vllm.v1.attention.ops.ultraquant.triton_dequant
¶
Full KV dequant for the UltraQuant cache format.
Used by continuation prefill: when a chunk brings many new query tokens on top of a long cached prefix, dequanting the prefix once and running a dense prefill kernel beats replaying the decode kernel per query token.
Reads FP4 codes plus UE8M0 group scales and writes K (Hadamard-rotated, as stored) and V (raw) into pre-allocated fp16/bf16 buffers.
Functions:
-
ultraquant_full_dequant_kv–Dequant
alloc_lencached positions intok_out/v_out.
_get_fp4_decode_table(device, dtype)
¶
16-entry FP4 E2M1 bit-pattern -> value table (cached per device).
Source code in vllm/v1/attention/ops/ultraquant/triton_dequant.py
ultraquant_full_dequant_kv(kv_cache, block_table, k_out, v_out, alloc_len)
¶
Dequant alloc_len cached positions into k_out / v_out.
k_out / v_out are [B, Hk, alloc_len, D] in fp16 or bf16.