vllm.v1.attention.ops.ultraquant.format
¶
FP4 codepoint + UE8M0 scale constants for the UltraQuant KV cache.
K/V codes are FP4 E2M1. Per-group scales are UE8M0 (one byte, a power of
two): s = 2^round(log2(c · absmax)) with c = 0.156. Decode
consumes Q as FP8 E4M3 so scaled F8F6F4 MFMA can run natively on CDNA4.
AoS slot layout (group_size=32), per (token, head): bytes [0 .. D/2) : K codes (FP4 nibbles, 2/byte) bytes [D/2 .. D/2 + Gk) : K scales (UE8M0, 1 byte × Gk groups) bytes [D/2 + Gk .. D + Gk) : V codes bytes [D + Gk .. D + 2*Gk) : V scales D=256 → Gk=8 → 272 B/slot. D=128 → Gk=4 → 136 B/slot.
No per-token norm-fold. No V rotation. K is Hadamard-rotated at store.
Functions:
-
get_constant_c–Return the fixed UltraQuant scale constant.
-
get_group_size–Return the fixed group size required by scaled MFMA.
-
k_codes_bytes–Bytes for one head's packed K codes.
-
k_scales_bytes–Bytes for one head's K scales (one E8M0 byte per group).
-
slot_size–Bytes per (token, head) slot.
-
ue8m0_decode–Decode a UE8M0 byte back to fp32. 0 → 0.0 (zero sentinel).
-
ue8m0_encode–Snap a positive fp32 scale
sto the nearest power of 2 and
Attributes:
-
DEFAULT_CONSTANT_C(float) –MSE-optimal scale constant;
s_raw = c * absmax, then UE8M0-snapped. -
FP4_MAX(float) –Largest representable FP4 magnitude.
-
GROUP_SIZE(int) –Elements per group; one E8M0 scale per group. Fixed at 32 because the
DEFAULT_CONSTANT_C = 0.156
module-attribute
¶
MSE-optimal scale constant; s_raw = c * absmax, then UE8M0-snapped.
FP4_MAX = 6.0
module-attribute
¶
Largest representable FP4 magnitude.
GROUP_SIZE = 32
module-attribute
¶
Elements per group; one E8M0 scale per group. Fixed at 32 because the AMD scaled F8F6F4 MFMA instruction also consumes one E8M0 scale per 32 elements along K. Other group sizes would force a fallback to plain MFMA + accumulator-side scale multiply (no hardware fast-path).
get_constant_c()
¶
get_group_size()
¶
k_codes_bytes(head_dim)
¶
k_scales_bytes(head_dim, group_size=None)
¶
slot_size(head_dim, group_size=None)
¶
Bytes per (token, head) slot.
Source code in vllm/v1/attention/ops/ultraquant/format.py
ue8m0_decode(byte)
¶
ue8m0_encode(s)
¶
Snap a positive fp32 scale s to the nearest power of 2 and
encode as a UE8M0 byte. s <= 0 encodes as 0 (zero sentinel).