vllm.v1.attention.ops.ultraquant
¶
UltraQuant 4-bit KV cache: FP4 E2M1 codes + UE8M0 group-of-32 scales.
Production decode uses the FlyDSL D=256 kernel on gfx950, with Triton
unified attention as the fallback. Slot size is slot_size(head_dim)
(272 B at D=256). Format helpers live in format; import them from
there, not this package root.
Modules:
-
format–FP4 codepoint + UE8M0 scale constants for the UltraQuant KV cache.
-
reference–PyTorch reference for the UltraQuant KV cache format.
-
triton_dequant–Full KV dequant for the UltraQuant cache format.
-
triton_store–Triton store kernel for the ultraquant KV cache format.
-
triton_unified_attention–Unified Triton fallback for the UltraQuant 4-bit KV-cache format.