vllm.models.deepseek_v41.common.ops.fused_compress_quant_cache
¶
V4.1 state saving/compression and independently schedulable cache insertion.
Functions:
-
fused_save_compress_norm–Pool each closed group into a normalized BF16 latent; save FP32 states.
-
rope_quant_insert–Apply GPT-J RoPE and publish a latent to the compressed KV cache.
-
save_ring_rows–Store FP32 [kv, score] rows to their ring slots, skipping slot -1.
_rope_quant_insert_mxfp8_kernel(latent, positions, cos_sin, cache, cache_slots, COS_STRIDE, CACHE_STRIDE, CACHE_BLOCK, COMPRESS_RATIO, SANITIZE_CACHE_NANS)
¶
V4.1 record: RoPE first, then MXFP8-quantize all 512 dims.
The RoPE dims are quantized here too, so unlike the V4 kernel the rotation has to happen before the scales are picked.
Source code in vllm/models/deepseek_v41/common/ops/fused_compress_quant_cache.py
_rope_quant_insert_nvfp4_kernel(latent, positions, cos_sin, cache, cache_slots, COS_STRIDE, CACHE_STRIDE, CACHE_BLOCK, COMPRESS_RATIO, SANITIZE_CACHE_NANS)
¶
V4.1 NVFP4 record: RoPE, then e2m1 with one e4m3 scale per 16 dims.
The scale is amax / 6 (6 is e2m1's largest magnitude) clamped to the
e4m3 range, with no per-tensor scale on top.
Source code in vllm/models/deepseek_v41/common/ops/fused_compress_quant_cache.py
fused_save_compress_norm(kv_score, positions, state_cache, slot_mapping, query_start_loc, token_to_req_indices, rms_norm_weight, rms_norm_eps, compress_ratio, latent_out, prev_rows=None, prev_row_indices=None)
¶
Pool each closed group into a normalized BF16 latent; save FP32 states.
The latent feeds the main-cache insert and the indexer K path, which the attention layer schedules on separate streams.
Ratio 2 keeps one ring block per request holding the open group's rows:
position p lives in row p % capacity and slot_mapping encodes
block * capacity + p % capacity. The grid has one program per request
followed by one per pair of packed tokens. A request program handles the
group that the chunk's first token closes with its predecessor's ring row,
then stores the chunk's last capacity rows to the ring; because the
same program does both, ring reads and writes never race. A pair program
handles the group that ends inside its pair, reading both rows from the
raw input. Ratio 1 has no ring and one program per token; slot_mapping
then only marks valid tokens.
Parameters:
-
(kv_score¶Tensor) –FP32 [tokens, 512] for CR1, [tokens, 1024] for CR2.
-
(positions¶Tensor) –Absolute positions of the packed request tokens.
-
(state_cache¶Tensor | None) –Ring FP32 [blocks, capacity, 1024] KV/score states (CR2).
-
(slot_mapping¶Tensor) –Ring slots (CR2) or main-cache slots (CR1).
-
(query_start_loc¶Tensor | None) –[num_reqs + 1] token offsets of each request's chunk.
-
(token_to_req_indices¶Tensor | None) –Request indices for the packed token rows.
-
(rms_norm_weight¶Tensor) –BF16 [512] normalization weight.
-
(rms_norm_eps¶float) –RMSNorm epsilon.
-
(compress_ratio¶int) –Group size, either 1 or 2.
-
(latent_out¶Tensor) –BF16 [tokens, 512], written only at valid group boundaries.
-
(prev_rows¶Tensor | None, default:None) –FP32 [rows, 1024] rows gathered from all PCP ranks (CR2). A chunk reads its predecessor here where
prev_row_indicesis set, and the ring is left tosave_ring_rows. -
(prev_row_indices¶Tensor | None, default:None) –[num_reqs] row in
prev_rowsholding the token before each chunk's first token, or -1 to read the ring.
Source code in vllm/models/deepseek_v41/common/ops/fused_compress_quant_cache.py
22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 | |
rope_quant_insert(latent, positions, cos_sin_cache, kv_cache, slot_mapping, compress_ratio, fp8_scale=None)
¶
Apply GPT-J RoPE and publish a latent to the compressed KV cache.
The BF16 latent supplies both NoPE quantization and RoPE input. It is read
only for valid slots at group boundaries. The cache dtype selects the
layout: uint8 is a paged FlashMLA layout, whose record the per-token
byte width names -- 584 B for V4 (576 value bytes and eight segregated
UE8M0 scale bytes, including one zero padding scale), 528 B for V4.1
(512 MXFP8 value bytes covering the RoPE dims too, then 16 UE8M0 scales of
32 dims each), or 288 B for V4.1 NVFP4 (256 bytes of e2m1 pairs then 32
e4m3 scales of 16 dims each), which only the compressed cache uses.
bfloat16 and float8_e4m3fn are the plain [448 NoPE | 64 RoPE] rows
read by FlashInfer, the latter scaled by the per-tensor fp8_scale.
Source code in vllm/models/deepseek_v41/common/ops/fused_compress_quant_cache.py
295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 | |
save_ring_rows(rows, slot_mapping, state_cache)
¶
Store FP32 [kv, score] rows to their ring slots, skipping slot -1.