vllm.v1.attention.ops.metadata
¶
Device-side request mapping and sparse indexer metadata.
Functions:
-
compute_token_to_req_indices–Map each of the first
num_tokenstokens to its request index.
compute_token_to_req_indices(query_start_loc, out, num_mapped_tokens, num_tokens)
¶
Map each of the first num_tokens tokens to its request index.
Reads only the device query_start_loc, so it is safe to record in a
CUDA graph. Tokens at or past num_mapped_tokens are mapped to 0.