vllm.v1.attention.backends.mla.prefill.selector
¶
Selector for MLA prefill backends.
This module provides functions for selecting the appropriate MLA prefill backend based on device capabilities and configuration.
Classes:
-
MLAPrefillSelectorConfig–Hashable configuration for MLA prefill backend selection.
Functions:
-
get_mla_prefill_backend–Select the MLA prefill backend based on configuration and device.
MLAPrefillSelectorConfig
¶
Bases: NamedTuple
Hashable configuration for MLA prefill backend selection.
This is analogous to AttentionSelectorConfig and contains model-specific configuration needed to select an MLA prefill backend, extracted from VllmConfig into a hashable form for caching.
Source code in vllm/v1/attention/backends/mla/prefill/selector.py
_auto_select_mla_prefill_backend(selector_config)
cached
¶
Auto-select the best available MLA prefill backend.
Parameters:
-
(selector_config¶MLAPrefillSelectorConfig) –Hashable configuration for backend selection.
Returns:
-
type[MLAPrefillBackend]–The selected prefill backend class.
Raises:
-
ValueError–If the platform has no valid MLA prefill backend.
Source code in vllm/v1/attention/backends/mla/prefill/selector.py
get_mla_prefill_backend(vllm_config)
¶
Select the MLA prefill backend based on configuration and device.
This function first checks for explicit user preferences via mla_prefill_backend in AttentionConfig, then falls back to automatic priority-based selection.
Parameters:
-
(vllm_config¶VllmConfig | None) –The vLLM configuration. May be None before model resolution, in which case defaults are used.
Returns:
-
type[MLAPrefillBackend]–The selected prefill backend class.