Skip to content

vllm.model_executor.layers.quantization.utils.humming

Humming quantization integration.

Modules:

  • activation –

    MoE activation expressions for Humming input processing.

  • linear –

    Prepare and execute Humming linear layers.

  • moe –

    Configure, prepare weights for, and assemble Humming MoE kernels.

  • priority –

    Humming kernel selection priority.

  • schema –

    Map Humming schemas and handle shared checkpoint quantization settings.

Functions:

convert_linear_layer_to_humming_standard(layer, name_map)

Rename/reshape a linear layer's quantized params (the canonical MPLinear layout: weight_packed int32 + weight_scale) into the parameter names and layout humming's weight schema expects (weight / weight_scale).

Source code in vllm/model_executor/layers/quantization/utils/humming/linear.py
def convert_linear_layer_to_humming_standard(
    layer: LinearBase, name_map: dict[str, str]
):
    """Rename/reshape a linear layer's quantized params (the canonical MPLinear
    layout: ``weight_packed`` int32 + ``weight_scale``) into the parameter names
    and layout humming's weight schema expects (``weight`` / ``weight_scale``)."""
    for name, checkpoint_name in name_map.items():
        tensor = getattr(layer, checkpoint_name)
        delattr(layer, checkpoint_name)

        if name == "weight":
            input_dim = getattr(tensor, "input_dim", 1)
            output_dim = getattr(tensor, "output_dim", 0)

            if input_dim == 0 and output_dim == 1:
                tensor = tensor.transpose(1, 0).contiguous()
            else:
                assert output_dim == 0 and input_dim == 1

            tensor = tensor.view(tensor.size(0), -1).view(torch.int32)
        elif name in ["weight_scale", "zero_point"]:
            if getattr(tensor, "output_dim", 0) == 1:
                tensor = tensor.transpose(0, 1).contiguous()
            if tensor.ndim == 1:
                tensor = tensor.unsqueeze(1)

            tensor = tensor.view(torch.int32) if name == "zero_point" else tensor

        if isinstance(tensor, torch.nn.Parameter):
            param = tensor
        else:
            param = torch.nn.Parameter(tensor, requires_grad=False)

        setattr(layer, name, param)

convert_to_humming_moe_kernel_format(layer, quant_config=None, sublayer_configs=None, weight_schema=None, input_schema=None, force_weight_schema=None, allow_input_schema_fallback=True)

Convert MoE weights from checkpoint format to Humming kernel format.

This function processes weights for each sublayer (w13, w2) by: 1. Converting from checkpoint format to humming format if needed 2. Force requanting if a different quantization schema is specified 3. Preparing layer metadata for the Humming kernel 4. Transforming weights for inference

Parameters:

  • layer

    (RoutedExperts) –

    The RoutedExperts layer containing weights to process

  • quant_config

    (dict | None, default: None ) –

    Optional quantization config dict. Required if weight_schema or input_schema are None. Used to build schemas via BaseWeightSchema.from_config().

  • sublayer_configs

    (dict[str, Any] | None, default: None ) –

    Optional configuration dict for each sublayer (w13, w2). Each config must have "shape_n" and "shape_k" keys. If None, configs are built from layer.moe_config properties.

  • weight_schema

    (Any | None, default: None ) –

    Optional initial weight quantization schema. If None, built from quant_config.

  • input_schema

    (Any | None, default: None ) –

    Optional initial input quantization schema. If None, built from quant_config or env vars.

  • force_weight_schema

    (Any | None, default: None ) –

    Optional schema to force requantization to

  • allow_input_schema_fallback

    (bool, default: True ) –

    Whether incompatible input schemas may be replaced.

Side effects
  • Modifies layer parameters in place
  • Sets layer.weight_schemas and layer.input_schemas
  • Sets layer.humming_configs for quant config construction
Source code in vllm/model_executor/layers/quantization/utils/humming/moe.py
def convert_to_humming_moe_kernel_format(
    layer: "RoutedExperts",
    quant_config: dict | None = None,
    sublayer_configs: dict[str, Any] | None = None,
    weight_schema: Any | None = None,
    input_schema: Any | None = None,
    force_weight_schema: Any | None = None,
    allow_input_schema_fallback: bool = True,
) -> dict[str, "LayerConfig"]:
    """Convert MoE weights from checkpoint format to Humming kernel format.

    This function processes weights for each sublayer (w13, w2) by:
    1. Converting from checkpoint format to humming format if needed
    2. Force requanting if a different quantization schema is specified
    3. Preparing layer metadata for the Humming kernel
    4. Transforming weights for inference

    Args:
        layer: The RoutedExperts layer containing weights to process
        quant_config: Optional quantization config dict. Required if weight_schema
                     or input_schema are None. Used to build schemas via
                     BaseWeightSchema.from_config().
        sublayer_configs: Optional configuration dict for each sublayer (w13, w2).
                         Each config must have "shape_n" and "shape_k" keys.
                         If None, configs are built from layer.moe_config properties.
        weight_schema: Optional initial weight quantization schema.
                      If None, built from quant_config.
        input_schema: Optional initial input quantization schema.
                     If None, built from quant_config or env vars.
        force_weight_schema: Optional schema to force requantization to
        allow_input_schema_fallback: Whether incompatible input schemas may be replaced.

    Side effects:
        - Modifies layer parameters in place
        - Sets layer.weight_schemas and layer.input_schemas
        - Sets layer.humming_configs for quant config construction

    """
    # Build schemas from quant_config if not provided
    has_bias = layer.moe_config.has_bias
    num_experts = layer.moe_config.num_local_experts
    param_dtype = layer.params_dtype

    if weight_schema is None or input_schema is None:
        if quant_config is None:
            raise ValueError(
                "Must provide either weight_schema/input_schema or quant_config"
            )

        from vllm.utils.humming import BaseWeightSchema, HummingInputSchema

        if weight_schema is None:
            weight_schema = BaseWeightSchema.from_config(quant_config)

        if input_schema is None:
            input_quant_config = (envs.VLLM_HUMMING_INPUT_QUANT_CONFIG or {}).copy()
            if humming_is_layer_skipped(input_quant_config, layer.layer_name):
                input_schema = HummingInputSchema()
            else:
                # TODO: read input_quant_config from quant_config
                input_quant_config = humming_schema.resolve_humming_layer_config(
                    input_quant_config, layer.layer_name
                )
                allow_input_schema_fallback = input_quant_config.pop(
                    "allow_fallback", False
                )
                input_schema = HummingInputSchema.from_config(input_quant_config)

    # Build sublayer configs from layer properties if not provided
    if sublayer_configs is None:
        is_gated = layer.moe_config.activation.is_gated
        intermediate_size = layer.moe_config.intermediate_size_per_partition
        sublayer_configs = {
            "w13": {
                "shape_n": intermediate_size * (2 if is_gated else 1),
                "shape_k": layer.moe_config.hidden_dim,
            },
            "w2": {
                "shape_n": layer.moe_config.hidden_dim,
                "shape_k": intermediate_size,
            },
        }

    layer.weight_schemas = {}
    layer.input_schemas = {}
    humming_configs = {}

    for sublayer_name, configs in sublayer_configs.items():
        final_weight_schema, final_input_schema, humming_config = (
            _process_single_sublayer(
                layer=layer,
                sublayer_name=sublayer_name,
                shape_n=configs["shape_n"],
                shape_k=configs["shape_k"],
                weight_schema=weight_schema,
                input_schema=input_schema,
                has_bias=has_bias,
                num_experts=num_experts,
                param_dtype=param_dtype,
                force_weight_schema=force_weight_schema,
                allow_input_schema_fallback=allow_input_schema_fallback,
            )
        )

        layer.weight_schemas[sublayer_name] = final_weight_schema
        layer.input_schemas[sublayer_name] = final_input_schema
        humming_configs[sublayer_name] = humming_config

    layer.humming_configs = humming_configs
    return humming_configs

prefers_humming(compute_capability=None)

Whether Humming outranks Marlin. A missing compute capability uses the current device.

Source code in vllm/model_executor/layers/quantization/utils/humming/priority.py
def prefers_humming(compute_capability: int | None = None) -> bool:
    """Whether Humming outranks Marlin. A missing compute capability uses the
    current device."""
    if compute_capability is None:
        from vllm.platforms import current_platform

        if current_platform.is_cuda():
            cc = current_platform.get_device_capability()
            compute_capability = cc.to_int() if cc is not None else None
    return compute_capability in _HUMMING_PREFERRED_CAPABILITIES

prioritize_humming(kernels, compute_capability=None)

Move Humming directly ahead of Marlin where Humming is preferred.

Every other entry keeps its relative order and the input list is not modified. Match kernel class names or MoE backend enum names.

Source code in vllm/model_executor/layers/quantization/utils/humming/priority.py
def prioritize_humming(
    kernels: list[_KernelT],
    compute_capability: int | None = None,
) -> list[_KernelT]:
    """Move Humming directly ahead of Marlin where Humming is preferred.

    Every other entry keeps its relative order and the input list is not
    modified. Match kernel class names or MoE backend enum names.
    """
    if not prefers_humming(compute_capability):
        return kernels

    names = [
        (kernel.name if isinstance(kernel, Enum) else kernel.__name__).lower()
        for kernel in kernels
    ]
    humming = next((i for i, name in enumerate(names) if "humming" in name), None)
    marlin = next((i for i, name in enumerate(names) if "marlin" in name), None)
    if humming is None or marlin is None or humming < marlin:
        return kernels
    kernels = kernels.copy()
    kernels.insert(marlin, kernels.pop(humming))
    return kernels

select_humming_moe_experts(config, weight_key, activation_key)

Select the primary Humming MoE Experts class Note: Shape-specific fallbacks may still occur at runtime.

Source code in vllm/model_executor/layers/quantization/utils/humming/moe.py
def select_humming_moe_experts(
    config: FusedMoEConfig,
    weight_key: QuantKey | None,
    activation_key: QuantKey | None,
) -> type[mk.FusedMoEExperts] | None:
    """Select the primary Humming MoE Experts class
    Note: Shape-specific fallbacks may still occur at runtime.
    """
    if not has_humming():
        return None

    # NOTE: the kernels are selected in the following order.
    AVAILABLE_EXPERTS: list[type[mk.FusedMoEExperts]] = [
        BatchedHummingGroupedExperts,
        HummingGroupedExperts,
        HummingIndexedExperts,
    ]

    # NOTE(rob): We need to peak into the P/F selection to determine
    # if we are using the batched or standard expert format, which
    # if not ideal. Once we unify TP + DP/EP, we can select P/F first.
    activation_format = (
        mk.FusedMoEActivationFormat.BatchedExperts
        if config.moe_parallel_config.use_batched_activation_format
        else mk.FusedMoEActivationFormat.Standard
    )

    def _make_log_backend(experts_cls: type[mk.FusedMoEExperts]):
        return f"Using {experts_cls.__name__} Humming MoE backend."

    def _make_log_unsupported(
        experts_cls: type[mk.FusedMoEExperts], reason: str | None
    ) -> str:
        if reason:
            return (
                f"Humming MoE experts {experts_cls.__name__} does not support the "
                f"deployment configuration since {reason}."
            )
        else:
            return (
                f"Humming MoE experts '{experts_cls.__name__}' does not support the "
                "deployment configuration."
            )

    for k_cls in AVAILABLE_EXPERTS:
        supported, reason = k_cls.is_supported_config(
            k_cls,
            config,
            weight_key,
            activation_key,
            activation_format,
        )
        if supported:
            logger.info_once(_make_log_backend(k_cls))
            return k_cls
        else:
            logger.debug_once(_make_log_unsupported(k_cls, reason))

    return None