Skip to content

vllm.entrypoints.generate.structured_decisions.protocol

Request and response models of the structured decisions API (/v1/systemone).

Classes:

  • ReadPromptRequest –

    Chat options for one read's prompt: the state and the question in the

ReadPromptRequest

Bases: OpenAIBaseModel

Chat options for one read's prompt: the state and the question in the user turn, ending at the generation prompt so the label is the reply's first token. Thinking is off unless the request turns it on, or the label would follow a thought rather than start the reply.

Source code in vllm/entrypoints/generate/structured_decisions/protocol.py
class ReadPromptRequest(OpenAIBaseModel):
    """Chat options for one read's prompt: the state and the question in the
    user turn, ending at the generation prompt so the label is the reply's
    first token. Thinking is off unless the request turns it on, or the label
    would follow a thought rather than start the reply."""

    chat_template_kwargs: dict[str, Any] | None = None

    def build_chat_params(
        self,
        default_template: str | None,
        default_template_content_format: ChatTemplateContentFormatOption,
    ) -> ChatParams:
        return ChatParams(
            chat_template=default_template,
            chat_template_content_format=default_template_content_format,
            chat_template_kwargs=merge_kwargs(
                merge_kwargs({"enable_thinking": False}, self.chat_template_kwargs),
                dict(add_generation_prompt=True, continue_final_message=False),
            ),
        )

    def build_tok_params(self, model_config: ModelConfig) -> TokenizeParams:
        return TokenizeParams(
            max_total_tokens=model_config.max_model_len,
            max_output_tokens=1,
            add_special_tokens=False,
        )