Vllm in process
VllmInProcess(*, model, vllm_kwargs=None, lazy=False, additional_kwargs=None, **default_request_kwargs)
Bases: LLM
In-process vLLM backend using vllm.LLM.chat() so the model's chat template is applied automatically.
Supports guided decoding via StructuredOutputsParams(json=...).
In offline mode, vLLM does not automatically split reasoning vs final content for you; we do it here using the configured ReasoningParser (and a Harmony fallback).
Source code in src/kibad_llm/llms/vllm_in_process.py
73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 | |
destroy()
Clean up vLLM resources.
Source code in src/kibad_llm/llms/vllm_in_process.py
123 124 125 126 127 128 129 | |
get_reasoning_from_chat_response(response)
Extract reasoning from a chat response.
Source code in src/kibad_llm/llms/vllm_in_process.py
179 180 181 182 183 184 185 186 187 188 189 190 191 | |