vLLM has released version 0.25.0, making Model Runner V2 the default execution path for dense models. The release matters because this is a change to the project’s standard serving route, wrapped in a notably broad set of additions across models, inference and deployment.
vLLM Watch analysis
What happened
The 11 July release contains 558 commits from 232 contributors. Alongside the default runner change, it adds new model support, a unified streaming parser engine for tool calls and reasoning, broader speculative decoding, and a long list of accelerator, quantisation and distributed-serving improvements.
Key findings
- Model Runner V2 becomes standard
Dense models now use the newer execution path by default, giving the project a clearer centre of gravity. - The model zoo gets wider
New entries include LLaVA-OneVision-2, Unlimited OCR, MOSS-Transcribe-Diarize, GLM-5 and DeepSeek-V3.2. - Parsing gets a proper engine
A unified streaming framework handles tool-call and reasoning parsers, including several newly supported model families. - Speculative decoding gets more flexible
Universal speculative decoding supports heterogeneous vocabularies, alongside new drafter and backend work. - Hardware coverage keeps spreading
The notes include updates for NVIDIA Blackwell, AMD ROCm, Intel XPU, CPUs, RISC-V and PowerPC. - Quantisation moves in more directions
The release expands support across several low-bit formats, FP8, NVFP4, Marlin and KV-cache paths.
Why it matters
For teams running vLLM in production, the default-path change is the headline. It suggests the newer runner has moved from promising work to the project’s expected operating model, while the surrounding changes target the practical frictions of serving different models on different hardware.
The breadth is useful, but it also means the upgrade is not one neat switch. Operators should test their specific model, backend and quantisation combination. Inference infrastructure remains a game of compatibility whack-a-mole, only now the mallet is distributed across several accelerators.
Our read
This is a substantial platform release, not a cosmetic version bump. Read the notes with your deployment matrix open, then stage the upgrade against the models and hardware you actually run.
What to watch
- Whether Model Runner V2 changes latency, throughput or stability for existing dense-model deployments.
- Which of the new model integrations are production-ready in your chosen backend.
- Real-world gains from the release’s speculative-decoding and kernel work.
- Follow-up fixes as users exercise the expanded hardware and quantisation combinations.
Discussion spark: If you run vLLM, which matters more for your next upgrade: the Model Runner V2 default, new model coverage, or hardware and quantisation support?
Sources and evidence
- Release v0.25.0 · vllm-project/vllm · GitHub (2 September 2026, 13:37 UTC)
This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.