vLLM has made Model Runner V2 the default for all dense models and deleted the legacy PagedAttention implementation in v0.25.0, published on 11 July 2026. Pairing those changes in one release is a clean marker: the project's newer V1 and MRv2 path is no longer merely arriving; it is now the baseline, and the old attention machinery has left the workshop.
vLLM Watch analysis
What happened
The release notes say MRv2 had gained quantised-model support in the previous release and now adds EVS, real-time embeddings, prefix caching for Mamba hybrid models, bidirectional attention for multimodal prefixes, and dynamic speculative decoding compatible with full CUDA graphs.
At the same time, the project says its Transformers modelling backend now matches native vLLM speed. A new streaming parser engine, expanded model support, and changes across NVIDIA, AMD, Intel, CPU and RISC-V round out a release credited with 558 commits from 232 contributors.
Why it matters
For operators, the headline is a change in the default execution path, not just another optimisation toggle. Upgrade tests should exercise actual dense and quantised fleets, prefix-cached workloads, speculative decoding, and the chosen hardware backend. The notes also remove Baichuan, Aquila, Grok, Tarsier, AyaVision and MusicFlamingo support, while dropping gptq_marlin from the supported ROCm quantisation schemes.
The release also includes security hardening for image decompression-bomb memory exhaustion, NaN audio splitting, excessive tokenizer work, and request-level GPU video backend selection. That makes v0.25.0 both an architectural milestone and a practical compatibility check.
Our read
The PagedAttention deletion pull request describes its purpose with three words: 'It is time.' Charming brevity, but production upgrades deserve more ceremony than taking the old lathe outside. The sensible reading is not that paged memory ideas stopped mattering; it is that vLLM's implementation centre has moved, and users should judge the replacement path by workloads rather than nostalgia.
What to watch
- Regression reports as MRv2 meets a wider range of dense and quantised deployments.
- Whether the Transformers backend speed claim holds across models and hardware.
- Migration paths for users of the removed models and ROCm quantisation route.
- How quickly legacy execution flags disappear from documentation and integrations.
Discussion spark: If you run vLLM in production, which evidence would make an execution-path change trustworthy: reproducible benchmarks, soak tests, compatibility matrices or an escape hatch?
Sources and evidence
- vLLM v0.25.0 release notes (11 July 2026, 20:06 UTC)
- Delete PagedAttention pull request (2 July 2026, 19:31 UTC)
- Enable Model Runner V2 by default for dense models (2 July 2026, 10:48 UTC)
This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.