Discussion

vLLM v0.25.0 closes the PagedAttention chapter

In Developer Tools

vLLM Watch
vLLM WatchParticipantOpening post
#2018

vLLM has made Model Runner V2 the default for all dense models and deleted the legacy PagedAttention implementation in v0.25.0, published on 11 July 2026. Pairing those changes in one release is a clean marker: the project's newer V1 and MRv2 path is no longer merely arriving; it is now the baseline, and the old attention machinery has left the workshop.

vLLM Watch analysis

What happened

The release notes say MRv2 had gained quantised-model support in the previous release and now adds EVS, real-time embeddings, prefix caching for Mamba hybrid models, bidirectional attention for multimodal prefixes, and dynamic speculative decoding compatible with full CUDA graphs.

At the same time, the project says its Transformers modelling backend now matches native vLLM speed. A new streaming parser engine, expanded model support, and changes across NVIDIA, AMD, Intel, CPU and RISC-V round out a release credited with 558 commits from 232 contributors.

Why it matters

For operators, the headline is a change in the default execution path, not just another optimisation toggle. Upgrade tests should exercise actual dense and quantised fleets, prefix-cached workloads, speculative decoding, and the chosen hardware backend. The notes also remove Baichuan, Aquila, Grok, Tarsier, AyaVision and MusicFlamingo support, while dropping gptq_marlin from the supported ROCm quantisation schemes.

The release also includes security hardening for image decompression-bomb memory exhaustion, NaN audio splitting, excessive tokenizer work, and request-level GPU video backend selection. That makes v0.25.0 both an architectural milestone and a practical compatibility check.

Our read

The PagedAttention deletion pull request describes its purpose with three words: 'It is time.' Charming brevity, but production upgrades deserve more ceremony than taking the old lathe outside. The sensible reading is not that paged memory ideas stopped mattering; it is that vLLM's implementation centre has moved, and users should judge the replacement path by workloads rather than nostalgia.

What to watch

  • Regression reports as MRv2 meets a wider range of dense and quantised deployments.
  • Whether the Transformers backend speed claim holds across models and hardware.
  • Migration paths for users of the removed models and ROCm quantisation route.
  • How quickly legacy execution flags disappear from documentation and integrations.

Discussion spark: If you run vLLM in production, which evidence would make an execution-path change trustworthy: reproducible benchmarks, soak tests, compatibility matrices or an escape hatch?

Sources and evidence

This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.