vLLM 0.31.0, released on 5 October, brings serving and performance changes for newer models, plus a GPU-resident weight cache designed to speed engine restarts. For operators, the gains come with several compatibility changes that make this one worth testing before upgrading production.
vLLM Watch analysis
What happened
The project’s release notes describe 717 commits from 307 contributors. Among the changes are new optimisations for DeepSeek-V4.1-Flash, an expanded preload tool for keeping post-quantised weights in GPU memory between engine restarts, and additional speculative-decoding options. The release also adds model capabilities and deployment options across CUDA, ROCm, XPU and CPU.
Our top picks
- A weight cache that survives engine restarts
The preload CLI launches a daemon to keep post-quantised weights in GPU memory, with added support for data parallelism and MTP draft models. - More DeepSeek-V4.1-Flash optimisations
The release adds fused attention and other serving changes, with FlashMLA and NVFP4 compressed KV cache set as the SM100 default. - More speculative-decoding routes
Model Runner V2 gains draft-model support, while new and updated drafters extend options for serving teams. - A cap on active requests
–max-num-active-seqs lets operators limit running requests separately from the maximum number of sequences. - Tighter handling of request-supplied multimodal settings
Per-request multimodal kwargs are rejected unless –trust-request-mm-kwargs is set, making the trust boundary an explicit configuration choice.
Why it matters
This is a substantial serving release: it offers more ways to manage restart time, memory pressure, scheduling and decoding for demanding models. The cache daemon could be especially useful where loading weights repeatedly is a deployment bottleneck; the release notes also describe experimental engine snapshots, so that restoration path is not yet the same as a routine production feature.
Upgrading is not just a matter of collecting the performance improvements. The notes list changed or removed settings, including the removal of tokenizermode="slow", a renamed Mamba cache option and changes to online quantisation configuration. Operators should check their launch flags and integrations before rolling forward.
Our read
There is real operator value here, particularly for teams serving DeepSeek models or trying to make restarts less of an occasion. Start with a representative workload, test the cache and scheduling changes, and treat the migration notes as part of the release rather than decorative reading.
What to watch
- Whether the preload cache delivers reliable restart improvements across real deployments.
- How the new scheduling and speculative-decoding options behave under varied workloads.
- Whether the changed flags and multimodal trust setting require migration work in existing services.
Discussion spark: Would you adopt the new weight-cache and serving controls early, or wait for them to prove themselves in production?
Sources and evidence
- Releases · vllm-project/vllm · GitHub (Publication date not supplied)
This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.