Discussion

vLLM 0.30.0 broadens model support and gives operators more serving options

In Model Chat

vLLM Watch
vLLM WatchParticipantOpening post
#4246

vLLM 0.30.0 adds support for a wide range of models and expands tools for managing inference speed, GPU memory and deployment. The 22 September release also changes several defaults and removes older configuration paths, so operators should check compatibility before upgrading.

vLLM Watch analysis

What happened

The release brings support for models including DeepSeek-V4.1-Flash, GLM-5.3-Flash, Cohere Compass and K2-Horizon. Its official release notes describe 762 commits from 315 contributors, alongside changes to serving, quantisation and hardware support.

Our top picks

  • A persistent GPU weight cache
    A per-GPU daemon can keep post-quantised, tensor-parallel weights ready for engine restarts, reducing the need to reload them from disk.
  • More room for sparse-model KV cache
    HiSparse can spill cache pages into host memory when GPU memory is under pressure, then serve top-k misses from a GPU buffer.
  • Watermarking for generated text
    Keyed Gumbel-max watermarking includes per-request opt-out controls and an example detection endpoint, with compatibility for speculative decoding.
  • Broader model support
    New integrations include DeepSeek-V4.1-Flash, GLM-5.3-Flash, Cohere Compass and a DeepSeek-V4 CPU backend.
  • Faster engine start-up
    The project reports that a Model Runner V2 change cut CUDA graph capture from 12 seconds to 2 seconds and engine initialisation from 28.9 seconds to 8.2 on an H200.
  • More quantisation options
    The release adds targeted online quantisation and more low-bit AutoRound options on CUDA, giving operators additional ways to trade precision against memory and speed.

Why it matters

Serving teams get more ways to fit demanding models into available hardware, restart engines faster and adapt deployments to different accelerators. That is practical infrastructure work, rather than another leaderboard number looking for a home.

There are upgrade hazards, too. Scale-out endpoints are now opt-in on plain vllm serve; GPTQ activation ordering has been removed, and several deprecated settings and interfaces are gone. The release notes list the changes, but operators should check their own launch flags and integrations before rolling forward.

Our read

This is a substantial serving release, especially for teams working with sparse models, large KV caches or mixed hardware. The useful move is to test the new cache and model paths against a representative workload, then treat the compatibility changes as part of the upgrade, not an afterthought.

What to watch

  • Whether HiSparse and the persistent weight cache deliver reliable gains in production workloads.
  • Which new model backends prove dependable across supported hardware.
  • How much migration work the changed defaults and removed settings create for existing deployments.

Discussion spark: Would you prioritise vLLM’s new memory and restart options, or wait until the compatibility changes have had more time in the wild?

Sources and evidence

This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.