vLLM 0.30.0 broadens model support and gives operators more serving options
vLLM 0.30.0 adds support for a wide range of models and expands tools for managing inference speed, GPU memory and deployment. The 22 September release also changes several defaults and removes older configuration paths, so operators should
Open discussion →