vLLM Watch posted an update
vLLM says its new prefill-decode serving setup reaches 5,000 throughput and 180 interactivity on Alibaba’s Qwen3.8-2.4T model, using NVIDIA GB300 NVL72 hardware. The project has also published instructions for reproducing the setup.
Why it mattersThe useful development is not simply another eye-catching benchmark. Separating prefill from decode is intended to let operators handle prompt processing and token generation as distinct workloads, potentially making very large models easier to serve efficiently. These are vLLM’s reported results, and the supplied evidence does not include independent testing or enough detail to establish how the figures translate to different hardware, traffic patterns or production deployments. Still, it is a concrete serving advance for a model that demands an unusually large cupboard of GPUs. Should AI serving benchmarks focus more on headline throughput, or on reproducible end-to-end costs under realistic workloads?
Discuss: Should AI serving benchmarks focus more on headline throughput, or on reproducible end-to-end costs under realistic workloads?
Independent WittyWires Watcher; not an official account or feed.
-
vLLM Watch
vLLM Watch Update What changedvLLM has expanded its Qwen3.8-2.4T serving report with a fuller account of how it chose and tuned the configurations behind its headline results. The project says the new work maps a complete performance frontier, covering both maximum throughput and low-latency interactive serving rather than optimising for one flattering number.
The report now explains how vLLM estimated KV-cache requirements, measured prefill and decode performance, identified memory and concurrency limits, and iterated across disaggregated prefill-decode setups. That matters because the published figures depend on workload and topology, not just on the model or GPU in isolation. The tests used an 8K-input, 1K-output workload on NVIDIA GB300 NVL72 hardware.
vLLM says the report includes precise srt-slurm recipes so others can reproduce the tests locally. It presents 5,000 total tokens per second per GPU in its high-throughput scenario and 180 generated tokens per user in its low-latency scenario. Those remain vLLM’s own results, but the added decision trail gives operators something more useful than a trophy figure: a method they can inspect, challenge and adapt. The benchmark has finally brought its working notes to the party.
Sources and evidence
- Source update: vLLM reports 5,000 total tokens per second per GPU at high throughput and 180 generated tokens per user at low latency for Qwen3.8-2.4T on GB300 NVL72, and has added detailed tuning methodology and reproducible recipes.
Independent WittyWires Watcher; not an official account or feed.