vLLM says it has made speculative decoding for the 2.8-trillion-parameter Kimi K3 practical to train, package and deploy. Its DSpark speculator reportedly lifts single-stream maths-reasoning interactivity from about 110 to 435 tokens per second per user, while delivering up to 3.5x higher output throughput at matched interactivity under concurrent load.
vLLM Watch analysis
What happened
The official vLLM blog post describes an extension of DFlash, a block-level speculative-decoding method. A lightweight draft model proposes several tokens, then the full model verifies them together. DSpark adds sequential correction, confidence estimates and hardware-aware scheduling to reduce the coherence problems that can make parallel drafts fall apart.
The released Kimi K3 speculator uses a five-layer, five-billion-parameter draft model and proposes eight tokens per decoding step. vLLM says it has also validated the approach on other target models, including Qwen3.6-35B-A3B and Gemma-4-31B-it.
Key findings
- Faster interactive generation
vLLM reports about 435 tokens per second per user, up from roughly 110, on its maths-reasoning test. - More throughput under load
The project reports up to 3.5x higher output throughput at matched interactivity. - Long-context headroom
On a 378,000-token LongBench-v2 prompt, vLLM reports up to 5.31 output tokens per decoding iteration. - Concurrent scaling
Increasing concurrency from one to 16 reportedly raises aggregate throughput from 177 to 683 tokens per second. - A real deployment recipe
vLLM provides a multi-node Docker configuration, with tensor parallelism across 16 GPUs and an eight-token DSpark setup.
Why it matters
Speculative decoding is meant to make a large model feel less large at serving time. The target model still checks the answer, but a smaller drafter can propose blocks of tokens in parallel, reducing how often the full model must work through generation one token at a time.
For operators, the important detail is that this is presented as infrastructure rather than a research demo. The recipe covers training, hidden-state extraction and multi-node serving, although the reported gains depend on the hardware, model, workload and acceptance behaviour described by vLLM.
Our read
This is a substantial serving upgrade if the reported numbers survive outside vLLM’s own test setup. The clever bit is not simply making a smaller drafter guess faster. DSpark tries to recover enough local coherence and scheduling awareness to make those guesses useful when the queue gets busy.
Operators running Kimi K3 should treat the post as a concrete starting point, then benchmark representative prompts and concurrency levels before changing production defaults. Four hundred tokens per second is a lovely headline. It is not a capacity plan.
What to watch
- Independent results for Kimi K3 DSpark on different GPU clusters.
- Whether the gains hold on coding, translation and long-context workloads beyond vLLM’s reported tests.
- How much extra cost the five-billion-parameter draft model adds to each serving configuration.
- Whether DSpark support broadens across open models and serving stacks.
Discussion spark: If you run large open models, would you trade extra draft-model complexity for this kind of serving speedup, and which workload would you benchmark first?
Sources and evidence
- How we trained the fastest DSpark for Kimi-K3 using GB300 NVL72 (15 September 2026, 00:00 UTC)
This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.