vLLM has substantially improved MiniMax M3 serving on AMD Instinct MI355X, with reported throughput gains of up to 4.45× per GPU in its benchmark checkpoints. The interesting part is the method: the team kept chasing whichever bottleneck surfaced next, from launch overhead and shared-expert work to sparse-attention metadata and speculative decoding.
vLLM Watch analysis
What happened
In the vLLM optimisation report, the project describes cumulative results from public SemiAnalysis InferenceX benchmarks. At concurrency 32, fixed-topology MXFP8 serving rose from 109.1 to 342.4 output tokens per second per GPU, or 3.14×, while median time to first token dropped from 1.46 to 0.67 seconds.
At concurrency 128, the same TP4/EP1 four-GPU path rose from 297.8 to 623.7 output tokens per second per GPU, or 2.09×. Median time to first token fell from 3.53 to 1.54 seconds, and mean time per output token fell from 100.7 to 48.8 milliseconds.
The headline 943.5 output tokens per second per GPU came from a later TP2/EP1 MXFP4 result. That is a denser deployment contract, not a clean fixed-topology speed comparison. A separate P/D-disaggregated configuration reached 6,370.5 total tokens per second per GPU at concurrency 512, with 1.32 seconds median time to first token.
Key findings
- Fuse the shared expert
Grouping the shared and routed experts removed a separate MLP path and intermediate traffic, improving output throughput by 30.2% at concurrency 1 in the reported PR tests. - Batch repeated sparse work
One workgroup can process all draft positions for a request, improving the MiniMax sparse-attention index kernel by up to 48.9% in the cited tests. - Reuse attention decisions
Sharing top-k block selections across adjacent sparse layers cut mean time per output token by about 10% at concurrency 1 and about 4% at high concurrency. - Move conversion out of serving
MXFP8 weight and scale reshuffling happens at model load, leaving the serving loop with a prepared layout. - Dispatch by real shapes
Separate prefill and decode tile choices improved TP8 8K/1K output throughput by 7.8% to 9.4% in the reported measurements. - Treat topology as a performance variable
TP2 and TP4 can select different attention backends, so changing tensor parallelism changes the operator path as well as collective size.
Why it matters
This is a useful reminder that frontier-model performance is often won in the plumbing. A sparse model can reduce arithmetic while adding metadata, routing and cache-management work; the best result depends on the shapes, concurrency and hardware path that actually run.
For operators, the practical lesson is to reproduce the benchmark contract before borrowing its number. The report explicitly distinguishes configured features from eligible and executed paths, and notes that its checkpoints are cumulative rather than neat one-change experiments.
Our read
This is strong engineering evidence and a more useful performance story than a single shiny throughput claim. Read the recipe as a map of optimisation ideas, then benchmark your own workload. “AMD support” is not one speed number, and TP4, TP2, standard decoding, EAGLE3 and disaggregated serving are different experiments.
What to watch
- Whether the reported gains reproduce outside the cited MiniMax M3 shapes and concurrency levels.
- Which sparse-attention and quantisation backends are selected automatically in future vLLM releases.
- Whether EAGLE3 gains remain worthwhile once draft-model cost and acceptance behaviour are included.
- Whether similar serving techniques transfer cleanly to other sparse mixture-of-experts models.
Discussion spark: Which matters more for your deployments: vLLM's raw throughput gains, its shape-aware dispatch approach, or the warning that benchmark topology changes the meaning of the number?
Sources and evidence
- Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X (10 September 2026, 00:00 UTC)
This is an independent WittyWires tracker and is not operated by, endorsed by, or affiliated with vLLM or the vLLM Project.