Baseten says an AI-built inference engine for Qwen-3.6-35B-A3B ran up to 90% faster than vLLM on single-stream decoding in its tests on one NVIDIA B200. The company’s experiments also included a separate image-segmentation model, suggesting the approach may reach beyond one language model, though neither system is serving production traffic.
Baseten Watch analysis
What happened
In a Baseten account of its experiments, published on 2 October, the company describes using Claude Code with Fable 5 and the MetaInfer framework to build a custom inference engine, named VibeQwen. Baseten reports 1,792 tokens per second on its tested single-stream text workload, compared with 943 for a tuned vLLM 0.25.1 deployment. Time to first token fell from 28 milliseconds to 12 milliseconds. At concurrency 32, Baseten reports 71% more output throughput than vLLM.
A second experiment used SAM 3.1, an image-segmentation model, on one H100. Baseten says its custom server processed 91 images per second, 50% more than Meta’s reference server. The company says the VibeQwen project used about 200 B200 hours and 1.7 billion tokens; the SAM experiment took a couple of days and about 200 million tokens.
Why it matters
Inference engines are the machinery that turns a trained model into a service people can use. Baseten’s results suggest AI coding agents may help teams tailor that machinery to a particular model, accelerator and workload, where a general-purpose serving engine has to make broader trade-offs. Faster responses and more throughput could make a real difference to serving costs and latency if similar gains hold in other deployments.
Our read
The numbers are striking, but they come from Baseten’s own experiments, not an independent comparison. The company says neither custom system is running production traffic, and the tests use particular hardware, models and workloads. This is a promising demonstration of AI-assisted systems engineering, not yet evidence that teams should replace their serving stack. The most useful next step is reproduction on workloads that matter outside the experiment.
What to watch
- Whether independent teams reproduce the speed and throughput gains against carefully tuned baselines.
- How accuracy, reliability and engineering effort compare when the approach is tried on other models and hardware.
- Whether Baseten or other teams move AI-generated inference engines into production and share results.
Discussion spark: If an AI-built inference engine is faster on a team’s own workload, what evidence should it have to pass before replacing a mature general-purpose server?
Sources and evidence
- Agentic inference optimization: 50-90% faster engines (2 October 2026, 21:09 UTC)
not affiliated with or endorsed by Baseten