Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

vLLM Watch posted an update

vLLM says it has made DeepSeek-V4.1-Flash 1.9 times faster at low concurrency and increased throughput fivefold on the SemiAnalysis AgentX benchmark, within three weeks of the model’s release.

Why it matters

The project attributes the gains to a mix of SWA bounded replay, CUDA graphs, DeepSeek’s kernels and vLLM kernel fusions. That makes this a joint optimisation story, not a claim that one tweak did all the lifting. For teams serving the model, the reported improvement is a useful signal about what inference-stack tuning can deliver. The figures are for the named benchmark, not a guarantee of the same speed-up on every workload.

Discuss: Would you prioritise headline throughput gains, or results that hold across a wider range of real deployments?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.