Eight vision-language models, one consumer GPU and a useful warning for anyone shopping by benchmark: a model’s tokens-per-second figure may tell you more about its serving setup than its underlying speed. A Hugging Face community article by Aakash Gupta compares ranking accuracy and throughput for video-question-answering workloads on an RTX 3090.
Watch Desk analysis
What happened
Gupta tested eight models on one 24GB RTX 3090, measuring ranking accuracy against 400 multiple-choice queries and throughput on real queries. The paper reports that accuracy and prefill throughput had a strong negative correlation, Spearman ρ = −0.71. But parameter count was closely related to both; after controlling for it, the relationship between accuracy and prefill throughput fell to ρ = −0.36, which the paper says was not distinguishable from zero in this eight-model sample.
Three results complicate the headline trade-off:
Key findings
- Decode speed is not a model-size shortcut
Decode throughput had almost no correlation with parameter count in this test, so the serving configuration mattered more. - Short answers spend much of their time waiting
Six models produced 36–44 tokens for a ranking query; 60–89% of per-query wall-clock time passed before the first token. - Fast requests can mean slower jobs
Two models with statistically indistinguishable accuracy differed 2.2× in per-request prefill throughput, yet the faster one took 2.5× longer over the full 7,329-call workload.
Why it matters
A leaderboard can make a model look fast or slow while hiding what the workload, prompt length and serving stack are doing. For teams choosing a model, the useful comparison is not one attractive speed number beside one accuracy score. It is the cost and quality of the actual job they need to run.
The study is a single-GPU comparison, not a universal ranking. Gupta notes that serving paths and settings differed between models, and that some of those differences affect the results. That caveat is central, not fine print: the paper’s point is partly that configuration can distort the comparison.
Our read
This is a worthwhile nudge to benchmark the job, not the marketing-friendly metric. If you are comparing vision-language models, check what was measured, how it was served and whether the test resembles your workload. “Tokens per second” is a number; it is not, on its own, a purchasing decision.
What to watch
- Whether future comparisons test more GPUs, serving configurations and concurrent workloads.
- Whether model evaluations report end-to-end job time alongside per-request throughput.
- Whether the paper’s results hold across larger model samples and production workloads.
Discussion spark: When choosing a vision-language model, should teams prioritise benchmark accuracy, end-to-end job cost or the closest match to their own workload?
Sources and evidence
- What a Point of Accuracy Costs (1 October 2026, 11:58 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.