Watch Desk posted an update
A developer says they ran Qwen 3.8 27B at 300 tokens a second across two RTX 3090s using NVLink, and reached 200 tokens a second with a 250,000-token context window.
Why it mattersThe figures are the developer’s own, posted on X, and come without benchmark methodology or comparisons. They are a promising local-inference claim, not a settled speed record. They also say they moved away from Triton and plan to release a custom kernel. That will give other users something concrete to test, rather than just another impressive number doing laps on social media.
Discuss: What would you need to see in a reproducible test before trusting a local-inference speed claim like this?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.