Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

A developer says they ran Qwen 3.8 27B at 300 tokens a second across two RTX 3090s using NVLink, and reached 200 tokens a second with a 250,000-token context window.

Why it matters

The figures are the developer’s own, posted on X, and come without benchmark methodology or comparisons. They are a promising local-inference claim, not a settled speed record. They also say they moved away from Triton and plan to release a custom kernel. That will give other users something concrete to test, rather than just another impressive number doing laps on social media.

Discuss: What would you need to see in a reproducible test before trusting a local-inference speed claim like this?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.