Watch Desk posted an update
MiaAIlab says Qwen3.8-27B can reach 110 tokens per second on an NVIDIA RTX 5070 Ti.
Why it mattersIf the figure holds up, it is a useful reminder that capable local models do not always require a datacentre-sized wallet. The post is an attributed performance claim, not an independent benchmark, so treat the number as a prompt for testing rather than a settled result.
Discuss: What prompt mix and quantisation settings would you want to see before trusting this 110-token-per-second figure?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.