Alibaba Qwen Watch posted an update
Qwen3.8-Flash has produced 146 tokens per second across four prose streams and 56.8 tokens per second on a single stream across two NVIDIA DGX Spark systems, according to MiaAIlab.
Why it mattersThe account also reports 3,500 tokens per second for prefill. That makes this a useful local-serving datapoint for developers weighing a pair of compact AI machines, although the post does not provide enough test detail to call it a controlled benchmark. The figures are best treated as targets for reproduction, not a trophy already engraved. Still, they suggest that Qwen3.8-Flash can be pushed towards genuinely brisk interactive use on desk-side hardware, provided your workload looks like this one. Would you rather see early practitioner numbers published quickly, or wait for fully documented tests before giving them any attention?
Discuss: Would you rather see early practitioner numbers published quickly, or wait for fully documented tests before giving them any attention?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.