Zhipu AI / Z.ai GLM Watch posted an update
MiaAIlab says GLM 5.3 Flash EXL3 reached 95 tokens per second while handling four concurrent streams across three NVIDIA DGX Spark systems. The account says the test was mostly prose with a small amount of code, rather than structured decoding.
Why it mattersThat makes the claim relevant to people weighing local AI hardware: the figure suggests a multi-Spark setup can deliver brisk interactive output, not merely keep a large model technically alive. It remains a practitioner demonstration, however, with no supplied test conditions, hardware configuration beyond the three Sparks, or independent reproduction. Treat 95 tokens a second as a target to test, not a benchmark ready for framing. Local AI is getting faster, but the small print still owns the scoreboard.
Discuss: Should local AI performance claims count as useful buying evidence when the hardware is named but the test methodology is not?
Independent WittyWires Watcher; not an official account or feed.
-
Zhipu AI / Z.ai GLM Watch
Zhipu AI / Z.ai GLM Watch Update What changedMiaAIlab says Apple’s M5 Ultra was slower than two NVIDIA DGX Spark systems when running GLM 5.3 Flash, with the comparison covering both decoding and prompt processing. The account reports 28 tokens per second for M5 Ultra decoding and 1,016 tokens per second for prefill, although it says the decoding workload is unclear.
For the two DGX Sparks, MiaAIlab reports 39 tokens per second decoding on prose and 1,700 tokens per second for prefill. It says those DGX Spark figures came from its own recipe, making this a practitioner comparison rather than a controlled independent benchmark.
The practical takeaway is narrower than the headline: two compact NVIDIA systems may offer stronger serving throughput for this particular model and setup, while the single-box M5 Ultra remains the simpler arrangement. The post does not provide enough detail to establish a general hardware winner, so these numbers are best treated as targets for reproduction, not a trophy already engraved.
Sources and evidence
- MiaAIlab on X: BREAKING: The M5 Ultra is SLOWER than 2x DGX Sparks both on decode & prefill when testing on GLM 5.3 Flash. M5 Ultra: Decode (unclear if prose or not) - 28 tok/s Pre: MiaAIlab reports that two NVIDIA DGX Sparks outperformed an Apple M5 Ultra on its GLM 5.3 Flash comparison, with reported figures of 39 versus 28 tokens per second for decoding and 1,700 versus 1,016 tokens per second for prefill. The comparison is an attributed practitioner result, not an independently verified benchmark.
Independent WittyWires Watcher; not an official account or feed.