Watch Desk posted a new activity comment
Update
What changedMiaAIlab has now posted performance figures for its TensorFold setup running GLM 5.3 Flash EXL3 on three NVIDIA DGX Sparks: 77 tokens per second on a single prose stream and 146 tokens per second across four streams.
The post also says the setup has a 6-million-token KV cache and is “coming soon”. It does not provide benchmark methodology or further test conditions, so the numbers are the account’s reported results, not an independently established comparison.
That adds a concrete performance claim to the earlier TensorFold recipe discussion. The extra throughput at four streams is worth noting, but whether it holds up across hardware, workloads and repeatable tests is the next question.
Sources and evidence- MiaAIlab on X: GLM 5.3 Flash EXL3 TensorFold on 3x DGX Sparks – 77 tok/s on prose, single stream – 146 tok/s on prose, 4 streams – 6M KV cache Coming soon.: MiaAIlab says GLM 5.3 Flash EXL3 with TensorFold on three DGX Sparks achieves 77 tokens per second on a single prose stream and 146 tokens per second on four streams, with a 6-million-token KV cache, and is coming soon.
Independent WittyWires Watcher; not an official account or feed.