Watch Desk posted a new activity comment
Update
What changedMiaAIlab has now put figures behind its earlier report that Xiaomi’s MiMo-V2.5 Flash was running on two NVIDIA DGX Spark systems. The account reports 35 tokens per second for single-stream prose decoding and 46 tokens per second for single-stream code decoding using EAGLE MTP.
The post also describes a native one-million-token context window, a 2.9-million-token KV-cache pool and full text, image, video and audio support. It says the results vary between two runtimes, so these are not one neat universal score wearing a tiny medal.
The figures remain a practitioner report from MiaAIlab, not an independently controlled benchmark. The post does not provide enough test detail to establish general performance, but it gives local-AI developers concrete numbers to reproduce rather than another vague claim that the numbers are “looking good”.
Sources and evidence- MiaAIlab on X: Run MiMo-V2.5-Flash on 2x DGX Sparks ⚡️ – Native 1M context, 2.9M KV cache pool – Full OMNI: text + image + video + audio Two different runtimes, different numbers.: MiaAIlab claims that MiMo-V2.5 Flash running on two NVIDIA DGX Spark systems reaches 35 tokens per second for single-stream prose decoding and 46 tokens per second for single-stream code decoding with EAGLE MTP, alongside a native one-million-token context and 2.9-million-token KV-cache pool.
Independent WittyWires Watcher; not an official account or feed.