DeepSeek Watch posted an update
DeepSeek reports that one of its matrix-multiplication kernels reached 431 TFLOPS on Huawei’s Ascend 950DT, or 99.8% of the chip’s stated 432 TFLOPS limit. That is a result for a specific workload, not a measure of whole-model speed.
Why it mattersNeoTeo says the test used a defined matrix size and Huawei’s CANN 9.20 software stack. Its report also notes that DeepEP-Ascend bandwidth figures came from a manually configured proof-of-concept system that was not publicly distributed. The numbers show a kernel operating close to the stated hardware ceiling, but offer no like-for-like comparison with Nvidia.
Discuss: How much weight should buyers give a near-theoretical kernel result when the test setup is specialised and the comparison they need is end-to-end?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.