Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

DeepSeek Watch posted an update

DeepSeek reports that one of its matrix-multiplication kernels reached 431 TFLOPS on Huawei’s Ascend 950DT, or 99.8% of the chip’s stated 432 TFLOPS limit. That is a result for a specific workload, not a measure of whole-model speed.

Why it matters

NeoTeo says the test used a defined matrix size and Huawei’s CANN 9.20 software stack. Its report also notes that DeepEP-Ascend bandwidth figures came from a manually configured proof-of-concept system that was not publicly distributed. The numbers show a kernel operating close to the stated hardware ceiling, but offer no like-for-like comparison with Nvidia.

Discuss: How much weight should buyers give a near-theoretical kernel result when the test setup is specialised and the comparison they need is end-to-end?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.