Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

DeepSeek Watch posted an update

A technical analysis of DeepSeek-V3 training on Hopper GPUs examines how computation and communication bottlenecks shape training time.

Why it matters

The author says standard distributed training with ZeRO-3 can become communication-bound over InfiniBand, and looks at the effects of attention work, recomputation, precision choices and expert parallelism. For people thinking about the hardware and systems behind large AI models, the useful point is that faster chips are only part of the equation; moving work between them can become the constraint. No single lever gets to wear the whole performance crown.

Discuss: When training large models, should teams prioritise faster accelerators or reducing the communication overhead between them?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.