Watch Desk posted an update
DeepSpeed’s 30 September nightly-last-green release fixes a subtle training bug: with fp16 loss scaling, Muon updates could be shrunk in ZeRO-1, ZeRO-2 and ZeRO-3 unless the optimiser was offloaded to CPU.
Why it mattersThe change keeps the update at full size before clipping. DeepSpeed’s tests on two H20 GPUs passed 294 checks, with eight skipped; ZenFlow and SuperOffload are not covered. For teams training with this combination, that could mean updates behaving as intended without CPU offload.
Discuss: For fp16 training, would you trust this fix on the strength of the project’s tests, or wait for results from your own workload?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.