The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. We present Dion3, a revision of Muon that targets this overhead at every level of the stack. Our Gram Newton-Schulz algorithm reduces the FLOP cost of orthogonalization, our CuteDSL kernels accelerate it by exploiting symmetry, and our megabatching strategy reduces communication overhead. Moreover, we propose a simple change to the update rule that cuts costs even further: selecting only a fraction of the momentum matrix's rows to orthogonalize at each step. This update rule improves on Dion (another "compressed" version of Muon), in both speed and performance. Overall, Dion3 matches or improves on the loss achieved by Muon but reduces optimizer step time by up to 6x. Dion3 is available via the dion package (https://github.com/microsoft/dion) as a drop-in replacement for Muon.
Dion3:全栈正交更新的 Muon 优化器改进版
AI 导读
Dion3 针对 Muon 优化器中立方时间 Newton-Schulz 正交化步骤的高开销,在堆栈各层进行优化:Gram Newton-Schulz 算法降低正交化 FLOP 成本,CuteDSL 内核利用对称性加速,megabatching 策略减少通信开销。
HuggingFace Daily Papers(社区热门论文)
54
AI 编辑部评分,满分 100Dion3:全栈正交更新的 Muon 优化器改进版
Dion3 针对 Muon 优化器中立方时间 Newton-Schulz 正交化步骤的高开销,在堆栈各层进行优化:Gram Newton-Schulz 算法降低正交化 FLOP 成本,CuteDSL 内核利用对称性加速,megabatching 策略减少通信开销。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org