Thomas Wolf· @Thom_Wolf · X·· 2 小时前AI 评分55
AI 导读
@classiclarryd 通报 Deven Pzak 将 NanoGPT 训练纪录从 67.6 秒压缩至 39.9 秒,PR 为 KellerJordan/modded-nanogpt#360。核心思路是在单个 flop 层面按价值取舍而非只优化 matmul,包括采样 softmax(约 8 秒)、ngram 嵌入的稀疏更新与优化器状态、分片稀疏通信、手写 64 维 head 的 flash attention、最后 300 步 EMA 与新优化器 Anvil2(约 1 秒)等。稀疏嵌入参数扩到 65B,贡献该 PR 约 25% 的收益;Thomas Wolf 转发并称 impressive,附上与 Deven 访谈整理的技术复盘 https://hyperstition.cc/training-nanogpt-in-39-9-seconds。
正文
impressive
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds在 X 查看被引用的帖子
来源:Thomas Wolf · x.com