jietang · @jietang · X·2026-09-10 10:10·33分钟前
AI 导读

唐杰指出确定模型最优规模并不简单,需权衡数据量、激活参数量、环境数量和目标推理成本,性能还受多种因素影响。其引用讨论称 Fable 约为 2-2.5T 参数而非 10T,Kimi K3 为 2.8T 参数、训练约用 20-30k Blackwell 等效算力,能力接近 Fable 5(非 5.1)。

jietang@jietang
28AI 编辑部评分,满分 100
2026-09-10 10:10· 33分钟前
AI 导读

唐杰指出确定模型最优规模并不简单,需权衡数据量、激活参数量、环境数量和目标推理成本,性能还受多种因素影响。其引用讨论称 Fable 约为 2-2.5T 参数而非 10T,Kimi K3 为 2.8T 参数、训练约用 20-30k Blackwell 等效算力,能力接近 Fable 5(非 5.1)。

Are you sure? Finding the optimal model size is tricky: data volume, active parameter count, the number of environments, and the target inference cost. Model performance also depends on many other factors, each introducing its own variability.

Charlie O'NeillFable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable...

来源:jietang· x.com