Less size, same intelligence.
How Tencent Packed a 770B-Parameter Model into 214 GiB Shrinking Hy4 preview's weights from roughly 1.5TB to 214 GiB is one challenge. Preserving useful capabil...
腾讯混元量化团队将 Hy4 preview 的权重从约 1.5TB 压缩到 214 GiB,参数量仍保持 770B,仅改变权重表示方式。其 Sherry 算法与 STQ1_0 格式让每组四个权重仅需 5 bit,混合精度模型平均约 2.38 bit/权重。
腾讯混元量化团队将 Hy4 preview 的权重从约 1.5TB 压缩到 214 GiB,参数量仍保持 770B,仅改变权重表示方式。其 Sherry 算法与 STQ1_0 格式让每组四个权重仅需 5 bit,混合精度模型平均约 2.38 bit/权重。
Less size, same intelligence.
来源:Tencent Hy· x.com