腾讯混元 Hy4-preview 压缩至 200GiB 精度几乎无损

Tencent Hy · @TencentHunyuan · X·2026-08-29 13:31·23天前
AI 导读

腾讯混元将 Hy4-preview(770B 参数、49B 激活、1M 上下文)从 1.5TB 压缩至约 200GiB GGUF 格式,通过 MIX-STQ1_0 混合量化按层分配位宽(最低 1.31-bit STQ1_0,最高 2.06-bit IQ2_XXS)。

Tencent Hy@TencentHunyuan
62AI 编辑部评分,满分 100

腾讯混元 Hy4-preview 压缩至 200GiB 精度几乎无损

2026-08-29 13:31· 23天前
AI 导读

腾讯混元将 Hy4-preview(770B 参数、49B 激活、1M 上下文)从 1.5TB 压缩至约 200GiB GGUF 格式,通过 MIX-STQ1_0 混合量化按层分配位宽(最低 1.31-bit STQ1_0,最高 2.06-bit IQ2_XXS)。

We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well !

Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.

Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5

See the details on HF : AngelSlim/Hy4-preview-GGUF

Weights & low-bit GGUFs 👇 https://huggingface.co/AngelSlim/Hy4-preview-GGUF

#LLM #Quantization #llamacpp #Hy

Tencent Hy🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. Mo...