We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning further enhance its overall performance. Hunyuan-A13B also introduces a dual-mode Chain-of-Thought framework that adapts reasoning depth to task complexity: fast thinking for routine queries and slow thinking for complex, multi-step problems. Evaluations show competitive performance across mathematics, science, programming, general language understanding, and agent tasks, often approaching that of much larger models. Its high inference throughput makes it suitable for latency-sensitive applications. We release Hunyuan-A13B to support open research and practical LLM deployment.
Hunyuan-A13B 技术报告发布:80B 总参数、13B 激活的开源 MoE 模型
AI 导读
腾讯混元发布开源 MoE 大语言模型 Hunyuan-A13B,总参数 80B、推理仅激活 13B,以平衡能力、效率与部署成本。模型在经严格筛选的 20T token 语料上预训练并加强 STEM 数据,再经 SFT 和大规模强化学习;引入双模式 Chain-of-Thought 框架,简单问题用快思考、复杂多步问题用慢思考。
HuggingFace Daily Papers(社区热门论文)
52
AI 编辑部评分,满分 100Hunyuan-A13B 技术报告发布:80B 总参数、13B 激活的开源 MoE 模型
腾讯混元发布开源 MoE 大语言模型 Hunyuan-A13B,总参数 80B、推理仅激活 13B,以平衡能力、效率与部署成本。模型在经严格筛选的 20T token 语料上预训练并加强 STEM 数据,再经 SFT 和大规模强化学习;引入双模式 Chain-of-Thought 框架,简单问题用快思考、复杂多步问题用慢思考。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org