MoE 强化学习中的专家空间探索:ESRL 框架让 Qwen3-30B-A3B 的 Pass@1 提升 3.2 个百分点

HuggingFace Daily Papers(社区热门论文)·2026-09-11 08:00·5天前
AI 导读

研究者提出 Expert-Space Exploration Reinforcement Learning(ESRL),一个显式探索 MoE 模型专家路由空间的架构感知框架,通过保留高置信度专家作为锚点、将随机路由限制在合理候选池内,并按 router 熵自适应调整扰动强度。

HuggingFace Daily Papers(社区热门论文)
42AI 编辑部评分,满分 100

MoE 强化学习中的专家空间探索:ESRL 框架让 Qwen3-30B-A3B 的 Pass@1 提升 3.2 个百分点

2026-09-11 08:00· 5天前
AI 导读

研究者提出 Expert-Space Exploration Reinforcement Learning(ESRL),一个显式探索 MoE 模型专家路由空间的架构感知框架,通过保留高置信度专家作为锚点、将随机路由限制在合理候选池内,并按 router 熵自适应调整扰动强度。

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the sparse computation paths that induce output distributions, expert selection offers an additional source of rollout diversity. Through empirical analysis, we find that perturbing expert routing effectively alters model output and increases rollout diversity, which is similar to increasing the decoding temperature.

However, direct perturbation can activate unsuitable experts and substantially degrade rollout quality. Motivated by these observations, we introduce Expert-Space Exploration Reinforcement Learning (ESRL), an architecture-aware framework that explicitly explores the expert-routing space of MoE models. ESRL preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths. The perturbation strength is further adapted according to router entropy to avoid over-perturbation.

To mitigate the routing mismatch introduced by perturbation, ESRL records the expert paths used during rollout and replays them during policy optimization. Experiments demonstrate that ESRL achieves the best performance across MoE backbones with top-K, top-1, and shared-expert routing, as well as across mathematics, science, and code tasks without additional sampling or computational cost. Specifically, ESRL on Qwen3-30B-A3B achieves the best among all compared methods, improving average Pass@1 and Pass@8 over GRPO by 3.2 and 4.5 percentage points, respectively.

Further analyses of expert utilization and training dynamics provide insights into how exploiting MoE-specific routing structure benefits RL training.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org