Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional agent-side warming up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to environment-side adaptation by constructing Feedback-Enriched Environments (FEEs). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs (1) stabilizes training dynamics by reducing entropy volatility, (2) facilitates proactive state-space exploration in difficult tasks, (3) ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and (4) identifies intra-group feedback consistency as a critical boundary for stable optimization.
环境作为脚手架:反馈增强环境如何引导长程任务中的自进化智能体
AI 导读
针对长程任务中强化学习奖励稀疏的问题,研究者提出环境侧适配范式,构建反馈增强环境(FEEs),在幕内探索与幕间演化的后期将反馈从动作引导转向观察丰富化。基于Qwen3多种规模模型及GRPO、GSPO、DAPO等RL算法在SciWorld和BFCL基准上的大规模实验显示,FEEs相较标准设置持续带来性能提升,并稳定训练动态、促进主动探索、将环境引导内化至策略权重。
HuggingFace Daily Papers(社区热门论文)
41
AI 编辑部评分,满分 100环境作为脚手架:反馈增强环境如何引导长程任务中的自进化智能体
针对长程任务中强化学习奖励稀疏的问题,研究者提出环境侧适配范式,构建反馈增强环境(FEEs),在幕内探索与幕间演化的后期将反馈从动作引导转向观察丰富化。基于Qwen3多种规模模型及GRPO、GSPO、DAPO等RL算法在SciWorld和BFCL基准上的大规模实验显示,FEEs相较标准设置持续带来性能提升,并稳定训练动态、促进主动探索、将环境引导内化至策略权重。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org