腾讯新论文:持续变难的环境优于共同进化,提升 Terminal-Bench 2.1 成绩

Rohan Paul · @rohanpaul_ai · X·2026-09-09 10:35·35分钟前
AI 导读

腾讯新论文表明,智能体持续进步时训练任务不能一成不变,持续变难的环境在 Terminal-Bench 2.1 上优于共同进化方法。合成终端任务终会过易,稳定解决后便不再提供有效 RL 信号。

Rohan Paul@rohanpaul_ai
43AI 编辑部评分,满分 100

腾讯新论文:持续变难的环境优于共同进化,提升 Terminal-Bench 2.1 成绩

2026-09-09 10:35· 35分钟前
AI 导读

腾讯新论文表明,智能体持续进步时训练任务不能一成不变,持续变难的环境在 Terminal-Bench 2.1 上优于共同进化方法。合成终端任务终会过易,稳定解决后便不再提供有效 RL 信号。

New Tencent paper shows, if an agent keeps improving, its training tasks cannot stay still: continuously harder environments produced better Terminal-Bench 2.1 results than co-evolution.

Synthetic terminal tasks eventually become too easy. Once the agent solves them reliably, they stop giving much useful RL signal.

Instead of waiting for the model to fail and then building new tasks around those failures, this paper evolves the tasks themselves. It gradually makes each environment less familiar, adds rarer required skills, or makes the job take more steps. Each harder version is checked and introduced as the agent improves.

That produced progressively harder environments when tested with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol.

On Terminal-Bench 2.1, Qwen3.6-27B reached 71.5% versus 62.9% with co-evolution, while Qwen3.6-35B-A3B reached 64.9% versus 55.1%.

来源:Rohan Paul· x.com