腾讯混元发布 EvolveScaler 信息演化基准

Tencent Hy · @TencentHunyuan · X·2026-09-15 14:34·1小时前
AI 导读

腾讯混元推出 EvolveScaler,将世界定义为可执行状态机再渲染成自然语言,模拟记录被撤回、修正、回填的信息演化场景。该基准含 117 个原型、159 个问题算子、5 个难度层级,单样本最多约 1200 个事件;14 个前沿模型在最难层级上 avg@5 中位数降至 11.3,而用它训练可在 8 个分布外基准上平均提升 5.25。

Tencent Hy@TencentHunyuan
45AI 编辑部评分,满分 100

腾讯混元发布 EvolveScaler 信息演化基准

2026-09-15 14:34· 1小时前
AI 导读

腾讯混元推出 EvolveScaler,将世界定义为可执行状态机再渲染成自然语言,模拟记录被撤回、修正、回填的信息演化场景。该基准含 117 个原型、159 个问题算子、5 个难度层级,单样本最多约 1200 个事件;14 个前沿模型在最难层级上 avg@5 中位数降至 11.3,而用它训练可在 8 个分布外基准上平均提升 5.25。

🚀 EvolveScaler is here.

Read a 40-day RPG log. Now answer one question: if you skip the mini-boss on Day 7, do you still beat the final boss?

The answer isn't in the log. You have to replay the world.

That's Information Evolution — records get retracted, corrected, backfilled. The world keeps changing after you read it.

So we build it backwards: define the world as an executable state machine, then render it into natural language. Code guarantees the logic. Language delivers the mess.

➡️ 117 prototypes. 159 question operators. 5 difficulty tiers. Up to ~1,200 events per sample.

➡️ 14 frontier models, hardest tier: median avg@5 falls to 11.3.

➡️ Train on it instead: +5.25 average across 8 out-of-distribution benchmarks.

Check out our paper and project page.

📚 Paper: https://arxiv.org/abs/2609.08435 🏠 Project Page: https://tencent-hunyuan.github.io/evolve-scaler/

来源:Tencent Hy· x.com