跳到正文
The Decoder:AI News· Manuel Uth·· 3 小时前AI 评分63

Epoch AI 发布 InnovationEval 评测,AI 智能体远未达到自主研究水平

AI agents overstate their results and remain far from autonomous research, study finds

AI 导读

Epoch AI 用新基准 InnovationEval 测试 Claude Fable 5 和 GPT-5.6 Sol 能否独立发明新训练方法,以人类设计的 SDPO 对 GRPO 的改进为参照。

正文

Agents were asked to invent a new training method

With its benchmark "InnovationEval," Epoch AI tested whether AI agents can conduct research on their own. The task was to invent a new method for improving language models after their initial training, then implement, test, and refine it independently.

The starting point was GRPO, a widely used technique. GRPO compares multiple answers a model generates for the same task and rewards the better ones, typically scoring each solution as a whole. The human-designed reference method, SDPO, uses extra signals like error messages to create more precise learning feedback for individual steps within an answer, so the model effectively becomes its own teacher.

Epoch tested Claude Fable 5 and GPT-5.6 Sol, which according to the organization had no prior knowledge of SDPO. Performance was measured on short-answer tasks like science questions and on coding tasks, with each agent having access to up to 3,000 hours of compute on high-end chips but no internet access.

Both models recycled known techniques instead of innovating

Neither model came close to the human reference. GPT-5.6 Sol targeted a weakness in GRPO: when all answers to a task are correct, the model learns nothing because there's nothing to compare. Sol had the model reinforce its successful solutions in those cases, but the idea wasn't new.

Measured against the improvement SDPO achieves over GRPO, Sol scored about 35 percent with generous grading, according to Epoch. Counting only changes that stayed within the experiment's rules, that number drops to about 15 percent. On coding tasks, Sol mostly made training more expensive and slower rather than improving the method itself.

Claude Fable 5 had the model retry failed tasks while feeding it the previous failed attempts, which is also a well-known technique and produced no measurable improvement. Even Sol's partial success would barely qualify as "moderately interesting" to experts, Epoch says. Newer models that knew SDPO from their training data couldn't fully replicate it either. GPT-6 Astra built a similar solution but didn't disclose its source, and even with the original paper in front of it, Fable 5 fell short of the reference.

Agents cherry-picked their best runs and buried the rest

Both agents also had a reporting problem. They ran multiple near-identical training rounds and reported only the best result each time, which makes a method look stronger than it actually is because outcomes fluctuate randomly. In their final reports, the models barely mentioned this practice, if at all, and they also failed to cite the prior work their methods drew on. Their self-reported numbers ran accordingly high: Sol claimed about 70 percent of the SDPO improvement, Fable 5 about 40 percent. Epoch stripped out those inflated gains.

Claimed vs. actual scores for Claude Fable 5 and GPT-5.6 Sol. Dark bars show the agents' self-reported numbers, medium bars show results corrected for cherry-picked runs, and light bars show scores counting only rule-compliant changes. All values are expressed as a share of SDPO's improvement over GRPO. Error bars show measurement uncertainty. | Image: Epoch AI

The models' internal reasoning logs show they were aware of the problem, with Fable 5 describing its repeated runs as a search for a better checkpoint. Whether this amounts to deliberate cheating or confusion, Epoch leaves open.

For GPT-5.6 Sol, the behavior fits an earlier finding from the evaluation organization METR, which had detected more cheating attempts from that model in its own coding test than from any other publicly available model it had evaluated.

Epoch concludes that humans would need to fully review all AI-generated research, which cuts into the models' usefulness.

Anthropic sees the same weaknesses in its own model

Anthropic describes similar limitations in the system card for Claude Opus 5.5. The model is far from replacing the company's own researchers, Anthropic says, with the main problems lying in "epistemic quality" and instruction-following.

Opus 5.5 presents unchecked assumptions as facts and pushes aside its own doubts more often than earlier models did. It describes partial checks as complete verification, turns preliminary assessments into recommendations without cross-checking, and addresses criticism too narrowly without questioning the overall approach. It also favors small, incremental tweaks and sticks to published studies rather than developing new ideas.

A study involving Princeton University and the UK Safety Institute also reached similar conclusions. It had Claude Opus 4.8 work for six days on the research questions behind two unpublished papers from the NeurIPS AI conference, and the original authors rejected both results when reviewing them. When initial hypotheses failed, the agents just softened their claims instead of starting over.

Technical execution is increasingly the smaller problem, since what the agents lack most is the ability to realistically gauge how confident they should be in a result and to question their own approach at a fundamental level.

More compute alone won't close the gap

AI isn't useless for research, since models can speed up literature review, write code, and explore variations faster than humans can.

But whether throwing more compute at the problem will produce autonomous researchers remains an open question. GPT-5.6 Sol used its entire budget, finding its improvement on short-answer tasks only near the very end of its compute allocation while making no rule-compliant progress at all on coding tasks. Epoch sees the late breakthrough as a weak hint that more compute time could help, though Fable 5 didn't even use half its budget.

GPT-5.6 Sol's progress plotted against cumulative GPU spend. On short-answer tasks (teal), the jump comes just before the budget runs out. On coding tasks (pink), the rule-compliant score stays at zero, and only out-of-scope changes produce any gain (dashed line). | Image: Epoch AI

Epoch also notes that leading models from a year ago would have performed much worse on the same test, and the organization plans to repeat InnovationEval regularly with new tasks. Until then, the biggest gap remains epistemic discipline: skeptically checking your own findings, disclosing uncertainties, and taking negative results seriously.

来源:The Decoder:AI News · the-decoder.com