# GPT-6 Astra 登顶 ErdosBench，OpenAI 称未刻意优化数学能力

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-09-10 21:45
- AIHOT 分数：63
- AIHOT 链接：https://aihot.news/items/cmtvlepvs0cxlronb4yp0h278
- 原文链接：https://the-decoder.com/gpt-6-astra-gives-mathematicians-a-breather-and-openai-says-thats-by-design

## AI 摘要

GPT-6 Astra 在 ulam.ai 的 ErdosBench 上以 3.23 分、解决 106 题（43 题完全解出、推翻 27 题）排名第一，超过 GPT-5.6 Sol 的 3.12 分和 78 题。

## 正文

OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says the company didn't prioritize math. That says something about where "AGI" actually stands.

Mathematicians can take a brief (likely very brief) breather after Astra's release. The field has been grappling with existential questions lately, and while the new model does push math research forward, it's not moving at the pace some had hoped for and others had feared.

On ulam.ai's ErdosBench, Astra took first place. The benchmark covers 226 open math problems inspired by the famous Erdős problems. Astra scored 3.23 and solved 106 problems, 43 of them completely. It disproved 27 more. It also disproved 27 others.

Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims. In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested," but the benchmark is far from saturated.

GPT-6 Astra leads the ErdosBench with a score of 3.23 and 106 solved problems, followed by GPT-5.6 Sol at 3.12 and 78 solved problems. | Image: ulam.ai

OpenAI chose not to optimize for math

OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement. In his essay "An Alien Mind," OpenAI chief scientist Jakub Pachocki writes, "[…] we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later."

That means the most capable math model right now isn't the result of targeted optimization. It's a byproduct of other priorities. OpenAI is pouring its resources into recursive self-improvement and securing future AI systems, since "we believe it is the only way to remain at the frontier of AI research moving forward," Pachocki writes. (Note: Since the essay was published, OpenAI has reportedly trained better internal math models, a complicated story, and the topics of RSI and AI safety have exploded.)

Pachocki's statement is interesting for another reason, too: it shows that hard optimization trade-offs are being made at the jagged frontier of AI development. A model that dominates math won't necessarily dominate everything else. That reminded me of a visualization by Cambridge researcher Adam Hunt, which contrasts two possible paths for AI. The "mainstream AGI" thesis assumes models improve gradually and broadly until they cover all human tasks.

The alternative scenario describes an increasingly "spiky" trajectory, with extreme strength in a few domains like coding and math but stagnation or even regression in areas like language quality, common sense, or social reasoning. That would give us a highly specialized model, not something most people would call Artificial General Intelligence. Of course, you can call anything AGI if you feel like it.

Two paths for AI development: gradual, broad improvement (top) versus extreme specialization with stagnating baseline capabilities (bottom). | Image: via Adam Hunt

Pachocki's point that OpenAI deliberately skipped math optimization for Astra is evidence for the "spiky" thesis. Even the leading AI lab can't push maximum progress in every direction at once. Capabilities grow where you optimize, and making math stronger means cutting back somewhere else. Behind the RSI priority, though, is the hope that the model will eventually make those optimizations itself, scaling faster across the board, including in math.

Math is wrestling with AI and with itself

Regardless of whether Astra is 5 or 50 percent better than its predecessor, the deeper question remains for a field that's grappling with a new world of compute power that helps solve problems but doesn't necessarily help understand them. Mathematician Terence Tao raised it at the 2026 International Congress of Mathematicians: if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload.

The critical task would then no longer be solving problems but deciding which results actually matter. Tao says math faces a crisis of its values and practices, one he compares to the foundational upheaval of the early 20th century.

The hardest problems in math remain unsolved for now, which should buy the discipline some time. The direction still seems set, though, even if some mathematicians don't think language models can deliver real breakthroughs because they lack human-like creativity. For similar reasons, there are also doubts about AI's potential for genuine self-improvement.
