Astra,OpenAI 正在内部测试的一款新模型,确实令人惊叹。这一点无可否认:
第一部分:最大的谬误
但与此同时,一大批人——其中不乏相当有分量的人物——正在四处宣扬一种关于其影响的、存在严重缺陷的论点。
看看你能不能找出这个谬误。以下是从众多例子中挑出的三个。


另一条推文甚至声称“这个物种刚刚跨过了一道单向门槛”;埃隆·马斯克则将其视为我们已抵达奇点的证据。
他们每个人犯的本质上都是同一个错误。
这个谬误有名字,叫作“合成谬误”。

维基百科上关于它的词条里满是绝佳的例子;这里只列开头几个。

这个谬误在当前案例中的表现是什么?就是认为一个擅长某类数学问题的系统,就擅长所有数学、擅长科学,甚至很可能擅长一切。“AGI 临近”群体一次又一次地犯着同一个逻辑谬误。每次有进展出现,我都会看到同样的错误。
这个谬误是这样运作的。
1. 有人假装所有认知都是同等的。(完全不符合事实。)
2. 每当 AI 在某种(花哨的)认知形式上取得成功,他们就希望你相信,AI 在一切形式上的成功都近在眼前。
你不需要是认知心理学家也能看出这个推论根本不成立。比如我们都知道,数学方面的专长并不能保证在所有领域都成为天才。一个擅长数学、物理或编程的人,可能在写作或理解人际关系上很吃力(反过来,一个伟大的作家也可能不擅长数学,等等)。某一领域的专长完全不能保证在所有甚至大多数领域都具备专长。
这正是霍华德·加德纳和罗伯特·斯腾伯格等人提出多元智能理论的原因,也是 SAT 将数学和语文分开考试的原因,等等。
Astra 看起来——我们还没有看到其方法论——在数学方面表现出色,或者至少在某些形式的数学上如此,但这并不意味着它能避免模型幻觉,或解决其他生成式 AI 系统存在的可靠性问题。这甚至不意味着它能可靠地阅读 PDF。也不意味着它将成为第一个能够遵守硬性规则的生成式 AI。(这应该让你感到恐惧。)
特别是,Astra 显然擅长某些问题;但这并不意味着它在难以形式化的问题上会表现出色,甚至不一定能胜任。这并不意味着它会神奇。这并不意味着它是 AGI 或 ASI 或任何类似的东西。
它*非常*令人印象深刻。但我看不出任何理由认为 Astra 是 AGI,更不用说 ASI 了。如果它能在我和 Miles Brundage 的 2024 年赌注中拿到 5/10 的分数,我会感到惊讶。
第二部分。有一个重要的、原则性的理由认为,数学上的成功是一个特例,不会像人们可能希望的那样广泛推广。
如果(a)这种谬误不是极其普遍,并且(b)我不认为有非常充分的理由认为这种谬误适用于此,我就不会花这么长的篇幅来讨论它。
数学成为这些模型大放异彩的领域并非偶然。数学适合两件事:验证(使用符号工具),以及大规模廉价生成的合成数据,在这些数据中你可以保证答案是正确的。编码也是如此——但一般来说并非如此。你可以生成任意多的数学事实;你无法模拟开放式的世界。你可以验证数学;你无法以同样的方式验证军事战略。
这并非新闻。我至少从 2025 年 1 月起就一直在指出这一点【第一行本应说编码和数学】,呼应了 Ernie Davis 和我在 2019 年关于围棋(以及为什么 AlphaGo 不会是万能灵药)所说的:

OpenAI 能做到 Astra 在数学方面所做的事情,是因为数学允许使用外部工具进行验证并创建合成数据。
这并不会让 Astra 显得不那么令人印象深刻,但正如十年前我们从 IBM 那次混乱且最终失败的尝试中学到的——它试图把赢得《危险边缘》的 Watson 改造成一台攻克癌症的机器——在一个领域取得成功并不能保证在所有领域都能成功。
OpenAI 在这里所处理的数学,与许多(也许是大多数)现实世界问题有着根本性的不同;在那些问题中,验证既不能保证解决方案,也无法让人几乎零成本地无限生成训练数据。
指望它成为万能溶剂,只能说明你并不理解这个基本事实。
第三部分:关于 Astra 还需要了解的七件事
昨天的推文和博客是营销,不是科学。这两者,连同随附的 249 页数学论文,都没有提供任何关于这一成就是如何实现的信息:
在一封评论本文初稿的电子邮件中,Ernie Davis 非常中肯地补充道:
要恰当地评估这些发现的意义,我们需要一些关键信息。第一:他们尝试了多少个猜想?如果 OpenAI 团队是从所有未解决的重要数学猜想空间中随机挑选出这 10 个猜想,而 Astra 全部解决了,那将令人惊叹;但这似乎完全不可能。如果他们挑选了 50 个他们认为 Astra 能成功的猜想,而它成功了 10 个,那仍然令人惊叹,但程度要低得多。如果他们在所有大约 1000 个未解决的 Erdős 猜想以及另外 10,000 个未解决猜想上运行了 Astra,那么它未能解决这些问题的结果,就是关于其作为数学推理者能力极限的重要信息。第二:OpenAI 大肆吹嘘说,完成这项工作总共只花了 2000 美元的计算成本。可以肯定的是,这很可能只包括 Astra 成功的那些猜想,而不包括它失败的。但是,考虑到参与这个项目的高薪数学家和计算机科学家的工资,成本又是多少呢?如果低于 20,000 美元,我会感到震惊;如果超过 200,000 美元,我也不会感到意外。
我们对此并不了解,而且照目前的发展态势来看,我们可能永远也无法了解。对于在2025年国际数学奥林匹克竞赛中夺得金牌的OpenAI系统,我们依然一无所知;而对Google的那个系统,我们掌握的信息也少得可怜。
Astra 在某些数学问题上表现出色,但要说数学已被“攻克”(X 平台上很多人似乎这么认为),这实在令人怀疑。它似乎只是擅长特定类型的数学问题。
但很可能并非所有数学问题它都擅长。以下是 Eric Weinstein 在 X 平台上列举的一些它可能不太擅长的数学类型示例。
Ernie Davis 还补充道:“自动形式化的问题——即将人类数学家撰写的数学内容转化为严格逻辑形式——似乎远未得到解决,尽管确实已取得进展。例如:过去两年里,数学家 Kevin Buzzard 一直在领导一个项目,用 Lean 语言将 Andrew Wiles 对费马大定理的证明形式化。可以肯定地说,没有任何 AI 系统有能力完成这项艰巨的任务。可以想见,如果它们能在无人协助的情况下做到这一点,它们早就抢先一步了;如果它们能将他数年的工作缩短为数周,他早就用上它们了。”
与我关于数学能力并不能保证通用性的总体预测一致,Astra 的证明写作水平似乎与其证明本身并不相称(这呼应了 Ernie Davis 和我一年前在此处的观察)”[点击查看完整推文]
我严重怀疑 Astra 能神奇地解决当前模型中所有无法可靠运行的问题。它能可靠地从任意 PDF 中提取数字吗?我表示怀疑。它能满足 Sabine Hossenfelder 希望 AI 为她的 YouTube 系列节目撰写脚本的愿望吗?[点击推文查看详情] 我也表示怀疑。
Astra 并非(如 Musk 所暗示的)奇点。无需多言。(不过可以看看我最近关于这个话题的文章;情况其实没什么变化。)
声称 Astra 的首次亮相“很可能是数学史上最重要的一天”是荒谬的,而一大群人却基于 Fable 某一次运行中的一句话(如果问两次,可能会给出不同答案)到处这么说。这里没有新理论,没有新技术,而且尽管它令人印象深刻,但它与微积分、代数、对数、概率、十进制系统、信息论、零的概念的发展完全不在一个量级上。Davis 表示:“Fable 引文中声称‘其中几项单独拿出来都足以成为这十年来的重大成果’是荒谬的……作为现实检验,自 1900 年以来,David Hilbert 的 23 个问题中已有 14 个得到解决——也就是平均每九年解决一个——而这些结果远未达到那个水平。你很容易就能列出一份自 1926 年以来被证明的 100 项重要得多的成果清单。”
我怀疑它不会直接带来癌症的治愈,也不会像一些兴奋的 AI“网红”所设想的那样,在“材料研究、能源生产、药物发现、一切领域”带来巨大进步。这也不意味着我们刚刚进入了“自动化科学发现的时代”。针对这些过度的说法……
……这篇刚刚发布,呼应了几天前我链接到的由等人撰写的新文章。
第四部分:总结
尽管昨天在 X 上遭到了巨大的反对(其中大部分是人身攻击,有些甚至涉及捏造,还有很多是彻头彻尾的谎言),我对 Astra 的总体看法仍然是我最初发布的那条:

可悲的是,Ernie Davis 和我在一年前就提出了很多关于如何正确检验新系统的相同观点:

我理解人们对 Astra 感到兴奋,但任何期待奇迹的人都很可能会失望。
Astra, a new model that OpenAI is testing internally, is amazing. No denying that:
Part I: The Biggest Fallacy
But at the same time, a whole raft of people, some fairly prominent, are running around making a deeply flawed argument about the implications.
See if you can spot the fallacy. Here are three examples among many.


Another tweet went so far as to claim that “the species just crossed a one-way threshold”; Elon Musk took it as evidence that we had reached The Singularity.
Each is making essentially the same error.
The fallacy has a name; it’s called the fallacy of composition.

The wiki on it is filled with great examples; here are just the first few.

What’s the manifestation of the fallacy in the current case? Thinking that a system that is great at a certain kind of math problem is great at all math, great at science or even quite possibly great at everything. The “AGI-is-near” community keeps committing the same logical fallacy over and over. Every time there’s an advance, I see the same error.
Here’s how the fallacy works.
1. Someone pretends that all cognition is created equally. (Totally untrue.)
2. Whenever AI achieves success on some form of (fancy) cognition, they want you to believe that success on all forms of AI is imminent.
You don’t have to be a cognitive psychologist to realize that this inference just doesn’t follow. We all know, for example, that expertise in math doesn’t guarantee genius in all domains. Someone who is great at math or physics or programming may struggle with writing or understanding human relationships (and conversely a great writer may be weak at math, etc). Expertise in one domain does not at all guarantee expertise in all or even most domains.
That’s precisely *why* people like Howard Gardner and Robert Sternberg developed multidimensional theories of intelligence, why the SAT tests math separately from verbal, etc.
Astra appears to be —we still haven’t seen the methodology—great at math, or at least some forms of math, but that does not mean that it will avoid hallucinations or solve the reliability problems other GenAI systems have. It doesn’t even mean it will be able to read PDFs reliably. And it doesn’t mean it will be the first generative AI to be able to obey hard rules, either. (Which should terrify you.)
In particular, Astra is obviously excellent at some problems; but that doesn’t mean it will be excellent or even competent at problems that are hard to formalize. It doesn’t mean it will be magic. It doesn’t mean it’s AGI or ASI or any of that.
It’s *very* impressive. But I see no reason whatsoever to think Astra is AGI let alone ASI. If it can score even a 5/10 on my 2024 bet with Miles Brundage I will be surprised.
Part II. There is an important, principled reason to think that success on math is a special case which will not generalize as much people might hope.
I wouldn’t go on about this fallacy at such length if (a) it wasn’t wildly common and (b) I didn’t think that there was very good reason to think that the fallacy applied here.
It’s not an accident that math is where these models are shining. Math lends itself to two things: verification (using symbolic tools), and massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct. The same applies to coding — but it is not true in general. You can generate as many math facts as you want; you can’t simulate the open-ended world. You can verify math; you can’t verify a military strategy in the same way.
This is not news. I have been pointing this out at least since January 2025 [first line should have said coding and math], echoing something Ernie Davis and I said about Go (and why AlphaGo would not be a panacea) in 2019:

OpenAI can do what Astra does in math because math allows for external tools to do verification and to create synthetic data.
That doesn’t make Astra less impressive, but as we learned a decade ago from IBM’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all.
The kind of math OpenAI is dealing with here is radically different from many (perhaps most) real-world problems in which verification neither guarantees solutions nor allows one to produce infinite training data effectively for free.
To expect it to be a universal solvent is to show you don’t understand that basic fact.
Part III: Seven more things to know about Astra
Yesterday's tweet and blog were marketing, not science. Neither of those nor the 249-page math article that went with them give any information about how this was accomplished:
In an email commenting on the first draft of this essay Ernie Davis added, quite rightly:
To properly evaluate the significance of these discoveries we would need some
critical pieces of information.
First: How many conjectures were attempted? If the OpenAI team picked these 10 conjectures at random from the space of all outstanding significant mathematical conjectures and Astra solved all 10, that would be amazing; but that seems altogether unlikely. If they cherry-picked 50 conjectures that they thought Astra would succeed on and it succeeded on 10, that’s still amazing, but significantly less so. If they ran Astra on all 1000 or so open Erdos conjectures and on 10,000 other open conjectures, then its failure to solve those is significant information on its limits as a mathematical reasoner.
Second: OpenAI brags loudly that this was done for a total computation cost of $2000. It seems a safe bet that this includes only the conjectures where Astra succeeded, not the ones where it failed. But what was the cost in terms of the salaries of the highly-paid mathematicians and computer scientists who worked on this project? I’d be astonished if it was less than $20,000 and would not be surprised if it was upward of $200,000.We don’t know this, and the way things are going, we may never know this. We still have zero information about the OpenAI system that achieved gold-medal performance at the 2025 International Mathematical Olympiad, and very little information about the Google system.
Astra is good at some math but it is very doubtful that math is “solved” (as a lot of people on X seemed to believe). It seems to be good at certain kinds of math
But quite possibly not all. Here are some examples from Eric Weinstein on X about some kinds of math it might be less good at
Ernie Davis also adds “The problem of autoformalization --- turning mathematics written by a human mathematician into a strictly logical form -- does not seem close to being solved, though certainly progress has been made. For example: For the last two years, mathematician Kevin Buzzard has been leading a project to formalize Andrew Wiles' proof of Fermat's Last Theorem in Lean. It is quite safe to say that no AI system is capable of carrying out that monumental task. Presumably if they could do it unassisted, they would scoop him, and if they could cut down his work from years to weeks, he would be using them.”
Consistent with my general prediction that facility in math doesn’t guarantee universality, it appears that Astra’s proofwriting is not on par with the proofs themselves (echoing something Ernie Davis and I observed here a year ago)” [click through for full tweet]
I seriously doubt that Astra will magically solve all the things that don’t work reliably in current models. Will it be able to reliably extract numbers from arbitrary PDFs? Doubt it. Will it solve Sabine Hossenfelder’s desire to have AI write scripts [click the tweet for more details] for her YouTube series? Doubt it.
Astra is not (as Musk suggested) the Singularity. Enough said. (But see my recent essay on the topic; nothing has really changed).
It is absurd to claim that the Astra debut is “plausibly the most significant day in the history of mathematics” as a bunch of people went around saying, based on a quote from one run of Fable (which might give different answers if asked twice). There was no new theory, no new techniques, and as impressive as it is, it is not on remotely on par with the development of calculus, algebra, logarithms, probability, the decimal system, information theory, the concept of zero. Says Davis “The claim quoted from Fable that "several of those would individually have been the result of the decade" is absurd … As a reality check, 14 of David Hilbert's 23 problems have been solved since 1900 --- that's one every nine years --- and these results are nowhere near that league. You could easily compile a list of 100 much more important results that have been proved since 1926.”
I doubt it will lead directly to a cure for cancer, or enormous advances in “materials research, energy production, drug discovery, Everything”, as some excitable AI “influencers” envision. Nor does this mean we just entered “the era of automated scientific discovery”. Against overwrought claims like these…
.. this just came out, echoing the new article by and others that I linked to a few days ago.
Part IV: Summary
Despite immense pushback on X yesterday (most of it ad hominem, some literally involving fabrication, a lot of it involving outright lies), my overall hot take on Astra remains what I first posted:

What’s sad about this is that Ernie Davis and I raised a lot of the same points about how one could properly examine new systems a year ago:

I get that people are excited about Astra, but anyone hoping for magic is likely to be disappointed.

