OpenAI 的一个内部推理模型推翻了匈牙利数学家保罗·埃尔多斯提出的所谓“单位距离猜想”。OpenAI 在公布这一结果的同时,还发布了一篇由九位外部数学家共同撰写的配套论文,他们对这一证明进行了验证、精简并给出了评注。
这个问题本身看似简单:在一张纸上放置一定数量的点。有多少对点之间的距离恰好为一个单位?1946年,埃尔多斯猜想,在一个略微倾斜的方形网格上进行简单排列,就已经接近最优解。这种排列产生的点对数量,其增长速度仅略快于点数本身。据数学家托马斯·布鲁姆称,埃尔多斯曾悬赏500美元寻求反例。根据标准参考书《离散几何研究问题》的说法,该问题被认为是“组合几何中可能最著名(且最容易解释)的问题”。
八十年后更优的构造
OpenAI 的模型发现了一种新的点排列方式,其产生的单位距离点对数量明显多于经典的方形网格。普林斯顿大学的威尔·萨温指出,每将点数翻倍,新构造带来的增益大约为百分之一。这听起来很小。但在上下文中,这意义重大,因为埃尔多斯的猜想认为几乎不可能实现任何此类增益。不过,这个问题并未完全解决:自1984年以来已知的一个理论上界,仍然远高于新构造所达到的水平。
引人注目的是所用工具的来源:并非几何学,而是代数数论。该模型没有使用经典的点网格,而是利用了复数系统,其内部对称性可以转化为特别密集的点阵模式。这些工具在数论领域已应用数十年。然而,参与其中的数学家们认为,将它们应用于平面几何中的一个基础问题,此前被认为是相当牵强的。
人类为何错过了这个解法
托马斯·布卢姆在合著论文的投稿中写道,要让人类找到这个解法,需要同时满足四个条件:必须在该问题上投入大量时间,敢于挑战埃尔德什的既定观点并实际尝试证伪,希望将原始构造转化到数域领域,并且对相当专业的类域论有足够了解。“AI 满足了所有这些条件,”布卢姆写道。它结合了“超人的耐心与对大量技术工具的熟悉程度”。
萨温补充了一个技术原因,说明为什么显而易见的推广方法会失败。自然的做法是选择一个扩展数系,然后观察其中越来越大的区块,本质上是在一个更复杂的数世界中放大原有的网格。根据萨温的说法,这只会回到埃尔德什的旧有界。该模型的关键技巧恰恰相反:它在每个数系内保持尺度不变,但每一步都切换到越来越丰富的数系。萨温写道,为什么这种特定的切换有效,对人类来说并不明显。
就在AI给出解法前一个月,布卢姆还在博客中将这个问题列为他的“埃尔德什十大难题”之一。他的动机是:一些观察者看到早期AI解决了更简单的埃尔德什问题后,便得出结论,认为这位数学家的所有问题都很简单。布卢姆想证明,许多埃尔德什问题催生了数十年的深度研究方法。
单位距离猜想是他列表中唯一的离散几何问题,正因为它“几十年来一直未被证明”。布卢姆指出,斯宾塞、塞迈雷迪和特罗特于1984年确立的上界在40多年里从未被改进过:“这个问题是一个很好的例子,说明尽管近年来离散几何领域取得了一些惊人成果,但我们距离理解甚至一些最基本的问题仍有很长的路要走。”他没想到AI仅仅一个月后就攻克了这个特定问题:“虽然我相信AI最终至少能在该列表中的几个问题上取得进展,但我没想到这会在仅仅一个月后就发生!”
数学界的反应
著名组合数学家诺加·阿隆称这一成果为“杰出成就”,并描述这一令人惊讶的发现是“其构造与分析以优雅而巧妙的方式运用了代数数论中相当精深的工具”。菲尔兹奖得主蒂莫西·高尔斯写道,如果这篇论文由人类提交给《数学年刊》并请求快速评审,“我会毫不犹豫地建议接受”。此前没有任何AI生成的证明能达到这一水平。高尔斯称其为“AI数学领域的一个里程碑”。
数论学家阿鲁尔·尚卡尔认为,这项工作表明当前的AI模型“已不仅仅是人类数学家的助手——它们能够产生独创性的天才想法,并将其付诸实现”。布鲁姆对此加以限定:该证明并未提供任何全新的几何工具,而一个完整的猜想证明很可能需要这类工具。但它表明“数论构造对于这类问题的解释力远超我们此前的预期”。他预计“未来几个月内,许多代数数论学家将密切关注离散几何中的其他未解问题”。
为何这次情况不同
近几个月来,AI系统已经解决或部分解决了一系列Erdős问题。由布鲁姆维护的网站erdosproblems.com收录了约1000个问题。据菲尔兹奖得主陶哲轩称,截至2025年9月,其中约380个已被解决。在2026年初的一段混乱时期里,又有大约50个问题被攻克,有些由人类完成,有些由AI完成,有些则是两者合作的结果。其中几个解答只有几页篇幅,或者相当于具有挑战性的课后习题水平。
这正是促使布鲁姆编制其前十名榜单的原因。他注意到,“不幸的是,我观察到一些数学家近来对Erdős问题越来越不屑一顾,或许是因为他们看到AI解决了这个网站上一些相当简单的问题,便错误地推广开来,认为Erdős提出的所有问题都只是奥林匹克竞赛级别的有趣小把戏。”
单位距离反证被明确归入不同类别,无论是在配套论文中还是在 OpenAI 看来都是如此。据 OpenAI 称,这是"人工智能首次自主解决一个数学子领域核心的著名开放问题"。Bloom 描述了自己的反应:当他得知这是一个反证时,他的巨大惊喜"略有减弱",而看到具体构造后,惊喜感进一步降低。
尽管如此,这一发现依然成立:与之前的 Erdős 解法不同,这并非一个易于理解的练习。这是一个被公认为八十年难题的问题,其上界自 1984 年以来从未改变,其解决方案需要借助一个遥远领域的工具。
Gowers 总结道:如果这项工作是由人类提交的,他会"毫不犹豫地"将其发表在《数学年刊》上。此前没有任何 AI 生成的证明能达到这一水平。
这一结果对数学本身意味着什么
参与其中的几位数学家利用配套论文反思了 AI 对其领域贡献的结构性影响。合著者 Daniel Litt 提出了一些令人不安的问题:为什么存在一些可以用相对简短、巧妙的论证来解决的著名问题?他的猜测是:要么是研究人员固守着次优的假设——比如 Erdős 本人认为自己的猜想是正确的——要么是解决方案需要用到相关领域大多数人并不熟悉的其他领域的思想。
"如果这些解释是正确的,它们应该让我们感到不安,"Litt 写道。"它们表明,追求专业化和各自为政的倾向虽然可以理解,却让我们付出了一些高质量科学的代价。"Litt 将人类的研究方式——研究者出于个人好奇心深入钻研少数几个问题——与当前 AI 系统性地遍历整个问题列表的模式进行了对比。这相当于"对数学问题投入的关注度得到了极大的扩展"。
高尔斯对自己的反应很坦诚。当他最初以为 AI 是证明了该猜想而非证伪时,他花了一整个晚上“调整我的世界观:如果 AI 能拿出这样的证明,那么数学家们可能很快就要全军覆没了。”第二天早上,当错误被澄清时,他感到“如释重负”。证伪可以被想象成耐心和反复试错的结果。而真正的证明则需要“深刻的洞察力”,那才是令人不安的。
在配套论文中,高尔斯提出了自己衡量证明难度的标准。他称之为“专家视角下的柯尔莫哥洛夫复杂度”——即专家为独立重构该证明所需的最短提示序列长度。他的初步判断是:AI 尚未在整体上超越人类,但在某些特定类型的问题上具有优势。它拥有“数学领域的百科全书式知识”,不太担心时间管理问题,因此可以“不惜代价地努力去证明那些看似不太可能成立的命题。”
即便如此,他表示进步不会停滞。很快就会出现“我们事后很难将其轻描淡写为‘比预期简单’”的 AI 解决方案。即使 AI 无法找到冗长复杂的证明,“我们很可能已经进入了一个人类在解决数学问题上很难与 AI 竞争的时代。”
布鲁姆持中间立场。对于他自己提出的测试问题——该证明是否让该领域对问题有了新的认识——他给出了“有保留的肯定”作为回答。数论构造方法对这些问题的解释力似乎超出了所有人的预期,而所需的数论知识也可能非常深奥。该领域的某些人可能会失望,因为该证明没有提供“强大的新几何工具”或意想不到的结构性结果——而一个完整的猜想证明很可能需要这些。这个解法,“事后看来”,是一个自然的推广,但“极其不平凡”。人类要发现它,需要四种罕见的巧合同时发生。
布卢姆这样描述该 AI 的优势:它兼具“超人级别的耐心”与“对海量技术工具的熟悉”,并执着地追寻“人类可能认为不值得花时间探索的路径”。他的看法是:“知识的边界非常崎岖,毫无疑问,未来数月乃至数年,我们将在数学的许多其他领域看到类似的成功——那些长期悬而未决的开放问题,将由 AI 通过揭示意想不到的联系、将现有技术工具推向极限而得到解决。”
人类与机器分工协作
OpenAI 同步发表的配套论文本身,就预示了未来 AI 与研究人员之间如何分工。据布卢姆称,模型生成的原始证明“完全有效”,但人类作者对其进行了“大幅改进”。只有索温的优化才得出了具体的改进度量。最终印在配套论文中的版本,比原始版本更简洁、更具普适性。
这第二步正是陶哲轩近期在斯坦福大学“数学未来研讨会”上演讲的焦点。陶哲轩认为,数学实践目前正经历“证明消化不良”:AI 系统生成和验证证明的速度越来越快,但人类的“消化”过程——即理解、阐释、将结果置于上下文中并在此基础上进一步推进——却跟不上。他判断一个解法是否真正完整的标准是:能否有人就此做一场演讲并回答提问?在单位距离反例证明这个案例中,九位著名数学家同意承担这项工作。这一标准能否大规模推广,则完全是另一个问题了。
An internal reasoning model from OpenAI has disproved the so-called unit distance conjecture posed by Hungarian mathematician Paul Erdős. OpenAI announced the result alongside a companion paper written by nine external mathematicians who verified, shortened, and commented on the proof.
The problem itself is deceptively simple: place a certain number of points on a sheet of paper. How many pairs of points can be exactly one unit apart? In 1946, Erdős conjectured that a simple arrangement on a slightly skewed square grid was already close to optimal. That arrangement produces a number of pairs that grows only barely faster than the number of points itself. According to mathematician Thomas Bloom, Erdős had offered $500 for a disproof. The problem is considered "possibly the best known (and simplest to explain) problem in combinatorial geometry," according to the standard reference Research Problems in Discrete Geometry.
A better construction after eight decades
OpenAI's model found a new point arrangement that produces noticeably more unit-distance pairs than the classic square grid. Will Sawin of Princeton University puts the gain at roughly one percent more pairs per doubling of the point count. That sounds small. In context, it's significant, as Erdős's conjecture said virtually no such gain was possible at all. The problem isn't fully solved, though: a theoretical upper bound known since 1984 still sits well above what the new construction achieves.
What's striking is where the tools came from: not geometry, but algebraic number theory. Instead of working with classical point grids, the model used complex number systems whose internal symmetries translate into especially dense point patterns. These tools have been standard in number theory for decades. Applying them to a basic problem in plane geometry, however, was considered far-fetched by the mathematicians involved.
Why humans missed the solution
Thomas Bloom writes in his contribution to the companion paper that four conditions had to line up for a human to have found this solution: you had to spend serious time on the problem, bet against Erdős's established opinion and actually attempt a disproof, want to translate the original construction into the world of number fields, and be sufficiently familiar with the fairly specialized class field theory. "The AI met all of these criteria," Bloom writes. It combines "superhuman levels of patience with familiarity with a vast array of technical machinery."
Sawin adds a technical reason why obvious generalizations failed. The natural approach would have been to pick one extended number system and look at bigger and bigger chunks of it, essentially inflating the old grid in a more complicated number world. According to Sawin, that just leads back to the old Erdős bound. The model's key trick was the opposite: it kept the scale fixed within each number system but switched to progressively richer number systems at every step. Why that particular switch works wasn't obvious to any human, Sawin writes.
Bloom had listed this problem just one month before the AI solution in a blog post as one of his "Top 10 Erdős Problems." His motivation: some observers had looked at earlier AI solutions to simpler Erdős problems and concluded that all of the mathematician's questions were trivial. Bloom wanted to show that many Erdős problems have spawned decades of deep methods.
The unit distance conjecture was the only discrete geometry problem on his list, precisely because it "has resisted proof for decades." Bloom pointed out that the upper bound established in 1984 by Spencer, Szemerédi, and Trotter hadn't been improved in over 40 years: "This problem serves as a great example that, despite some spectacular results in recent years in discrete geometry, we are still a long way from understanding even some of the most basic questions." He didn't expect an AI to crack this particular problem just one month later: "While I believed that AI would make some progress on at least a couple of the problems in that list eventually, I did not expect this to happen just one month later!"
Reactions from the math community
Noga Alon, one of the leading combinatorialists, calls the result an "outstanding achievement" and describes the surprising finding as a "construction and its analysis apply fairly sophisticated tools from algebraic number theory in an elegant and clever way." Fields Medalist Tim Gowers writes that if a human had submitted the paper to the Annals of Mathematics and asked for a quick assessment, "I would have recommended acceptance without any hesitation." No previous AI-generated proof has come close. Gowers calls it "a milestone in AI mathematics."
Number theorist Arul Shankar sees the work as evidence that current AI models "go beyond just helpers to human mathematicians - they are capable of having original ingenious ideas, and then carrying them out to fruition." Bloom qualifies that: the proof doesn't deliver any fundamentally new geometric tools, the kind a complete proof of the conjecture would likely require. But it shows that "there is a lot more that number theoretic constructions have to say about these sorts of questions than we suspected." He expects "many algebraic number theorists will be taking a close look at other open problems in discrete geometry in the coming months."
Why this case is different
AI systems had already solved or partially solved a whole series of Erdős problems in recent months. The platform erdosproblems.com, maintained by Bloom, catalogs around 1,000 problems. According to Fields Medalist Terence Tao, about 380 of them were solved by September 2025. During a chaotic stretch in early 2026, roughly 50 more fell, some by humans, some by AI, some by a mix. Several of those solutions fit on a few pages or were at the level of challenging homework exercises.
That's exactly what drove Bloom to compile his top-10 list. He noticed that he "have, unfortunately, seen some mathematicians grow dismissive of Erdős problems recently, perhaps because they have seen reports of AI solving problems on this site that turned out to be quite simple, and wrongly generalised this to assume that all problems posed by Erdős are amusing novelties, of the level of olympiad problems."
The unit distance disproof is explicitly placed in a different category, both in the companion paper and by OpenAI. According to OpenAI, this is "the first time that a prominent open problem, central to a subfield of mathematics, has been solved autonomously by AI." Bloom describes his own reaction: his big surprise was "dampened slightly" when he learned it was a disproof, and dampened further when he saw the construction.
Still, the finding stands: unlike the previous Erdős solutions, this isn't an accessible exercise. It's a problem considered hard for eight decades, with an upper bound unchanged since 1984, whose solution required tools from a distant field.
Gowers sums it up: had a human submitted the work, he would have accepted it for the Annals of Mathematics "without any hesitation." No prior AI-generated proof has come close.
What the result says about math itself
Several of the mathematicians involved use the companion paper to reflect on the structural consequences of AI contributions to their field. Co-author Daniel Litt asks uncomfortable questions: why do famous problems exist that can be solved with a relatively short, clever argument? His guess: either researchers cling to suboptimal assumptions—like Erdős's own belief that his conjecture was correct—or the solution demands ideas from areas most people in the relevant field don't know well.
"These explanations, if correct, should cause us some discomfort," Litt writes. "They suggest that incentives towards specialization and silo-ing, though understandable, have cost us some high-quality science." Litt contrasts the human approach, where a researcher digs deep into a few questions out of personal curiosity, with the current AI mode of systematically working through entire problem lists. That amounts to "a vast expansion of the attention aimed at mathematical problems."
Gowers is candid about his own reaction. When he initially assumed that the AI had proved the conjecture rather than disproved it, he spent the evening "adjusting my world view: if AI could come up with a proof like that, then maybe it would be all over for mathematicians very soon." The next morning, when the mistake was cleared up, it was "a big relief." A disproof can be imagined as the result of patience and trial and error. A real proof would have required "deep insight" and that would have been unsettling.
In the companion paper, Gowers develops his own measure for proof difficulty. He calls it "Kolmogorov complexity modulo experts" - the length of the shortest sequence of hints an expert would need to reconstruct the proof independently. His tentative take: AI isn't broadly better than humans yet, but it has advantages on certain problem types. It has "encyclopaedic knowledge of mathematics," worries less about time management, and can therefore "afford to try quite hard to prove statements that seem unlikely to be true."
Even so, he says progress won't plateau. There will soon be AI solutions "that we will find hard to explain away as easier than expected with hindsight." Even if AI can't find long, complex proofs, "we have still probably entered an era where it will become very difficult for humans to compete with AI at solving mathematical problems."
Bloom takes a middle position. To his own test question—whether the proof taught the field something new about the problem—he answers with a "moderated yes." Number-theoretic constructions apparently have more to say about these kinds of questions than anyone suspected, and the required number theory can run very deep. Some in the field may be disappointed that the proof doesn't deliver "powerful new geometric tools" or unexpected structural results, the kind a full proof of the conjecture would likely need. The solution is, "with the benefit of hindsight," a natural generalization, but "highly non-trivial." It took four rare coincidences for a human to have found it.
Bloom describes the AI's strength this way: it combines "superhuman levels of patience" with "familiarity with a vast array of technical machinery" and stubbornly pursues "paths that a human may have dismissed as not worth their time to explore." His outlook: "The frontiers of knowledge are very spiky, and no doubt the coming months and years will see similar successes in many other areas of mathematics, where long-standing open problems are resolved by an AI revealing unexpected connections and pushing the existing technical machinery to its limit."
Humans and machines split the work
The companion paper OpenAI published is itself a preview of how labor might be divided between AI and researchers going forward. The original proof generated by the model was, according to Bloom, "completely valid," but the human authors "significantly improved" it. Only Sawin's refinement produced the concrete measure of improvement. The version printed in the companion paper is shorter and more general than the original.
That second step was recently the focus of a talk by Tao at the Future of Mathematics Symposium at Stanford. Tao argues that mathematical practice is currently experiencing "proof indigestion": AI systems generate and verify proofs faster and faster, but human digestion, meaning understanding, explaining, contextualizing, and building on results, can't keep up. His bar for whether a solution is truly complete: can someone give a talk about it and answer questions? In the case of the unit distance disproof, nine prominent mathematicians agreed to do exactly that work. Whether that standard can scale is another question entirely.