OpenAI 展示了 AI 能够攻克数学领域长期悬而未决的问题。专家们对由此带来的可能性感到兴奋——同时也担忧这对他们所在领域接下来意味着什么。
过去一年里,数学家 James Maynard 花了大量时间“自我反思”。作为牛津大学教授、享有盛誉的菲尔兹奖得主,Maynard 告诉 The Verge,随着这门传统上进展缓慢的学科匆忙适应 AI,他一直在思考自己领域的未来。
在我们交谈的几天前,OpenAI 披露它已给出 10 个长期悬而未决的数学问题的解答,其中一些问题困扰了学术界数十年。就像生成式 AI 被用于生成文本和图像,或在科学和医学领域提出构想一样,这项技术从其训练所用的大量材料中学习模式和关联,并利用它们创造出新的东西。应用于数学时,这可能意味着以新的方式组合已知的结果、方法和工具来攻克一个问题,有时在不同领域之间建立联系,或让埋没在学术文献中的概念重新浮现。
对于 Maynard 和其他接受 The Verge 采访的数学家来说,这一消息让他们对本领域走向的复杂情绪更加翻涌。对于加速数学发现的前景,人们有着显而易见的兴奋——但也有忧虑,某些情况下甚至是绝望,因为他们担心这对那些毕生致力于这一追求的人,以及将追随他们的未来一代代数学家意味着什么。很少有人怀疑,一场深刻的剧变已经在进行之中。
OpenAI 使用一款名为 Astra 的未发布先进模型所解决的问题,横跨了广泛的数学领域,从高度抽象的问题到具有实际意义的问题。其中一项突破涉及球体在三维以上空间中能有多紧密地堆积,这个问题与数据能被多高效地编码和传输相关。另一项突破推进了纠错码的极限,纠错码有助于从含噪信号中恢复信息。第三项突破解决了两个长期悬而未决的问题:复杂连通网络在结构性模式出现之前能变得多复杂。清单上的其他成果还攻克了量子博弈论中的问题,以及在后量子网络安全所用技术中具有意义的高维网格内目标搜索问题。
其中最引人注目的成果之一,涉及非 sofic 群的存在性——这类无限数学结构,粗略地说,无法用有限结构来近似。这类结构究竟是否存在,几十年来一直是一个悬而未决的问题。它引人注目还有另一个原因:围绕功劳应归于 OpenAI 的 AI 多少、又应归于其近期工作所依托的人类数学家多少,存在一场争议。
剑桥大学数学家 Francesco Fournier-Facio 告诉The Verge,他和该领域的其他人认为,OpenAI 最初的公告淡化了研究人员 Andreas Thom 和 Gábor Kun 的贡献,而他们近期的研究为这一成果奠定了基础。
当 OpenAI 首次发布其公告时,它表示分享的是“针对那些长期悬而未决、且主要结果至少十年、多数情况下远不止十年毫无进展的问题所取得的成果”。它后来修改了这一说法,改为分享“各项成果,每一项都解决了一个长期悬而未决的问题,或在其上取得了实质性进展”。该页面没有更正说明,也没有对这一改动做出任何解释。
Kun 是匈牙利阿尔弗雷德·雷尼数学研究所的研究员,他告诉The Verge,OpenAI 在发布前不久通过电子邮件联系了他,向他分享了其发现。他认为最初公告中那种笼统的措辞“相当滑稽”,尤其是所附的更详细研究论文“明确表示它建立在我 2016 年和 2019 年的成果之上”,后者是与 Thom 合著的。“这相当草率,”Kun 说。
发布之后,OpenAI 再次联系了 Kun。他拒绝分享完整的邮件往来,但向The Verge朗读了他所说的其中一段摘录,其中一位 OpenAI 数学家写道,那段措辞本意是指该合集里的其他成果。“它并非意在暗示这个问题没有任何进展,”他说,邮件中这样写道。“我们当然同意,这一论证关键依赖于你的工作。”这位数学家还补充说,他们会要求修改公告的措辞。
OpenAI 发言人 Laurance Fauconnet 向 The Verge 证实,该帖子随后进行了更新。“我们更新了措辞,以更好地反映这些结果所基于的先前研究。尽管非 sofic 群是否存在这一问题数十年来一直悬而未决,但我们的 sofic 群证明依赖于更近期发表的重要数学工作,我们希望确保这些贡献得到恰当的认可。”
Kun 表示,他怀疑其他结果中是否也存在类似的疏漏,而那些结果超出了他的专业领域。
更广泛地评估 OpenAI 的这些结果并非易事。数学已经变得如此专业化,以至于很少有研究者具备充分的能力去审视 Astra 所触及的所有领域。该公司发布了超过 250 页论文来阐述这些解答,另有 60 多页描述“这些想法是如何汇聚而成的”,并用 Lean(用于验证数学证明的软件)对每一项结果进行了认证。尽管 The Verge 采访的许多数学家表示,他们个人无法评估其中部分甚至全部解答,但所有人都表示,普遍共识是 OpenAI 的这一成就背后似乎确有实打实的分量。
这些问题——OpenAI 声称由其“下一代主要模型”Astra 的内部版本解决——绝非等闲。Maynard 表示,它们是数学家和计算机科学家曾投入大量时间思考、却屡屡未能解决的那类问题。
“人们普遍觉得,[解决]这 10 个问题中的任何一个都能让你在学术界找到一份工作,”伦敦数学科学研究所研究员何杨辉(Yang-Hui He)说。他刚从韩国一场为期四周的 AI 与数学研究会议回来,他说,许多人感到过去六个月里出现了一种“相变”,AI 带来了真正而有意义的进展。
今年 5 月,OpenAI 震惊了数学界,它宣布,一个未具名的内部模型破解了 Paul Erdős 提出的一个猜想,而这个猜想在近一个世纪的大部分时间里一直困扰着数学家。7 月,哈佛大学数学家 Levent Alpöge 发推称,Anthropic 的 Claude Fable 5 用一个微小的反例推翻了棘手的雅可比猜想,颠覆了数十年来试图证明其为真的努力。这些只是数学家们真正在意的问题被 AI 攻克的最新例子。
Maynard 回忆说,直到最近,情况并非总是如此。他说,除了一些值得注意的例外,AI 在数学领域的突破往往伴随着大量宣传,而所涉及的问题却常常很少受到研究者的认真关注。这种变化的速度让许多研究者措手不及,整个领域正忙于思考如何应对,一些人则在怀疑,它是否还能以任何接近当前形态的方式继续存续下去。
许多接受 The Verge 采访的数学家似乎仍在消化自己对这一切的看法,同时表达了对变化速度的惊讶,甚至是震惊。兴奋之情溢于言表,但他说,从韩国的那场会议以及更广泛的领域来看,他的印象是,领域内许多人在淡化近期进展的重要性,以求对事情推进的速度“保持冷静”。
金钱是许多这类担忧的核心。“目前还不太清楚我们的大学是否愿意为我们的定理支付那么高的费用,”苏格兰圣安德鲁斯大学的教授 Colva Roney-Dougal 说。“数学通常是一门廉价的学科,”她说,指了指一块写满她研究内容的白色书写板。“大多数时候我甚至懒得去申请研究经费。我不需要。我只要继续做我的工作就行。”
即便按照 AI 的标准这些成本相对不高,这套体系也可能崩溃。OpenAI 估计,按照其 Sol 模型当前的 API 价格,生成 Astra 的 10 个解答大约要花费 2,000 美元的 token 费用。接受 The Verge 采访的研究人员表示,真实成本很可能要高得多,具体取决于在成功之前尝试了多少问题和多少次尝试。当被问及最终名单是如何汇总而成时,OpenAI 没有详细说明。但对于一个习惯于靠微薄预算运作的领域来说,即便是公开标出的价格也可能过于高昂。Roney-Dougal 担心,规模较小、财力较弱的机构的研究人员可能会被完全挡在某些研究之外。
还有一种更广泛的不安,即商业利益正日益侵入一个基本上一直公开运作的领域。尽管数学已变得更加依赖计算,但研究人员所依赖的许多工具都是开源且可免费获取的。而 OpenAI 和 Anthropic 等公司最强大的模型是专有的,访问权限由这些公司严格控制。虽然两家公司都有项目为学术研究人员提供免费访问权限,但这种访问远非普遍,而且The Verge采访的研究人员中,很少有人能够使用最先进的系统。Maynard 和 Roney-Dougal 都表示希望开放权重模型最终能够弥合这一差距,让数学家无需依赖少数几家大型 AI 公司就能获得工具。
访问权限并不是唯一的担忧。The Verge采访的多位研究人员质疑,AI 公司的价值观是否与数学界的价值观一致。他们说,这些公司有产品要卖,还要为即将 IPO 所对应的巨额估值提供支撑,因此它们有充分的动机去炒作和夸大自己的贡献,同时贬低这些成果所依赖的人类学术工作。
“这是目前我最不满的一点,”Roney-Dougal 说。“他们把我们的学科当成了广告游乐场。”
这些担忧并不局限于 The Verge 采访过的研究人员。今年 6 月,数学家们发布了 《莱顿宣言》,这是一套关于在数学中负责任地使用 AI 的原则,已获得国际数学联盟的认可,并有超过 3,400 人签署。它敦促政策制定者、政府、媒体和其他群体不要轻信那些“夸大其产品能力”的公司所制造的“炒作”。人们担心的是,夸大的说法将对该领域产生切实后果,让资助方和政府相信人类数学家的必要性比实际上更低。
OpenAI 被指责在其 10 项 Astra 进展中正是这么做的,尤其是因为它淡化了 Kun 和 Thom 的贡献。Kun 说他的“感受相当矛盾”。看到自己的工作被证明有助于解决一个重要问题,这令人欣慰,尽管他希望完成它的人是自己;这也给他带来了原本可能永远不会得到的关注,以及大量的祝贺。他对祝贺者的玩笑是:他将成为一个“非常著名的失业者”。
Fournier-Facio 的感受则远没有那么矛盾。他说:“大多数人只会看看 OpenAI 的公告,然后照单全收。”他认为,很少有人有时间、专业知识或意愿去翻阅数百页的技术论文,弄清标题背后的人类劳动,尤其是那些需要快速出稿的记者或政策制定者。“
他们只是在选择对自己最有利的叙事,”他说。“说一个 AI 系统独立想出了人类 10 年来毫无进展的东西,要唬人得多。说这是在过去 10 年人类想法的基础上,用一种聪明的方式把它们组合起来,就没那么抓眼球了。”
Fournier-Facio 承认,数学界在如何分配功劳这件事上,本身就存在不公平。“走完最后一步的人拿走了大部分功劳,对吧?”他把数学研究比作建造一座金字塔,在最终放下最后一块石头的人之下,是几代人积累的工作。他说,如果解决这个问题的是人类,也许就不会有这么多人坚持要确保 Kun 和 Thom 得到署名。但如今是 AI 迈出了最后一步,他担心其他所有人都会被错误地“视为无用”。
更糟的是:“在人类的情况下,放下最后一块石头的人未必会抢走所有建造金字塔者的饭碗。”
当数学家们思考 AI 将如何影响下一代研究者时,这种恐惧尤其令人不安。牛津大学教授 Andras Juhasz 表示,大语言模型开始解决的许多问题,恰恰是研究生“通常会研究”的那类问题。它们并不总是特别引人注目,但钻研这些问题对于培养成为一名成功研究者所需的技能、直觉和习惯至关重要。苏黎世联邦理工学院研究员 Johannes Schmitt 表示,他们现在面临的风险是,被别人用 AI 抢先一步解决掉。
这种影响已经开始显现。Juhasz 表示,随着 AI 系统越来越有能力完成项目作业,它对于本科生来说已经“成为一种有问题的评估形式”。这些工具或许能帮助学生更快得到答案,但可能损害通过挣扎求索才能获得的整体数学理解。与此同时,Maynard 已经在担忧如何让学生的研究项目面向未来。“如果一篇可发表论文的标准是 AI 做不到的事情,尤其是当博士通常要读四年时,你要找的不是一个 AI 现在做不到的问题,而是 AI 四年后做不到的问题。”
这种不确定性已经在固化为绝望情绪,尤其是在职业生涯早期的数学家中。“与我交谈过的许多人真的、真的非常害怕,”Fournier-Facio 说。他向 The Verge 指出了两篇绝望的文章,出自在网上流传的研究生之手,质疑自己在该领域的前途。Roney-Dougal 说,她在社交媒体上也看到了类似的疑虑在蔓延,评论大意是“我不确定现在该做什么。我不知道自己继续下去还有什么意义。”Roney-Dougal 本人没那么悲观:“我对此没那么沮丧,”她说。“但很多人是的。”
问题不止于 AI 能解出多少道题。数学的进步不是靠把问题从清单上一条条划掉来实现的。一个解答最重要的意义可能在于它引出了什么。最好的成果能开辟出全新的问题、技巧,甚至研究领域。AI 能否做到这一点,尚待观察。反过来说,现在就断言落入 AI 之手的问题不过是唾手可得的低垂果实,也为时过早。要充分理解这些成果贡献了什么、它们最终被证明有多重要,还需要时间。
到目前为止,Schmitt 并不信服。他说,他几乎没有看到证据表明 AI 的重大贡献能像这样开辟出新的思考与探索领域,这引出了一个令人不安的前景:这项事业正变得越来越失衡。“我们可能正走向一种有些失衡的局面,大量优秀而有趣的问题被 AI 智能体横扫,”他说,而这些解答并不会反过来为研究者催生出富有成果的新方向。
并非所有人都如此不安。伦敦研究所的 He 对接下来会发生什么出奇地乐观。他表示,从事传统学术数学研究的人数出现一定程度的收缩,甚至可能是一个“健康的方向”,并指出数学专业毕业生本就面临严峻的就业市场。另一些人则看到了 AI 拓宽研究参与人群的潜力,让那些无法接触到集中在精英机构的专业知识的本科生和研究人员也能发展自己的想法,并有可能产出此前可能遥不可及的、可发表的研究成果。
而对于资深研究人员来说,与此同时,将更多常规工作交给 AI,可以腾出时间专注于数学中更具创造性和概念性的部分。
尽管有这些恐惧与希望,没有人知道这一切将走向何方。AI 发展得太快,而它所产出的数学成果还太新,无法判断其影响可能是什么。几位研究人员担心,在任何人还来不及理解 AI 究竟意味着什么之前,关于 AI 可能成为什么的种种宣称,就可能让这个领域朝着更糟的方向被重塑。
当被问及他是否认为 Astra 的任何成果配得上菲尔兹奖——数学界最高荣誉之一,也是他本人四年前获得的奖项——Maynard 表示,他的初步印象是这些成果还不够格。
“我看过的那些,确实有着非常令人印象深刻的成果的气质,”他说,但未必是那种在概念层面提升数学的成果。他表示,就目前而言,他看到的更多是现有技术和方法被进一步推进并以巧妙方式相互连接的证据,而非某种深刻的新数学思想。
“但我觉得现在下结论也为时过早,”他说,“有时候需要时间和视角才能意识到,‘哦,这里确实有一个根本性的新想法。’”
以目前的发展速度,AI 可能不会给数学家们太多时间去弄清楚这一点。
OpenAI showed that AI can tackle long-standing problems in mathematics. Experts are excited about the possibilities — and worried about what comes next for their field.
Mathematician James Maynard has spent a lot of time this past year “soul searching.” A professor at the University of Oxford and winner of the prestigious Fields Medal, Maynard told The Verge he’s been grappling with the future of his field as the traditionally slow-moving discipline hurries to adapt to AI.
Days before we spoke, OpenAI revealed it had produced the solutions to 10 long-standing mathematics problems, some of which had confounded academics for decades. Like generative AI used to produce text and images or propose ideas in science and medicine, the technology learns patterns and connections from the vast amount of material it’s trained on and uses them to create something new. Applied to mathematics, that can mean combining known results, methods, and tools in new ways to attack a problem, sometimes drawing links between disparate fields or resurfacing concepts buried in academic literature.
For Maynard and other mathematicians The Verge spoke to, the announcement has added to a complex swirl of emotions about where their field is headed. There is palpable excitement at the prospect of accelerating mathematical discovery — but also apprehension, and in some cases despair, about what this could mean for the people who have dedicated their lives to the pursuit, and for the generations of future mathematicians who will follow them. Few doubt that a profound upheaval is already underway.
The problems OpenAI solved, using an advanced unreleased model known as Astra, spanned a wide range of mathematical fields, from the highly abstract to questions with practical implications. One breakthrough concerned how tightly spheres can be packed in more than three dimensions, a problem linked to how efficiently data can be encoded and transmitted. Another pushed the limits of error-correcting codes, which can help recover information from noisy signals. A third resolved two long-standing questions about how complex connected networks can become before structural patterns emerge. Other results on the list tackled problems in quantum game theory and the search for targets inside high-dimensional grids, with implications for techniques used in post-quantum cybersecurity.
One of the most attention-grabbing results concerned the existence of non-sofic groups, infinite mathematical structures that, roughly speaking, cannot be approximated by finite ones. Whether such structures existed at all had remained an open question for decades. It was attention-grabbing for another reason, too: a dispute over how much credit belonged to OpenAI’s AI, and how much to the human mathematicians whose recent work it built upon.
Francesco Fournier-Facio, a mathematician at the University of Cambridge, told The Verge he and others working in that area believed OpenAI’s original announcement minimized the contributions of researchers Andreas Thom and Gábor Kun, whose recent works helped lay the groundwork for the result.
When OpenAI first published its announcement, it said it was sharing “results to problems that have been open and have seen no progress on the main result for at least a decade, and in most cases much longer.” It later changed this to say it was sharing “results, each of which resolves or makes substantial progress on a long-standing open problem.” The page contains no correction note or explanation for the change.
Kun, a researcher at the Alfréd Rényi Institute of Mathematics in Hungary, told The Verge OpenAI reached out to him by email shortly before publishing to share its findings. He found the sweeping language in the original announcement “rather comical,” particularly as the more detailed research paper attached “clearly said that it builds on my results from 2016 and 2019,” the latter coauthored with Thom. “It’s rather sloppy,” Kun said.
After publication, OpenAI contacted Kun again. He declined to share the email exchange in full but read The Verge what he said was an excerpt in which an OpenAI mathematician wrote that the language had been intended to refer to other results in the collection. “It was not intended to suggest that there had been no progress on this problem,” the email read, he said. “We certainly agree that the argument relies crucially on your work.” The mathematician added that they would ask for the wording of the announcement to be revised.
OpenAI spokesperson Laurance Fauconnet confirmed to The Verge that the post was subsequently updated. “We updated the language to better reflect the prior research these results build upon. Although the question of whether non-sofic groups exist had remained open for decades, our sofic group proof relies on important mathematical work published more recently, and we wanted to ensure those contributions were properly acknowledged.”
Kun said he wondered whether similar oversights might have been made in the other results, which were outside his areas of expertise.
Assessing OpenAI’s results more broadly is complicated. Mathematics has become so specialized that few researchers are fully equipped to scrutinize all of the fields Astra touched upon. The company released more than 250 pages of papers laying out the solutions, along with 60 more pages describing “how the ideas came together,” and certified each result with Lean, software for verifying mathematical proofs. While many of the mathematicians The Verge spoke to said they couldn’t personally evaluate some or even any of the solutions themselves, all said there is a broad consensus that there seems to be real weight behind OpenAI’s achievement.
The problems, which OpenAI claimed were solved by an internal version of its “next major model,” Astra, were not trivial. Maynard said they were the kind of questions mathematicians and computer scientists had spent serious time thinking about, and repeatedly failed to resolve.
“There’s a general feeling that [solving] one of these 10 problems would get you a job in academia,” said Yang-Hui He, a fellow at the London Institute for Mathematical Sciences. He had just returned from a four-week AI and mathematics research conference in South Korea, where he said many felt there had been something of a “phase transition” over the past six months, with AI producing genuine and meaningful advances.
In May, OpenAI stunned mathematicians when it announced that an unnamed internal model had cracked a conjecture by Paul Erdős that had eluded mathematicians for the better part of a century. In July, Harvard mathematician Levent Alpöge tweeted that Anthropic’s Claude Fable 5 had disproved the fiendish Jacobian conjecture with a tiny counterexample, overturning decades of efforts to prove it was true. They are among the latest examples of problems mathematicians seriously care about falling to AI.
Maynard recalled that, until recently, this was not always the case. With some notable exceptions, he said AI breakthroughs in mathematics generated a lot of publicity while involving problems that had often attracted little serious attention from researchers. The speed at which this is changing has caught many researchers off guard, leaving the field scrambling to work out how to respond, and some wondering whether it can continue to survive in anything like its current form.
Many of the mathematicians The Verge spoke to seemed to still be working out what they thought about it all, while expressing surprise, even shock, at the speed of change. There was plenty of excitement, but He said his impression from the conference in South Korea, and from the field more broadly, is that many in the field are downplaying the significance of recent advances in a bid to “keep calm” about how quickly things are moving.
Money is at the heart of many of these concerns. “It’s not quite clear whether our universities are going to be willing to pay that much for our theorems,” said Colva Roney-Dougal, a professor at the University of St Andrews in Scotland. “Maths is a cheap discipline typically,” she said, gesturing to a whiteboard covered in her work. “Most of the time I don’t even bother getting a research grant. I don’t need one. I just get on with my job.”
That system could break down even if costs are relatively modest by AI standards. OpenAI estimates that generating Astra’s 10 solutions would have cost around $2,000 in tokens at the current API prices for its Sol model. Researchers The Verge spoke to said the true cost was likely considerably higher, depending on how many problems and attempts preceded the successful ones. OpenAI did not elaborate when asked about how the final list was assembled. But for a field used to operating on a shoestring budget, even the advertised price could be too much. Roney-Dougal fears researchers at smaller and less wealthy institutions could be locked out of some research entirely.
There is also a broader unease about the growing intrusion of commercial interests into a field that has largely operated in the open. Even as mathematics has become more computational, many tools researchers rely on are open source and freely available. The most capable models from companies like OpenAI and Anthropic are proprietary and access is tightly controlled by the companies. While both have programs offering free access to academic researchers, that access is far from universal, and few of the researchers The Verge spoke to had been able to use the most sophisticated systems. Both Maynard and Roney-Dougal expressed hope that open-weight models could eventually close that gap, giving mathematicians access to tools without having to rely on a handful of big AI companies.
Access isn’t the only concern. Several researchers The Verge spoke with questioned whether the values of AI companies align with those of the mathematical community. With products to sell and enormous valuations for impending IPOs to justify, companies have every incentive to hype and exaggerate their contributions, they said, while underselling the human scholarship those results rely on.
“That’s the bit I’m most unhappy about at the moment,” Roney-Dougal said. “They’re treating our discipline as an advertising playground.”
Those concerns extend beyond the researchers The Verge spoke to. In June, mathematicians published the Leiden Declaration, a set of principles for the responsible use of AI in mathematics that has been endorsed by the International Mathematical Union and signed by more than 3,400 people. It urges policymakers, governments, the media, and other groups to not buy into “the hype” created by companies who “overstate the capabilities of their products.” The fear is that exaggerated claims will have real consequences for the field, convincing funders and governments that human mathematicians are less necessary than they actually are.
OpenAI has been accused of doing just that with its 10 Astra advances, not least through its glossing over the contributions of Kun and Thom. Kun said his “feelings are quite ambivalent.” It was gratifying to see his work prove useful in resolving a significant problem, even if he wished he had been the one to finish it, and it brought him attention he otherwise might never have received, along with plenty of congratulations. His joke to well-wishers: He will be a “very famous unemployed” person.
Fournier-Facio feels considerably less ambivalent. “Most people will just look at the OpenAI announcement and take it at face value,” he said. Few people, he argued, will have the time, expertise, or inclination to dig through hundreds of pages of technical papers to understand the human work behind the headline, particularly journalists or policymakers working quickly. “They’re just choosing the narrative that benefits them most,” he said. “It’s a lot more impressive to say that an AI system came up independently with something that humans have done nothing on for 10 years. It’s a lot less sexy to say that this is a kind of building on ideas from the past 10 years from humans and combining them in a clever way.”
There is already an element of unfairness in how mathematics assigns credit, Fournier-Facio acknowledged. “The person that does the last step gets most of the credit, right?” He likened mathematical research to building a pyramid, with generations of work accumulated beneath the person who finally places the last stone. Maybe there wouldn’t have been this much insistence on making sure Kun and Thom were credited if it were a human solving this problem, he said. But with AI taking that final step, he worries that everyone else will wrongly “be seen as useless.”
Worse still: “In the case of humans, the human that puts the last stone in is not necessarily going to steal the job of all the people that built the pyramid.”
That fear is especially troubling when mathematicians consider how AI will impact the next generation of researchers. Many of the problems LLMs are beginning to solve are precisely the kind that graduate students “typically work on,” said Oxford professor Andras Juhasz. They are not always particularly flashy, but working through them is crucial for developing the skills, intuition, and habits needed to become successful researchers. They now risk being scooped by someone swooping in and solving it with AI, said ETH Zurich researcher Johannes Schmitt.
The effects are already beginning to be felt. Project work has “become a problematic form of assessment” for undergraduates as AI systems become more capable of completing it, Juhasz said. The tools may help students get answers quicker, but risk harming the overall mathematical understanding that comes through struggle. Maynard, meanwhile, is already worrying about how to futureproof research projects for students. “If the standard for a publishable paper is something that an AI can’t do, particularly when a PhD is typically four years, you’re not trying to come up with a problem that AI can’t do now. It’s AI in four years’ time.”
This uncertainty is already hardening into despondency, particularly among mathematicians early on in their careers. “Many of the people that I talked with are really, really scared,” said Fournier-Facio. He pointed The Verge to two despairing essays from graduate students circulating online questioning their futures in the field. Roney-Dougal said she had seen similar doubts spreading across social media, with comments along the lines of “I’m not sure what I should be doing now. I don’t know if there’s any point in me carrying on.” Roney-Dougal herself is less pessimistic: “I’m not that depressed about it,” she said. “But plenty of people are.”
The issue goes beyond how many problems AI can solve. Mathematics does not progress by simply checking problems off a list. A solution can matter most for what comes next. The best results can open up entirely new questions, techniques, or even fields of research. It remains to be seen whether AI can do that. Conversely, it is too early to dismiss the problems falling to AI as simply low-hanging fruit. Fully understanding what these results contribute, and how important they ultimately prove to be, will take time.
So far, Schmitt is unconvinced. He said he has seen little evidence of major AI contributions opening up new areas of thought and inquiry like this, raising the unsettling prospect of an increasingly lopsided endeavor. “We might be heading for a somewhat imbalanced situation, where lots of good and interesting problems get mowed down by AI agents,” he said, without those solutions feeding back into the creation of fruitful new directions for researchers to pursue.
Not everyone is so unsettled. The London Institute’s He was strikingly optimistic about what comes next. He said some contraction in the number of people pursuing traditional academic mathematics could even be a “healthy direction,” pointing to an already brutal job market for mathematics graduates. Others saw the potential for AI to broaden who gets to participate in research, allowing undergraduates and researchers without access to expertise concentrated at elite institutions to develop their ideas and potentially produce publishable work that might previously have been beyond their reach. For established researchers, meanwhile, handing off more routine work to AI could free up time to focus on more creative and conceptual parts of mathematics.
For all the fears and hopes, nobody knows where this is going. AI is moving too fast, and the mathematics it is producing is still too fresh to judge what its impact may be. Several researchers worried that the field could be reshaped for the worse by claims about what AI could become before anyone has had time to understand what it actually means.
When asked whether he thought any of Astra’s results were worthy of a Fields Medal, one of mathematics’ highest honors and an award he himself received four years ago, Maynard said his initial impression was that they fell short.
“The ones that I’ve looked at, they have the flavor of being very impressive results,” he said, but not necessarily ones that elevate mathematics at a conceptual level. For now, he said he sees more evidence of existing techniques and methods being pushed further and connected in clever ways than of some profound new mathematical ideas.
“But I think it’s also too early to say,” he said. “Sometimes it takes time and perspective to realize, ‘Oh, here is really a fundamental new idea.’”
At the speed it is going, AI may not give mathematicians much time to figure it out.