纽约大学数学教授特里斯坦·巴克马斯特周二宣布了三项证明,其中一项初步成果涉及理论数学中一个重大未解问题。这些发现是与 Anthropic 数学家莱文特·阿尔波格合作完成,并同时使用了 Codex 和 Claude AI 模型。这些成果本身意义重大——但同时也伴随着一场围绕 OpenAI 试图解决同一问题的不寻常争议。
“这件事还有另一面,”巴克马斯特在宣布这些证明的声明中写道,“说实话,我非常希望自己不必为此操心。”
就在巴克马斯特和阿尔波格即将完成他们的成果时,他们得知“关于我们研究进展的信息已被传递给 OpenAI”。当他们就此联系 OpenAI 时,对方告知 OpenAI 已经完整证明了这一核心问题。但当他们追问 OpenAI 何时开始研究该问题、以及其中涉及多少人工输入时,对方的回答变得越发含糊其辞。
“后来发现,原来有一个完整的团队一直在攻克这个问题,”巴克马斯特说,“而且动用了极其庞大的算力……最终,大家一致认为,第一条提示词是在过去几天内发出的,当时关于我们工作的信息已经传到了OpenAI。”
如果此事属实,那就意味着OpenAI团队已经确信巴克马斯特和阿尔珀格的思路是正确的,并决定利用自身在算力资源上的巨大优势,率先达成形式化证明。
负责领导 OpenAI 数学研究的 Sebastian Bubeck 表示,这些说法“虚假且具有煽动性”。
“需要澄清的是,我是遵循学术规范参与讨论的,我对事情发展到这一步感到失望,”他在 Buckmaster 发表声明后的一篇帖子中写道。“了解我的人都知道,学术标准对我来说是最重要的。”Bubeck 表示,他很快会发表一份更完整的声明。
这场争议的核心是“纳维-斯托克斯方程的存在性与光滑性”问题,这是七个“千禧年大奖难题”之一——这是一组重要的未解数学难题,克莱数学研究所为每个问题的首位解答者(个人或团队)提供 100 万美元奖金。纳维-斯托克斯方程在流体力学中被广泛使用,但在理论层面却鲜有深入理解。若有人能解答该问题,将代表着人类对数学物理整体认知的一次重大飞跃。
尽管这个问题在数学家中被广泛研究,但 Buckmaster 及其合作者所采用的具体策略却远非普遍做法。因此,Buckmaster 认为 OpenAI 最终在同一时间采用相同方法,这一点十分可疑。
“通过光滑外力来解决这个克莱问题的路径,即 Fefferman 对问题陈述中的选项 c 和 d,是 Luis 和 Diego 开辟的路线,也是 Levent 和我悄悄选择要攻克的方向,”Buckmaster 写道。“据我所知,几乎没有人在这条路线上工作,”他继续说道。“这不是把问题陈述丢给一个模型、几天之内就能得出的方向。”
虽然 Alpöge 受雇于 Anthropic,但他并非代表该公司开展这项研究。因此,两人在研究中使用了一系列不同的模型,主要依赖 OpenAI 的 Codex。即便如此,Alpöge 与竞争对手实验室的关联似乎仍是 OpenAI 的一个痛点,Buckmaster 声称 Bubeck 曾要求他移除对 Alpöge 的致谢,作为提议的折中方案的一部分。
当 Buckmaster 推动将争议公之于众时,他说 Bubeck 的回应是:“你为什么要毁掉自己的职业生涯?”Buckmaster 表示,当他反驳时,Bubeck 接着说道:“如果你不希望我客气,那我也不必客气。”
Buckmaster 还提出了担忧,因为他在构建该项目时大量使用了 Codex,他工作中的信息可能已经为 OpenAI 自己解决该问题的努力提供了参考。OpenAI 保留使用 Codex 交互数据训练模型的权利,不过用户可以选择退出。如果 OpenAI 团队使用了基于 Buckmaster 自身 Codex 交互数据训练的模型,那么当该模型面对类似问题时,确实有可能复现出他的工作成果。
OpenAI 未回应就此可能性提出的置评请求。
无论如何,这个问题很可能会重新点燃关于 AI 在数学研究中作用以及 OpenAI 具体动机的持续辩论。就 Buckmaster 本人而言,他似乎认为最好的答案是将尽可能多的研究信息公之于众。
“我没有看过 OpenAI 的证据,”巴克马斯特写道。“我不知道他们的模型做了什么,也不知道是怎么做的。我不知道我们的数据是否被使用。我没有指控任何人任何事。我只是陈述我所被告知的内容、时间,以及向我提出的建议。我之所以说出来,是因为如果不这样做,就等于任由一连串的公告说出一些我知道是虚假的话。”
NYU mathematics professor Tristan Buckmaster announced three proofs on Tuesday with a preliminary finding on one of the major unsolved problems in theoretical mathematics. The findings, made in collaboration with Anthropic mathematician Levent Alpöge and using both Codex and Claude AI models, are significant in themselves — but they’re also accompanied by an unusual controversy surrounding OpenAI’s attempts to solve the same problem.
“There is another part of this story,” Buckmaster wrote in his statement announcing the proofs, “and one that, honestly, I very much wish I did not have to be concerned with.”
While Buckmaster and Alpöge were finalizing their results, they learned that “information about our progress had been passed to OpenAI.” When they contacted OpenAI about this, they were told that OpenAI had already achieved a full proof of the central problem. But when they asked followup questions about when OpenAI had begun its research into the problem and how much human input was involved, the answers became more evasive.
“It emerged that an entire team had been working on the problem,” Buckmaster says,” and that an insane amount of compute had been used…. Eventually, it was agreed that [the first prompt] had been sent in the past few days, after information about our work had reached OpenAI.”
If true, that would suggest the OpenAI team had become convinced that Buckmaster and Alpöge’s approach was the right one, and decided to use its material advantage in computing resources to reach a formal proof first.
Sebastian Bubeck, who leads OpenAI’s mathematical research, says those claims are “false and inflammatory.”
“To clarify, I came into the discussion following academic norms, and I’m disappointed that it has come to this,” he wrote in a post after Buckmaster’s statement. “Anyone who knows me knows that academic standards are of the highest importance to me.” Bubeck said he would make a more complete statement soon.
The dispute centers on the “Navier-Stokes existence and smoothness” problem, one of the seven “Millennium Prize problems” — a set of major unsolved math problems, each carrying a $1 million bounty from Clay Mathematics Institute for the first person or group to provide a solution. The Navier-Stokes equations are widely used in fluid mechanics but poorly understood in theoretical terms. A solution would represent a significant advance in the collective understanding of mathematical physics.
Although the problem is widely pursued among mathematicians, the specific tactic taken by Buckmaster and his collaborator is far less common. As a result, Buckmaster found it suspicious that OpenAI ended up taking the same approach at the same time.
“The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack,” Buckmaster wrote. “Almost nobody else I know of was working on it,” he continued. “It is not the direction one arrives at in a few days by giving a model the problem statement.”
While Alpöge is employed by Anthropic, he was not conducting this research on the company’s behalf. As a result, the duo used a mix of models, relying primarily on OpenAI’s Codex in their work. Even so, Alpöge’s affiliation with a rival lab seems to have been a sore point for OpenAI, and Buckmaster alleges that Bubeck asked him to remove Alpöge’s credit as part of a proposed compromise.
When Buckmaster pushed to make the dispute public, he says that Bubeck replied: “Why would you ruin your career?” Buckmaster says that when he pushed back, Bubeck followed up with: “If you don’t want me to be nice, then I don’t have to be nice.”
Buckmaster also raised concerns that, because he used Codex extensively in assembling the project, information from his work could have informed OpenAI’s own efforts to solve the problem. OpenAI reserves the right to train models on Codex interactions, although users are able to opt-out. If the OpenAI team used a model trained on Buckmaster’s own Codex interactions, it’s plausible that it could have regurgitated his work when faced with a similar problem.
OpenAI did not respond to a request for comment on this possibility.
Regardless, the issue is likely to reignite the ongoing debate about AI’s role in mathematical research, and OpenAI’s specific incentives. For his part, Buckmaster seems to believe the best answer is to get as much information about the research out into the public eye.
“I have not seen OpenAI’s proof,” Buckmaster wrote. “I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me. I am stating it because the alternative is to let a sequence of announcements say something I know to be false.”