纽约大学数学教授 Tristan Buckmaster 周二宣布了三项证明,其中一项初步成果涉及理论数学中一个长期未解的重大难题。这些成果是与 Anthropic 数学家 Levent Alpöge 合作、并借助 Codex 和 Claude AI 模型完成的,本身意义重大——但与此同时,围绕 OpenAI 试图解决同一问题的努力,也引发了一场不同寻常的争议。
“这件事还有另一面,”Buckmaster 在宣布这些证明的声明中写道,“而且说实话,我非常希望自己不必为此操心。”据该声明称,OpenAI 的一项并行工作在其成果公开之前就建立在他们研究的基础上,由此引发了一连串学术竞争与相互矛盾的主张。
在 Buckmaster 发表声明后不久,OpenAI 便发布了关于 Navier–Stokes 方程存在性与光滑性问题的完整证明,而 Buckmaster 的成果正是朝着这一方向迈出的步骤。据 OpenAI 称,该证明是由一个尚未发布的下一代模型发现的,该模型在过去一周内攻克了多个不同的未解难题。总体而言,这项为期一周的工作消耗了 3000 亿个输出 token——若按当前 Astra 费率计费,相当于价值 2250 万美元的计算量。
纳维-斯托克斯方程的存在性与光滑性问题,是七大千禧年大奖难题之一——这是一组尚未解决的重要数学难题,每道题由克莱数学研究所为第一个或第一组提供解答的人悬赏 100 万美元。纳维-斯托克斯方程在流体力学中被广泛使用,但在理论层面却鲜有深入理解。若有人能解决该问题,将代表着人类对数学物理整体认知的一次重大飞跃。
当 Buckmaster 和 Alpöge 正在完善他们自己的研究成果时,他们得知“关于我们研究进展的信息已被传递给了 OpenAI”。当他们联系 OpenAI 时,对方告诉他们,OpenAI 已经取得了该核心问题的完整证明。但当他们追问 OpenAI 是何时开始研究这个问题、其中涉及多少人工输入时,回答就变得含糊其辞了。
“原来有一整个团队一直在研究这个问题,”Buckmaster 说,“而且动用了极其惊人的算力……最终,对方承认,[第一个提示词]是在过去几天内发出的,是在关于我们工作的信息传到 OpenAI 之后。”
如果情况属实,那就意味着 OpenAI 团队已经确信 Buckmaster 和 Alpöge 的研究思路是正确的,并决定利用其在计算资源上的巨大优势,率先取得正式证明。
OpenAI 的帖子证实了这一时间线的大部分内容,特别提到最新一轮研究始于 9 月 1 日,起因是听闻有两道千禧年大奖难题已被解决的传言。此外,该帖子也证实了与 Buckmaster 和 Alpöge 之间正在进行的对话。
尽管这个问题在数学家中被广泛研究,但巴克马斯特及其合作者所采取的具体策略却远非普遍。因此,巴克马斯特对 OpenAI 竟然在同一时间采用了相同的方法感到可疑。
“通过光滑力场通往克莱问题(即费弗曼问题陈述中的选项 c 和 d)的路径,是路易斯和迭戈开辟的路线,也是莱文特和我悄悄选择进攻的路线,”巴克马斯特写道。“据我所知,几乎没有人在这条路线上工作,”他继续说道。“这不是把问题陈述丢给模型几天就能想到的方向。”
虽然阿尔珀格受雇于 Anthropic,但他并非代表公司进行这项研究。因此,两人使用了多种模型的组合,主要依赖 OpenAI 的 Codex 开展工作。即便如此,阿尔珀格与竞争对手实验室的关系似乎成了 OpenAI 的一个痛点,巴克马斯特声称,布贝克曾要求他移除对阿尔珀格的致谢,作为提议的妥协方案的一部分。
当巴克马斯特推动将争议公开时,他说布贝克的回应是:“你为什么要毁掉自己的职业生涯?”巴克马斯特说,当他反驳时,布贝克接着说道:“如果你不想让我客气,那我也不必客气。”
巴克马斯特还提出了担忧:由于他在组装该项目时大量使用了 Codex,他工作中的信息可能已经为 OpenAI 自己解决该问题的努力提供了参考。OpenAI 保留使用 Codex 交互数据训练模型的权利,尽管用户可以选择退出。如果 OpenAI 团队使用了基于巴克马斯特自己的 Codex 交互数据训练的模型,那么当面对类似问题时,该模型很可能复述了他的工作成果。
OpenAI 在自己的帖子中淡化了复述内容可能参与其中的可能性。“我们(研究人员和智能体)在他们公开发布之前,没有通过任何方式看到他们的任何工作——特别是,我们没有访问任何特定用户数据来解决这个问题,”帖子中写道。“虽然可能性不大,但我们不能排除从他们使用我们产品中得出的去标识化数据帮助改进了我们的模型。然而,我们的证明有显著差异,甚至在欧拉案例中,所证明的精确结果也是不同的(受迫与不受迫)。”
无论如何,这个问题很可能会重新点燃关于 AI 在数学研究中的角色以及 OpenAI 具体动机的持续争论。就巴克马斯特而言,他似乎认为最好的答案是尽可能多地将研究信息公之于众。
更新(美国东部时间下午 2:35):纳入了 OpenAI 发布 Navier-Stokes 结果的细节。
NYU mathematics professor Tristan Buckmaster announced three proofs on Tuesday with a preliminary finding on one of the major unsolved problems in theoretical mathematics. The findings, made in collaboration with Anthropic mathematician Levent Alpöge and using both Codex and Claude AI models, are significant in themselves — but they’re also accompanied by an unusual controversy surrounding OpenAI’s attempts to solve the same problem.
“There is another part of this story,” Buckmaster wrote in his statement announcing the proofs, “and one that, honestly, I very much wish I did not have to be concerned with.” According to the statement, a parallel effort by OpenAI built on their work before it became public, leading to a tangle of academic rivalries and conflicting claims.
Shortly after the Buckmaster’s statement, OpenAI published a full proof of the Navier–Stokes existence and smoothness problem, which Buckmaster’s findings had taken steps towards. According to OpenAI, the proof was discovered by an unreleased next-generation model, which has tackled a range of different unsolved problems over the past week. All told, the week-long effort consumed 300 billion output tokens — $22.5 million worth of compute, if charged at current Astra rates.
The Navier-Stokes existence and smoothness problem is one of the seven Millennium Prize problems — a set of major unsolved math problems, each carrying a $1 million bounty from Clay Mathematics Institute for the first person or group to provide a solution. The Navier-Stokes equations are widely used in fluid mechanics but poorly understood in theoretical terms. A solution would represent a significant advance in the collective understanding of mathematical physics.
While Buckmaster and Alpöge were finalizing their own results, they learned that “information about our progress had been passed to OpenAI.” When they contacted OpenAI, they were told that OpenAI had already achieved a full proof of the central problem. But when they asked follow-up questions about when OpenAI had begun its research into the problem and how much human input was involved, the answers became more evasive.
“It emerged that an entire team had been working on the problem,” Buckmaster said, “and that an insane amount of compute had been used…. Eventually, it was agreed that [the first prompt] had been sent in the past few days, after information about our work had reached OpenAI.”
If true, that would suggest the OpenAI team had become convinced that Buckmaster and Alpöge’s approach was the right one, and decided to use its material advantage in computing resources to reach a formal proof first.
OpenAI’s post confirms much of this timeline, specifically saying that the latest effort began on September 1, inspired by rumors that two Millenium Prize problem had been solved. Additionally, the post confirms the ongoing conversations with Buckmaster and Alpöge.
Although the problem is widely pursued among mathematicians, the specific tactic taken by Buckmaster and his collaborator is far less common. As a result, Buckmaster found it suspicious that OpenAI ended up taking the same approach at the same time.
“The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack,” Buckmaster wrote. “Almost nobody else I know of was working on it,” he continued. “It is not the direction one arrives at in a few days by giving a model the problem statement.”
While Alpöge is employed by Anthropic, he was not conducting this research on the company’s behalf. As a result, the duo used a mix of models, relying primarily on OpenAI’s Codex in their work. Even so, Alpöge’s affiliation with a rival lab seems to have been a sore point for OpenAI, and Buckmaster alleges that Bubeck asked him to remove Alpöge’s credit as part of a proposed compromise.
When Buckmaster pushed to make the dispute public, he says that Bubeck replied: “Why would you ruin your career?” Buckmaster says that when he pushed back, Bubeck followed up with: “If you don’t want me to be nice, then I don’t have to be nice.”
Buckmaster also raised concerns that, because he used Codex extensively in assembling the project, information from his work could have informed OpenAI’s own efforts to solve the problem. OpenAI reserves the right to train models on Codex interactions, although users are able to opt-out. If the OpenAI team used a model trained on Buckmaster’s own Codex interactions, it’s plausible that it could have regurgitated his work when faced with a similar problem.
In its own post, OpenAI downplayed the possibility that regurgitation could have been involved. “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem,” the post reads. “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).”
Regardless, the issue is likely to reignite the ongoing debate about AI’s role in mathematical research, and OpenAI’s specific incentives. For his part, Buckmaster seems to believe the best answer is to get as much information about the research out into the public eye.
Update 2:35p.m. ET: Incorporated details from OpenAI’s release of the Navier-Stokes result.