OpenAI 取得了令人瞩目的成果,他们使用一款未发布的模型,为
纳维-斯托克斯存在性与光滑性问题
——七大
千禧年大奖难题
之一——给出了解答。该问题自 2000 年 5 月 24 日起悬赏 100 万美元。
这一发现多少被一些暗箱操作的指控所掩盖。指控来自纽约大学数学教授 Tristan Buckmaster,他一直在与 Levent Alpöge 合作研究相关问题。Levent Alpöge 是一位卓有成就的数学家,目前在 Anthropic 工作。
Tristan 的投诉附带着一份仓促发表的版本,其中包含他们自己的研究成果。这份 PDF 描述了事件经过。最简短的版本是:Tristan 和 Levent 花了将近一年时间研究这个问题,大量使用 Claude 和 Codex(主要是 GPT-5.6 Sol),然后在 8 月 15 日取得了突破。数学界的流言蜚语开始发酵,Tristan 和 Levent 听说 OpenAI 听说 Anthropic 已经解决了“一个重大的开放问题”,于是他们主动联系,得知 OpenAI 有一个团队正在用类似的方法研究一个相关问题。引用 Tristan 的话:
我问他们第一次发送提示词是什么时候。这个问题OpenAI方面有一段时间没有直接回答。最终他们承认是在过去几天内发送的,是在了解到我们工作的相关信息之后。
我问该模型是否基于我们在Codex中的会话进行过训练,或者能否访问这些会话——整个项目期间我们一直把所有草稿放在Codex里。我得到的答复是,该模型不会查阅用户数据。我又追问了训练方面的问题,但没有得到回答。
事情从那时起变得更加复杂。OpenAI团队提出可以等Tristan先发布,或者由他作为作者撰写一篇关于他们成果的论文,但他们明确表示,由于Levent所在雇主与OpenAI存在竞争关系,Levent 不会被列为共同作者。
以下是OpenAI对自己工作的描述:
9月1日(周二),我们听到传闻说有两道千禧年大奖难题已被解决。受这些传闻以及我们内部模型性能阶跃式提升的鼓舞,我们启动了一项工作,在所有未解决的千禧年大奖难题以及其他几个高影响力问题上对该模型进行评估。[...]
智能体于9月5日(周六)得出解答,距首批智能体启动约88小时。Lean形式化与验证又额外花费了17小时,由GPT‑6 Astra完成。
在所有尝试求解的问题中,智能体共发送了 490 万条消息,使用了约 3000 亿个输出 token。在解决纳维-斯托克斯问题的过程中,智能体发送了 270 万条消息,使用了约 1300 亿个输出 token。
(我们不知道他们使用的内部模型的成本结构,但按 GPT-6 Astra 的公开 API 价格计算,3000 亿个输出 token 的成本将达到 1500 万美元。)
以下是他们对 Tristan 和 Levent 工作的看法(强调为我所加):
我们的工作始于 9 月 1 日,此前我们听到一个传言,后来意识到该传言与 Anthropic 员工 Levent Alpöge 以及纽约大学数学教授 Tristan Buckmaster 有关。在我们完成整个项目并通过 Lean 验证(9 月 6 日)之后,由于相信传言称他们也已得到纳维-斯托克斯问题的解,我们主动联系了他们,提议同时发布我们的结果,并在联合公告中承认他们的优先权。[……]
我们(研究人员和智能体)在他们公开发布之前,没有通过任何途径看到他们的任何工作——特别是,我们没有访问任何特定用户数据来解决这个问题。虽然可能性不大,但我们无法排除从他们使用我们产品中获得的去标识化数据有助于改进我们的模型。不过,我们的证明存在显著差异,甚至在欧拉方程情形下,所证明的具体结果也有所不同(有外力 vs 无外力)。
我对这件事的理解是:OpenAI 听说有人用大语言模型解决了千禧年大奖难题,便把这当作展示其最新模型实力的机会,而没有过多考虑抢先于一个团队会带来怎样的观感——这个团队近一年来一直在用 OpenAI 自家的模型攻克这道题。
这种情况似乎正与当下计算机安全领域发生的事如出一辙。Anil Madhavapeddy 最近指出,如今,仅仅是一个漏洞的传闻就足以让人找到安全漏洞,因为一旦有人知道某款软件存在未修补的漏洞,他们就可以给自己的智能体布置任务去把它找出来。数学领域现在是否也面临同样的情况?仅仅知道某个问题存在一个未公开的解法,就可能引发数百万美元的 LLM 算力投入,只为抢先一步。
这也凸显了我在这一切运作方式上长期存在的一个困惑。当 AI 实验室说我的数据被“用于改进模型性能”时,这到底意味着什么?
关于这一点,我以前最喜欢问的两个假设性问题分别是:
- 如果我在运行 Codex,而我的某个 API 密钥不小心被卷入了上下文之中,那么将来别人向模型索要 API 密钥时,把我这个密钥吐出来的概率有多大?(我曾问过 OpenAI 的某个人,他们把这称为“反刍”问题,并向我保证他们费了很大力气来防止这种情况……但不愿描述具体做法。)
- 如果我与 ChatGPT 头脑风暴,探讨我公司潜在的新发展方向,那么这些信息在六个月后被我的一位竞争对手问出“X 公司下一步可能会做什么”时泄露出去的概率有多大?
我对此新偏好的假设是:
- 如果我使用 ChatGPT 帮我部分解决一个千禧年大奖难题,那么我的工作影响训练、使得后来的模型帮助其他人先解决该问题的概率有多大?
Impressive result from OpenAI, who used an unreleased model to produce a resolution to
the Navier–Stokes existence and smoothness problem
, one of the seven
Millennium Prize Problems
that have been subject to a $1,000,000 prize since May 24th, 2000.
The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic.
Tristan's complaint accompanied a hastily published version of their own results. Here's the PDF describing what happened. The very short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan:
I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would not be invited as a co-author due to OpenAI's competitive relationship with his employer.
Here's how OpenAI described their work:
On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...]
The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.
Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.
(We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost $15,000,000.)
Here's where they provide their perspective on Tristan and Levent's work (emphasis mine):
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...]
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year.
This situation appears to mirror what's happening in the world of computer security right now. Anil Madhavapeddy recently pointed out that Just a rumour of a bug is enough to find a security exploit these days, because if someone knows that some software has an unpatched vulnerability, they can set their agents the task of finding it. Is the same now true of mathematics? Just knowing that there is an unpublished solution to a problem might trigger millions of dollars in LLM spending to get there first.
This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", what does that actually mean?
My two favourite hypothetical questions regarding this used to be:
- If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
- If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?
My new preferred hypothetical for this is:
- If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?