在这篇客座文章中,物理学家兼科学作家 Matt von Hippel 分享了他向 AI 公司发起一项挑战后发生的事情,该挑战涉及他此前所在的理论物理子领域中的一个问题。

你发起一项挑战,结果一个月后就被人攻克,这种事并不常见。但我们正身处一个不寻常的时代。
让我自我介绍一下:我是 Matt von Hippel。我曾是一名理论物理学家;如今我是一名科学作家。一直以来,我还是一名博主,每周在 4gravitons.com 上撰写关于物理学以及从事物理学研究的人的文章。
越来越多的时候,写物理学博客就意味着写 AI。这是个问题,因为我绝对算不上 AI 专家。我确实涉猎过一些。我大概比你家奶奶懂得多。但大多数时候,我不得不退后一步,去信任专家。而令人沮丧的是,专家们意见不一!我听过一些聪明、见多识广的人,他们确信 AI 距离超级智能只有几年之遥,而超级智能将能够做出真正令人恐惧的事情。我也听过一些同样聪明、同样见多识广的人,他们同样确信,基于 LLM 的 AI 已接近天花板,像 Claude 这样的模型甚至无法在物理学中做出令人印象深刻的工作,更不用说征服世界了。
我一直不愿做出自己的预测。在形成观点之前,我想先看到 LLM 在一个我熟悉的领域取得进展——一个我知道很难做到的事情,因为我自己曾尝试过类似的事情。
除此之外,我还想看看大语言模型能否完成一件我认为在计算上很难的事情。诚然,大语言模型在数学上已经取得了令人瞩目的进展,而且仅这个月就可能已经改变了许多人的看法。但数学上的进步来自新的想法,而想法是神秘的东西:在找到之前,人永远无法确切知道它们有多难找。计算感觉更实在。我想看到大语言模型去应对一个看似遥不可及的挑战——不是因为在原则上研究人员不知道怎么做,而是因为做这件事似乎需要比研究人员合理可及范围更多的计算机和时间。我想看看那些研究人员是否错了:一个更聪明的人工研究者,能否用同样的计算机,仍然把问题解决掉。
于是,我提出了一个挑战:
“如果 AI 公司想打动像我这样的人(或者就此而言,吓到我们),那它们就需要攻克我过去所在的领域。证明 AI 能够利用一名学者可获得的那些计算机资源,解决散射振幅领域某个重大的未解问题。证明一个所有人都以为会成为问题的计算极限,实际上并不重要。给我们算出 N=8 超引力到七圈,或者 N=4 超对称杨-米尔斯到九圈。”
简而言之:AI 能否解决我从前所在的理论粒子物理子领域中的一个前沿问题?而且它能否在有限的预算下做到?
这个挑战
我原来的领域是理论粒子物理学的一个分支,叫做振幅学。当其他粒子物理学家预测新粒子时,他们会确保自己能够完成计算来检验这些预测。他们计算被称为散射振幅的公式,这些公式让物理学家能够利用亚原子粒子的动量和能量,来计算它们以特定方式发生反应的概率。
如果物理学家能够对这些反应做出更精确的预测,他们就可以检验大型强子对撞机等实验的结果是否与这些预测相符。不一致之处可能成为新理论的证据,而新理论或许能解释物理学中一些长期悬而未决的重大谜团,比如暗物质的本质,或者宇宙中物质与反物质之间的平衡。
这些散射振幅公式很难计算,难到物理学家几乎总是使用近似方法。他们做部分计算,在特定数量的“圈”处截断,圈数是衡量粒子间相互作用被允许复杂到何种程度的指标。计算中包含的“圈”越多,就越接近真实答案,而计算在算力上也就越难完成。
在实践中,大多数散射振幅公式只计算到了两圈。少数达到了三圈。你在粒子物理学中可能听说过的最精确的预测用到了五.
振幅学家希望做得更好。他们开发实验性的新技术,并在特殊的“玩具模型”理论上进行测试。通过在计算更容易的玩具模型上尝试这些技术,而非真实世界中更具挑战性的粒子,振幅学家可以对新技术进行压力测试,看看它们能走多远。
我为其中两个玩具模型发布了挑战。Anthropic 的人选择攻克的那个,是要在一个名为 N=4 super Yang-Mills 的特定玩具模型理论中做到九圈。
“Yang-Mills”是一类理论的技术名称,这类理论解释了我们周围世界的大部分现象。自然界四种基本力中的三种:电磁力、把原子核束缚在一起的强核力,以及导致香蕉之类东西发生放射性衰变的弱核力,都是 Yang-Mills 理论。
“N=4 super”中的“super”来自超对称。物理学家曾推测,每个粒子都有一个“超对称伙伴”,一种电荷相同但类型不同的粒子,把电子这类物质粒子与光子这类力粒子对应起来。他们一度乐观地认为,通过这些更熟悉粒子的尚未发现的伙伴粒子,可以解释暗物质。那些推测用的是“N=1”超对称。而在“N=4”中,每个粒子有 四个超对称伙伴,而不只是一个。
粒子数量如此之多,使这个理论非常不现实。N=4 super Yang-Mills 并不被用来解释暗物质,也不被用来解释现实世界中的任何东西。相反,振幅研究者用它来磨炼自己的技术,因为 N=4 悖论般地更容易计算。不同粒子之间的微妙平衡意味着只需要某些特定变量组合,从而简化了计算。
这些计算是用一种名为“自举”(bootstrap)的实验性技术完成的,而这项技术最终出人意料地非常适合借助 AI 来使用。要对一个振幅进行自举,你不需要考虑每一种可能的粒子相互作用。你只需要大致知道答案应该长什么样,并用一种专门的字母表在计算机文件中记录下每一种可能性。
然后你开始逐一核对你所知道的一切:来自其他计算技术的预测、答案必须遵守的规则、与相关问题的联系——那些问题中答案更容易找到。这有点像数独,你一开始拿到一个填满所有可能数字的网格,然后随着推理逐步把它们划掉。最终,你希望发现只有一种可能性满足所有校验,同时还剩有足够多的校验来确保你没有犯错。
这意味着 Lance 已经做好了充分准备,可以去检验是否有人把下一个振幅公式——九圈的结果——交到了他手上。那将会是一个有趣的答案,不仅因为它能验证自举技术,还因为它是复杂度达到那么多圈的振幅中一个罕见的例子,一个本身就值得研究的答案。
但他没有算出它,这个领域里的其他人也没有。他找到八圈答案的方式本就已经有些间接,是通过一个出人意料的联系通向另一个不同但相关的公式,称为形状因子(form-factor),这是一种涉及不同粒子的部分振幅,结果发现它计算起来稍微容易一些。他原本预期找到下一圈会更加间接,可能是通过另一种不同的 AI 方法。如果人们认为只要把通常的自举方法再跑一圈就有可能做到,那早就有人去做了。
然后人们就做到了
显然,Anthropic 有人读了我的博客。
八月底,Anthropic 的两位物理学家 Liam Fitzpatrick 和 Siddharth Mishra-Sharma 联系了我,说他们解决了我那篇博文中提出的挑战之一。在跟 Lance 核实了结果之后,他们向我讲解了他们的解法。
本着这项挑战的精神,他们并没有动用价值数百万美元的计算资源。他们用的是 Fable 5.1,运行在 Claude Science 之中——这是一个科学家可以付费使用的平台。Claude Science 就是业内人士所说的“harness”(驾驭框架),一种利用 Claude LLM 配合结构化规则和提示词的程序,以获得更稳健、更具科学实用性的行为表现。
显然,在询问 Claude 它最有可能解决哪个问题之后,他们给了它一个简单的提示词:
“问题是计算平面 N=4 SYM 中九圈水平的六粒子(六边形)振幅。”
从那之后,他们就只是不断告诉它继续做下去,评论诸如:
“我要去睡觉了,接下来几个小时都不在。继续做这件事,直到我让你停。每 4-6 小时给我一次进展汇报。”
Claude 最终用两种不同的方式完成了计算:原始的 bootstrap 方法,以及间接的形状因子方法。无论哪种方式,终端用户都需要花费大约一到两千美元,主要成本来自长时间运行 Claude 的开销。其中 bootstrap 计算使用 Python 编程语言和 SymPy 包完成,花费了大约 100 美元的预算,相当于让 96 个 CPU 运行一周。
让 96 个 CPU 运行一周,在我十年前做这类工作时可能会觉得开销很大,但如今如果你有充分的理由,这已经相当实惠了。
事实证明,人类距离这个结果也并不太远。在我收到 Anthropic 消息几天后,我们收到了宋禾的来信,他是北京中国科学院的一位振幅学家。宋禾的团队已经得到了大部分结果。他们借助了一些基于 GPT-6 的 AI 辅助,但不是 Anthropic 所用的那种一次性、几乎无需人类参与的方式。
这里每个人都很友好,这让人松了一口气。人类——Lance、宋禾及他们的合作者——将发表这些结果,花时间解释和分析它们,以造福未来的研究者。Claude 的角色目前已经完成了。
那么,问题解决了吗?
我设置这个挑战,是因为我想更好地了解当前 AI 能做什么,以及它从这里能走向何方。那么我学到了什么?
我本以为这可能是一个机会,让我们看到 AI 以一种出人意料的方式突破某个计算壁垒。结果它做的事情,人类原来也能做到。Claude 使用了已知方法,只是比人们此前尝试过的多投入了一些算力。它可能因为使用了 Python 而获得了助力,而不是 Maple(Lance 最喜欢的数学程序)或 Mathematica(我最喜欢的),而且它可能使用了比我们好得多的软件工程实践,但也没有聪明到超乎常理的地步。
我最大的收获是,那里的低垂果实比你预想的要多。即便一个目标简单且定义明确,有时在专家看来,它也会显得比实际难实现得多。多年来一直有计算机科学背景的人告诉我,振幅研究者只要雇几个程序员就能取得大得多的进展。他们应该感到自己得到了平反。
同样值得注意的是,Claude Science 一次就完成了这件事,而且没有任何比“继续做下去”更精细的科学监督。这些计算既挑剔又混乱。如果我花一周时间、用 96 个 CPU 来做这类计算,那我几乎肯定会最终用掉两周:几乎可以保证我第一次尝试就会搞砸某个地方。
我不知道 Claude 在内部过程中犯了多少错误,但整个执行框架在没有任何外部合作者输入的情况下把它带到了终点。到了这个地步,我不确定这还让我感到意外。但如果你之所以不知道它能做到这一点,是因为你仍然认为 AI 太容易出错、根本没法用,那么这应该是你的收获:它现在能可靠地做这类事情了。
情况显然进展得很快。在三月,AI 完成物理项目的方式还像个学生:任务规模较小,需要大量手把手指导,而且错误频出。相比之下,这是一次真正的前沿计算,通常是由振幅领域的顶尖专家来处理的那类工作。虽然也有可能这只是一个对 AI 友好得多的问题,但我认为不仅仅是如此:我认为这项技术确实变强了。
我能把这一点推广到什么程度?这我不确定。
这些玩具模型理论往往是小型子社区关注的重点。而现实世界中的振幅计算是一个更广阔的领域,有许多团队在相互竞争、争夺前沿。那里的低垂果实可能更少。但我不会对此抱有指望。我认识一些从事这些计算的人,他们越来越多地在编码中使用 AI。如果人们还没有在检验 AI 科研工具链能否一次性完成那里的前沿计算,那他们应该去检验(而且他们应该有一套检验结果的方案)。如果在合理预算内还能再挤出一圈,我也不会太惊讶。
然后这就成了整个社区需要讨论的问题:新的前沿在哪里,接下来需要弄清楚什么?与数学中的许多问题不同,振幅不仅仅是新方法的训练场。它有一个目标,就是让预测足够精确,以便与即将到来的实验进行比较。这个领域距离那个目标还有多近?
不过,比这更宽泛地说,我并没有真正得到一个答案。
我带着好奇心投入这件事,不仅想知道 AI 今天在研究中能做什么,也想知道它的未来。当你读到 LLM 出现之前那些关于超级智能的预测时,它们往往会提出一些看似荒诞的风险。人们想象 AI 能够模拟人来预测其反应并加以操纵,或者能够从第一性原理出发,弄清楚如何制造出灭绝物种的病毒或吞噬世界的纳米技术。
而对这些风险的常见反驳是,它们把智能——那个模糊而神秘的新想法之源——与算力混为一谈。批评者认为,即便新建一批数据中心,也不具备完成上述任何任务所需的算力,那些不过是科幻未来的噩梦,短期内不会到来。
对于这些批评者,我并不觉得自己有了更好的答案。我多少了解到 AI 现在能做什么:它能以合理的预算,在我以前的领域完成有意义的工作,而且几乎可以完全自主地完成。但我原本希望看到更奇特的东西——用于计算本身的新方法,带着意想不到的能力。我原本希望窥见未来的一角,让我在关于超级智能的争论中能有依据地发表意见。我想知道 AI 能把算力极限推进到多远……而我觉得我在这里学到的只是:对于极限在哪里,我过去太天真了。
附记:被机器抢先一步是什么感觉?
作者:Lance Dixon,SLAC 国家加速器实验室与斯坦福大学粒子物理与天体物理学教授,他核查了 Claude 的九圈结果。
我认识的大多数理论物理学家都意识到,当前的大语言模型时代将彻底改变我们思考物理学的方式。问题只在于:它什么时候才会真正让人切身感受到冲击?对我来说,这一天发生在 9 月 1 日,当时 Anthropic 的 Liam Fitzpatrick 和 Siddharth Mishra-Sharma 告诉我,Claude 计算出了平面 N=4 超杨-米尔斯理论中的九圈 MHV 六粒子振幅,并请我验证其结果。
我不打算解释上一句话中所有的技术术语;Matt 已在上文介绍了背景。我确实需要提到,实际上有两个相关的对象,即“振幅”和我们称之为“形状因子”的东西。每一个都对应一个圈数:一圈、两圈、三圈,以此类推。在计算上,每一个圈阶都比前一个更难,即使已经找到了许多简化计算的技巧也是如此。此外,在相同圈阶下,形状因子比振幅更容易处理。2023 年,Andy Liu 和我展示了如何利用形状因子以及一种我们称之为对跖对偶的奇特对称性,来得到八圈的振幅。
自 2023 年以来,我和合作者们一直着眼于利用我们 2023 年的思路推进到九圈,先是形状因子,然后是振幅。我原以为直接计算振幅会太难。所以 Claude 能直接做到这一点,确实让我印象非常深刻。倒不是因为这有多大的计算量,而是因为整个构造非常脆弱:如果你在计算方案中犯了任何错误,一切就会像失败的舒芙蕾一样轰然坍塌,而你只能困惑于原因何在(并去调试)。
此外,这个构造有太多细节枯燥到无法在论文中完整记录。所以 Claude 不得不从零开始开发所有这些代码。
从九圈振幅回到形状因子相对容易,而对我来说,主要通过这种方式来验证结果也更方便。这意味着在过去两周里,我一直在验证一个结果——九圈形状因子,这是我们团队花了几年时间努力攻关的目标。而一台机器解决了一个我认为直接求解过于困难的问题。这让我个人感到困扰吗?这让人心力交瘁吗?
不,有两个原因。其一,我们团队已经在开展一项计划,使用定制 Transformer 模型来预测更高圈数,而我们的口号之一就是:“我们拥有所有工具来验证机器提供的任何候选解。”Claude 是一种不同类型的 Transformer 模型,可能比我们的定制模型大上百万倍。
但当然,我们说过我们能验证 AI 模型给出的任何结果,所以我们能够也应该这样做。第二个原因是,如果你看看 Claude 是如何解决这个问题的,它使用了我和合作者多年来开发的所有方法,并且(也许是给我们面子)用我们已经建立好的相同格式呈现了解法。
所以当我在验证 Claude 的结果时,Claude 也在验证我们之前的所有工作。事实上,我敢断言,Claude 对我们 2019 年和 2023 年论文的理解比任何人类都更透彻,除了我的合著者之外。
在我写完这些之后,Song He 告诉我,他的团队也计算出了九圈振幅中被称为 symbol 的那部分。(不知道为什么,人们似乎就是喜欢告诉我他们在九圈上的成功。)Song 的团队使用了 AI(GPT-6)来帮助他们计算一些约束条件,但没有用于整体框架。所以现在,我在两周之内被一台机器抢了先,也被人类加一台机器抢了先。
回到 Claude 的计算:在我看来,一个大语言模型能够执行我们所列出的复杂流程中的所有步骤,并组织起这样的计算算力,这本身就是一项相当了不起的成就。但更触及灵魂的时刻将会到来——当大语言模型开始先于人类提出新的物理原理和洞见之时。
补充材料
披露声明
Anthropic 邀请 Matt von Hippel 撰写本文,并为其投入的时间支付了报酬。Anthropic 员工对草稿提供了反馈;内容与观点均为其本人所有。Lance Dixon 独立验证了该结果,并获得了 Claude 使用额度。
In this guest post, physicist and science writer Matt von Hippel shares what happened when he issued a challenge to AI companies regarding a problem in his former subfield of theoretical physics.

It’s not often that you issue a challenge, only to see it beaten a month later. But we’re living in unusual times.
Let me introduce myself: I’m Matt von Hippel. I used to be a theoretical physicist; these days I’m a science writer. Throughout, I’ve been a blogger, writing weekly at 4gravitons.com about physics and the people who do it.
More and more, blogging about physics has meant blogging about AI. That’s a problem, because I’m definitely not an AI expert. I’ve dabbled in it, sure. I probably know more than your grandma. But I mostly have to step back and trust the experts. And frustratingly, the experts disagree! I’ve heard from smart, well-informed people who are confident that AI is a few years away from superintelligence, and that superintelligence will be capable of truly terrifying things. And I’ve heard from smart, well-informed people who are equally confident that LLM-based AI is close to a ceiling, that models like Claude won’t even be able to do impressive work in physics, let alone conquer the world.
I’ve been reluctant to make my own predictions. Before forming an opinion, I wanted to see an LLM make progress on something familiar, something I knew was hard to do because I’d tried to do something similar myself.
In addition to that, I wanted to see an LLM do something that I expected to be computationally hard. LLMs have made impressive strides in math, certainly, and this month alone has likely changed many peoples’ minds. But progress in math comes from new ideas, and ideas are mysterious things: one never quite knows how hard they are to find until they’re found. Computation felt more solid. I wanted to see an LLM tackle a challenge that seemed out of reach not because researchers didn’t know how to do it in principle, but because doing it seemed like the kind of thing that would take more computers and time than the researchers reasonably had access to. I wanted to see if those researchers were wrong: if a smarter, artificial researcher could use the same computers, and solve the problem anyway.
So, I issued a challenge:
“If AI companies want to impress people like me (or scare us, for that matter), then they need to tackle my old field. Show that an AI can take the kinds of computer resources an academic has access to, and solve one of the scattering amplitudes field’s big outstanding problems. Show that a computational limit everyone expected to be a problem doesn’t actually matter. Give us N=8 supergravity to seven loops, or N=4 super Yang-Mills to nine loops.”
In short: can AI solve a frontier problem in my former subfield of theoretical particle physics? And can it do it on a budget?
The challenge
My old field is a branch of theoretical particle physics called amplitudeology. When other particle physicists predict new particles, they make sure they can do the calculations to test those predictions. They compute formulas called scattering amplitudes, which let physicists use the momenta and energies of subatomic particles to calculate how likely they are to react in particular ways. If physicists can make more accurate predictions for these reactions, they can check whether results from experiments like the Large Hadron Collider match those predictions. A mismatch could be evidence for a new theory, one that could explain some of physics’ big lingering mysteries, like the nature of dark matter, or the balance between matter and antimatter in the universe.
These scattering amplitude formulas are hard to compute, so hard that physicists almost always use approximations. They do partial calculations, cut off at a specific number of “loops,” a measure of how complicated interactions between particles are allowed to get. The more “loops” they include in their calculations, the closer they get to the real answer, and the harder, computationally, the calculation is to do.
In practice, most scattering amplitude formulas have only been calculated to two loops. A few have three. The most precise prediction in particle physics you might have heard of used five.
Amplitudeologists want to do better. They develop experimental new techniques, and test them on special “toy model” theories. By trying the technique with a toy model where the calculation is easier, rather than the more challenging particles of the real world, amplitudeologists can stress-test the new methods and see how far they can go.
I posted challenges for two of those toy models. The one the folks at Anthropic chose to tackle was to go up to nine loops with a particular toy model theory, called N=4 super Yang-Mills.
“Yang-Mills” is a technical name for a type of theory that explains most of the world around us. Three of the four fundamental forces of nature: electromagnetism, the strong nuclear force that holds the nuclei of atoms together, and the weak nuclear force that causes radioactive decay in things like bananas, are all Yang-Mills theories.
The “N=4 super” comes from supersymmetry. Physicists have speculated that each particle has a “supersymmetric partner,” a particle with the same charge, but of a different type, matching matter particles like electrons to force particles like photons. At one time they were optimistic these particles could explain dark matter, via undiscovered partners of more familiar particles. Those speculations used “N=1” supersymmetry. In “N=4,” each particle has four supersymmetric partners, not just one.
That surfeit of particles makes the theory very unrealistic. N=4 super Yang-Mills isn’t used as an explanation for dark matter, or for anything in the real world. Instead, amplitudeologists use it to hone their techniques, because N=4 is paradoxically easier to calculate with. The delicate balance between the different particles means only certain combinations of variables are needed, streamlining calculations.
These calculations were done with an experimental technique called a bootstrap, which ended up bizarrely well-suited for use of AI. To bootstrap an amplitude, you don’t have to take into account every possible particle interaction. You just need to know roughly what the answer ought to look like, keeping track of every possibility in computer files in a specialized alphabet. Then you start checking everything you know: predictions from other calculation techniques, rules the answer has to obey, links to related problems where the answer was easier to find. It’s a bit like Sudoku, where you begin with a grid with all possible numbers, then cross them out as you go. In the end, you’re hoping to find that only one possibility satisfies all the checks, while having enough checks left over to make sure you didn’t make a mistake.
That meant that Lance was already well set up to check if someone had handed him the next amplitude formula, with nine loops. It would be an interesting answer, not just as a validation of the bootstrap technique, but as a rare example of an amplitude with that many loops of complexity, an answer that could be worth studying in its own right.
But he hadn’t computed it, and neither had anyone else in the field. The way he found the eight-loop answer was already a bit indirect, via a surprising link to a different but related formula called a form-factor, a kind of partial amplitude involving different particles that turns out to be a bit easier to calculate. He was expecting to find the next loop even more indirectly, potentially by a different kind of AI method. If people thought it was possible to just run the usual bootstrap method for one more loop, someone would have done it.
Then people did it
Apparently, there are folks at Anthropic who read my blog.
At the end of August, Liam Fitzpatrick and Siddharth Mishra-Sharma, two physicists at Anthropic, reached out to me to say they had tackled one of the challenges in my post. After verifying the result with Lance, they talked me through how they got it.
True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use. Claude Science is what folks in the biz call a “harness,” a program that uses the Claude LLM with structured rules and prompts in order to get more robust and scientifically useful behavior.
Apparently, after asking Claude which problem it was most likely to be able to tackle, they gave it a simple prompt:
“The problem is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.”
From there, they just kept telling it to keep going, with comments like:
“I'm going to sleep and won't be available for another several hours. Keep working on this until I tell you to stop. Give me updates every 4-6 hours.”
Claude ended up doing the calculation two different ways: the original bootstrap, and the indirect form-factor approach. Either approach would have cost an end-user around one or two thousand dollars, mostly due to the expense of running Claude for so long. The bootstrap calculation, done with the Python programming language with package SymPy, took around $100 of the budget, corresponding to running 96 CPUs for a week.
Running 96 CPUs for a week might have felt like a lot when I was doing this kind of work ten years ago, but it’s pretty affordable now if you have a good reason.
As it turned out, the result wasn’t all that far away for humans either. A few days after I heard from Anthropic, we heard from Song He, an amplitudeologist at the Chinese Academy of Sciences in Beijing. Song’s group had already gotten the majority of the result. They’d used some AI assistance, based on GPT-6, but not the kind of one-shot almost human-less approach Anthropic used.
Everyone has been friendly here, which is a bit of a relief. The humans, Lance and Song and their collaborators, will get to publish the results, taking time to explain them and analyze them for the benefit of future researchers. Claude’s role is done, for now.
So, problem solved?
I set my challenge because I wanted a better sense of what current AI can do, and where it could go from here. So what have I learned?
I’d thought this could be a chance to see AI overcome a computational barrier in a surprising way. Instead, it did something it turned out humans were also able to do. Claude used known methods, with a bit more compute than people had tried to use before. It may have gotten a boost from using Python, and not Maple (Lance’s favorite program for math) or Mathematica (mine), and it may have used much better software engineering practices than we would have, but not super-intelligently so.
My biggest takeaway is that there is more low-hanging fruit out there than you’d expect. Even when a goal is simple and well-defined, sometimes it’s going to look much less achievable to experts than it actually is. There are people with a computer science background who’ve been telling me for years that amplitudeologists could make a lot more progress just by hiring a few programmers. They should feel vindicated.
It’s also noteworthy that Claude Science accomplished this in one shot, without any scientific oversight more sophisticated than “keep going.” These are finicky, messy calculations. If I’d used a week of time on 96 CPUs to do this kind of calculation, then I’d almost certainly end up using two weeks: it’s practically guaranteed I’d screw up something on the first try. I don’t know how many mistakes Claude made internally on the way, but the harness got it to the end without an outside collaborator’s input. I’m not sure that surprises me, at this point. But if you didn’t know it could do that because you’re still thinking of AI as so error-prone that it’s unusable, then this should be your takeaway: It can do this kind of thing reliably now.
Things definitely seem to be moving fast. In March, AI was accomplishing physics projects like a student: smaller-scale tasks with a lot of hand-holding and mistakes. In contrast, this is a real frontier calculation, the kind of thing normally tackled by the top experts in amplitudes. While it’s possible that this is just a much more AI-friendly problem, I don’t think it’s just that: I think the technology has genuinely gotten better.
How far can I generalize this? That I’m not sure of.
These toy model theories tend to be the focus of small sub-communities. The real-world amplitudes calculations are a wider field, with many groups trying to beat each other to the frontier. It’s possible there’s less low-hanging fruit there. But I wouldn’t count on it. I know people who work on those calculations have been increasingly using AI for coding. If people aren’t already checking whether AI science harnesses can one-shot frontier calculations there, they ought to (and they ought to have a plan for how to check the results). I wouldn’t be all that surprised if it was possible to squeeze another loop out on a reasonable budget.
Then it becomes a question for the community to discuss: where is the new frontier, and what needs to be figured out next? Unlike many problems in mathematics, amplitudes aren’t just a training ground for new methods. There’s a goal, to make predictions precise enough to compare with upcoming experiments. How much closer is the field to that goal?
More broadly than that, though, I didn’t really get an answer.
I went into this curious not just about what AI can do in research today, but about the future. When you read predictions about superintelligence from the days before LLMs, they often propose fantastical-seeming risks. People imagined AI that could simulate people to predict their reactions and manipulate them, or figure out how to build a species-ending virus or world-devouring nanotech from first principles. And the usual objection to these risks is that they conflated intelligence, the vague and mysterious source of new ideas, with computational power. Critics argued that even a fleet of new datacenters wouldn’t have the computational power to do any of those tasks, that they were nightmares of a sci-fi future that wasn’t coming any time soon.
I don’t feel like I have a better answer for those critics. I learned a bit about what AI can do now, that it can do work that matters in my old field on a reasonable budget, and do it pretty much autonomously to boot. But I’d hoped to see something stranger, new methods for the calculation itself with unexpected power. I’d hoped to get a glimpse of the future, something that would give me an informed opinion in debates about superintelligence. I wanted to know how far AI could push computational limits… and I feel like what I learned here is just that I was too naïve about where the limit was.
An addendum: How does it feel to be scooped by a machine?
By Lance Dixon, Professor of Particle Physics and Astrophysics at SLAC National Accelerator Laboratory and Stanford University, who checked Claude's nine-loop result.
Most theoretical physicists I know recognize that the current era of large language models is going to completely transform the way we think about physics. The question was just: when was it going to really hit home? For me, it happened on September 1, when Liam Fitzpatrick and Siddharth Mishra-Sharma at Anthropic told me that Claude had computed the nine-loop MHV six-particle amplitude in planar N=4 super Yang-Mills, and asked me to validate its result.
I'm not going to explain all the technical terms in that last sentence; Matt has covered the background above. I do need to mention that there are really two related objects, the "amplitude" and something we call the “form factor.” Each has an associated number of loops: one, two, three, and so on. Every loop order is harder than the previous one, computationally, even after finding lots of tricks to make things easier. Also, the form factor is easier than the amplitude at the same loop order. In 2023 Andy Liu and I showed how to use the form factor and a weird symmetry we call antipodal duality to get the amplitude at eight loops.
Since 2023, my collaborators and I have eyed getting to nine loops, first for the form factor and then for the amplitude, using our 2023 idea. I thought it would be too hard to do the amplitude directly. So I was really quite impressed that Claude could do it directly. Not so much because it was a big computational task, but because the whole setup is very fragile: if you make any mistake at all in the computational recipe, it all crashes down like a failed soufflé, and you are left to wonder why (and debug). Also, there are so many details of the construction that are too boring to document fully in a publication. So Claude had to develop all that code from scratch.
From the nine-loop amplitude it is relatively easy to go back to the form factor, and it was easier for me to validate the result mostly that way. That meant that for the last two weeks I've been validating a result, the nine-loop form factor, that our team had been working toward for a couple of years. And a machine had solved a problem that I thought was too hard to do directly. Does that bother me personally? Is it soul-crushing?
No, for two reasons. One is that our team already had a campaign to use custom transformer models to predict higher loops, and part of our slogan was: “We have all the tools to validate any candidate solution a machine would provide us.” Claude is a different kind of transformer model, probably over a million times bigger than our custom one. But sure, we said we could validate any result an AI model would give us, so we can and should do it. The second reason is that, if you look at how Claude solved the problem, it used all the methods my collaborators and I developed over the years, and it presented the solution (maybe as a favor to us) in the same format we had already set up. So while I'm validating Claude's result, Claude is validating all of our previous work. In fact, I would assert that Claude understands our 2019 and 2023 papers better than any human, aside from my co-authors.
After I wrote this, Song He told me that his group had also computed the piece of the nine-loop amplitude called the symbol. (People just seem to like to tell me about their nine-loop successes, for whatever reason.) Song's group used AI (GPT-6) to help them compute some of the constraints, but not for the overall framework. So now I've been scooped by both a machine and by humans plus a machine, within two weeks.
Going back to the Claude computation: it's quite a triumph, in my opinion, for a large language model to execute all of the steps in the complicated recipe we laid out, and to organize the computational horsepower. But the more soul-searching moments will come when large language models start to come up with new physical principles and insights before humans.
Additional material
- The full nine-loop result, in the format used for the earlier loop orders;
- The concurrent nine-loop result by Song He, Jirong Jing, and Xiang Li.
Disclosure
Anthropic invited Matt von Hippel to write this post and compensated him for his time. Anthropic staff gave feedback on drafts; the content and opinions are his own. Lance Dixon validated the result independently and received Claude usage credits.