Nathan 和 Florian 坐下来讨论开放模型领域发生的一切。继上周 Kimi K3 发布之后,感觉一切都在加速——美中地缘政治、开放与闭源模型的经济学、AI 前沿的安全问题,等等。
章节:
00:00 欢迎与背景介绍
04:38 与 Kimi K3 共处 / 使用 Kimi K3
08:53 GLM 5.2 持续扮演的角色
12:47 中国模型为何如此出色?
17:41 数据、环境,以及中国实验室巡礼
19:47 中国厂商综述:Qwen、DeepSeek、MiniMax……
24:08 美国的开放模型生态
30:25 前沿与近前沿,以及反对禁令的网络安全论据
34:58 知识蒸馏与 Ben Thompson 之争
44:12 预测与前沿梯队排名
48:36 总结
可在 Apple Podcasts、Spotify 以及你获取播客的任何平台收听。其他 Interconnects 访谈,请见此处。
00:00:06 Nathan Lambert:好的,欢迎回到 Interconnects。我们正在做季度开放模型综述,基本上就是我们调侃或解释——不是调侃——为什么这么多关于知识蒸馏的观点都很糟糕,并理解当前形势。我想上周四就是 Kimi K3 发布的日子。我认为在不久的将来我们会看到更多更多。这似乎相当不可避免。比如上周末,Xi 发表了他的讲话,直接承诺将开放和开源作为一种战略。那并不是一份详细的形势布局。
Qwen 宣布他们的下一个大模型将采用开放权重,这是一个很大的变化。我觉得要聊的东西实在太多了。我觉得 Flo 你其实已经在讲性能差距和知识蒸馏的代价了。所以我们大概可以从那里开始,然后我一边聊一边有个小清单,我们可以随时过一遍那些话题以及我写的那篇博客,它们都非常细致入微。所以我觉得我们有聊不完的东西。那就继续你刚才那番吐槽吧。
00:01:17 Florian Brand:是的,我觉得,或者说每次模型发布时最大的话题,至少每次开放模型发布时,就是它落后闭源前沿多少、多少个月。嗯,人们喜欢给这件事安上一个明确、嗯、确定性的数字,而这真的真的非常混乱,因为如今我们有这么多不同的评测基准提供方,还有、嗯、这么多不同的评测基准,以至于每个网站——我在这件事上也不是无辜的——嗯,都会拿出自己偏爱的评测基准,来证明当前模型或新发布的模型处于前沿,然后另一方就会拿出另一个评测基准来反驳,说哦它其实落后一年左右之类的。嗯,而且这很大程度上似乎都取决于那个问题:开放模型到底落后多少个月。
00:02:26 Nathan Lambert: 是的。所以我的挑衅性观点是,有些基准其实与人们正在做的事情有相当合理的相关性,这就是智能体编程和智能体计算机使用任务,而有些基准则与长尾相关,我认为这正是 Claude 和 GPT 如此有价值的地方。但如果问题是 Claude Code 和 Codex 现在的市场是什么,如果答案是软件工程,那么模型在这方面落后比如几个月,这可能是一件非常非常大的事,然后我怀疑这个模型会没问题——免责声明,模型权重还没发布,据说是在 7 月 20 日、27 日,很多讨论都会建立在它们会发布的假设之上。
但人们可以对这个模型做后训练,很可能在人们想要的许多这类细分领域中匹配 Opus 和 GPT。我认为,我的意思是,我们对后训练开源模型行业有着不同的视角,但在让这些模型针对特定高价值任务进行微调方面,有极其大量的兴奋和进展,这在历史上是通过 Qwen 和 GLM 的混合来实现的,而 GLM 5.2 真正加速了这一点。
我很好奇第一个发出博客文章的人,比如“我们在我们的任务上微调了 Kimi K3”,因为我敢打赌你能获得巨大的收益。我想,甚至你使用 Kimi K3 也比我多,但我的直觉是,它会是一个有点粗糙边缘的后训练,仅仅因为它的规模扩展有多大,而这通常意味着仍然可以从中提取出大量性能。
不。
00:04:03 Florian Brand: 是的。跑起来、跑起来、跑起来,尤其是后训练,那个会超级难,因为你需要大概一个节点的 B300 才能加载权重,从规模上来说这简直疯狂。所以大概需要一些时间,而且我听说需要大量的工程工作,才能真正把它弄到可以微调的状态。嗯,但是你们想聊的是实际使用这个模型,比如你真的注册了那个编程套餐并且用了它。所以把它推出来是很好的背景信息。
00:04:38 Nathan Lambert: 是的。所以我在发布后第二天左右注册了 200 美元的套餐,嗯,那是他们最大的套餐,和其他家类似,不过他们好像也有 40 美元和 100 美元的。嗯,但最大的套餐有 100 万上下文,我觉得,或者至少感觉它在 API 请求方面也有一些优先级,因为网上有太多人说他们不断遇到 API 错误,而到目前为止,如果我要这么说的话,我一直过得还挺好的。嗯,就模型能力而言,在某些方面,除了前端确实非常好之外,它在某些方面真的很出彩,表现很出色。
嗯,甚至超出了我的预期,即便是像一些研究任务这样的东西,比如我在 Interconnects 这边,我们现在已经有超过一年的开放模型数据了,嗯,我让前沿模型提出一些我们以前没做过的有趣分析,因为我们自己做分析并把它发表出来,嗯,我让它们全都来做点新的东西,基本上就是让我惊喜一下。而很多模型,或者说前沿模型,嗯,或者说基本上所有模型,都会抓住我们做过的东西,重做数据分析那部分,然后搞一些奇怪的冷门部分。
呃,Kimi K3 做了一些更有意思的事情,嗯,我明确让它去抓取 Reddit,嗯,然后它找到了一些我甚至都没考虑过的 subreddit,接着发现,比如说,Reddit 上的讨论要比下载量通常起飞的时间早一两个月,嗯,更近期一些,或者说他们发现,他们能在下载量通常起飞之前一两个月就发现那些有意思的模型,比如他们都盯上了 Qwen,然后人们就会下载更多的 Qwen 模型,这类分析是开创性的,嗯,但这是 Kimi 让我感到惊讶的地方,相比其他所有前沿模型,嗯。
一个简单的问题,比如你能把它用于你所做的大部分核心工作吗,就你拥有的经验而言,你有一堆你倾向于做的事情的分布,其中大多数是用 Codex 做的,我觉得你是 Codex 派而不是 Claude 派,那你觉得这个模型能胜任的那部分,维恩图重叠的比例大概是多少
00:07:24 Florian Brand:呃,这真的取决于我给它多大的自由度,比如,我现在在 Kimi K3 上看到的最大问题是,我正在做,呃,我们在 Prime Intellect 做的那个框架,呃,那是我工作的地方,而我在 Kimi 上发现的主要问题是它的代码要简单得多,呃,这让它可读性强了很多,呃,但它漏掉了一些 Codex 能搞定的事情,或者说,我们说的是 56、55,尤其是 54 会达到那些水平。所以,我会说 Kimi K3 在这类任务上大概是 54-55 的水平。
但如果我喜欢,我会读代码,然后说好吧,这代码真不错,接着我用 Codex 过一遍,它会发现各种小众的边缘情况,在这些情况下它表现不佳,但用于监督运行或者跑一些实验,它其实真的挺可用。嗯,对于一些其他小众的事情,你直接让它跑就行。
唯一的缺点是——但这也是因为 API 在用户量上完全被挤爆了,而且他们的服务器在中国。实际耗时比 GPT 显著显著地高。但我想说,如果我要硬推着把它用进我的日常工作流,我会变慢,但不会慢到让我说“好吧,这没法用”的程度。
00:08:53 Nathan Lambert: 那这和 GLM 5.2 相比如何?因为在我看来 GLM 5.2 还是一个正在展开的故事,就像我在旧金山到处晃悠,人们会说,是啊,我在我的智能体编程和/或工作流的这部分里真的在用这个。嗯,你怎么看——我感觉你是不是也属于用 GLM 的那一拨,还是说
00:09:20 Florian Brand: 是的。比如你会在哪里……我也用过、现在也还在用 GLM,主要是因为我们有一个内部端点,速度非常快,而且我们还有——或者说在那之前,我也用过一个 API,速度大概是每秒 200 或 300 个 token 吧。嗯,如果你能以足够好的水平快速完成大量任务,你就会直接用那个模型,而不是去 Codex 里先选较小的模型,再选合适的推理强度,再选 fast 模式——我就直接用 GLM,得到同样的结果,而且它还挺不错的,能力上绝对算是 Sonnet 级别的,对于很多清理类任务、那些纯粹是苦力活的任务,它真的管用。我想说,你大概可以用 Kimi K3 作为主智能体、GLM 负责子智能体工作,在大量工作上走得非常远。
00:10:24 Nathan Lambert: Kimi 这次发布以及这些模型的规模,有一点相当不同。我认为这些开放模型要真正得到优化并在各家推理服务商那里可用,还需要更长一点时间。比如 GLM 5.2 已经相当快了,但一来我们还没拿到权重,二来我不觉得它的采用铺开速度会像 500B、700B 的 MoE 那样快。我觉得那边会有更多问题,这是一个非常不同的局面——过去中国模型完成 RL 训练后,会在几小时到几天、也许一周内发布开放权重,然后整个生态几乎立刻就懂得该怎么做了。
我认为,在开放权重模型的下一阶段,基础设施层面的提升会大得多,我觉得我们必须把这一点考虑进去,就像闭源实验室在发布模型之前会在幕后做这些事一样。所以这就像是在某种程度上操纵时间差,可能要多等一个月,人们才能真正进行后训练,并大规模地把 Kimi 用于自己的工作流。
我们作为开放权重的拥趸总爱说,哦,只有在闭源模型可用的时候,你才能利用这个时间差,但现在开放模型也有类似的动态了,就像 Kimi 的 API 完全崩了。需求太多,供给不足。所以并不是说这个模型会立刻扩散开来,我只是、我只是在从性能时间差的角度思考这件事
00:11:54 Florian Brand:因为那确实是真的,但另一方面,开放生态在过去几个月里已经相当专业化了。呃,比如在你们最初发布的时候,它们都会带着一些提前拿到权重的合作伙伴一起推出。这些天里,vLLM 的补丁在模型发布前几天甚至前几周就已经出来了,这和一年前完全不同,那时候基本上就是权重一放出来,模型制作者就说,行吧,你们自己想办法搞定。所以我预计第一天的普遍可用性会相当不错,然后各家提供商之间的竞赛就开始了,大家都开始优化,以争取越来越高的速度,因为这太有面子了。
00:12:47 Nathan Lambert: 是的。好,有两个方向可以聊。我们为什么认为中国模型能做到这么好?我想我写过,在我们的 Discord 里和 Epoch 的 JSD 有过一场辩论,我觉得那场辩论非常好,我在我的文章里有一节,我正逐渐倾向于认为中国实验室的资本效率更高,你可以把资本转化为算力、数据和人才,从而让模型变得更好,而这真的——我认为如果这确实是一种结构性优势,那就极其重要,不管原因是什么。我认为原因可能是,他们的人才无论其教育体系如何,都更擅长训练去解决那些能让 LLM 变得更好的问题。
也可能只是在中国,所有算力、人才以及其他一切的成本都低得多,不管是补贴,还是平均薪资更低。但随着我们在模型迭代上不断推进,这是一件非常重大的事。如果下一代模型对 Anthropic 来说要花 $10 billion,但对 Kimi 来说只要 $4 billion,那这可能非常巨大,但目前还不清楚为什么会这样。
比如,我记得 Big Eagle,那位 Kimi 工程师,回复了我关于这件事的推文,他说这有帮助是因为我们并不是在试图推进前沿,只是在努力追赶,这真的可能是一种心态问题,即中国实验室的目标是如何被界定的,从而让他们构建这些模型的成本低得多。
但过去一年里,我们问了很多类似“中国模型会不会掉队”的问题。我原本以为,由于训练这种资本密集型的特性,闭源和开源模型之间的差距会拉大,但看起来情况正朝着相反的方向发展,这很难解读,不过你是否同意,这些实验室——也就是中国的实验室——保持跟进的程度比我们预期的要更高一些?为什么?
00:14:43 Florian Brand:嗯,我其实看了我们基于去年回顾对今年做出的预测,我们基本上说的是差距会保持在几个月之内。嗯,所以这个预测似乎大体上成立。嗯,对我们来说幸运的是,我们没有给出具体数字,不管是 3 个月、6 个月还是 9 个月。所以在这方面我们是安全的。嗯,但我觉得,我们在中国和这些人交谈时,双方共同的总体感受是,他们——研究人员本身就是两三百人的团队,全都是二十多岁的中段,全都只想让一个模型变得真正出色,他们似乎不做任何支线任务。
他们似乎不做任何偏离这些事情的事。嗯,至于算力,这对我们来说是一个很难回答的问题,尤其是现在这些中国的芯片正在上线。我还认为,芯片——我认为芯片走私在过去大概 6 到 9 个月里大幅增加,或者说那些被走私的芯片开始上线了。
00:15:58 Nathan Lambert:走私是规避出口限制的一个笼统说法。如果这些芯片在马来西亚而且他们在用,我我认为那也算类似情况,我认为这在过去六到九个月里大幅增加了,这在一定程度上就是那个的结果,而且你说但是我只是想把这一点提出来,就是我的确认为他们现在拥有的算力比他们训练上一代模型时要多得多。
00:16:24 Florian Brand:是的。就像我们就像或者只是为了提供背景,两周前我想 LongCat 发布了他们的模型,他们呃声称而且我我们知道这很可能是真的呃,是完全在中国芯片上训练的。呃他们没有公开说明是哪些芯片,但人们推测是呃是华为的一些呃昇腾,嗯随着国内产量提升,你可以他们很可能大部分用于或者它们被用于训练,但它们在推理方面尤其有用,而推理也是训练中很大的一部分。
所以他们很可能在训练部分混合使用了一些呃 Nvidia 和其他芯片,然后在推理部分使用越来越大的比例,在这个阶段这确实非常重要。所以我认为他们的总算力在增加,而且他们实际上并没有很多用户。所以他们不需要为 10 亿用户提供算力,就像 ChatGPT 必须做的那样,也不需要为数百或数千家企业提供算力,就像 Anthropic 必须做的那样,因为他们没有那种量级的呃付费客户。
00:17:41 Nathan Lambert: 是的。而且我认为即便是那些付费客户,至少在企业侧,当你在支持这些东西的时候,公司里就是会有时间和闲聊。即便你不是做研究的,即便这不在你的工作范围内,它确实会改变公司的注意力。如果 SSI 推出一个优秀的模型,那将最终验证分心确实是个问题,但这是题外话,我们可以等等再说。我觉得也有关于数据和环境产业开始在那里出现的传闻。
你还记得任何具体的例子吗?因为我们在中国的时候,他们似乎对外部数据的利用少得令人震惊。所以就在几个月后——我们四月去的,然后仅仅几个月后的七月,我们就听到了一些关于中国新公司的事情,以及他们想要购买数据之类的。这是一个有趣的时间线,展示了情况是如何变化的。
00:18:40 Florian Brand: 而且我会对他们实际告诉我们的内容加上误差范围。
00:18:44 Nathan Lambert: 那是因为时间上太接近了,我也不知道。
00:18:49 Florian Brand: 是的。那那那可能那可能是真的。呃,但这类事情很难精确判断。不过我我想说,看起来购买外部数据正成为一个更重要的因素。呃,如果开放模型能以折扣价买到同样的数据——因为它们买数据环境的时间更晚——这将帮助它们赶上闭源模型。但这确实是一个因素。至于这个因素有多大,我们不知道。我们没有任何公开的洞察,而且我怀疑我们也不会从任何人那里获得这些洞察,呃,真的,呃,所以这肯定是为什么我们能够追赶或提升他们模型分数的原因之一。
00:19:47 Nathan Lambert: 好,来盘点一下其他中国模型厂商。我们已经聊了 Kimi,聊了智谱 / GLM。我觉得很快会有更多非常好的 GLM 模型出现。他们可能会叫它 GLM 5.5 之类的。呃,通义千问(Qwen),我们聊过他们即将推出的最大模型。我要说的是,通义千问最大的模型,相对于他们小模型的出色表现,在性能上往往没有获得同样的绝对排名,这可能是专注方向的代价。我觉得这和云厂商有关。这几乎就像是,如果你眯着眼看,几乎就像 Google。
这就好比 Qwen 背后有阿里巴巴,他们在这里有太多机会,而让开发者通过这些小模型与阿里巴巴 Qwen 建立关联,对他们的云来说是一个巨大的机会,我认为他们正在大获成功。但他们的大模型一直不如小模型那么出色。所以我不指望他们的模型会像 Kimi K3 或 GLM 5.2 那样具有突破性。
我预计它会在新闻中被当作重大开源模型来报道,就像中国开源领域的重磅名字丢出一颗巨型炸弹一样,但我不认为它会像一条新闻故事那样持续发酵。嗯,DeepSeek,你可以接着说,随便谁补充都行。
00:21:01 Florian Brand:有意思的是——不知道你对此关注多少——他们有一个端点,你可以用来访问预览版本,而且他们每天都在更新这个端点,所以他们的迭代周期非常快,因为我们能在所有这些嗯 Twitter 嗯基准测试上看到进展。所以很多这些 SVG 相关的东西和 three.js 之类的所有视觉生成任务,模型在过去几天里进步了很多。所以他们摸索出了某种快速反馈机制,嗯,其他公司也有类似的东西。呃,我们知道这一点,或者说 Cursor 有很多博客讲他们如何快速迭代。嗯,但他们似乎在持续上传新的 checkpoint 并使其可用。
00:21:47 Nathan Lambert:嗯,但我同意。我猜这像是他们最终 RL 训练过程中的一个时间门控。就好比在他们的 RL 训练结束时它仍在略微改进,而他们只是走个过场打个勾而已。
00:22:03 Nathan Lambert: 好。Qwen DeepSeek V4 据说会推出一个预览版。嗯,关于 DeepSeek V4,我觉得有意思的一点是,flash 模型实际上要受欢迎得多,也就是他们那个更小的版本,对很多人来说它似乎是一匹绝对的主力干将。所以我认为那才是他们值得关注的模型。我不指望 V4 Pro 会是一个戏剧性的突破。这跟任何类似情况都一样,比如小米要是很快发布一款新的 MiMo Pro 模型。我不指望它的发布会有那么大的冲击力,但它很可能会是一款非常扎实的模型。只是这很难说。他们仍然是一个相当新的入局者。MiniMax,我觉得,在玩一个不一样的游戏。我不认为 MiniMax 在追逐那种 Kimi/GLM 月之暗面式的通往 AGI 的氛围。
00:22:46 Florian Brand: 哦,那我可不同意这一点。
00:22:49 Nathan Lambert: 你觉得,你觉得 MiniMax 还在这场局里吗?
00:22:52 Florian Brand: 是的,我我我我觉得他们确实感受到了这种压力,尤其是因为他们和 GLM 一样是一家上市公司,如果你看看股价表现,RIP,那些股票在过去几天里的表现,嗯,这这这这似乎造成了巨大的差别,而有趣的部分将是,呃,许可证,因为他们已经多次更改许可证,呃,变得越来越严格,嗯,如果现在在习,呃,讲话之后,态度又再次转变的话。
呃,MiniMax 是否会回到完全开放的许可证,这会很有意思。另一个有意思的点是,K3 会采用哪个许可证,因为他们说过会开源,但我认为他们并没有就实际叠加在上面的许可证做出任何承诺。
00:23:45 Nathan Lambert:是的。我的意思是,这正是关键所在。是的,我们我们会看到的。嗯,Ling、美团、LongCat 有点类似,都是非常强的模型,可能内部从中获得了大量价值。但它们并没有同样的开发者突破。嗯,所以这大概是七到八家中国实验室。我可能漏掉了一些。另外我们也可以聊聊美国实验室。除此之外,Gemini 3.6 flash 发布了。看起来还行。就像是就像是它只是一个小小的提升。它更快了。它没那么啰嗦了,但好像也没什么大不了的。我们会停止我们会停止分享这个。嗯,这就是 Gemini 从我们这里得到的提及量。
但我确实认为值得稍微聊聊美国生态。我认为有一些新兴玩家。Thinking Machines 发布了他们的第一个模型。我和他们中的一些人聊过。他们非常支持弄清楚如何用 Tinker 做出一个可微调的模型,我认为这是一个我真心真心推荐给大多数开放模型构建者的研究领域。
我认为如果你能在那里获得心智份额,你就会获得大规模采用,因为关键更多在于能针对真实任务进行微调,而不是拥有最好的最好的数字。嗯,所以这就是他们的 Inkling 模型,一个一万亿参数的模型,得分还算不错,但并非前沿水平。
我认为,有点像 DeepSeek V4 那样,他们打算发布一个更小的版本,总参数量大约只有四分之一,但性能非常非常出色。如果 Inkling small preview 在几周内推出,我确实认为那会是一个被广泛使用的模型。
它的规模很适合做自动化任务和特定领域任务,可能不像 Kimi 和 GLM 5.2 那样属于通用智能体类型,但我觉得这非常契合他们的业务。嗯,我知道还有一些其他的——我想说,美国那些较小的玩家似乎也还不错,比如 Arcee 今年早些时候发布了他们的模型,目前仍在稳步推进。
Poolside 已经开始发布一些模型了。
他们在过去几个月里已经拿到了几个,而且看起来还准备在此基础上发布更多模型。所以他们确实——Reflection 永远处于"模型即将推出"的阵营,如果他们真的致力于开源,那对他们来说确实有必要拿出一些模型、一些代码或者别的什么东西,这样他们才能开始让开发者飞轮转起来。
这需要付出很多——把模型发布出来真的很难。比如我和 Thinking Machines 的一些人聊过,感觉就像是,哦,要真正做成这件事工作量很大,我觉得是这样。嗯,Nvidia 也在稳步推进。我觉得他们现在已经是稳定的玩家了。
他们一直在发布模型。他们很快就会发布更多。他们发布了很多数据。我正在"逼"他们,想让他们发布 Qwen 风格的小模型,就像 Gemma 那样。
Gemma 只有这些像 Qwen 竞品一样超级受欢迎的模型。嗯,Gemma 模型有点——它们在规模上、或者在对应规模的架构之类的东西上都有点杂乱无章,但 Gemma 模型在采用率上真的真的能跟 Qwen 模型匹敌。嗯,我不确定它们在研究上是否同样易用,这可能需要一段时间。
可能需要多次迭代。就像现在很多语言模型研究都是围绕小型 Qwen 模型和基于 Qwen 的模型来设计的,所以这需要一段时间。就像人们非常熟悉怎么用这些模型,以及怎么用它们得出研究成果。所以我希望 Gemma 继续推出,并能在这个细分领域里竞争。
我不知道我在这里漏掉了谁。
00:27:22 Florian Brand:不,我认为两者都是大玩家。呃,这个领域正变得越来越广。呃,就模型创造者而言,像去年,除了 Gemma 3 和嗯 GPT-OSS,我们还有别的发布吗?
00:27:41 Nathan Lambert:GPT-OSS 2 会很猛,而且显然,显然 Nemotron 也是。嗯,哦,我还想到年初的 Llama 4,但呃,我不想让那个被遗忘,不过我们确实看到,越来越多的玩家正在加入,并以一种真正令人难以置信的速度推出模型。
00:27:59 Florian Brand:比如 Poolside 在过去两三个月里已经发布了三到四个模型。呃,而且他们似乎已经找到了某种方法,能够相当稳定地持续产出模型。嗯,这也是我们在开源侧看到的情况。我们正在谈论 GLM,我觉得他们模型发布的迭代周期现在已经在 1 到 2 个月之间,每一次新的迭代都变得越来越好,这与闭源实验室的做法非常相似,就像我们现在每隔大约 6 周就会得到一个全新的 GPT、一个全新的 Claude。呃,所以就拥有足够好的流水线来发布越来越强的模型而言,他们必须做到,或者说开源生态确实已经搞明白了,或者看起来已经搞明白了。
00:28:59 Nathan Lambert:是的,我同意。这很有前景,但同时也挺好笑的是,美国生态圈刚开始发布一些模型,然后你就看到 Xi 上了麦,还有这两个模型。就好像真的很难追上,因为训练出人们真正会使用的模型需要大量的机构级专业能力。而我认为,这正是那些正在发布模型的美国公司现在所意识到的——这些可不只是刷榜刷出来的、蒸馏出来的知识产权窃取模型。
这些模型确实相当出色,人们会在自己的内部交易基准上拿它们做对比,然后看看在可量化的指标上击败它们有多难。我觉得,我从美国做交易模型的一些人那里感受到了这种情绪,就是——我认为人们应该在大规模化和可微调性上做创新,试着去利用这个离家很近的潜在市场,但与此同时,每家公司面临的压力都太大了,都想要发布一个能号称是前沿的模型。
我觉得投资者对这么多参与者都抱有这种期待,所以他们其实是在做一件相当困难的事,而接下来一年美中之间的平衡会如何演变,将会很有意思。
00:30:25 Florian Brand: 是的,我觉得,或者说总体而言,我和很多其他人都谈过整个生态系统,这也是你在一开始就提到过的。我觉得我们正看到模型能力之间越来越明显的分化——对于很多任务来说,能力已经足够好了,比如很多编程任务,当前的前沿模型无论是开源还是闭源都已经足够好。嗯,在这里,改进感觉越来越不那么重要了。
但如果我们把目光投向最前沿的前沿,也就是发现新的数学证明、找到新的疗法、发现新药物、发明新事物,那似乎完全是另一头猛兽,而且很可能在相当长一段时间内都会被最前沿所主导。那么最大的问题就变成了:就潜在可触达市场而言,这究竟有多重要,以及这会在多大程度上成为焦点。
我认为,或者说我的基本判断是,我们正看到前沿在不断收窄。我们在 Mythos 用于网络安全、GPT……或者用于生物技术时已经看到了这一点——那些模型不会对所有人开放,而且如果我们考虑到有关 Anthropic 如今正在孵化或创建一些内部实验室来研发药物的报道,或许连外部合作伙伴也无法获得。
所以,最前沿对所有人来说都是不可触及的,而接近前沿的能力则正变得越来越商品化,这带来了许多不同的影响,尤其是当你考虑到网络安全之类的事情时。两三天前 Hugging Face 有一份报告称,他们发现有某个智能体试图入侵他们的系统。
他们试图用 GPT 和 Claude 来分析它,却无法做到,因为所有的护栏都把它们拦住了。于是他们不得不使用 GLM,一个能力较弱的模型,但它在应对这类防御性操作时没有护栏。他们不得不使用一个更差的模型来保护自己或分析数据,这是一种糟糕的处境——如今美国公司因为封闭的前沿对它们不可触及,而不得不依赖能力较弱的模型。
00:33:10 Nathan Lambert:是的。而且我认为这实际上是不采取任何行动的最佳论据之一。就好比,如果世界其他地方都能使用这些开放模型,而我们却禁止美国公司使用它们,那么美国公司的防御能力与全世界攻击者之间的差距就会不断扩大——在网络攻防方面——我们可以争论在当前能力水平下,网络方面的直接风险到底有多大,但如果你从结构上把局面设定成防御方不会随着时间推移变得更强,而攻击方却可以,那看起来就是——网络风险为什么会变得更加真实?而且要说清楚,那将会是:如果你禁止美国公司使用最好的中国开放权重模型。
而这种禁令很可能是一种影子禁令,也就是以法律行动或惩罚的法律威胁为手段,却不明确说明具体通过什么路径来实施。目前有很多关于这方面的讨论。我不太确定我们会不会对此有大量评论,但很明显,华盛顿正在试探各种不同的方式来限制美国使用最好的中国开放权重模型。
我认为这是某种恐慌煽动的下游产物。我们也会过渡到知识蒸馏的问题。就好比,所有这些来自美国主流 AI 媒体叙事的说法,都在把中国模型指向为窃取知识产权、或者具有危险性、或者与中国政府——一个威权政府——有关联。
而这一切似乎都在汇聚到这样一个时刻:人们对在 AI 上采取行动产生了兴趣,却又不太清楚该从哪里下手。于是可能会对所谓的“敌人”动用一种粗糙的手段,接下来我们就可以过渡到知识蒸馏这个话题。我觉得这方面有很多讨论。最近 Ben Thompson 终于就知识蒸馏发表了看法。我认为 Ben 大概是科技界阅读量最高的博客之一(Stratechery)。我觉得这场辩论——让我们看看该从哪里开始这场辩论。核心问题大概是:知识蒸馏到底有多大帮助,以及你该对此做些什么?我一直持这样的观点:随着中国模型越来越接近前沿,并且交易机制转向 RL,知识蒸馏的影响力正随着时间推移变得越来越小。知识蒸馏通常发生的方式是,中国实验室黑入 API。黑入这个词可能有点重,但他们会越狱 Claude 和 GPT 的 API,以提取推理 token。
当你拿到带有工具调用的推理 token 时,那就是完美的 SFT 数据和/或中期训练数据,可以用来训练基础模型,从而在一个重要领域里植入一些智能体行为。而在此之后,后训练的核心部分就是在智能体领域进行大规模 RL,以此推动前沿,以及他们今天所做的一切;随着 SFT 在以往几代中变得不那么普遍,RL 只会变得越来越流行。
过去你只要扩大 SFT 的规模,就能非常接近前沿,而如果真能说从 Claude 或 GPT 那里拿到一百万条智能体 rollout,把它作为你的 SFT 数据集并在其上训练,那才会真正产生巨大影响。
我认为在往年,这本来能在更大程度上帮助你抵达前沿。Ben Thompson 说过一番让我非常恼火的话,他极力宣称,随着你做 RL,知识蒸馏正变得愈发有影响力。他是在他那篇《谁害怕中国模型》的文章里这么说的。我们可以在下面放上链接。那是一篇公开文章。然后他还上了自己的播客巡演。他也有自己的播客,说了同样的东西。而我认为非常重要的是要指出,在 RL 阶段进行知识蒸馏要困难得多。
他所说的是,在 RL 期间可以使用的那种评分模型,本质上就是你可以让一个模型检查一次 rollout 的智能体轨迹,并对不同部分打分,判断它是否完成了奖励、采取了哪些动作。而他暗示,中国的实验室正在用 Fable 和 GPT 5.6 以及最强的模型,在 RL 中实际执行这种监督。
问题在于,大规模的 RL 运行意味着数以百万计、数以百万计的 rollout。我记得 Thinking Machines 的博客文章提到,他们最终的 RL 运行大约有 2000 万到 4000 万次之类的。所以要在像 Fable 或 GPT 5.6 这样的 API 上做这件事,会贵得离谱,而且很可能还会成为时间瓶颈,因为这些模型相当慢,而且坦白说,相比使用你自己定制的更强大的模型,它甚至可能不会给你带来性能提升,还有很多类似这样的问题。
所以,我只是觉得,那种认为蒸馏之所以更有帮助是因为强化学习正变得越来越普遍的观点,并没有我们今天所拥有的文献作为依据。这让我很为难,因为 Ben 的文章还得出结论说,我们应该把禁止蒸馏的服务条款定为非法,我有点想支持他这个激进结论,让美国公司进行蒸馏合法化,但我无法支持任何我认为基于不实信息、或者说被误导的信息而得出的结论。
所以,我也是 Ben 的粉丝。如果你是 Ben 的粉丝,也能在这件事上推他一把,我真的觉得你应该这么做,因为大概还有一期播客。他要在什么上面录?他什么时候录 Sharp Tech?周四。
我们得行动起来,让他更正记录,因为我也不知道。我我觉得特别烦的是,科技界最有影响力的人试图在我们关于蒸馏的观点上充当盟友,而他的观点却是我们什么都不该做。嗯,但这很难。就好像他的影响力如此之广,以至于这现在就要成为我们不得不去驳斥的既定共识了,我猜这总比……我也不知道,其实不,这没帮助,因为他在说蒸馏更重要,这意味着那些对此感到害怕的人会拿这一点作为论据,说我们应该采取行动,即便他们不会这么做,因为他们大概不会同意他的结论。
我也不知道。这就是我的吐槽。Ben,你错了。
00:39:16 Florian Brand: 是的,我确实认为区分这些阶段很重要。尤其是,毫无疑问它被用于 SFT 阶段,也就是后训练的第一个阶段或其中一个阶段,这也正是模型习得自身举止方式的地方——这就是为什么模型会说“哦,我是 Claude”,因为它们是在 SFT 阶段学到这一点的,人格就是在这个阶段形成的。但强大的能力是在 RL 阶段获得的,那才是花钱的地方,也是你需要一个足够快的评判器的地方——最好情况下它直接跑在同样的 GPU 上,或者离你的 GPU 非常近,用一个规模较小或足够快的模型,这样你就不会被它拖成瓶颈。
嗯,就影响而言,也很难说更好的模型生成的 SFT 数据相比更差的模型能带来多大的影响或多大的提升。所以,如果你能拿到来自最新 Claude 模型的 1000 万 token,对比来自落后两代的开放模型,在保持阶段不变、预训练阶段也不变的情况下,这究竟能带来多大的提升,是一个悬而未决的问题,我认为我们不会在论文里看到答案,因为那样你就得展示你的 SFT 能力和越狱能力了。
00:40:48 Nathan Lambert:但我想在这一点上再深入强调一下,关于生成 SFT 推理轨迹,已经有相当多的文献了,无论是其中最知名的,还是他们所做的那条开源工作线,Open Thoughts 3和 Open Thoughts Agent 在某种程度上一直是过去几年里扩展推理 SFT 的基础性工作,而每当有人重新审视这个问题时,他们都没有找到那个答案,即在你所在领域中性能最强的模型就是最适合做 SFT 的教师模型。人们试过,我试过,很多人都试过,这个想法太简单了,就好比最先进的开放 SFT 数据集是建立在 QwQ-32B 之上的,一个古老的推理模型之类的。
为什么我们不能直接从 GLM 5.2 生成补全,在其上做 SFT 并改进模型呢?我们不知道。这就像这项研究,很多人试过,但它并不是一个已有答案的研究问题。可能存在某种情况,比如基座模型、中期训练与 Qwen 太接近了。因此就像很难打破。
你必须重做中期训练。我认为你必须为推理重做中期训练。我认为推理中期训练和推理 SFT 是如此紧密地交织在一起,几乎为它们使用不同的词都没有意义。这可能就是问题所在。但文献甚至不知道如何扩展,就像如果我有一个神奇的 API,能给我来自 Claude/Gemini 的推理轨迹。
我其实不知道,我拿它去微调一个 OLMo 模型,是否会让 OLMo 变得更聪明。
这是最疯狂、最悬而未决的研究问题之一。而这恰恰让知识蒸馏这件事变得特别好笑,因为情况是这样的:没错,我认为中国的实验室确实在用 Opus 这样的强模型来生成一些 SFT 数据,但他们同时也在创新。我当时想,我真希望他们能告诉我们怎么把这玩意儿搞成。
我认为这个范式是这样的:OpenAI 和 Anthropic 找到一个他们做得极其出色的细分领域,然后中国的实验室可以从那里拿到一些样本,用来启动他们的数据引擎,而这就是你能在某个特定领域上领先几个月的地方。
但像数学、代码以及 Terminal-Bench 这类核心领域上的爬山式优化,他们做的其实是同一件事,而这件事难就难在:要生成那些对当前模型来说很难的提示词——也就是带有环境的难题——并且要能提供真实的、非奖励黑客式的学习行为。而这正是当前前沿数据研究的样子。而且,生成这些难题真的很难。我确信中国的实验室也在做同样的事情。我不知道。这就是我的一通吐槽。我有点忘了我们对话的上下文了。
00:43:23 Florian Brand:不,不,我同意。或者总结一下,是的,SFT 或知识蒸馏确实有一定效果。是的,它能给他们带来提升,但没有人们所希望或似乎认为的那么大。
00:43:39 Nathan Lambert:我觉得这是对这场对话一个很好的总结。这种说法也变得有点令人厌烦,因为它说所有这些开放模型之所以好,只是因为它们在搞知识蒸馏。这绝对不是事实,因为如果真是这样的话,每个人都能轻易地通过用 GLM 或 K3 的数据做知识蒸馏来赶上它们。但我们并没有看到,或者说不会仅凭 SFT 就实现这一点。
00:44:12 Florian Brand:是的,我同意。你有什么预测,或者还有更多想聊的话题吗?
00:44:18 Nathan Lambert:说到预测,我觉得我们——或者说我重新回顾了我们去年的预测,基本上就是说一切都会像前一年那样继续下去。我们当时预测会看到更大的模型,超过 2 万亿参数,事实确实如此,而我认为今年在模型规模方面不会出现更大幅度的爆发。我们可能会看到某个模型的总参数量略超 3 万亿,但我不指望今年会出现 5 万亿或 10 万亿参数的模型,那会让我非常意外。然后去年那份清单——我们能重新做一遍吗,不用把整个都做完。
00:45:03 Florian Brand:哦,当然可以。
00:45:03 Nathan Lambert:这就是我们在 2025 年底所处的位置,现在你会把谁列入前沿?嗯,是 Kimi,是智谱,DeepSeek 如今有点难对付。我觉得它们会算是接近的竞争者。所以我会把 DeepSeek 和通义千问列为与 Kimi 和智谱同属前沿的接近竞争者。你觉得还有其他人配得上接近竞争者吗?因为在那之后,值得关注的和更低的就有太多了。
00:45:38 Florian Brand:我我觉得到年底我们会看到 MiniMax 带来的一个惊喜。我觉得我们会看到一个大模型,像是真正的大模型,不是 M3 那种规模,而是万亿参数以上,它会在输出方面让我们惊讶,与 MiniMax 之前的表现相比,所以我仍然会把它们列为到年底的接近竞争者。
00:45:54 Nathan Lambert:你觉得到年底会有任何美国公司进入接近竞争者之列吗?Nemotron,我不觉得我会把它放在那里。Thinking Machines 更接近,尤其是如果那个更小的模型真的取得突破的话。但我觉得我暂时还不会把它们放在那里。Reflection 据说只想在拥有前沿模型时才发布。但那么问题就是,我们能拿到它吗?比如我们觉得会有任何美国公司进入这个行列吗?大致来说,到年底我们的前五名会是什么样?所以前五名还是同样的,只是重新洗牌了。
00:46:36 Florian Brand:我会说,它们确实有可能非常接近。嗯,这也取决于我们认为什么才算接近。比如我觉得 Nemotron 和 Thinking Machines 将会发布的模型,可以作为非常好的基础模型,供你针对自己的领域进行微调,这并不意味着它们能像前沿模型那样直接使用,但它们具有如此大的实用价值,嗯,我会把它们列为势均力敌的竞争者,因为你只需要找到自己的数据,嗯,把模型推向正确的方向。
00:47:12 Nathan Lambert:嗯,我当时想的是,到那时我们会让这成为一个六家公司的群体,再加上一家美国公司,比如如果我们在十一月下旬做这件事,我猜会有一家美国公司——很可能是 Nvidia thinky 或者 Reflection mo——做出一些东西,让我们可以说有一家美国公司进入了这个顶级集群,这将是相当长一段时间以来的第一次。
00:47:43 Florian Brand:是的,我我我觉得这是现实的。我我的一张外卡是腾讯,我觉得我们可能会在年底前看到一些东西。嗯,他们有了新的领导层。嗯,他们这次以 Apache 许可证发布了他们的混元模型。等等,腾讯一直都有这些自定义许可证,不允许英国和韩国的任何人以及整个欧盟使用他们的模型,还有可接受使用政策等等。嗯,而有了混元和他们的新领导层,他们得到了一个约 2500 亿参数的非常能打的模型。嗯,我觉得到年底我们可能会看到一个大模型发布,会让那些不关注这个生态的人感到惊讶。
00:48:36 Nathan Lambert: 是的,我也确信我们还会遇到一些意外。AI 尤其是开放模型,向来如此。它非常非常难以预测。好,我觉得这是个不错的收尾点。我们大概真的应该每季度做一次这个。并不难,大家也会喜欢。嗯,很高兴见到你,我们很快再聊。希望很快能线下见面。
00:49:02 Florian Brand: peace。
Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — geopolitics of US v China, economics of open vs. closed models, security at the frontier of AI, and so on.
Chapters:
00:00 Welcome & context
04:38 Living with / using Kimi K3
08:53 GLM 5.2’s continued role
12:47 How are the Chinese models this good?
17:41 Data, environments, and a tour of the Chinese labs
19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax…
24:08 The US open-model ecosystem
30:25 Frontier vs. near-frontier, and the cybersecurity case against bans
34:58 Distillation and the Ben Thompson debate
44:12 Predictions and a frontier tier list
48:36 Wrap-up
Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.
00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We’re doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn’t a detailed layout state of affairs.
Qwen announced their next big model is going to be open weight, which is a big change of things. I think there’s just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.
00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I’m not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it’s actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.
00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it’s it’s like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren’t out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.
But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there’s a lot of performance that could still be extracted from it. No.
00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and uh a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. Um but people you want to talk about using the model like you actually signed up for the the coding program and used it. So like getting this out there is good context.
00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so uh for the $200 plan uh which is their biggest one similar to to all the others but they have um like I think $40 and $100 as well. Uh but the biggest plan has uh 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people um online are saying that they hit API errors constantly and so far I’ve been I’ve been uh pretty well off if I’m uh going to say that. Um and in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.
Um even my um expectations even with things like uh some research tasks like I have or at interconnects we now have over a year of data on on open models um and I ask the frontier models to come up with some interesting analysis which we haven’t done before in uh because we do our own analysis and have this published uh and I asked them all right do something new and um surprise me, basically. And a lot of the models or the frontier models u or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.
Uh Kimi K3 did some more interesting things um I’ve told it explicitly to scrape Reddit um and then it found uh some subreddits I haven’t even considered and then found out for example that the Reddit discussions are um one or two months in uh more recent or they found they find the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking um but it is something that Kimi surprised me at compared to to all the other frontier um models.
A simple question like can you use this for most of the core work you do in terms of like the exp you you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine
00:07:24 Florian Brand: uh it’s really depends on how much leeway I give it like, the big thing I have seen with Kimi K3 right now I’m I’m working on u the framework we are doing at uh at Prime Intellect, where I work, and the main thing I found with Kimi is its code is a lot simpler uh which makes it way more readable uh but it misses some things that Codex just or like we’re talking 56, 55 and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kind of tasks.
But if I like I read the code and I say all right that’s really good code and then I give it a pass over with with Codex and it finds all these niche niche cases where it doesn’t excel but for supervising runs or for running uh some experiments it is actually really usable. Um and for some other niche things like you can just let it run. The one downside is but that’s also because the API is completely swamped in terms of users and it their servers are in China. The wall clock time is significantly significantly higher than GPT. But I would say like if I was to to push it and use it in my daily workflow, I would be slower, but I wouldn’t be slowed down by so much that I would say, “All right, that’s unusable.”
00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion where like I would go bop around SF and people are like yeah I genuinely use this for this part of my like agentic coding and/or workflow. Um, how do you like I feel like were you in that camp using GLM at all or
00:09:20 Florian Brand: Yeah. like where do you I also use used and use uh GLM mostly because we have an internal endpoint which is really fast and we have or or before that I I also used an API which had I don’t know 200 or 300 tokens per second. Um and if you can do a lot of task at a good enough level like really fast you just use that model compared to going to Codex then selecting the lesser model then selecting the right reasoning effort then selecting fast like I just use GLM get the same result and uh and it’s uh pretty fine like it it definitely is Sonnet-ish level in terms of capabilities and for a lot of cleanup task for a task that just is grunt work. It really works. Like I I would say you could probably go really far for a lot of the work uh with Kimi K3 as the main agent and GLM for for sub agent work.
00:10:24 Nathan Lambert: Something that’s pretty different with Kimi’s announcement and the scale of models this is. I think it’ll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don’t have the weights yet, and then two, like I don’t think it’s going to be as fast of a roll out on adoption as the like 500B, 700B MoE. like I I there’s going to be more problems there which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week and then like immediately the ecosystem kind of knew how to do this.
I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap
00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.
00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.
It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.
But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?
00:14:43 Florian Brand: Well, I I actually looked at our uh predictions for uh for this year based on our last year’s recap and we basically said that the gap will stay with within a few months. Uh so that prediction seems to largely hold. Um luckily for us, we didn’t put a concrete number whether it’s 3 months, 6 months or 9 months. So we are safe on that side. Um but I think like the general thing we both felt when we were in China and talking to these people like they are like the researchers themselves are teams of two or 300 people all mid20s and all just want one model to be really good like they don’t seem to do any side quests.
They don’t seem to do anything that uh deviates from from these things. And um they might or in terms of compute which is a really hard question for for us to answer especially as uh these Chinese uh chips are now coming online. We have I also think chips I think chip smuggling increased substantially in the last like 6 to 9 months or the chips that have been smuggled started to become online.
00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they’re using them, I I count that similar and I think that that has massively increased in the last six to nine months, this is the partially the result of that and and you’re saying but I just wanted to put that out there of like I do think that they have a lot more compute though than they did when they were training the previous generation of models.
00:16:24 Florian Brand: Yeah. like we like or just for for context two weeks ago I think LongCat released their model which they uh claim and I we know that it is very likely true uh is trained entirely on uh on Chinese chips. uh they didn’t specify publicly which ones but people speculate that it’s uh that it’s some uh Ascends from Huawei um and as the domestic production ramps up and you can they’re probably used most or they are used for for training but they are especially useful for inference which is a huge part of training as well.
So they probably use some mix of uh of Nvidia and other chips for the training part and then an increasingly larger part for the inference part during which during the stage is is really important. So I think their overall compute is increasing and also they don’t actually have a lot of users. So they don’t need to power 1 billion users like ChatGPT has to do, hundreds or thousands of enterprises like Anthropic has to do because they don’t have that magnitude of uh of of paying customers.
00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re like not a research, even if it’s not in your job, it like does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are are a problem, but that’s an aside that we we can wait on. I think the there’s also rumblings of the data and environments industry starting to appear there.
Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seem to utilize external data. So just a few months hearing a whole bunch of a month months after our trip we went in April and then just months later in July, we’re are hearing a few things of like new companies in China and them wanting to buy data and things. And that is uh like a funny timeline of how that changes.
00:18:40 Florian Brand: And I would put error bars on what they actually told us.
00:18:44 Nathan Lambert: And that cuz it’s like so close in time that I don’t know.
00:18:49 Florian Brand: Yeah. That that that might that might be true. Uh but like those things are hard to to pinpoint. I but I would say it it seems like the buying of external data is becoming more of a factor. Um which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because um they buy the the data environments later. But it is it is a factor. How big of a factor like we don’t know. we don’t have any public insights and I doubt that we will get those insights uh from from anyone b uh really uh so that’s definitely one of the parts why um why we are able to to catch up or improve their their model scores.
00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We’ve talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um Qwen, we talked about their biggest model coming. Qwen’s biggest models I will say have tended to relative to the excellence of their small models not had the same like absolute ranking in performance which is a probably a cost of focus. I think it goes with a cloud companies. It’s it’s almost like it’s if you squint it’s almost like Google.
It’s like Qwen has Alibaba has so much opportunity here and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they’re succeeding wildly. But their big models have always not been as excellent as their small models. So I don’t expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open quite as the open bottle name in China drops giant bottle but I don’t think it will be as sustained as a um news story um DeepSeek you can go if chime in whatever
00:21:01 Florian Brand: the the interesting thing is don’t know how how much you follow this but they are have or they have an endpoint which you can use for a preview version and they’ve updated this endpoint daily so they have some really fast iteration cycle because we the we progress in all these um Twitter um benchmarks. So a lot of these SVG things and three.js like all these visual generation tasks the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism um which other companies have as well. Uh we we know this or cursor has a lot of blogs about this how they iterate really fast. Um but they seem to continuously upload new checkpoints and make them available.
00:21:47 Nathan Lambert: Um but I agree. I’m guessing it’s like a time gated within their final RL run. It’s like still slightly improving at the end of their RL run and they’re just like checking the box.
00:22:03 Nathan Lambert: Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um the thing about DeepSeek V4 I think is that the flash model is actually way more popular which is their smaller which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don’t expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don’t expect it to be as big of a drop but it would probably be a very solid model. It’s just like it’s hard to know. They’re still a pretty new entrance. MiniMax, I think, is playing a different game. I don’t think MiniMax is chasing this um Kimi/GLM moonshot to AGI type vibe.
00:22:46 Florian Brand: Oh, I would, I would disagree there.
00:22:49 Nathan Lambert: You think, Do you think MiniMax is still in this?
00:22:52 Florian Brand: Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they’ve changed the license a lot uh to be more and more restrictive and um if there’s now a change of heart again after the Xi, uh, speech.
Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It’s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don’t think they have done any commitments in terms of the actual license where you put on top.
00:23:45 Nathan Lambert: Yeah. I mean that’s it’s super important is the thing. Yeah, we we’ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren’t don’t have the same developer breakthrough. Um, so what that’s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It’s like it’s like it’s it’s a tiny bump. It’s faster. It’s less of a yapper, but like doesn’t really matter. We’re going to stop we’ll stop sharing this. Um that’s that’s the amount of mention that Gemini gets for us.
But I do think it’s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking machines released their first model. I’ve talked to some of them. they’re very on board for figuring out this how to make a fine-tunable model with Tinker and I think that’s a research area that I really really recommend for most of the open model builders. I think if you can get mind share there you will get massive adoption because it’s more about being fine-tunable for real tasks than it is about having that be best best numbers. Um, so this was their Inkling model which is a one trillion parameter which has like decent but not frontier scores.
I think kind of like DeepSeek V4 they’re going to they’re planning to release a smaller which is like a quarter of the size in total parameters which has really really good performance and if Inkling small preview comes out in a few weeks I do think that that will be a really used model. It’s a good size for kind of automating tasks and kind of domain specific tasks and might not be a like general agent type thing like Kimi and GLM 5.2 but I think that suits their business really well. Um I know that there’s some other the I would say like the smaller players in the US seem well like Arcee released their models earlier this year still chugging along. Poolside has started releasing some models.
They’ve gotten a few in the last few months and seem poised to release more models on top of that. So they’re really going Reflection is perpetually in the model coming soon camp and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they’re really committed to open source. It just takes a lot this it’s hard to get the models out. Like I talked to some people at Thinking Machines and it’s like kind of like oh that’s a lot of it’s a lot of work to actually do this I think. And um Nvidia chugging along. I think they’re at the stable player at this point. They’re keeping to release models. They’ll release more soon. They release a lot of data. I’m bullying them to try to get them to release Qwen style small models, which is like Gemma.
Gemma only has these like Qwen competitor models that are super popular. Um, the Gemma models are a little they’re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really really matching the Qwen models in terms of adoption. Um, I’m not sure they’re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that like it takes a while. Like people know how to use these models really well and if with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don’t know any anyone that I missed here.
00:27:22 Florian Brand: No, I think both are the big players. Uh it’s, it is becoming broader. Uh in terms of model creators like last year, did we have any release aside from Gemma 3 and um GPT-OSS?
00:27:41 Nathan Lambert: was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don’t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.
00:27:59 Florian Brand: like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that’s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.
00:28:59 Nathan Lambert: Yeah, I agree. It’s it’s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It’s just like it’s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.
These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I’ve I’ve picked this sentiment up from a few people in the US trading models and it is just like there’s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they’re kind of trying to do a a pretty hard thing and it’ll be interesting how the next year unfolds for the US China balance.
00:30:25 Florian Brand: Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that’s also something you’ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.
But if we look at the frontiers frontier, so finding new math proofs, finding uh new uh cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter uh in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won’t be accessible for everyone um and maybe not even external partners if we consider the reports that Anthropic is now spawning or or creating some internal labs to develop drugs.
Um so the very frontier is inaccessible for everyone and then the near frontier capabilities is becoming more and more commoditized um which has a lot of different implications especially if you think about things like uh cyber security. There was that report from Hugging Face two or three days ago that they had some agent trying to to hack their system. um and they tried to analyze it with GPT and with Claude but were unable to because all the guardrails blocked them. So they had to use GLM, a lesser capable model, but it had no guardrails for this kind of defensive action. And they had to use a worse model to defend themselves or to analyze the data, which is a horrible state to be in that we have US-based companies now relying on lesser models because the closed frontier is inaccessible to them.
00:33:10 Nathan Lambert: Yeah. And I think this is actually one of the best arguments for not doing anything. It’s like if the rest of the world has access to these open models and we ban them for the companies in the US to use and it’s just like a growing disparity between US companies ability to defend and the attackers all over the world in terms of cyber and we could debate like how much of an immediate risks the cyber stuff is at the current capability levels but if you’re setting it up structurally so that the defenders get don’t get better over time and the attackers can like that seems like when the Why would cyber risk become more real? And that would to be very clear that would be if you ban the best Chinese openweight models from being used at companies in the US.
And this ban would likely be a kind of shadow ban, which is the threat of legal threat of legal action or punishment without it being clear on exactly what the pathway to do it is. And there are a lot of talks about this right now. I don’t like like I don’t know if we’re going to have a ton to say about this, but it’s clear that DC is flirting with different ways of restricting the best Chinese openweight models in the US. This is I think downstream of some fear-mongering. We’ll transition into the distillation question too. It’s like all these things from the primary AI media narrative in the US that is pointing towards Chinese models as stealing IP or being dangerous or being affiliated with the Chinese government, an authoritarian government.
And it’s like all these things are leading up to this moment of interest in taking action on AI and then not really knowing where to do it. So potentially taking a crude instrument to the like quote unquote enemy and we could transition into distillation. I think there’s a lot of discussion on it. Most recently Ben Thompson finally chimed in on distillation. I think Ben is probably one of the is probably the highest read blog in tech (Stratechery). I think that the the debate let’s see where do we even start the debate. The core question is like how much does distillation help and what should you do about it? I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the trading regime shifts to RL. The way that distillation tends to happen is that the Chinese labs hack the APIs. Hack is like maybe a strong word, but they jailbreak the APIs of Claude and GPT to extract the reasoning tokens.
When you have the reasoning tokens with the tool calls, that is perfect SFT data and or mid-training data to train the base model with to seed some agentic behaviors in an important domain. And now after that the core part of post-training is to do large-scale RL in agentic domains to so like push the frontier and everything that they’re doing today and RL is only becoming more prevalent with this as SFT becomes less prevalent in previous generations you could get very close to the frontier just by scaling up SFT and that would be what really impactful if you could say take a million agentic rollouts from Claude or GPT have that be your SFT set and train on it.
I think in previous years that would have done a lot more to get you to the frontier. What Ben Thompson has said which made me really annoyed is that he very strongly proclaimed that distillation is getting more impactful as you do RL. He did this in his article who’s afraid of Chinese models. We can link it below. It’s a public one. And then he was also on his own podcast tour. He has also podcast as well saying the same things. And I think it’s really important to say that distillation during the RL stage is a lot harder.
What he said was that the kind of grading models that can be used during RL, which is essentially you can have a model check over the agentic trajectory of a roll out and grade different parts on if it completed the reward, what actions it took. And he’s insinuating that the Chinese labs are using Fable and GPT 5.6 and the strongest models to actually do this supervision in RL. The problem is that big RL runs are millions and millions of rollouts. I think Thinking Machines blog post had like 20 to 40 million or something for their final RL run. So to do this on an API like Fable or GPT 5.6 would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift versus using your own tailored greater model or and many things like this.
And so I just think the argument that distillation is helping more because RL is becoming more prevalent is not grounded in literature that we have today. This is tough for me because Ben’s article also concludes that we should like make terms of service disallowing distillation illegal, which I kind I like want to support his radical conclusion to make distillation legal for US companies, but I can’t support any conclusion that I think is on um infactual mis like misguided information. So, I’m also a fan of Ben. If you’re a fan of Ben and could also nudge him on this, I would you really should because there’s probably one more podcast. What is he going to record it on? Like when does he record Sharp Tech? Thursday.
We We got to get on and get him to correct the record because I I don’t know. I I find it so annoying that the most prominent voice in tech is trying to be an ally for our point of view on distillation is that we should do nothing. Um but like it’s hard. It’s like he has such wide reach that this is now going to be the status quo that we have to debunk which I guess it’s a better status quo than I don’t know actually no it’s not helpful because he’s saying that distillation is more important which means the people who are afraid about that are going to use that as a data point to say that we should take action even if they don’t because they probably won’t agree with his conclusions. I don’t know. That was my rant. Ben, you’re wrong.
00:39:16 Florian Brand: Yeah, I-I do think it is important to to differentiate these phases. Um and especially like there is no doubt that it is used during the SFT stage which is the first stage of or one of the stages for post-training and that’s also where the model picks up its manners like that’s why the models say oh I am Claude because they learn this during the SFT stage that’s where uh this personality is formed but the strong capabilities come during the RL stage which is where the money is spent which is where you need to have a fast enough judge which in the best case just runs in at the same GPUs or very close to your GPUs with smallish or with a fast enough model so you can uh are not bottlenecked by this.
Um and it in terms of impact it is also very hard to say how much impact or how much of a boost the better model SFT data gives you versus a lesser model. Um so if you are able to to have 10 million tokens from the latest Claude model versus two generations behind open model how much of a boost that really gives you if you keep the stage right uh the same and the pre-training stage the same is an open question which I don’t think we will see answered in a paper because then you have to showcase your uh SFT and your jailbreaking capabilities
00:40:48 Nathan Lambert: but I I wanted to double down on this like there’s been a good amount of literature on generating SFT reasoning traces whether it’s the most prominent ones have been opens line of work they did Open Thoughts 3 and Open Thoughts Agent have kind of been the foundational like scaling reasoning SFT works in the last few years and whenever somebody revisits this question they have not found the answer that the strongest model on performance in your domain is the best teacher for SFT people have try I’ve tried many people have tried the idea is so simple is like the state-of-the-art open SFT data set is built on QwQ-32B like an ancient reasoning model or something.
Why can we not just generate completions from GLM 5.2 do SFT on it and improve the model? We don’t know. It’s like the research so many people have tried and it is not an answered research question. There might be something like the base model the mid-training is too close to Qwen. So therefore it’s like hard to break. You have to redo the mid training. I think you have to redo the mid-training for reasoning. I think reasoning mid-training and reasoning SFT are so closely intertwined. It almost doesn’t make sense to have different words for them. That could be the issue. But the literature doesn’t even know how to ext like if I had a magical API that gave me reasoning traces from Claude/Gemini. I actually don’t know if me like fine-tuning an OLMo model on that would make OLMo smarter.
It’s one of the most wild unanswered research questions. And this is just makes the distillation thing so funny where it’s like yes the Chinese labs I think are using strong models like Opus for some SFT data but they’re also innovating. I was like I would love to them to tell us how to make this freaking work. And I think it’s the the paradigm I think is like open AI and anthropic find a niche domain that they do so well at and then the Chinese labs can get some samples there to kind of bootstrap their data engine and that’s where you will gain you will gain a few months on a specific domain.
But a hill climbing on these core domains like math and code and like Terminal-Bench like they’re just doing the same thing which is like so hard to generate prompts which are problems with environments that are hard for the current models and provide real nonreward hacking um learning behavior. And like that is what frontier data research looks like right now. And it is like it’s hard to generate these hard problems. And I’m sure the Chinese labs are doing the same the same things. And I don’t I don’t know. That’s that’s my rant. I’m kind of lost the context of our conversation.
00:43:23 Florian Brand: No, no, I would I would agree. Or to to to recap, yeah, SFT or distillation has some effect. Yeah, it gives them a boost, but not that much uh as people would like or or seem to think it gives.
00:43:39 Nathan Lambert: I think that’s that’s a good good summary of the of the of the conversation. It also becomes kind of tiresome because it says uh that open all all these open models are just good because they are distilling. um which definitely isn’t the case cuz if if it were the case, everyone would would be easily able to catch up to a GLM or to a K3 um by using its data for distillation. But we have not or we won’t see this from SFT alone.
00:44:12 Florian Brand: Yeah, I agree. Do you have any predictions or or more topics you want to get to?
00:44:18 Nathan Lambert: Um in terms of predictions, I think we are or I I revisited uh ours from from last year and it basically said everything will continue uh like it did uh the previous year. Uh we predicted that we will see bigger models uh up over two trillion parameters which it did and I don’t think we will see a much bigger explosion in terms of model size this year. we might see something or some model a bit bigger than three trillion parameters uh total but I don’t expect a five or 10 trillion parameter model and we open this year that would really surprise me um then list from last year can we redo this we don’t have to do the whole thing
00:45:03 Florian Brand: oh sure
00:45:03 Nathan Lambert: this is where we were at the end of 2025 who do you put in frontier now well it is Kimi and it is Zhipu DeepSeek is kind of a hard nut these days. Like I think they would be in close competitors. So I would put DeepSeek and Qwen the one as close competitors with Kimi and Zhipu as Frontier. Do you think anyone else would deserve close competitor? Cuz after that noteworthy and below like there’s so many.
00:45:38 Florian Brand: I I think we will see a surprise from MiniMax by end of the year. I think we will see a big model which like a really big model not uh M3 size but trillion parameters plus which will surprise us in terms of uh the outputs of MiniMax compared to before uh so I would still put them at close competitors by the end of the year
00:45:54 Nathan Lambert: do you think any US companies will be in the closed competitors by end of the year Nemotron I don’t think I would put there Thinking Machines closer especially if the smaller model really breaks through. But I don’t think I would put them there yet. Reflection is supposedly like only wants to release if they have a model that’s frontier. But then the question is will we get it? Like do we think that any US companies will get into this what is roughly like our top five by the end of the year? So the top five are the same but reshuffled.
00:46:36 Florian Brand: I would say it is possible uh that they are really close. Um it it also depends on what we think matters for closeness. Like I think uh Nemotron and um uh Thinking Machines will release models which act as really good base to be fine-tuned for your domain which doesn’t mean they are usable like a frontier model but they have so much utility uh that I would put them into close competitors because you would just need to find your data and uh to push the model into the right direction.
00:47:12 Nathan Lambert: Um as a I was going to think that we would make this a group of six with a US company by then like if we do this in late November I would guess that a US company pro most likely Nvidia thinky or Reflection mo does stuff that gets us to say that there is an American company in this like top cluster which would be a first time for a while.
00:47:43 Florian Brand: Yeah, I I I think that is realistic. My my one wild card is Tencent, which I think we might see something by end of the year. Uh they got some new leadership. Uh they released their Hunyuan model under Apache this time. Wait, so Tencent always had these custom licenses which disallowed anyone in the UK and South Korea and the entirety of the EU to to use their their model and also had acceptance use policy and so on. Um, and with Hunyuan and their new leadership, they got a really competent model at 250ish billion parameters. Um and I think by end of the year we might see a big model release which will surprise the people not following the ecosystem.
00:48:36 Nathan Lambert: Yeah, I I am also sure we will be in for some surprises. This is always the thing with AI and especially open models. It’s very very unpredictable. Okay, I I think this is a good place to stop. We probably should really do this quarterly. It’s not that hard and people will enjoy it. Um, but good to see you and we’ll talk soon. Hopefully in person soon.
00:49:02 Florian Brand: Peace.