杂务说明:我目前在外旅行,所以这篇帖子无法录制配音。编者注——我在邮件发出后,补充了关于中国数据行业的第 5 点。
今天,Z.ai 发布了他们的 GLM-5.3 模型,目前仅在编程套餐中提供,即将上线其 API,并将在两周后上线 Hugging Face(开放权重)。这个模型看起来非常出色,分数提升幅度相当惊人。在许多基准测试中,该模型已经超越了月之暗面的 Kimi K3,在某些基准上甚至超过了 Claude Fable 5 或 GPT-5.6-Sol。

以下是更完整的对比:

这使得该模型基本处于智能体编程基准的前沿,而参数量仅约 750B——只有 Kimi K3 的三分之一!Z.ai 的博客文章相当直白,开头就是一句大胆的话:
我们为 GLM-5.3 所做的全部工作,就是扩展后训练。
GLM-5.3 与 GLM-5.2 使用相同的基础模型,但后训练大幅扩展。如果做一个粗略的概括,Z.ai 似乎在後训练方面有优势,而 Kimi 更像是预训练的杰作。这次发布之后,有很多讨论在问:中国怎么能跟得这么紧?这么小的模型怎么能媲美美国领先的公开模型?这些结果是真的吗?
最简单的解释是,Z.ai 非常擅长他们做的事情——值得记住的是,他们研究这条模型路线的时间比业内几乎任何人都长。以下是 GLM 模型的简要历史。
智谱AI成立——2019年
GLM(通用语言模型)——2021年3月——由清华大学数据挖掘/知识工程组 THUDM 发布。权重
GLM-130B——2022年8月——规模化版本。GLM-130B 至 GLM-4 的技术报告——权重
ChatGLM——2023年3月14日——首个聊天版本。权重
ChatGLM2——2023年6月25日——权重
ChatGLM3——2023年10月27日——权重
GLM-4——2024年1月16日——更名为 GLM;随后于6月发布开放权重的 GLM-4-9B。权重
GLM-5——2026年2月11日——最新主要代际。权重
GLM 5.2 于今年 6 月 22 日发布,当时引起了不小的轰动——发布数周后,我经常听到我认识的 AI 研究人员说,他们仍在使用这个模型,因为它的速度快(有些人在内部集群上部署该模型,以获得比公开服务更快的速度)和简洁(作为一个没有回滚等机制的模型,在开发前沿 AI 系统时尤其好用)。GLM-5.2 整体上确实配得上这股热度。
我自己也一直在经历某种类似的否认心态,心想“他们到底是怎么做到这一点的?这些模型肯定没有看起来那么好。”令人有些不适的是,美国公司拥有如此压倒性的资源领先优势,却似乎无法在能力上拉开差距。常见的解释是知识蒸馏,我对此写过很多,但我认为这不是主要因素。关于这一点,最近有一篇论文展示了从前沿模型中提取推理轨迹的简单方法——这正是中国实验室绝对可以大规模使用的东西。我不明白为什么美国实验室没有更快地修补这种行为;相反,他们跑去向政府寻求政策帮助。这在我看来说不通。
Z.ai 的博客文章直截了当,与以强化学习为主导的训练体系相吻合。他们说他们使用了“更多的环境、更多样化的任务,以及在这些任务上投入了更多的训练算力。”你不可能简单“蒸馏”出强化学习环境、大规模运行这些环境的基础设施,或者将它们有效混合的算法。
那么,如果不是蒸馏,中国实验室是怎么做到的呢?他们是在刷榜吗?一个被广泛接受的“刷榜”定义是让模型专注于测试集,从而使真实世界表现与纸面分数产生显著差异。决定因素更多是宏观层面的,而非技术层面的(是的,技术细节当然重要,但不同实验室之间的技术细节更难区分):
Z.ai 的发布周期很可能以天计算,而非像 OpenAI 或 Anthropic 那样以月计算。OpenAI 和 Anthropic 拥有远比 Z.ai 和月之暗面更优秀的内部模型,这几乎是板上钉钉的事。尽管如此,这些美国公司往往需要数月时间才能将模型公之于众,这在前沿模型的采用决策上极大地利好中国实验室。简而言之——中国实验室把美国实验室用于发布前测试的时间全部用来在基准测试上继续攀爬(SpaceXAI 在这方面很可能与中国实验室的做法更为接近)。由于进展速度如此之快,这很可能是决定中国实验室能否留在前沿的最大因素。迄今为止,这对美国实验室来说在经济上尚可接受,因为它们的模型需求依然旺盛。
随着构建大语言模型的实验室内部模型自我改进循环不断加速,如果这些反馈循环中的任何一个需要用户数据,这种更快的发布周期可能会极大地利好中国实验室,让它们的产品在下一个性能大幅超越的模型问世之前拥有更长的生命周期,从而削弱对其模型的需求。
这些显然就是业内许多人担忧的竞赛动态。在如此多的实验室都在领先能力范围内构建前沿模型的情况下,很难看到这种局面在短期内有所缓和。
没错,Z.ai 可能比 OpenAI 或 Anthropic 更看重公开基准测试。这些基准测试,例如在 Artificial Analysis Intelligence Index 或类似聚合平台上取得高分,会直接影响它们的股价。它们在很多方面都需要这样做,才能持续融资并维持团队士气,因为作为挑战美国巨头的草根逆袭者,这是一个绝佳的故事。微妙的基准优化(benchmaxxing)并不一定源于绝望或类似的压力。这在数量惊人的实验室中已是行业标准。许多公司的数据采集策略就是在自己落后的基准测试上购买数据。
Z.ai 并没有为了刷榜而把 GLM-5.3 搞到“烤焦”的程度(至少不是故意的,而且他们会去核查这一点)。眼下每一家实验室都在处理扩展强化学习时遇到的棘手问题。Anthropic 的 Opus 5 和 Sonnet 5 模型尽管基准测试分数高得惊人,但口碑却相当两极分化。整个行业都处于同一条船上,所以有些模型权重用起来就是比另一些顺手,但他们在发布博客里公布的基准分数是实打实的。
GLM-5.3 很可能比 Claude Fable 或 GPT Sol 更偏向窄众模型。GPT-5.2 发布时,除了智能体编程之外,评价也是褒贬不一。与此同时,OpenAI 和 Anthropic 支撑着规模极其庞大的企业,这些企业对模型有数不清的用例。这就是作为处于采用曲线更早期阶段的公司的好处——你可以瞄准最有价值的用例。在后训练阶段,稍微少在意一些东西,会让最终模型的组装容易得多。我这话稍微有点夸张,因为据报道 Z.ai 凭借强劲的本地化部署业务,年经常性收入(ARR)已经达到 10 亿美元。
同样,旗舰版 GLM 模型也一直不具备视觉能力。纯文本模式确实有助于 Z.ai 拿到更有竞争力的分数,但这也是一个竞争更激烈的领域。另一端的例子则是像 Inkling-Small 这样的模型,它被设计成全模态(omnimodal)模型。
(补充)强化学习数据产业正在中国起飞。我们关注的许多消息源和传闻渠道都在不断提到,中国的数据产业正在起飞——很大程度上是由美国数据公司向中国模型实验室销售所推动的。这可能表现为中国实验室购买许多美国前沿实验室所使用的同类强化学习环境,并更早地发布下游经过强化学习训练的模型。我们对这个市场的规模和影响力仍然存在很大的误差范围,但它无疑正变得日益重要。
Z.ai 是一家技术极为精湛的大语言模型机构——其算力利用效率很可能远超 OpenAI 和 Anthropic。这一点值得反复强调。这帮人非常擅长自己所做的事情。该公司与清华大学关系极为密切,而清华汇聚了许多中国最顶尖的计算机科学家。这个庞大且充满干劲的人才库,对他们的成功而言,与对任何西方同行一样,都是核心要素。
总体来看,他们用 GLM 系列模型所执行的似乎是一个非常完善的策略。恭喜发布!我很期待权重公开,这样我就能做更深入的测试了(我倾向于使用 Fireworks 或 Baseten 这类美国开源权重推理服务)。
这是向着强大网络能力在经济领域不可避免的普及所迈出的又一步。Z.ai 已经承认了这一点,并表示:
GLM-5.3 是我们迄今为止在网络安全任务上能力最强的模型。它在漏洞发现、漏洞利用分析以及复杂的多步骤安全任务方面带来了显著提升。这些能力可以帮助防御方更早识别弱点、验证风险并加速修复。
这些能力同时也带来了明确的双重用途风险。因此,我们采取分阶段发布的策略。选定的安全合作伙伴将首先在受控环境中评估 GLM-5.3。随后将提供更广泛的访问权限和 API 服务。一旦必要的安全评估和发布准备工作完成,我们将发布 GLM-5.3 的完整模型权重。
他们接着承认,他们正在通过请求分类器和思维链监控(在模型对齐之上)来监控其平台上的推理过程。关键在于细节,而目前尚不清楚每家 AI 实验室在这方面的执行水平如何。能力的扩散取决于最低的共同标准。
归根结底,当真正的开放权重模型来临时,这类安全措施几乎无关紧要。如果不是 GLM-5.3,那也会是另一个模型。具备这些能力的模型规模正在随时间缩小,变得更容易修改和部署(可能没有安全防护)。智谱(Z.ai)做了一些正确的事情,包括推动更多漏洞发现和主动管理,但任何单一公司都远无法独自应对这一问题。
我们需要由政府或行业联盟主导的产业级指导,立即为这一转变在所有软件领域做好准备。
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out.
Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks’ time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI’s Kimi K3 and on some it’s surpassed Claude Fable 5 or GPT-5.6-Sol.

Here’s a more complete comparison:

This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence:
Scaling post-training is all we did for GLM-5.3.
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.
Zhipu AI Founded – 2019
GLM (General Language Model) — March 2021 — released by THUDM, Tsinghua University’s Data Mining / Knowledge Engineering group. Weights
GLM-130B — August 2022 — Scaled version. Technical report for GLM-130B through GLM-4 — Weights
ChatGLM — March 14, 2023 — first chat version. Weights
ChatGLM2 — June 25, 2023 — Weights
ChatGLM3 — October 27, 2023 — Weights
GLM-4 — January 16, 2024 — rebranded as just GLM; open-weight GLM-4-9B followed in June. Weights
GLM-5 — February 11, 2026 — latest major generation. Weights
GLM 5.2, released on June 22 of this year, was a big deal – weeks after the release, I regularly heard from AI researchers I know who still used the model due to its speed (some deploy the model on internal clusters for faster speeds than public offerings) and simplicity (as a model with no rollbacks, etc., when working on frontier AI systems). GLM-5.2 altogether stood up to the hype.
I’ve been going through some of the same denial myself, thinking “how do they keep doing this? Surely the models aren’t as good as they look.” There’s something a bit off-putting with how the American companies have such a commanding resource lead, but can’t seem to pull away in capabilities. The common answer is distillation, which I’ve writtenat length about, but I deem not to be the major factor. On that note, there was a recent paper that showed simple methods for extracting the reasoning traces from frontier models – this is the sort of thing that Chinese labs could definitely use at scale. I’m confused why the labs in the U.S. haven’t patched this behavior faster; instead they’re running to the government asking for policy help. It doesn’t add up for me.
Z.ai’s blog is direct and matches with an RL-dominated training regime. They say they used “more environments, more diverse tasks, and more compute spent training on them.” One does not simply “distill” RL environments, infrastructure to run them at scale, or algorithms to mix them together effectively.
So, how do the Chinese labs do it if not distillation? Are they benchmaxxing? An accepted definition of benchmaxxing is focusing the model on the test sets, such that the real-world performance meaningfully differs from the on-paper scores. The determining factors are much more big picture than technical (yes, the technical details definitely matter, but are harder to differentiate from lab to lab):
The time to release for Z.ai is likely days, not months as with OpenAI or Anthropic. It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies tend to take months to release their models to the public, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks (SpaceXAI is likely far closer to the Chinese labs here). With the pace of progress being so fast, this is likely the largest determining factor of why Chinese labs stay at the frontier. This, so far, has been economically acceptable for the American labs, as they’ve still had massive demand for their models.
As model self-improvement loops ramp up within the labs building LLMs, if any of these feedback loops require user data, this faster release cycle could massively favor the Chinese labs, giving their offerings longer lifespans before the next vastly superior model comes out, undercutting demand for their models.
These are very clearly the race dynamics that many in the industry worry about. With so many labs building frontier models in the envelope of leading capabilities, it is hard to see this abating in the near future.Yes, Z.ai probably cares slightly more about public benchmarks than OpenAI or Anthropic. These benchmarks, e.g. scoring highly on the Artificial Analysis Intelligence Index, or similar aggregators, have a very direct impact on their stock price. They in many ways need to do this to keep raising capital and maintain team morale, as being the scrappy underdog matching American giants is a wonderful story.
Subtle benchmaxxing does not need to come out of desperation or any similar pressures. It’s the industry standard across a remarkable number of labs. Many companies’ data acquisition strategy is to buy data on the benchmarks they’re behind on.Z.ai is not benchmaxxing to the point where GLM-5.3 is fried (at least not intentionally, and they’ll check for it). Every lab is dealing with the rough edges of scaling RL right now. Anthropic’s Opus 5 and Sonnet 5 models have very mixed reputations, despite the incredible benchmark scores. Everyone in the industry is in the same boat, so some model weights end up being easier to use than others, but the benchmark scores in their release blogs are the real deal.
GLM-5.3 is likely a narrower model than Claude Fable or GPT Sol. When GPT-5.2 was released, it had mixed reviews outside of agentic coding. At the same time, OpenAI and Anthropic support very large businesses with countless use-cases for their models. This is a benefit of being a company earlier in their adoption curve – you can target the most valuable use-cases. Within post-training, caring about a bit less will make assembling the final model far easier.
I’m overstating this a bit, as Z.ai reportedly reached $1B of ARR on the back of a strong on-premises deployment business.Similarly, the flagship GLM models have not had visual capabilities. Being text-only definitely helps Z.ai get more competitive scores, but it is a more competitive space. On the other side of things are models like Inkling-Small, which is designed to be omnimodal.
(ADDED) The RL data industry is taking off in China. Many sources and rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier labs, and releasing the downstream RL’d model sooner. We still have large error bars on the scale and impact of this market, but it is certainly becoming important.
Z.ai is an extremely skilled LLM organization – one that is likely far more compute efficient than OpenAI / Anthropic. This needs repeating. These folks are very good at what they do. The company has very close ties to Tsinghua University, which is home to many of the best Chinese computer scientists. This abundant, eager talent pool is as central to their success as it is for any Western counterpart.
Altogether, it seems like a perfectly good strategy they’re executing with the GLM line of models. Congrats on the release! I’m excited for the weights to be out so I can do more extended testing (I tend to use American open-weight inference services like Fireworks or Baseten).
This is another step towards the inevitable proliferation of very strong cyber capabilities across the economy. Z.ai has acknowledged this, saying:
GLM-5.3 is our most capable model to date for cybersecurity tasks. It delivers substantial improvements in vulnerability discovery, exploit analysis, and complex multistep security tasks. These capabilities can help defenders identify weaknesses earlier, validate risks, and accelerate remediation.
They also create clear dual-use risks. We are therefore taking a staged approach to release. Selected security partners will first evaluate GLM-5.3 in controlled settings. Broader access and API availability will follow. Once the necessary safety evaluations and release preparations are complete, we will publish GLM-5.3’s complete model weights.
They go on to acknowledge how they’re monitoring inference on their platforms via a request classifier and chain of thought monitoring (on top of model alignment). The devil is in the details here, and it is unclear the level of execution every AI lab will have here. The capability diffusion is determined by the lowest common denominator.
At the end of the day, this type of safety barely matters when true open-weights are coming. If not GLM-5.3, then another model. The size of the models with these capabilities is reducing over time, becoming easier to modify and deploy (potentially without safeguards). Z.ai does some of the right things, including pushing for more vulnerability discovery and proactive management, but any single company is far from being able to handle this on their own.
We need industrial-scale guidance led by the government or industry coalitions to immediately prepare for this transition across all software.