# Z.ai 称 GLM-5.3 驱动的 Infra Agent 在 100，000+ 中国芯片上优化 GLM-5.3-Flash 推理，吞吐达基线 3 倍

- 来源：X.PIN (@thexpin)
- 发布时间：2026-09-18 19:35
- AIHOT 分数：52
- AIHOT 链接：https://aihot.news/items/cmu6wllox0e7crowk5crz6o0z
- 原文链接：https://x.com/thexpin/status/2100911462357102694

## AI 摘要

Z.ai 宣布其 GLM-5.3 驱动的 Infra Agent 为 GLM-5.3-Flash 设计、调试并优化了生产推理基础设施，覆盖 100,000+ 中国芯片。不到两周吞吐达到初始基线的 3 倍，硬件效率和单 token 成本据称与主流 Nvidia GPU 相当。团队称之为递归自我改进，即模型改进运行自身的基础设施。

## 正文

http://Z.ai says its GLM-5.3-powered Infra Agent designed, debugged and optimized production inference infrastructure for GLM-5.3-Flash across 100,000+ Chinese chips.

In under two weeks, throughput reached 3× the initial baseline, with hardware efficiency and per-token costs reportedly comparable to mainstream Nvidia GPUs.

The team calls it recursive self-improvement: a model improving the infrastructure that runs it.
