X.PIN · @thexpin · X·2026-09-18 19:35·3小时前
AI 导读

Z.ai 宣布其 GLM-5.3 驱动的 Infra Agent 为 GLM-5.3-Flash 设计、调试并优化了生产推理基础设施,覆盖 100,000+ 中国芯片。不到两周吞吐达到初始基线的 3 倍,硬件效率和单 token 成本据称与主流 Nvidia GPU 相当。团队称之为递归自我改进,即模型改进运行自身的基础设施。

X.PIN@thexpin
52AI 编辑部评分,满分 100
2026-09-18 19:35· 3小时前
AI 导读

Z.ai 宣布其 GLM-5.3 驱动的 Infra Agent 为 GLM-5.3-Flash 设计、调试并优化了生产推理基础设施,覆盖 100,000+ 中国芯片。不到两周吞吐达到初始基线的 3 倍,硬件效率和单 token 成本据称与主流 Nvidia GPU 相当。团队称之为递归自我改进,即模型改进运行自身的基础设施。

http://Z.ai says its GLM-5.3-powered Infra Agent designed, debugged and optimized production inference infrastructure for GLM-5.3-Flash across 100,000+ Chinese chips.

In under two weeks, throughput reached 3× the initial baseline, with hardware efficiency and per-token costs reportedly comparable to mainstream Nvidia GPUs.

The team calls it recursive self-improvement: a model improving the infrastructure that runs it.