http://Z.ai says its GLM-5.3-powered Infra Agent designed, debugged and optimized production inference infrastructure for GLM-5.3-Flash across 100,000+ Chinese chips.
In under two weeks, throughput reached 3× the initial baseline, with hardware efficiency and per-token costs reportedly comparable to mainstream Nvidia GPUs.
The team calls it recursive self-improvement: a model improving the infrastructure that runs it.