ComputeLabs Research

Zhipu AI’s GLM-5.3-Flash used domestic accelerators; its Encode–Prefill–Decode architecture tripled serving performance versus the same-hardware baseline.

· ComputeLabs Research · from the August 26, 2026 edition

Before release, Zhipu AI tested GLM-5.3-Flash anonymously as Ox-Alpha, known in the Chinese community as Niulai (牛来), on OpenCode and OpenRouter. Zhipu said Ox-Alpha became the week’s most-used model and set call-volume records on both platforms.

All request traffic during the test was served by domestic accelerator chips connected through Zhipu’s self-developed high-bandwidth interconnect network. Zhipu did not identify the accelerator models in its official statement; a separate report named Huawei, Moore Threads and Hygon as possible suppliers but said Zhipu did not comment on that information.

At the cluster level, Zhipu used a production Encode–Prefill–Decode architecture. It separated multimodal encoding, prompt prefill and token-by-token decoding into independently scheduled and independently scalable worker pools.

Zhipu said end-to-end serving performance was three times its initial baseline on the same hardware. The company also claimed that hardware efficiency and per-token cost had reached levels comparable with mainstream NVIDIA GPUs, but it did not provide the underlying benchmark values or accelerator model specifications.

Additional reporting

  • Zhipu AI’s GLM-5.3-Flash
  • Encode–Prefill–Decode

All 21 stories from August 26, 2026