Zhipu AI unveils GLM-5.3-Flash, claims record usage on domestic chip cluster

Zhipu AI revealed that its previously code-named Ox Alpha model is actually GLM-5.3-Flash, which ran on 100,000 domestically produced chips. The model processed 62 trillion tokens during a stealth trial, setting a usage record on OpenRouter. Shares of Zhipu AI rose over 12% in Hong Kong following the announcement.
The model's performance on OpenRouter, where it captured roughly a third of weekly token volume, signals growing demand for open-weight systems in coding applications. Zhipu's Hong Kong share price surge reflects investor confidence in domestic chip infrastructure.
The deployment directly addresses Beijing's strategic push to reduce dependence on Nvidia processors amid US export restrictions. By demonstrating that a 100,000-chip domestic cluster can handle global-scale inference workloads, Zhipu positions Chinese hardware as a viable alternative for large AI deployments.
The successful scaling of AI on domestic chips could reshape the global AI supply chain, potentially accelerating China's technological self-sufficiency and reducing its vulnerability to US export controls. Developers worldwide may gain access to alternative open-weight models, fostering competition in the AI marketplace. However, reliance on Chinese infrastructure could raise concerns among Western enterprises about data governance and regulatory alignment, potentially creating new divisions in how AI services are sourced and deployed internationally.