27 August 2026
Chinese lab Z AI reveals mystery model as GLM-5.3-Flash
First reported
AI Business, TLDR AI and 1 other ran this on , all on the same day.
- Z AI disclosed that Ox Alpha, an anonymous model that ranked highly on OpenRouter, is their new GLM-5.3-Flash.
- GLM-5.3-Flash uses a mixture-of-experts architecture, a technique where only part of the model activates per query, enabling cheap inference.
- The model costs roughly one-tenth as much to run as competing alternatives and performed comparably to Claude Opus on coding tasks.
- Z AI ran the model entirely on Chinese-made chips rather than US hardware, potentially addressing China's reliance on foreign semiconductor supply.
Where they differ
TLDR AIfocused on the model's technical efficiency and benchmark performance.
The Rundown AIemphasized the geopolitical significance of avoiding US chips and breaking into cost-competitive territory against Western models.
What each one reported
Z.ai disclosed that the anonymous ox-alpha model was GLM-5.3-Flash, a 320B-parameter mixture-of-experts model with 18B active parameters that approached Claude Opus 4.8 on coding and agentic benchmarks. The newsletter emphasizes the model's architecture for extreme efficiency and ultra-low-cost inference, served entirely on Chinese AI chips.
Chinese AI lab Z AI confirmed that the anonymous Ox Alpha model that dominated OpenRouter rankings is its new GLM-5.3-Flash, available with open weights at a tenth the cost of comparable rivals. The newsletter emphasizes that the model's entire free usage week ran on Chinese-made chips, suggesting Z AI may have solved a key bottleneck in China's AI development.
Reported by AI Business