27 August 2026
Z.ai open-sources large language model at reduced cost
First reported
Deep Learning Weekly ran this on .
- Z.ai released GLM-5.3-Flash, a model that processes text and images with a context window of 1 million tokens [amount of text it can hold at once].
- The model costs one-tenth as much to run as the full GLM-5.3, making frontier-scale AI [cutting-edge, powerful] models cheaper to access.
- GLM-5.3-Flash uses a mixture-of-experts architecture [activates different specialized sub-models for different tasks], with 320 billion total parameters but only uses 18 billion actively.
How it was covered
Deep Learning WeeklyEditorial team
Z.ai released GLM-5.3-Flash, a 320B/18B mixture-of-experts model with 1M-token multimodal context at one-tenth the cost of the full GLM-5.3. The newsletter emphasises the significant price reduction making frontier-scale models more accessible.