27 August 2026

Z.ai open-sources large language model at reduced cost

First reported

Deep Learning Weekly ran this on .

  • Z.ai released GLM-5.3-Flash, a model that processes text and images with a context window of 1 million tokens [amount of text it can hold at once].
  • The model costs one-tenth as much to run as the full GLM-5.3, making frontier-scale AI [cutting-edge, powerful] models cheaper to access.
  • GLM-5.3-Flash uses a mixture-of-experts architecture [activates different specialized sub-models for different tasks], with 320 billion total parameters but only uses 18 billion actively.

How it was covered

Deep Learning WeeklyEditorial team

Z.ai released GLM-5.3-Flash, a 320B/18B mixture-of-experts model with 1M-token multimodal context at one-tenth the cost of the full GLM-5.3. The newsletter emphasises the significant price reduction making frontier-scale models more accessible.