GLM-5.3-Flash
GLM-5.3-Flash is an open-weight language model from Z.ai (Zhipu AI), released on August 26, 2026. It is a mixture-of-experts model with 320B total parameters, of which 18B are active for each token. The model accepts text and image input.
- Developer
- Z.ai (Zhipu AI)
- Released
- August 26, 2026
- Total parameters
- 320B
- Active parameters per token
- 18B
- Architecture
- MoE
- Input
- Text, image
- License
- MIT
- Commercial use
- Yes
- Official weights, GB
- 328.3 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under the MIT license, which allows commercial use.
The official weights on Hugging Face take about 328 GB in FP8. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 320B
- Active parameters, billions
- 18B
- License type
- MIT
- License conditions
- MIT: keep the copyright and permission notice.
- Weights precision
- FP8
- Official quantized or alternative versions
- zai-org/GLM-5.3-Flash-BF16 (full precision)
- License text
- huggingface.co
- Sources
Page URL Hugging Face model card https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/README.md Hugging Face weights index https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/model.safetensors.index.json GitHub file https://raw.githubusercontent.com/zai-org/GLM-5/main/README.md z.ai announcement https://z.ai/blog/glm-5.3-flash