Model•active
GLM-5.3-Flash
GLM-5.3-Flash is Z.ai's 320B-parameter open-weight efficiency model with 18B active parameters, a 1,048,576-token context window, and a hybrid sparse/linear-attention multimodal architecture.
Official siteVerified Oct 5, 2026
Overview
Flash starts from a newly trained base model and combines sparse attention, linear attention, mixture-of-experts routing, and multimodal pre-training. The official processor configuration includes image and video processing.
Model specifications
- Family
- GLM-5.3
- Context
- 1,048,576 tokens
- Input
- Text, Image, Video
- Output
- Text
- Open weights
- Yes
- Capabilities
- Reasoning, Vision