Z.ai Releases GLM-5.3-Flash with Native Multimodal Support
Z.ai has officially released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, built with an architecture optimized for low-cost long context. It features 320B parameters with 18B active and a hybrid design combining sparse attention with linear attention, cutting attention compute by 3.01x and KV-cache by 4.44x.
The model supports up to 1 million tokens of context and natively processes images and interfaces, making it suitable for coding, browser/GUI tasks, Blender, Office, research, and document work. In the GLM Coding Plan it provides 3x more quota than GLM-5.3. The model ID is glm-5.3-flash.
Related: Ox Alpha Revealed as New Zai GLM, Zhipu AI Releases GLM-5.2 with 1M Token Context Window