Z.ai launches GLM-5.3-Flash under MIT license
Z.ai is pitching GLM-5.3-Flash as a cheaper multimodal GLM-5 model with a 1M-token context window and open MIT weights.
The MoE model has 320B total parameters and 18B active parameters, and Z.ai says it beats GLM-5.2 on reported coding and agent benchmarks at one-tenth the price. The company reports lower attention compute and KV cache use than GLM-5.3, with IndexPool used to manage long-context latency and memory. Visual reasoning is a core release point, covering rendered interfaces, games, 3D output, documents, spreadsheets, presentations and dashboards. The model was previously seen as ox-alpha on OpenCode and OpenRouter, and is now available through GLM Coding Plan, ZCode, Hugging Face and local deployment stacks including SGLang and vLLM. TestingCatalog's note
The MoE model has 320B total parameters and 18B active parameters, and Z.ai says it beats GLM-5.2 on reported coding and agent benchmarks at one-tenth the price. The company reports lower attention compute and KV cache use than GLM-5.3, with IndexPool used to manage long-context latency and memory. Visual reasoning is a core release point, covering rendered interfaces, games, 3D output, documents, spreadsheets, presentations and dashboards. The model was previously seen as ox-alpha on OpenCode and OpenRouter, and is now available through GLM Coding Plan, ZCode, Hugging Face and local deployment stacks including SGLang and vLLM. TestingCatalog's note
score 7