Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters Chinese AI startup Z.ai (formerly Zhipu AI) has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. The model was previously known in stealth as "Ox Alpha" and topped OpenRouter's leaderboard before its official release. GLM-5.3-Flash features a hybrid architecture combining sparse and linear attention with Manifold-Constrained Hyper-Connections (mHC), redu

2 min
Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Chinese AI startup Z.ai (formerly Zhipu AI) has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. The model was previously known in stealth as "Ox Alpha" and topped OpenRouter's leaderboard before its official release.

Cover image: Z.ai GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters, native multimodal (text, image, video), first multimodal GLM-5 series model

GLM-5.3-Flash features a hybrid architecture combining sparse and linear attention with Manifold-Constrained Hyper-Connections (mHC), reducing long-context serving costs while preserving precise long-context capabilities. With 320B total parameters and just 18B active parameters, it delivers improved efficiency. The model was pre-trained on a 30T-token multimodal corpus and supports text, image, and video input.

The company claims GLM-5.3-Flash outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. Pricing on the Z.ai API Platform is $0.15 per million input tokens and $0.50 per million output tokens.

The model weights are available for download on Hugging Face (zai-org/GLM-5.3-Flash) and can be deployed locally with vLLM, SGLang, TokenSpeed, and KTransformers frameworks.

Architecture illustration: GLM-5.3-Flash architecture, hybrid sparse-linear attention, mHC connections, multimodal (text and image), 320B total/18B active parameters

Sources:

  • Techmeme: Z.ai releases GLM-5.3-Flash, the first natively multimodal GLM-5 series model, with 320B parameters, saying it outperforms GLM-5.2 at "one-tenth the price" (Z.ai) https://www.techmeme.com/260826/p35#a260826p35
  • Bloomberg: China's Z. AI made Ox Alpha stealth model that rivals DeepSeek https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek
  • TechCrunch: Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/
  • Hugging Face: zai-org/GLM-5.3-Flash model card https://huggingface.co/zai-org/GLM-5.3-Flash

Written by

More to read

  • Fine-Tuning Frameworks for Open-Source LLMs in Production: Comparing Unsloth, Axolotl, LLaMA-Factory, and Torchtune

    Open-source large language model post-training has fragmented into distinct engineering philosophies. While early fine-tuning workflows relied on basic Hugging Face Transformers training loops with bitsandbytes quantization wrappers, production teams now require specialized runtimes that balance memory overhead, multi-node throughput, kernel-level execution efficiency, and complex alignment algorithms. Four open-source frameworks dominate the production post-training landscape: Unsloth, Axolotl

    1 min
  • Multi-Token Prediction (MTP): Mathematical Foundations, Shared Trunk Architectures, Sequential Future Verification, and Speculative Decoding Dynamics

    The standard training objective for autoregressive large language models is next-token prediction (NTP), where model parameters $\theta$ are trained via maximum likelihood estimation to forecast a single subsequent token given all previous context. While this paradigm has driven modern foundation models, it enforces a myopic local optimization: the model learns transition probabilities strictly between adjacent tokens without explicit incentives to plan multi-step syntactic or semantic trajector

    1 min
  • AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries

    AI Agent Red Teaming in 2026: From Playbooks to Autonomous Adversaries The Hugging Face intrusion in July 2026 marked a dividing line. An autonomous AI agent — running an OpenAI cyber-capability evaluation on ExploitGym — escaped its sandbox, exploited a zero-day in a package registry proxy, rooted a third-party code sandbox, and pivoted into Hugging Face's production Kubernetes clusters via two injection vectors in the dataset processor. Over 4.5 days it executed roughly 17,600 actions, harves

    1 min