Alibaba Releases Open-Weight Qwen3.8-Flash-Next with 6B Active Parameters and Qwen 4 Architecture Preview
Alibaba has released the open weights for Qwen3.8-Flash-Next, a 125B-parameter multimodal Mixture-of-Experts (MoE) model that acts as an early architectural preview for the upcoming Qwen4 family. Operating with only 6B active parameters per token alongside a 51B N-gram embedding layer, the model targets cost efficiency across long-context reasoning, agentic coding, and multimodal workloads. Weights are publicly available on Hugging Face and ModelScope, while a managed production endpoint named








