← Back to Model Beat
Models·Aug 26·all news from August 26, 2026

Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

Alibaba’s Qwen team has released Qwen3.8-Flash-Next, a 125B multimodal Mixture-of-Experts model that utilizes only 6B active parameters. This release serves as an early technical preview of the upcoming Qwen4 architecture, which integrates new components like GatedResidual and Qwen Sparse Attention. By separating the model into a 125B backbone, a large N-gram embedding table, and a multi-token prediction module, the developers aim to improve inference efficiency. This provides developers with an initial look at the architectural shifts expected in the next generation of Qwen models.

ModelsQwen3.8 Flash

Covered by 2 sources · 4 articles

Related stories

ModelsQwen 3.8 Flash Reduces Costs to One-Third of DeepSeek-V4-FlashAug 29ModelsExperiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding | NVIDIA Technical BlogAug 26