Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
Alibaba’s Qwen team has released Qwen3.8-Flash-Next, a 125B multimodal Mixture-of-Experts model that utilizes only 6B active parameters. This release serves as an early technical preview of the upcoming Qwen4 architecture, which integrates new components like GatedResidual and Qwen Sparse Attention. By separating the model into a 125B backbone, a large N-gram embedding table, and a multi-token prediction module, the developers aim to improve inference efficiency. This provides developers with an initial look at the architectural shifts expected in the next generation of Qwen models.
ModelsQwen3.8 Flash
Covered by 2 sources · 4 articles
- MMarkTechPost↗Asif RazzaqAug 26
- MMarkTechPost↗Asif RazzaqAug 26
- GGitHub Releases↗Aug 26
- GGitHub Releases↗Aug 26