← Back to Model Beat
Models·Jun 17·all news from June 17, 2026

MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget

AI startup MiniMax has introduced a sparse attention mechanism called MSA that improves processing efficiency for large language models. By using an index branch to select only the most relevant data blocks, the system reduces computational requirements during attention operations by more than 28 times compared to standard methods. This approach allows models to handle million-token context windows with significantly lower resource consumption while maintaining performance levels seen in existing state-of-the-art architectures.

Covered by 1 source

Related stories

ModelsFrom Chatbots to Collaborators: AI’s Next EraJun 15 · 39 sourcesModelsPredicting model behavior before release by simulating deploymentJun 16 · 3 sourcesModelsZ.ai Launches GLM-5.2 With a Usable 1M-Token Context, Two Thinking-Effort Levels, and No Benchmarks at LaunchJun 14 · 10 sourcesModelsGoogle’s Gemini-Powered AI Home Speaker Goes on Sale June 25Jun 17 · 5 sources