MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget
AI startup MiniMax has introduced a sparse attention mechanism called MSA that improves processing efficiency for large language models. By using an index branch to select only the most relevant data blocks, the system reduces computational requirements during attention operations by more than 28 times compared to standard methods. This approach allows models to handle million-token context windows with significantly lower resource consumption while maintaining performance levels seen in existing state-of-the-art architectures.
Covered by 1 source
- MMarkTechPost↗Asif RazzaqJun 17