AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
AMD has released Instella-MoE-16B-A3B, an open-source mixture-of-experts language model trained entirely on the company's Instinct MI300X and MI325X hardware. By providing weights from every training phase, AMD offers developers transparent insight into the model's development process and its use of Gated MLA and FarSkip-Collective architectures. This release serves as a functional demonstration of the firm's data center hardware capabilities for training efficient models with 2.8 billion active parameters.
Covered by 1 source
- MMarkTechPost↗Asif RazzaqAug 1