Motif 3: Technical Report
Researchers have introduced Motif 3, a decoder-only Mixture-of-Experts language model featuring 314 billion total parameters. The architecture utilizes fine-grained sparsity by selecting eight experts from a pool of 384 for every token processed, resulting in 13.2 billion active parameters per inference step.
ModelsMotif-3
Covered by 1 source
- AarXiv CS.AI↗Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki RyuAug 11