Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
Thinking Machines Lab has released Inkling-Small, an open weights multimodal model that utilizes a mixture-of-experts architecture with 276 billion total parameters and 12 billion active parameters. The model matches the performance of the larger Inkling while requiring significantly less hardware, as its NVFP4 checkpoint can run on a single NVIDIA B300 GPU.
Covered by 1 source
- MMarkTechPost↗Asif Razzaq3d ago