FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution
Researchers from UC Berkeley and MIT have introduced FreeToken, an open-source inference engine designed to optimize Mixture-of-Experts models for use on consumer-grade hardware. By utilizing dynamic scheduling and improved weight management, the software increases decoding speeds and overall performance for these complex models. This development potentially lowers the technical barriers for running large AI architectures on standard personal computing devices.
Covered by 1 source
- IInfoQ AI↗Olimpiu PopAug 29