← Back to Model Beat
Hardware·Aug 23·all news from August 23, 2026

Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU

FreeToken is a new serving engine designed to run large Mixture-of-Experts models locally by balancing data between a workstation GPU and CPU memory. By optimizing how cache misses are handled, the engine enables the execution of the 753B GLM-5.2 model on consumer-grade hardware.

Covered by 1 source

Related stories

HardwareNvidia to Pay AI Startup Poolside a $6 Billion License, Newcomer SaysAug 20 · 22 sourcesHardwareJalapeño’s first results show industry-leading speed and efficiency in AI inferenceAug 25 · 9 sourcesHardwareNvidia Must Prove It Can Be Tomorrow's AI PlatformAug 25 · 29 sourcesHardwareAnthropic to Pay Nscale $45 Billion for AI Computing PowerAug 26 · 4 sources