← Back to Model Beat
Models·Jul 9·all news from July 9, 2026

NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput

Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B, a smaller, compressed version of its previous model that reduces total parameters from 120.7 billion to 75.3 billion. By utilizing a technique called iterative puzzle to optimize hardware efficiency, the model achieves double the server throughput compared to its predecessor. This release highlights a broader industry shift toward making large language models more computationally efficient for enterprise deployment.

Covered by 1 source · 2 articles

Related stories

ModelsGPT-5.6 is now the preferred model in Microsoft 365 CopilotJul 7 · 24 sourcesModelsMistral AI Releases Robotics Model to Support Physical AI PushJul 8 · 43 sourcesModelsIntroducing GPT-LiveJul 8 · 13 sourcesModelsDeepseek is designing its own AI chipJul 5 · 62 sources