NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput
Nvidia has released Nemotron-Labs-3-Puzzle-75B-A9B, a smaller, compressed version of its previous model that reduces total parameters from 120.7 billion to 75.3 billion. By utilizing a technique called iterative puzzle to optimize hardware efficiency, the model achieves double the server throughput compared to its predecessor. This release highlights a broader industry shift toward making large language models more computationally efficient for enterprise deployment.
Covered by 1 source · 2 articles
- MMarkTechPost↗Asif RazzaqJul 9
- MMarkTechPost↗Asif RazzaqJul 9