← Back to Model Beat
Hardware·Aug 25·all news from August 25, 2026

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated

Nvidia has moved its Groq 3 LPX inference chip into full production, reporting speeds of 3,400 tokens per second when running Gemma 4 31B. While this performance is four times faster than the competing hardware from Cerebras, the comparison is complex because Nvidia requires a cluster of 64 accelerators to reach these speeds. In contrast, Cerebras achieves its output using only one or two units, highlighting different architectural approaches to scaling inference performance.

Covered by 1 source

Related stories

HardwarePreviewing the Model Hardware StandardAug 27 · 10 sourcesHardwareJalapeño’s first results show industry-leading speed and efficiency in AI inferenceAug 25 · 9 sourcesHardwareNvidia Must Prove It Can Be Tomorrow's AI PlatformAug 25 · 29 sourcesHardwareNvidia-Backed Lambda Inks $1 Billion Private Debt for Chip DealAug 27 · 5 sources