Nvidia is shipping its Groq 3 LPX inference chip at 3,400 tokens per second, claiming four times Cerebras' speed. The benchmark requires at least 64 accelerators versus Cerebras' one or two, so real-world comparison depends on scale and MoE efficiency. AI developers should weigh cost and architecture fit before choosing inference hardware.
Opening Kapyn…