Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence (The Register)
The benchmark figure sits inside Nvidia’s broader claim that Groq 3 LPX is now in full production.
Techmeme’s entry says Nvidia reported 3,400-plus tokens per second for Groq 3 LPX racks on Gemma 4 31B with a 100,000-token input. The note also says Nebius is the first customer and SpaceX will deploy Nvidia’s Vera CPUs. A linked Bluesky post from The Register’s Tobias Mann cautions that Gemma 4 31B is an idealized case and that scaling efficiency to MoE models is the harder test. Techmeme's note
Techmeme’s entry says Nvidia reported 3,400-plus tokens per second for Groq 3 LPX racks on Gemma 4 31B with a 100,000-token input. The note also says Nebius is the first customer and SpaceX will deploy Nvidia’s Vera CPUs. A linked Bluesky post from The Register’s Tobias Mann cautions that Gemma 4 31B is an idealized case and that scaling efficiency to MoE models is the harder test. Techmeme's note
score 6