0
Applied AI·August 24, 2026·1 min read

Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence

Share

3,400 tokens/sec on a 100,000-token Gemma 4 31B prompt is Nvidia signaling that long-context, low-latency inference is moving into production territory. If your workloads are bottlenecked on context length or response time, it’s time to re-run your infra and vendor benchmarks with LPU-class options in the mix.