ComputeLabs Research
NVIDIA’s Groq 3 LPX inference accelerator entered full production, reaching 3,400 output tokens per second.
· ComputeLabs Research · from the August 24, 2026 edition
NVIDIA announced that Groq 3 LPX, an interactive artificial-intelligence inference accelerator, had entered full production. A separate update reported performance of 3,400 output tokens per second.
Groq 3 LPX is described as an extension of the NVIDIA Vera Rubin platform and is intended to provide fast token generation for highly responsive agentic systems. NVIDIA also positions it as the interactive inference component of the Vera Rubin platform.
The product is an inference accelerator, not a CPU, server, or standalone data-center capacity figure. The 3,400-token-per-second figure describes output-token generation performance; the supplied source excerpt does not specify the model, precision, batch size, context length, or complete benchmarking methodology.
Sources
- HPCwire: NVIDIA Groq 3 LPX Enters Full Production for Agentic AI Inference(hpcwire.com)
- NVIDIA Blog: With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents(blogs.nvidia.com)
- NVIDIA Developer Blog: How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin(developer.nvidia.com)

