ComputeLabs Research

NVIDIA’s Groq 3 LPX inference accelerator entered full production, reaching 3,400 output tokens per second.

· ComputeLabs Research · from the August 24, 2026 edition

NVIDIA announced that Groq 3 LPX, an interactive artificial-intelligence inference accelerator, had entered full production. A separate update reported performance of 3,400 output tokens per second.

Groq 3 LPX is described as an extension of the NVIDIA Vera Rubin platform and is intended to provide fast token generation for highly responsive agentic systems. NVIDIA also positions it as the interactive inference component of the Vera Rubin platform.

The product is an inference accelerator, not a CPU, server, or standalone data-center capacity figure. The 3,400-token-per-second figure describes output-token generation performance; the supplied source excerpt does not specify the model, precision, batch size, context length, or complete benchmarking methodology.

All 22 stories from August 24, 2026