Nvidia Ships Groq 3 LPX at 3,400 Tokens a Second
Announced at Hot Chips 2026, the dedicated inference accelerator born from Nvidia’s $20B Groq deal is now in full production, hitting 3,400 output tokens per second on Gemma 4 31B. Nebius is the first AI cloud to deploy it.
Nvidia announced at Hot Chips 2026 on August 24 that Groq 3 LPX — its first accelerator purpose-built for inference rather than training — has entered full production as an extension of the Vera Rubin data-center platform. The pitch is latency rather than raw throughput: the target is agentic workloads, where a model's answer is only one step in a long chain and every extra second compounds.
In benchmarking by Artificial Analysis, Groq 3 LPX produced 3,400 output tokens per second running the open-weight Gemma 4 31B model with a 100,000-token context window. Nvidia claims the system is four times more responsive than the nearest alternative platform on latency-sensitive work, and says multistep agentic coding tasks that once ran for hours can finish in minutes.
The design leans on rack-scale integration rather than a single standout chip. Nvidia describes “extreme codesign across seven chips and five purpose-built racks,” pairing the accelerators with BlueField-4 DPUs, Vera CPU racks, BlueField-4 STX storage and Spectrum-6 SPX Ethernet. A full rack-scale deployment links up to 256 accelerators over the company's high-bandwidth interconnects so that GPUs and LPUs present to software as a single inference engine.
The technology arrived by acquisition. Nvidia licensed Groq's inference stack for $20 billion in December 2025 and hired founder Jonathan Ross and president Sunny Madra as part of the deal — a defensive move against a startup whose entire premise was that Nvidia's training-optimised GPUs were the wrong shape for serving tokens.
Nebius is the first AI cloud to adopt the part, folding it into its Nebius Token Factory inference platform, with racks going live later this year alongside Vera CPUs and Rubin GPUs. Groq itself, still operating as an inference provider, plans early adoption. Chief executive Jensen Huang said the technology “transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”
Inference-specific silicon has been a crowded pitch for years, with Cerebras and others selling speed itself as the product. What changes here is who is making it: the company that already owns the training market is now defending the serving market too, and it paid $20 billion for the sharpest argument against itself in order to do it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.