AI

Nvidia Puts Groq 3 LPX Rack Into Full Production, Hits 3,400 Tokens Per Second

Nvidia's $20 billion Groq acquisition became a shipping product on August 24, with the LPX rack now in volume production at neocloud Nebius and delivering 4x the inference throughput of the nearest competitor, Nvidia executives said.

M
By Marcus Chen AI Correspondent
August 25, 2026 / 6 min read

Nvidia on August 24 moved its Groq 3 LPX inference rack into full production, marking the commercialization of technology from the company's largest-ever acquisition. Senior director Dion Harris told reporters that the LPX rack, which combines 256 Groq 3 chips per chassis, has begun shipping to neocloud provider Nebius and will be online at Nebius data centers later this year, according to CNBC.

Performance Numbers

In an Artificial Analysis benchmark disclosed Monday, an LPX rack running Google's Gemma 4 31B open model delivered 3,400 output tokens per second on a 100,000-token input sequence, which Nvidia and analysts at The Register said is 4x faster than the nearest competing platform. The speed advantage is targeted at agentic AI workloads, where each token latency compounds across multi-step reasoning chains. Nvidia frames the LPX rack as an extension of the Vera Rubin NVL72 platform: Vera CPUs and Rubin GPUs handle prefill and bulk inference, while the LPX takes the latency-sensitive decode phase.

The $20 Billion Bet

Nvidia acquired Groq's assets in December for $20 billion, the largest purchase in Nvidia's history. Groq's LPU architecture relies on 500 MB of on-die SRAM per chip, which delivers memory bandwidth orders of magnitude higher than HBM4 stacks. The trade-off is total capacity: each LPU holds 576x less memory than a top-spec Rubin GPU, so production deployments must spread models across many chips. A single LPX rack uses 256 LPUs for 128 GB of high-bandwidth SRAM, and multiple racks can be ganged together for larger models.

The Competitive Picture

Nvidia's comparison targets Cerebras, which publicly serves Gemma 4 31B at 882 tokens per second on its CS-3 systems. AMD announced at its Advancing AI event last month that it will integrate Helios rack-scale GPU systems with Cerebras WSE-3T accelerators for a heterogeneous inference architecture similar to Nvidia's NVL72-plus-LPX configuration. OpenAI's new Ultrafast mode, powered by Cerebras, currently promises 750 tokens per second. Nvidia's 4x claim was independently corroborated by Artificial Analysis on the Gemma 4 31B workload.

What to Watch Through Year-End

Three checkpoints follow. Nebius Token Factory is expected to begin serving LPX-accelerated inference to developers in Q4 2026, the first commercial deployment of Groq silicon under Nvidia ownership. Nvidia's earnings call on Wednesday will provide the first disclosure of Groq-related revenue and gross-margin contribution. And SpaceXAI's announcement that NVIDIA Vera CPUs will power its next-generation agentic AI stack is the first major customer win disclosed for the broader Vera Rubin platform, signaling that the company's bet on heterogeneous inference is finding design wins beyond the LPX rack itself.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.