Nvidia entered full production on Groq 3 LPX on Aug. 24, 2026, announcing the milestone at Hot Chips in Palo Alto. The accelerator is an extension of the Vera Rubin NVL72 platform and targets agentic AI workloads—systems that chain hundreds or thousands of inference steps together to complete complex tasks in real time.
In Artificial Analysis benchmarking, Groq 3 LPX delivered 3,400 output tokens per second running Gemma 4 31B with a 100,000-token context window. Nvidia describes that result as the fastest ever recorded for the model. Token generation speed is the binding constraint in agentic workloads: each reasoning step waits on the prior one, so latency compounds across long task chains in ways that do not affect single-turn inference.
Nvidia reports up to four times faster responsiveness versus the nearest alternative platform for agentic and other latency-sensitive workloads. The company did not name the competing platform in its announcement.
Nebius will be the first AI cloud to deploy Groq 3 LPX commercially, running it through its Nebius Token Factory production inference platform. Groq—the AI cloud provider, separate from the chip design term—plans to be among the earliest additional adopters. These are planned deployments, not completed ones.
The hardware integrates into Nvidia's Vera Rubin AI factory architecture alongside BlueField-4 data processing units, Vera CPU racks, storage systems and Spectrum-6 Ethernet. Groq 3 LPX is not a standalone card a customer drops into an existing server. It only makes economic sense inside the full Vera Rubin NVL72 factory build, which keeps the upgrade cycle captive to Nvidia's infrastructure ecosystem.
Agentic AI applications generate token volumes that dwarf traditional chatbot deployments. A single autonomous agent completing a multi-step research or coding task can consume millions of tokens in one session. At that scale, the cost per token and wall-clock time per task become the primary operating variables for cloud providers selling inference capacity. A fourfold responsiveness gain at 3,400 tokens per second per rack is a pricing and margin argument as much as a performance one.
Nebius is a logical first mover. The company is building inference-as-a-service capacity and brands Token Factory specifically around high-throughput output. Deploying Groq 3 LPX first gives Nebius a window to market the fastest available inference for open-source agentic models before the same hardware becomes available to competitors.
The Gemma 4 31B model used in benchmarking is an open-source agentic model. That choice matters: open-source models are the primary workload on independent AI cloud platforms like Nebius and Groq rather than on hyperscaler clouds, which tend to run proprietary models. The benchmark result is directly relevant to the customers Nvidia is selling Groq 3 LPX to.
Nvidia shares closed at $210.25 on Aug. 24, down 2.1 percent on the day. The stock's movement does not reflect a market that reads product announcements as near-term revenue events. Nvidia's own historical data shows AI-tagged announcements have diverged from positive sentiment in four of five matched cases, with an average next-day move of negative 1.15 percent across that set. The Groq 3 LPX announcement fits that pattern: full production on a new accelerator signals a product cycle opening, not a quarter closing.
Competitive context includes inference-chip competition from custom silicon at hyperscalers—Google's TPUs, Amazon's Trainium and Inferentia lines—and from startups positioning purpose-built inference hardware as a cheaper alternative to Nvidia's stack. Nvidia's answer with Groq 3 LPX is not price but raw throughput: if agentic workloads require tens of millions of tokens per session and the alternative platforms cannot match 3,400 tokens per second at 100,000-token context, the economics of running those workloads on Nvidia hardware close the gap regardless of list price differences.
Commercial validation comes next. Nebius deploying Token Factory on Groq 3 LPX will produce real-world throughput data outside Nvidia's controlled benchmark environment. If the 3,400 tokens-per-second figure holds at scale under production load diversity, it becomes a selling point Nvidia can use across the rest of the Vera Rubin NVL72 customer base. If it does not, the benchmark stands as a ceiling rather than a floor.
