NVIDIA ships Groq 3 LPX into production
NVIDIA put Groq 3 LPX into full production on 24 August, an inference accelerator purpose-built for the long token streams that AI agents generate as they call tools, run code and iterate across thousands of steps (NVIDIA Newsroom; press release). The chip extends NVIDIA’s Vera Rubin rack-scale system and targets the bottleneck that turns multi-agent loops from snappy to sluggish: how fast the next token comes back to the model.
A specialist chip for the agent-loop bottleneck
That bottleneck matters more than it used to. Modern agents don’t answer in one pass — they read files, run scripts, inspect outputs and write again. Each round trip is a fresh inference, and latency compounds across long workflows. NVIDIA’s pitch is that making each inference faster is the cleanest way to make agents feel responsive, and that the rest of the stack — CPU, network, storage — needs to be built around the same goal.
In an Artificial Analysis benchmark, the platform hit what NVIDIA calls the fastest result yet for an open-weight agentic model on a long context window: roughly four times faster than the nearest alternative for latency-sensitive workloads (NVIDIA Blog). Coding tasks that ran for hours, the company says, finish in minutes.
Nebius, an Amsterdam-headquartered AI cloud with European data centres, is the first to deploy Groq 3 LPX into production. It slots into the company’s Token Factory inference platform and reaches customers through the same APIs they already use, so the change shows up as a latency drop rather than a migration project. Groq, the inference-only cloud that licensed its LPU silicon to NVIDIA, plans to be among the next adopters.
Jensen Huang, NVIDIA’s founder and CEO, framed the launch in workload-fit terms. Inference is the growth engine of AI.
Nebius CTO Danila Shtan put the customer-facing benefit more plainly: every step of an agent’s loop feels instant — through the same API developers are already using, with no migration to a new stack.
How the rack is wired together
The rack is wired together so the networking silicon, storage chips, orchestration CPUs and token-generation accelerator each do one job and do it fast. The main GPUs handle the heavy context work; the new chip focuses on decode — generating the next token as quickly as possible. A rack-scale deployment strings many accelerators together through direct chip-to-chip links, turning the rack into one deterministic inference engine. NVIDIA describes the design as “extreme codesign” across its full rack stack (NVIDIA Blog).
Networking is the other half of the bet. CoreWeave has put the multiplane switch fabric into production, splitting each server’s connection across multiple independent planes so the fabric scales to very large GPU counts without the third network tier traditional Ethernet fabrics eventually require. SpaceXAI is also committing to the platform, planning to build its next-generation architecture around the orchestration CPUs for scheduling, tool use and data processing.
3,400output tokens per second on a 100,000-token agentic workload — NVIDIA’s headline number for the launch, framed as roughly four times the nearest rival.
What it means for UK teams
This isn’t something a UK small team will buy or rent directly. Like Blackwell before it, Groq 3 LPX reaches readers only through the AI services they already pay for — coding assistants, agent platforms, hosted inference. Three shifts are worth watching:
- The cost curve for long agent runs. When inference gets roughly four times faster on the same hardware, the cloud bill for a 30-minute agent loop shrinks in proportion. That changes the maths on which workflows are worth automating at all — and it lands on the line items UK teams are already disputing with finance.
- UK and European cloud access. Nebius is the named first mover, and the closer the cloud sits, the more palatable data-residency conversations become for regulated buyers. Watch for additional European adopters rather than another US-only launch.
- Open-model agent performance. The headline benchmark ran on an open-weight model. If open-weight agents pick up most of the latency win, the closed-vs-open gap narrows for buyers who want to keep their options open — and the cost story in the first bullet improves further.
The deeper pattern is that NVIDIA’s agent-era pitch is no longer about raw GPU speed — it’s about removing the bottlenecks around the GPU. The new chip handles token generation, the networking silicon handles the fabric, the orchestration CPUs handle scheduling, the storage and security chips handle the rest. Agents, unlike training runs, punish latency at every layer, and NVIDIA is betting the only way to win is to optimise every layer together.
Sources & quotes
Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →
- NVIDIA Advances Vera Rubin Inference With New LPX and CPX Platforms (NVIDIA Blog, 24 Aug 2026)
- NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI (NVIDIA Newsroom, 24 Aug 2026)
- NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI (NVIDIA Investor Relations, 24 Aug 2026)


