Nvidia said its Groq 3 LPX accelerator is in full production and extends the Vera Rubin NVL72 platform with a decode engine built specifically for agentic inference. On an Artificial Analysis benchmark running the open Gemma 4 31B model, the setup delivered 3,400 output tokens per second at a 100,000-token context window, which the company said is 4x the nearest competing platform. Nebius is the first AI cloud to adopt Groq 3 LPX, and SpaceXAI committed to building its next-generation architecture around Vera CPUs.
The pitch is that inference is now the bottleneck, not training. Agentic systems generate tokens one at a time across long chains of tool calls, so decode latency compounds. A rack-scale Groq 3 LPX deployment fits 256 LP30 accelerators connected through direct chip-to-chip links, working alongside Rubin GPUs that handle the heavier context-processing side of the workload.
The split-role design is the point. Rubin GPUs churn through the large-context prefill; LPX handles the latency-sensitive decode. Nvidia's argument is that combining the two eliminates the traditional throughput-versus-speed tradeoff that has forced operators to choose between fast responses and cheap tokens. CoreWeave has deployed Spectrum-X Multiplane in production to knit Vera Rubin racks together at scale, and Nebius is folding the LPX into its existing Token Factory service.
Key facts
- 01Nvidia Groq 3 LPX enters full production, hitting 3,400 output tokens per second on Gemma 4 31B at 100,000-token context — 4x the nearest alternative.
- 02A rack-scale Groq 3 LPX deployment packs 256 LP30 accelerators linked through direct chip-to-chip interconnects.
- 03Spectrum-X Multiplane scales to 512,000 GPUs on a two-tier Ethernet fabric, with 1.6x better networking performance than off-the-shelf Ethernet.
- 04In an eight-plane topology, a single plane failure retains 90% of bandwidth with hardware recovery 11x faster than software load balancing.
- 05Nebius is the first AI cloud to adopt Groq 3 LPX; SpaceXAI committed to Vera CPUs; CoreWeave has deployed Spectrum-X Multiplane in production.
On the networking side, Spectrum-X Multiplane is the piece that lets AI factories grow past today's largest clusters without adding a third network tier. Each server's connection is split into parallel planes, each running a lightweight two-tier fabric. Nvidia says the design scales to 512,000 GPUs with 1.6x better AI networking performance than off-the-shelf Ethernet.
“Agentic AI is creating a new performance challenge: decode latency.”— NVIDIA, company statement
The reliability numbers are the interesting part. In an eight-plane topology, if one plane fails, the network holds roughly 90% of its total bandwidth, and the ConnectX-9 SuperNIC handles rerouting in hardware — 11x faster than software-based multiplane load balancing. Nvidia translates that into 1.6x higher AI factory output on identical hardware footprints.
The underlying silicon is the Spectrum-X SN6000 series switch built on a 102.4Tb/s Spectrum-6 Ethernet ASIC, paired with ConnectX-9 SuperNICs pushing up to 1,600Gb/s per port. This is the same broader Nvidia-networking strategy that showed up earlier this year when the company opened NVLink Fusion to let custom XPUs plug into its rack-scale systems — a shift from selling GPUs to selling the entire factory floor.
SpaceXAI's adoption is the boldest partnership on the slate. The company plans to deploy Vera CPUs to handle the CPU-heavy work of agentic AI — orchestration, tool use, code execution, data processing and simulation — across data centers on Earth and, per the release, orbital satellites. That last piece echoes recent moves like Starcloud's $250M raise for orbital data centers; the space-based-inference thesis is quietly picking up commercial commitments.
The economic story Nvidia is telling is straightforward: agentic AI generates more tokens per user session than chat, context windows are getting longer, and the marginal token has to be cheap and fast simultaneously. Nvidia's phrase for this is the "token factory," and the codesigned Rubin-plus-LPX architecture is meant to be the reference implementation.
The obvious question is whether the 4x claim holds outside Gemma 4 31B and the specific 100,000-token benchmark condition. Model-specific optimizations are common in vendor benchmarks, and the closest genuine competitor here is the actual Groq LPU line, which has built its business on deterministic low-latency decode. Nvidia has not published head-to-head numbers against Groq's own hardware, only against an unnamed "nearest alternative platform." Real deployment data from Nebius and CoreWeave will be more informative than launch-day figures.
Nvidia's ability to sell entire factory architectures — compute, networking, and now decode-specialized inference silicon — is what makes it hard to dislodge, and Groq 3 LPX is the piece that closes the loop on agentic workloads specifically. If Nebius's Token Factory numbers look anything like the benchmark, expect every serious inference cloud to be forced into the same rack shape within a year. That would consolidate more of the AI infrastructure margin at Nvidia and squeeze the standalone-accelerator startups whose entire pitch was faster decode.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




