Skip to main content
Live
Main content

Nvidia ships Groq 3 LPX to accelerate Vera Rubin agentic inference

The new accelerator hits 3,400 output tokens per second on 100,000-token contexts, 4x the nearest platform, with Nebius as first cloud adopter.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Nvidia said its Groq 3 LPX accelerator is in full production and extends the Vera Rubin NVL72 platform with a decode engine built specifically for agentic inference. On an Artificial Analysis benchmark running the open Gemma 4 31B model, the setup delivered 3,400 output tokens per second at a 100,000-token context window, which the company said is 4x the nearest competing platform. Nebius is the first AI cloud to adopt Groq 3 LPX, and SpaceXAI committed to building its next-generation architecture around Vera CPUs.

The pitch is that inference is now the bottleneck, not training. Agentic systems generate tokens one at a time across long chains of tool calls, so decode latency compounds. A rack-scale Groq 3 LPX deployment fits 256 LP30 accelerators connected through direct chip-to-chip links, working alongside Rubin GPUs that handle the heavier context-processing side of the workload.

The split-role design is the point. Rubin GPUs churn through the large-context prefill; LPX handles the latency-sensitive decode. Nvidia's argument is that combining the two eliminates the traditional throughput-versus-speed tradeoff that has forced operators to choose between fast responses and cheap tokens. CoreWeave has deployed Spectrum-X Multiplane in production to knit Vera Rubin racks together at scale, and Nebius is folding the LPX into its existing Token Factory service.

Key facts

  • 01Nvidia Groq 3 LPX enters full production, hitting 3,400 output tokens per second on Gemma 4 31B at 100,000-token context — 4x the nearest alternative.
  • 02A rack-scale Groq 3 LPX deployment packs 256 LP30 accelerators linked through direct chip-to-chip interconnects.
  • 03Spectrum-X Multiplane scales to 512,000 GPUs on a two-tier Ethernet fabric, with 1.6x better networking performance than off-the-shelf Ethernet.
  • 04In an eight-plane topology, a single plane failure retains 90% of bandwidth with hardware recovery 11x faster than software load balancing.
  • 05Nebius is the first AI cloud to adopt Groq 3 LPX; SpaceXAI committed to Vera CPUs; CoreWeave has deployed Spectrum-X Multiplane in production.

On the networking side, Spectrum-X Multiplane is the piece that lets AI factories grow past today's largest clusters without adding a third network tier. Each server's connection is split into parallel planes, each running a lightweight two-tier fabric. Nvidia says the design scales to 512,000 GPUs with 1.6x better AI networking performance than off-the-shelf Ethernet.

Agentic AI is creating a new performance challenge: decode latency.
NVIDIA, company statement

The reliability numbers are the interesting part. In an eight-plane topology, if one plane fails, the network holds roughly 90% of its total bandwidth, and the ConnectX-9 SuperNIC handles rerouting in hardware — 11x faster than software-based multiplane load balancing. Nvidia translates that into 1.6x higher AI factory output on identical hardware footprints.

The underlying silicon is the Spectrum-X SN6000 series switch built on a 102.4Tb/s Spectrum-6 Ethernet ASIC, paired with ConnectX-9 SuperNICs pushing up to 1,600Gb/s per port. This is the same broader Nvidia-networking strategy that showed up earlier this year when the company opened NVLink Fusion to let custom XPUs plug into its rack-scale systems — a shift from selling GPUs to selling the entire factory floor.

SpaceXAI's adoption is the boldest partnership on the slate. The company plans to deploy Vera CPUs to handle the CPU-heavy work of agentic AI — orchestration, tool use, code execution, data processing and simulation — across data centers on Earth and, per the release, orbital satellites. That last piece echoes recent moves like Starcloud's $250M raise for orbital data centers; the space-based-inference thesis is quietly picking up commercial commitments.

The economic story Nvidia is telling is straightforward: agentic AI generates more tokens per user session than chat, context windows are getting longer, and the marginal token has to be cheap and fast simultaneously. Nvidia's phrase for this is the "token factory," and the codesigned Rubin-plus-LPX architecture is meant to be the reference implementation.

Related · from this week
Nvidia pitches performance per watt as the AI factory's decisive metric
Jaeden Schafer · 5 min read →

The obvious question is whether the 4x claim holds outside Gemma 4 31B and the specific 100,000-token benchmark condition. Model-specific optimizations are common in vendor benchmarks, and the closest genuine competitor here is the actual Groq LPU line, which has built its business on deterministic low-latency decode. Nvidia has not published head-to-head numbers against Groq's own hardware, only against an unnamed "nearest alternative platform." Real deployment data from Nebius and CoreWeave will be more informative than launch-day figures.

Nvidia's ability to sell entire factory architectures — compute, networking, and now decode-specialized inference silicon — is what makes it hard to dislodge, and Groq 3 LPX is the piece that closes the loop on agentic workloads specifically. If Nebius's Token Factory numbers look anything like the benchmark, expect every serious inference cloud to be forced into the same rack shape within a year. That would consolidate more of the AI infrastructure margin at Nvidia and squeeze the standalone-accelerator startups whose entire pitch was faster decode.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia pitches performance per watt as the AI factory's decisive metric

GB300 NVL72 delivers up to 25x performance per watt over Hopper as power becomes the binding constraint on inference economics.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Spectrum-X's MRC protocol via OCP after OpenAI, Microsoft, Oracle deployments

The Multipath Reliable Connection transport, proven on Blackwell training runs, is now an open spec through the Open Compute Project.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia opens Spectrum-X's MRC protocol to the industry after OpenAI, Microsoft, Oracle deployments

The Multipath Reliable Connection protocol, co-developed with OpenAI and Microsoft, reroutes failed network paths in microseconds across gigascale GPU fabrics.

Jaeden Schafer5 min read