Skip to main content
Live
Main content

Nvidia Blackwell sweeps MLPerf Training 6.0 across all seven benchmarks

GB300 NVL72 delivers 1.6x faster training than GB200 NVL72, and CoreWeave hits a 2.02-minute time-to-train on DeepSeek-V3 671B.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Nvidia swept the MLPerf Training 6.0 benchmark suite, posting the fastest time-to-train on all 7 workloads and the largest-scale submission on record at 8,192 GPUs. It was the only platform with results across every benchmark in the round, and its new GB300 NVL72 rack-scale system trained up to 1.6x faster than the GB200 NVL72 at matched scale. The result reinforces Nvidia's grip on frontier training just as mixture-of-experts architectures move to the center of the field.

MLPerf Training 6.0 added two new MoE pretraining workloads — DeepSeek-V3 671B and GPT-OSS-20B — alongside Llama 3.1 405B and the rest of the suite. Microsoft Azure trained Llama 3.1 405B to the reference quality target in 7.07 minutes across 8,192 GB200 NVL72 GPUs, the fastest submission for that benchmark. CoreWeave took the DeepSeek-V3 671B record, reaching target quality in 2.02 minutes on 8,192 GB300 NVL72 GPUs connected over Nvidia's Spectrum-X Ethernet fabric.

The 1.6x generational gain from GB200 to GB300 comes from three places. Blackwell Ultra adds higher compute density via NVFP4, a low-precision training format Nvidia has been pushing across model sizes. It also carries more memory and a higher sustained power ceiling, letting the GPU hold peak throughput longer on long-running pretraining jobs.

Key facts

  • 01Nvidia was the only vendor to submit results on all 7 MLPerf Training 6.0 benchmarks and posted the fastest time-to-train on each.
  • 02GB300 NVL72 trained up to 1.6x faster than GB200 NVL72 at the same scale, driven by NVFP4 compute and a higher power ceiling.
  • 03Microsoft Azure trained Llama 3.1 405B to target quality in 7.07 minutes across 8,192 GB200 NVL72 GPUs.
  • 04CoreWeave hit a 2.02-minute time-to-train on DeepSeek-V3 671B using 8,192 GB300 NVL72 GPUs over Spectrum-X Ethernet.
  • 0519 ecosystem partners submitted results, including Azure, CoreWeave, Google Cloud, Dell, HPE, Cisco and Lambda.

Within each rack-scale system, fifth-generation Nvidia NVLink Switches tie 72 GPUs into a single high-bandwidth pool of compute and memory. That topology matters disproportionately for MoE training, where tokens have to be routed across GPUs to reach the right expert subnetwork. The same all-to-all communication pattern that makes MoE inference hard makes MoE training brutal at scale, and NVLink bandwidth is what keeps it from collapsing.

Nvidia also used NVFP4 training methods to push performance on both small- and large-scale pretraining as well as fine-tuning, while clearing MLPerf's accuracy bar. The company recently used NVFP4 to pretrain its own 550-billion-parameter Nemotron 3 Ultra model, a real-world stress test for the format outside the benchmark suite.

Scale-out networking ran on two complementary platforms — Nvidia Quantum InfiniBand and Spectrum-X Ethernet — giving operators a choice depending on their data center build. On DeepSeek-V3 671B, the largest MoE model in the suite, Nvidia scaled to 8,192 GPUs with GB200 NVL72 systems, the largest Blackwell-based MLPerf Training submission to date. The Llama 3.1 405B run scaled to 5,120 GPUs on GB200 NVL72 hardware before Azure pushed it to 8,192.

Reliability got a separate pitch. Nvidia screens each GPU across 30-plus manufacturing test stages before deployment, then leans on its Reliability, Availability and Serviceability Engine and self-healing routing to avoid mid-job failures. When something does break, the Nvidia Resiliency Extension — NVRx — detects underperforming nodes, restarts from a recent checkpoint instead of the full job, and at the network layer Spectrum-X reroutes around failed links in milliseconds. On training runs that span weeks across hundreds of thousands of GPUs, that resiliency stack is what separates a finished model from a stalled one.

19 ecosystem partners submitted results this round, including Microsoft Azure, CoreWeave, Google Cloud, Dell Technologies, Hewlett Packard Enterprise, Cisco, Lambda, Nebius, Fujitsu, ASUSTeK, Giga Computing, Inventec, Quanta Cloud Computing, Netweb Technologies and Krai. Several of them carried customer-facing results alongside the benchmark numbers. Cohere trained its North agentic AI platform 3x faster on GB200 NVL72. Thinking Machines Lab reported 2x faster training and serving on GB300 NVL72 versus the prior generation on Google Cloud. Nebius helped Higgsfield cut model training time by 30%, supporting a platform that now serves 22 million users and generates more than 6 million pieces of AI content per day.

Related · from this week
Nvidia's Vera Rubin NVL72 hits production with 10x tokens per megawatt
Jaeden Schafer · 5 min read →

Two caveats deserve weight. MLPerf submissions are tuned, expert-run configurations — real-world training jobs rarely match benchmark numbers, and the 2.02-minute and 7.07-minute results are time-to-target on fixed reference setups, not end-to-end model development cycles. And the benchmark didn't include AMD's MI355X or any non-Nvidia rack-scale submissions at the 8,192-GPU tier, so the sweep is partly a measure of who chose to show up. That said, when the headline frontier-lab buyers — Azure, CoreWeave, Google Cloud — all submit on Blackwell, the competitive signal is hard to ignore.

Nvidia's moat in training is no longer just the GPU; it's the rack, the switch fabric, the precision format, the resiliency stack and the partner network all engineered to a single benchmark target. For AI labs deciding where to spend the next several billion dollars on training capacity, MLPerf Training 6.0 makes the calculus uncomfortable for anyone outside the Blackwell ecosystem. The MoE results in particular matter — every frontier model coming out of the major labs is moving toward sparse, expert-routed architectures, and Nvidia just demonstrated it can train one to target quality in two minutes flat.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
News

Nvidia's Vera Rubin NVL72 hits production with 10x tokens per megawatt

CoreWeave's DeepSeek-R1 benchmark shows a 10x throughput-per-watt gain over Grace Blackwell, with racks now live at four major clouds.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia pitches performance per watt as the AI factory's decisive metric

GB300 NVL72 delivers up to 25x performance per watt over Hopper as power becomes the binding constraint on inference economics.

Jaeden Schafer5 min read
Nvidia logo
Models

NVIDIA's inference software stack cuts DeepSeek V4 token costs 5x in one month

Baseten, Cognition, Together AI and Cursor are riding compounding software gains on Blackwell GPUs as inference economics shift to cost per token.

Jaeden Schafer5 min read