Skip to main content
Live
Main content

Nvidia ships Nemotron 3.5 Lightning, a 30B open MoE for local agents

The new open-weights model runs 4x faster than class rivals and slots into RTX PCs, DGX Spark, and Jetson for always-on agentic workloads.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Nvidia expanded its Nemotron 3 lineup on August 11 with Nemotron 3.5 Lightning, a 30B parameter mixture-of-experts model released under open weights and aimed squarely at always-on local agents. The company claims up to 4x faster token generation and 30% faster time to completion than other open models in its class, positioning the release as the runtime piece of a broader local-AI push it is rolling out throughout August.

The model is designed to be fine-tuned on-device for narrow tasks — writing in a preferred tone, handling a specific coding style, or operating inside a vertical like photography or 3D design. That specialization is the thesis: a smaller, faster expert-routed model that lives on the user's hardware, wired into local files and tools, rather than a general-purpose frontier model called over the network for every step.

Nemotron 3.5 Lightning runs on Nvidia RTX PCs, DGX Spark and OEM GB10 systems, and Jetson at the low end, and scales up through RTX PRO workstations, DGX Station, and GB300 deskside systems into data centers and cloud. Nvidia listed Acer, ASUS, Dell Technologies, Exxact, GIGABYTE, HP, Lenovo, MSI, and Supermicro as Blackwell-system partners shipping compatible hardware — a hardware surface area few competitors can match.

Key facts

  • 01Nvidia released Nemotron 3.5 Lightning, a 30B parameter mixture-of-experts model with open weights, on August 11, 2026.
  • 02The model delivers up to 4x faster token generation and 30% faster time to completion versus open models in its class.
  • 03Nvidia's NeMo Switchyard router cut benchmark completion cost to roughly one-third of running Opus 4.8 alone by routing steps across models.
  • 04Day-one deployment support comes from vLLM, Ollama, llama.cpp, LM Studio, and Unsloth across NVFP4 and GGUF formats.
  • 05The model runs across RTX PCs, DGX Spark, Jetson, RTX PRO workstations, and GB300 deskside systems from Acer, ASUS, Dell, HP, Lenovo, and others.

The runtime story is deliberately open. Nvidia worked with vLLM, Ollama, llama.cpp, and LM Studio for local deployment, and offered the weights in both NVFP4 and GGUF formats. Unsloth is providing day-one quantized builds through Unsloth Studio. That collection covers essentially every serious local-inference stack in circulation, cutting the usual gap between a vendor announcement and community usability down to zero.

Alongside the model, Nvidia released NeMo Switchyard, an open-source routing library that directs each step of an agent workflow to whichever model — local or remote, Nvidia or otherwise — best fits the task on accuracy, speed, and cost. Internal benchmarks showed Switchyard maintained frontier-level task completion while cutting benchmark completion cost to roughly one-third of running Opus 4.8 alone. The library is on GitHub.

The economic pitch matters. Enterprise agent workloads chain many model calls together, and token costs on top-tier closed models compound quickly across multi-step tasks. Routing cheap, fast local calls through Nemotron 3.5 Lightning for the easy steps and reserving frontier models for the hard ones is the emerging pattern for keeping agentic AI affordable at scale, and Nvidia is now shipping both halves of that stack.

Distribution is broad. Nemotron 3.5 Lightning is available through OpenRouter, on build.nvidia.com as an NIM microservice, and through Nvidia Cloud Partners and third-party post-training and inference platforms. That mirrors the deployment surface for a closed frontier model while keeping the weights themselves open for fine-tuning — a middle-ground positioning Nvidia has been steadily building out across the Nemotron family.

The competitive frame is worth naming. Meta's Llama has been the default open-weights option for on-device agents, and Moonshot's Kimi K3, which we covered last week, sharpened pressure on closed labs. Nvidia entering with a hardware-optimized 30B MoE plus a router that can arbitrage across providers is a different play — the company is not trying to be the best model, it is trying to be the substrate everyone else's models run on.

Related · from this week
Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for agent workloads
Jaeden Schafer · 5 min read →

The obvious caveat: Nvidia has not published head-to-head accuracy benchmarks against Llama, Qwen, or the current Kimi release in the announcement itself, and the 4x throughput and 30% latency numbers are relative to an unnamed "class" of open models. Independent evaluation on standard reasoning and coding benchmarks will decide whether Nemotron 3.5 Lightning is genuinely competitive on quality or only on speed. Community testing across the vLLM and Ollama channels will surface that within days.

For Nvidia, the strategic logic is straightforward: every developer who fine-tunes Nemotron 3.5 Lightning on an RTX PC or a DGX Spark is a developer whose agent stack is anchored on Nvidia silicon by default. Pairing an open model with a router that can call anything makes the platform, not the model, the lock-in. That is a more durable position than trying to out-scale OpenAI or Anthropic on frontier training runs, and it is where the local-AI market appears to be heading.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for agent workloads

The 30B mixture-of-experts model runs 4x faster on output, while the routing library cuts task cost to a third of Opus 4.8.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia Nemotron 3 Ultra hits closed-model parity at 10x lower cost on LangChain

LangChain tuned its Deep Agents harness for Nemotron 3 Ultra, matching top closed models on business tasks without retraining.

Jaeden Schafer5 min read
Hugging Face's Delangue: half the Fortune 500 now runs on open source AI
Business

Hugging Face's Delangue: half the Fortune 500 now runs on open source AI

The CEO argues frontier API costs push companies to open models as they scale, and warns a handful of firms could otherwise control everything.

Jaeden Schafer5 min read