Nvidia expanded its Nemotron 3 lineup on August 11 with Nemotron 3.5 Lightning, a 30B parameter mixture-of-experts model released under open weights and aimed squarely at always-on local agents. The company claims up to 4x faster token generation and 30% faster time to completion than other open models in its class, positioning the release as the runtime piece of a broader local-AI push it is rolling out throughout August.
The model is designed to be fine-tuned on-device for narrow tasks — writing in a preferred tone, handling a specific coding style, or operating inside a vertical like photography or 3D design. That specialization is the thesis: a smaller, faster expert-routed model that lives on the user's hardware, wired into local files and tools, rather than a general-purpose frontier model called over the network for every step.
Nemotron 3.5 Lightning runs on Nvidia RTX PCs, DGX Spark and OEM GB10 systems, and Jetson at the low end, and scales up through RTX PRO workstations, DGX Station, and GB300 deskside systems into data centers and cloud. Nvidia listed Acer, ASUS, Dell Technologies, Exxact, GIGABYTE, HP, Lenovo, MSI, and Supermicro as Blackwell-system partners shipping compatible hardware — a hardware surface area few competitors can match.
Key facts
- 01Nvidia released Nemotron 3.5 Lightning, a 30B parameter mixture-of-experts model with open weights, on August 11, 2026.
- 02The model delivers up to 4x faster token generation and 30% faster time to completion versus open models in its class.
- 03Nvidia's NeMo Switchyard router cut benchmark completion cost to roughly one-third of running Opus 4.8 alone by routing steps across models.
- 04Day-one deployment support comes from vLLM, Ollama, llama.cpp, LM Studio, and Unsloth across NVFP4 and GGUF formats.
- 05The model runs across RTX PCs, DGX Spark, Jetson, RTX PRO workstations, and GB300 deskside systems from Acer, ASUS, Dell, HP, Lenovo, and others.
The runtime story is deliberately open. Nvidia worked with vLLM, Ollama, llama.cpp, and LM Studio for local deployment, and offered the weights in both NVFP4 and GGUF formats. Unsloth is providing day-one quantized builds through Unsloth Studio. That collection covers essentially every serious local-inference stack in circulation, cutting the usual gap between a vendor announcement and community usability down to zero.
Alongside the model, Nvidia released NeMo Switchyard, an open-source routing library that directs each step of an agent workflow to whichever model — local or remote, Nvidia or otherwise — best fits the task on accuracy, speed, and cost. Internal benchmarks showed Switchyard maintained frontier-level task completion while cutting benchmark completion cost to roughly one-third of running Opus 4.8 alone. The library is on GitHub.
The economic pitch matters. Enterprise agent workloads chain many model calls together, and token costs on top-tier closed models compound quickly across multi-step tasks. Routing cheap, fast local calls through Nemotron 3.5 Lightning for the easy steps and reserving frontier models for the hard ones is the emerging pattern for keeping agentic AI affordable at scale, and Nvidia is now shipping both halves of that stack.
Distribution is broad. Nemotron 3.5 Lightning is available through OpenRouter, on build.nvidia.com as an NIM microservice, and through Nvidia Cloud Partners and third-party post-training and inference platforms. That mirrors the deployment surface for a closed frontier model while keeping the weights themselves open for fine-tuning — a middle-ground positioning Nvidia has been steadily building out across the Nemotron family.
The competitive frame is worth naming. Meta's Llama has been the default open-weights option for on-device agents, and Moonshot's Kimi K3, which we covered last week, sharpened pressure on closed labs. Nvidia entering with a hardware-optimized 30B MoE plus a router that can arbitrage across providers is a different play — the company is not trying to be the best model, it is trying to be the substrate everyone else's models run on.
The obvious caveat: Nvidia has not published head-to-head accuracy benchmarks against Llama, Qwen, or the current Kimi release in the announcement itself, and the 4x throughput and 30% latency numbers are relative to an unnamed "class" of open models. Independent evaluation on standard reasoning and coding benchmarks will decide whether Nemotron 3.5 Lightning is genuinely competitive on quality or only on speed. Community testing across the vLLM and Ollama channels will surface that within days.
For Nvidia, the strategic logic is straightforward: every developer who fine-tunes Nemotron 3.5 Lightning on an RTX PC or a DGX Spark is a developer whose agent stack is anchored on Nvidia silicon by default. Pairing an open model with a router that can call anything makes the platform, not the model, the lock-in. That is a more durable position than trying to out-scale OpenAI or Anthropic on frontier training runs, and it is where the local-AI market appears to be heading.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




