Skip to main content
Live
Main content

Hermes Agent hits 140,000 GitHub stars, runs on NVIDIA RTX and DGX Spark

Nous Research's local-first agent framework is now the most-used agent on OpenRouter, paired with Alibaba's new Qwen 3.6 models.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Nous Research's Hermes Agent has crossed 140,000 GitHub stars in under three months and, as of last week, is the most-used agent in the world by OpenRouter's traffic measure. The framework is provider- and model-agnostic, built to run locally on NVIDIA RTX PCs, NVIDIA RTX PRO workstations, and the NVIDIA DGX Spark, which ships with 128GB of unified memory and 1 petaflop of AI performance. NVIDIA detailed the pairing in a May 13, 2026 post tied to its RTX AI Garage program.

The pitch is reliability and self-improvement, two traits that have eluded most agent frameworks. Hermes writes and refines its own skills, saving learnings from feedback so it can adapt over time, and treats sub-agents as short-lived, isolated workers with focused context windows. That last design choice matters for local deployment: smaller context windows mean smaller models can drive the agent without losing the plot mid-task.

The local angle gets sharper with Alibaba's new Qwen 3.6 series. The Qwen 3.6 35B model runs on roughly 20GB of memory while outperforming the previous-generation 120B-parameter models, which required 70GB+ to host. The denser Qwen 3.6 27B variant matches the accuracy of 400B-parameter models like Qwen 3.5 397B at one-sixteenth the size, which is the kind of compression that turns a workstation into a credible agent host.

Key facts

  • 01Hermes Agent crossed 140,000 GitHub stars in under three months and is now the most-used agent on OpenRouter.
  • 02Qwen 3.6 35B runs on roughly 20GB of memory while surpassing 120B-parameter models that need 70GB+.
  • 03NVIDIA DGX Spark ships with 128GB unified memory and 1 petaflop of AI performance.
  • 04NVIDIA RTX PRO GPUs deliver up to 3x faster token generation running Qwen 3.6 via llama.cpp.
  • 05Google Gemma 4 26B and 31B models are now available as NVFP4 checkpoints on NVIDIA Blackwell GPUs.

On RTX PRO hardware, NVIDIA cites up to 3x faster token generation running Qwen 3.6 models through llama.cpp. Hermes Agent ships with out-of-the-box support for LM Studio and Ollama in addition to llama.cpp, so the path from clone to running agent is short. NVIDIA Tensor Cores handle the inference acceleration, which is what lets multistep tasks resolve in seconds rather than minutes.

The Qwen 3.6 27B matches the accuracy of 400 billion-parameter models like Qwen 3.5 397B while being one-sixteenth the size, running on roughly 20GB of memory.
Jaeden Schafer

Hermes follows OpenClaw as the second open-source agentic framework to find serious community traction this year. Nous Research is positioning Hermes as an active orchestration layer rather than a thin wrapper around a chat API, with curated skills, tools, and plug-ins that the team stress-tests before shipping. The company says the framework "just works — even with 30 billion-parameter-class local models — without the constant debugging that most other agent frameworks require."

DGX Spark is the hardware NVIDIA wants paired with this workload. The compact standalone machine can run 120B-parameter mixture-of-experts models continuously, but the Qwen 3.6 35B model delivers comparable intelligence in a leaner footprint, leaving headroom for concurrent workloads. For an always-on agent that plans, executes, and refines itself in the background, that headroom is the difference between a demo and a daily driver.

The broader RTX AI Garage update bundles a few adjacent moves. Google's Gemma 4 26B and 31B models are now available as NVFP4 checkpoints optimized for NVIDIA Blackwell GPUs, and pairing them with Google's Multi-Token Prediction drafters delivers up to 3x faster inference at identical output quality. Mistral Medium version 3.5, released in April, added llama.cpp and Ollama compatibility for RTX PRO and DGX Spark systems.

NVIDIA also rolled out NemoClaw, an open-source stack that hardens OpenClaw experiences on NVIDIA devices and adds local-model support. NemoClaw now runs on Windows Subsystem for Linux, which opens the door for developers on Microsoft's platform without forcing a dual-boot. NVIDIA is pointing developers to a step-by-step playbook for getting NemoClaw running on DGX Spark.

Related · from this week
Nvidia ships Nemotron 3.5 Lightning, a 30B open MoE for local agents
Jaeden Schafer · 5 min read →

The skeptical read: agent frameworks have a short half-life. OpenClaw was the breakout of the prior cycle, and Hermes is the breakout of this one — neither guarantees the next. GitHub stars and OpenRouter traffic measure attention more than retention, and the test for Hermes is whether self-evolving skills actually compound over months of real use, or whether users hit the same brittleness wall that has dogged every prior generation of agents.

For NVIDIA, the strategic point is that local agents are becoming a credible category, and DGX Spark plus RTX PRO is the hardware configuration NVIDIA wants developers to default to when they design for it. The Qwen 3.6 efficiency gains matter here: every doubling of capability-per-gigabyte expands the addressable market for on-device agents and shrinks the case for routing every task to a cloud endpoint. If Hermes retains its lead, NVIDIA's local-AI flywheel gets another turn — and the cloud-hosted agent platforms lose some of their pricing power.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Tools

Nvidia logo
Models

Nvidia ships Nemotron 3.5 Lightning, a 30B open MoE for local agents

The new open-weights model runs 4x faster than class rivals and slots into RTX PCs, DGX Spark, and Jetson for always-on agentic workloads.

Jaeden Schafer5 min read
Nvidia logo
News

Nvidia unveils RTX Spark, a Windows PC class built for local AI agents

RTX Spark packs 1 petaflop of AI compute and 128GB of unified memory, with shipping set for this fall alongside an OpenShell runtime for Windows.

Jaeden Schafer5 min read
Nvidia logo
Tools

Nvidia's GeForce NOW adds cross-store library sync, cuts annual pricing

Ultimate tier drops $70 for 12 months as cloud service ties together Steam, GOG, Ubisoft+, EA app and Xbox libraries.

Jaeden Schafer4 min read