Nous Research's Hermes Agent has crossed 140,000 GitHub stars in under three months and, as of last week, is the most-used agent in the world by OpenRouter's traffic measure. The framework is provider- and model-agnostic, built to run locally on NVIDIA RTX PCs, NVIDIA RTX PRO workstations, and the NVIDIA DGX Spark, which ships with 128GB of unified memory and 1 petaflop of AI performance. NVIDIA detailed the pairing in a May 13, 2026 post tied to its RTX AI Garage program.
The pitch is reliability and self-improvement, two traits that have eluded most agent frameworks. Hermes writes and refines its own skills, saving learnings from feedback so it can adapt over time, and treats sub-agents as short-lived, isolated workers with focused context windows. That last design choice matters for local deployment: smaller context windows mean smaller models can drive the agent without losing the plot mid-task.
The local angle gets sharper with Alibaba's new Qwen 3.6 series. The Qwen 3.6 35B model runs on roughly 20GB of memory while outperforming the previous-generation 120B-parameter models, which required 70GB+ to host. The denser Qwen 3.6 27B variant matches the accuracy of 400B-parameter models like Qwen 3.5 397B at one-sixteenth the size, which is the kind of compression that turns a workstation into a credible agent host.
Key facts
- 01Hermes Agent crossed 140,000 GitHub stars in under three months and is now the most-used agent on OpenRouter.
- 02Qwen 3.6 35B runs on roughly 20GB of memory while surpassing 120B-parameter models that need 70GB+.
- 03NVIDIA DGX Spark ships with 128GB unified memory and 1 petaflop of AI performance.
- 04NVIDIA RTX PRO GPUs deliver up to 3x faster token generation running Qwen 3.6 via llama.cpp.
- 05Google Gemma 4 26B and 31B models are now available as NVFP4 checkpoints on NVIDIA Blackwell GPUs.
On RTX PRO hardware, NVIDIA cites up to 3x faster token generation running Qwen 3.6 models through llama.cpp. Hermes Agent ships with out-of-the-box support for LM Studio and Ollama in addition to llama.cpp, so the path from clone to running agent is short. NVIDIA Tensor Cores handle the inference acceleration, which is what lets multistep tasks resolve in seconds rather than minutes.
“The Qwen 3.6 27B matches the accuracy of 400 billion-parameter models like Qwen 3.5 397B while being one-sixteenth the size, running on roughly 20GB of memory.”— Jaeden Schafer
Hermes follows OpenClaw as the second open-source agentic framework to find serious community traction this year. Nous Research is positioning Hermes as an active orchestration layer rather than a thin wrapper around a chat API, with curated skills, tools, and plug-ins that the team stress-tests before shipping. The company says the framework "just works — even with 30 billion-parameter-class local models — without the constant debugging that most other agent frameworks require."
DGX Spark is the hardware NVIDIA wants paired with this workload. The compact standalone machine can run 120B-parameter mixture-of-experts models continuously, but the Qwen 3.6 35B model delivers comparable intelligence in a leaner footprint, leaving headroom for concurrent workloads. For an always-on agent that plans, executes, and refines itself in the background, that headroom is the difference between a demo and a daily driver.
The broader RTX AI Garage update bundles a few adjacent moves. Google's Gemma 4 26B and 31B models are now available as NVFP4 checkpoints optimized for NVIDIA Blackwell GPUs, and pairing them with Google's Multi-Token Prediction drafters delivers up to 3x faster inference at identical output quality. Mistral Medium version 3.5, released in April, added llama.cpp and Ollama compatibility for RTX PRO and DGX Spark systems.
NVIDIA also rolled out NemoClaw, an open-source stack that hardens OpenClaw experiences on NVIDIA devices and adds local-model support. NemoClaw now runs on Windows Subsystem for Linux, which opens the door for developers on Microsoft's platform without forcing a dual-boot. NVIDIA is pointing developers to a step-by-step playbook for getting NemoClaw running on DGX Spark.
The skeptical read: agent frameworks have a short half-life. OpenClaw was the breakout of the prior cycle, and Hermes is the breakout of this one — neither guarantees the next. GitHub stars and OpenRouter traffic measure attention more than retention, and the test for Hermes is whether self-evolving skills actually compound over months of real use, or whether users hit the same brittleness wall that has dogged every prior generation of agents.
For NVIDIA, the strategic point is that local agents are becoming a credible category, and DGX Spark plus RTX PRO is the hardware configuration NVIDIA wants developers to default to when they design for it. The Qwen 3.6 efficiency gains matter here: every doubling of capability-per-gigabyte expands the addressable market for on-device agents and shrinks the case for routing every task to a cloud endpoint. If Hermes retains its lead, NVIDIA's local-AI flywheel gets another turn — and the cloud-hosted agent platforms lose some of their pricing power.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




