Skip to main content
Live
Main content

Skild AI's S1 robot learns new factory tasks from a single video

The startup hit a $100M revenue run rate 10 months after first deployment and is now installing Blackwell systems at Foxconn with Nvidia.

Jaeden Schafer
Editor in Chief · · 5 min read
Nvidia logo

Skild AI's new S1 robot foundation model can learn a previously unseen factory task from a single video demonstration, and the startup has hit a $100 million annual revenue run rate just 10 months after its first commercial deployment. The model, launched last week, treats a video of a human performing a task as a prompt and executes the work on hardware without retraining or updating weights. Skild built S1 on Nvidia infrastructure and has already signed more than 60 deployment partnerships across manufacturing, logistics, inspection, security and food preparation.

The technique, called in-context learning, targets the biggest cost in industrial robotics: reprogramming. Fixed-function robots typically require fresh datasets, retraining and revalidation every time a product changes or a line is reconfigured. S1 skips that loop. An operator records a demonstration, the model infers the intent, objects and sequence, and the robot performs the task — often one that was not covered in pretraining.

Skild says S1 can handle unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee and kit assembly. In one plant-potting test, the team went from recording the demo to autonomous execution in 11 minutes. The model adjusts when objects move, recovers from errors, and stitches skills into sequences it was never explicitly programmed to run.

Key facts

  • 01Skild AI's S1 model learns unseen tasks up to 10 minutes long from a single video demo, without weight updates or task-specific post-training.
  • 02S1 hit 66% per-step success on new multistep tasks versus 9% for a comparable system — a more than sevenfold improvement.
  • 03Skild reached a $100M annual revenue run rate 10 months after first commercial deployment, with 60+ deployment partnerships.
  • 04One short video demo is roughly equivalent to 380 hands-on training examples, which would take a person 50–100 hours to collect.
  • 05Skild, Nvidia and Foxconn are deploying the Skild Brain on dual-arm robots assembling Nvidia Blackwell systems, including a 16-screw workflow.

The benchmark numbers are what make the release notable. On new multistep tasks, S1 succeeded at roughly 66% of each step, compared with 9% for a similar AI system — a more than sevenfold gap. Skild also estimates that one short video demonstration carries about as much signal as 380 hand-collected training examples, which would take a human operator 50 to 100 hours to gather.

The commercial work is already on factory floors. Skild, Nvidia and Foxconn are jointly deploying the Skild Brain on dual-arm manipulators building Nvidia Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, drives 16 screws and adapts when the scene diverges from the plan. That is precise, contact-aware manipulation with sequence tracking and recovery — the hardest surface for general-purpose robots to hold reliably.

The Nvidia stack shows up at every stage. Cosmos world foundation models diversify training data and turn raw video into structured descriptions, while Cosmos Curator handles annotation and filtering at scale. Isaac Sim and Omniverse libraries generate physically based virtual environments for edge-case testing, and Isaac Lab wraps reinforcement learning around the Newton physics engine to close the sim-to-real gap. Nsight profiles training and TensorRT optimizes inference for real-time robot control.

Skild and Nvidia are also co-developing new GPU-accelerated simulation solvers for how robots grip and manipulate solid objects, which they plan to release to all developers through Newton. That is a meaningful contribution to the open robotics stack: contact-rich simulation has been one of the slowest, most brittle parts of the pipeline, and faster solvers directly compound how quickly any lab can train dexterous policies.

The strategic question is whether one-shot video learning holds up outside curated demos. A 66% per-step success rate is a step-change against baselines, but for a 20-step assembly the compounded probability of an end-to-end clean run is still modest, which is why Skild leans on error recovery and skill composition rather than claiming perfection. Long-horizon reliability, safety certification on factory floors and the economics of hardware-agnostic deployment remain the open work.

Related · from this week
Nvidia unveils Jetson Thor T3000 and T2000 for mainstream robotics
Jaeden Schafer · 5 min read →

Skild's trajectory — $100M ARR in under a year, 60-plus deployments and a Foxconn partnership building the chips that power the next round of AI — reframes what a robotics foundation model company can be. The playbook of a shared robot brain trained across embodiments, priced against reprogramming budgets rather than robot units, is the fastest path yet to putting adaptable manipulation into the tasks that have resisted automation for decades. If in-context learning from video generalizes, the bottleneck in industrial robotics stops being data collection and starts being how fast an operator can hit record.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Nvidia logo
Models

Nvidia unveils Jetson Thor T3000 and T2000 for mainstream robotics

The Blackwell-based T3000 delivers 865 FP4 teraflops in half the size of the T5000; both modules ship Q1 2027.

Jaeden Schafer5 min read
Nvidia logo
Models

Nvidia and Hugging Face push Isaac GR00T 1.7 into LeRobot for open robotics

The integration connects 3M robotics developers to 16M AI builders, with Cosmos 3 world models coming next to Hugging Face's open library.

Jaeden Schafer5 min read
Nvidia logo
Models

NVIDIA ships Omniverse skills to train vision AI agents on synthetic defect data

Roboflow hit 95% average precision on Corning fiber defects using just 8 real images plus synthetic data from NVIDIA's new skill.

Jaeden Schafer5 min read