Skild AI's new S1 robot foundation model can learn a previously unseen factory task from a single video demonstration, and the startup has hit a $100 million annual revenue run rate just 10 months after its first commercial deployment. The model, launched last week, treats a video of a human performing a task as a prompt and executes the work on hardware without retraining or updating weights. Skild built S1 on Nvidia infrastructure and has already signed more than 60 deployment partnerships across manufacturing, logistics, inspection, security and food preparation.
The technique, called in-context learning, targets the biggest cost in industrial robotics: reprogramming. Fixed-function robots typically require fresh datasets, retraining and revalidation every time a product changes or a line is reconfigured. S1 skips that loop. An operator records a demonstration, the model infers the intent, objects and sequence, and the robot performs the task — often one that was not covered in pretraining.
Skild says S1 can handle unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee and kit assembly. In one plant-potting test, the team went from recording the demo to autonomous execution in 11 minutes. The model adjusts when objects move, recovers from errors, and stitches skills into sequences it was never explicitly programmed to run.
Key facts
- 01Skild AI's S1 model learns unseen tasks up to 10 minutes long from a single video demo, without weight updates or task-specific post-training.
- 02S1 hit 66% per-step success on new multistep tasks versus 9% for a comparable system — a more than sevenfold improvement.
- 03Skild reached a $100M annual revenue run rate 10 months after first commercial deployment, with 60+ deployment partnerships.
- 04One short video demo is roughly equivalent to 380 hands-on training examples, which would take a person 50–100 hours to collect.
- 05Skild, Nvidia and Foxconn are deploying the Skild Brain on dual-arm robots assembling Nvidia Blackwell systems, including a 16-screw workflow.
The benchmark numbers are what make the release notable. On new multistep tasks, S1 succeeded at roughly 66% of each step, compared with 9% for a similar AI system — a more than sevenfold gap. Skild also estimates that one short video demonstration carries about as much signal as 380 hand-collected training examples, which would take a human operator 50 to 100 hours to gather.
The commercial work is already on factory floors. Skild, Nvidia and Foxconn are jointly deploying the Skild Brain on dual-arm manipulators building Nvidia Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, drives 16 screws and adapts when the scene diverges from the plan. That is precise, contact-aware manipulation with sequence tracking and recovery — the hardest surface for general-purpose robots to hold reliably.
The Nvidia stack shows up at every stage. Cosmos world foundation models diversify training data and turn raw video into structured descriptions, while Cosmos Curator handles annotation and filtering at scale. Isaac Sim and Omniverse libraries generate physically based virtual environments for edge-case testing, and Isaac Lab wraps reinforcement learning around the Newton physics engine to close the sim-to-real gap. Nsight profiles training and TensorRT optimizes inference for real-time robot control.
Skild and Nvidia are also co-developing new GPU-accelerated simulation solvers for how robots grip and manipulate solid objects, which they plan to release to all developers through Newton. That is a meaningful contribution to the open robotics stack: contact-rich simulation has been one of the slowest, most brittle parts of the pipeline, and faster solvers directly compound how quickly any lab can train dexterous policies.
The strategic question is whether one-shot video learning holds up outside curated demos. A 66% per-step success rate is a step-change against baselines, but for a 20-step assembly the compounded probability of an end-to-end clean run is still modest, which is why Skild leans on error recovery and skill composition rather than claiming perfection. Long-horizon reliability, safety certification on factory floors and the economics of hardware-agnostic deployment remain the open work.
Skild's trajectory — $100M ARR in under a year, 60-plus deployments and a Foxconn partnership building the chips that power the next round of AI — reframes what a robotics foundation model company can be. The playbook of a shared robot brain trained across embodiments, priced against reprogramming budgets rather than robot units, is the fastest path yet to putting adaptable manipulation into the tasks that have resisted automation for decades. If in-context learning from video generalizes, the bottleneck in industrial robotics stops being data collection and starts being how fast an operator can hit record.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




