Show the robot once, then ask it to do the job

Skild AI’s S1 robot foundation model is designed to learn a new, long sequence of actions from a video demonstration. An operator records the task; the system interprets the objects and steps, then maps them to actions for the robot in front of it. The idea is to avoid collecting a new dataset and retraining the model every time a factory changes a process.

The NVIDIA source video shows the system around everyday manipulation tasks, including food preparation. The more legible reported result is a plant-potting trial: Skild says its team went from recording the demonstration to autonomous execution on hardware in 11 minutes. That is a concrete setup-time claim, not evidence that any factory task can be taught in eleven minutes.

Long tasks are the real test

S1 is intended to handle unfamiliar tasks lasting up to ten minutes and involving dozens of manipulation steps. NVIDIA and Skild list examples including potting a plant, making pancakes, brewing pour-over coffee and assembling a kit. Those are useful demonstrations because they require the robot to sequence actions instead of repeating one isolated motion.

The companies say the model can adjust when objects move, recover from errors and compose skills it was not explicitly programmed to perform. In factory use, those behaviors matter more than a polished single pick-and-place clip: the robot has to continue when an object or sequence is a little different from the demonstration.

The vendor-reported score deserves context

For new multistep tasks, Skild reports about 66% success at each step, compared with 9% for a similar AI system. It also estimates that one short video can provide information comparable to roughly 380 hands-on examples. Both numbers are reported by Skild and NVIDIA; the post does not give an independent replication or enough detail to generalize them to every task.

Per-step reliability compounds across a long task. Even a system that succeeds on two-thirds of individual steps needs recovery behavior if it is expected to finish a sequence reliably. The reported ability to adapt and recover is therefore as important as the headline success rate.

From kitchen demos to production

Skild, NVIDIA and Foxconn are also working on a dual-arm assembly workflow for NVIDIA Blackwell systems. NVIDIA describes a robot installing a busbar and limit block, fastening 16 screws and adapting to disturbances. That task combines precision, contact-aware motion and recovery—closer to the constraints of manufacturing than a static tabletop demonstration.

The source report describes the model, evaluation and deployment partners: https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/. For robotics, the key question is whether learning from video can reduce setup time without sacrificing repeatability, safety and recovery when a real production line varies.

Explore the original source ↗

Source published September 10, 2026. Coverage is based on the maker’s announcement and demonstration.