Robots usually need new demonstrations, fine-tuning and validation every time their job changes. Skild AI’s S1 is testing a different model: show the robot one video of a task it has never performed, then use that video as the prompt without changing the model’s weights.
In Skild AI’s S1 research release, the company describes a robotic foundation model built around in-context learning. Instead of treating each new manipulation job as another training project, S1 receives a visual demonstration and maps the demonstrated intent, objects and sequence onto the robot and environment in front of it.
Skild showed the model publicly in its S1 launch demonstration on X, where the company presented the one-video prompting approach for unfamiliar physical tasks.
The video is the prompt
For a conventional robot-learning pipeline, adapting to a new task often means collecting a new dataset and running another training or fine-tuning cycle. That makes robot training data one of the most important—and expensive—inputs in physical AI.
S1 changes the adaptation step. The model is pretrained so that the demonstration itself becomes part of the context at inference time. A person can record the desired behavior from a different viewpoint and even a different embodiment; S1 then has to infer the goal and translate it into actions available to the robot.
That is closer to prompting a language model with an example than programming a traditional industrial robot. The weights stay fixed. The new information arrives through context.
The hard examples last up to 10 minutes
Skild’s most interesting demonstrations are not single pick-and-place motions. The company reports unfamiliar tasks lasting as long as 10 minutes and spanning dozens of manipulation steps, including potting a plant, cooking a pancake, making pour-over coffee and assembling a kit.
Long-horizon work matters because errors compound. A robot has to remember what has already happened, recognize when the scene has changed, sequence multiple skills and recover when an intermediate step does not go exactly as demonstrated.
In the plant-potting example, Skild says the interval from recording the human demonstration to autonomous execution on hardware was 11 minutes. That is a radically different deployment loop from collecting hours of task-specific demonstrations before every change.
66% versus 9% needs the right context
On Skild’s internal benchmark for unseen long-horizon tasks, the in-context model reached a 66% average cumulative per-step success rate after pretraining at the company’s largest tested scale. A language-conditioned comparison model trained on the same data reached 9%.
That does not mean S1 autonomously completed 66% of entire 10-minute jobs from start to finish. Skild’s methodology grades steps cumulatively and uses human intervention to recover from failures so later steps can still be evaluated. The number is therefore a per-step research metric, not an end-to-end production reliability rate.
NVIDIA’s independent technical account of the collaboration reports the same 66% versus 9% comparison and says Skild estimates that one short video prompt can provide roughly the adaptation value of about 380 hands-on training examples, which could otherwise require 50–100 hours of manual collection.
Better data still matters
One-shot prompting does not eliminate the need for large-scale pretraining. S1 only has useful context learning because the underlying model has already learned from diverse robotics experience. Skild says it combines teleoperation, egocentric human video, simulation and deployment data rather than betting on one source.
That connects directly to newer motion-capture training systems designed to record precise human movement at scale. Better demonstrations can improve what models learn before they ever receive a new task prompt.
Skild has also explored simulation-based robot learning, including very long simulated training horizons for locomotion. S1 extends the same broader strategy into manipulation: pretrain a general system broadly enough that adaptation can happen from context instead of another gradient update.
Why this matters on a factory floor
Industrial environments change constantly. Fixtures move, products change, bins arrive in different positions and previously scripted sequences stop matching reality. A robot that needs a full retraining project for every variation remains expensive to redeploy even if its hardware is capable.
A successful in-context robot could make the human demonstration itself the programming interface: show the task, verify the behavior, then deploy. That would move part of robotics engineering away from writing or retraining a specialist policy for every workflow and toward maintaining a general model with strong validation and safety boundaries.
S1 is not proof that this problem is solved. Its headline results are company-reported research results, the evaluation uses recovery interventions, and Skild has not shown that every unfamiliar factory process can be learned safely from one video. But the capability being tested is important: whether a general robot policy can learn a genuinely new physical sequence at deployment time without changing its weights.
The next evidence to watch is production evidence—complete-task reliability, failure rates, recovery behavior, safety validation and how performance changes after weeks or months of real operating drift. If those metrics hold up, one-video prompting could change the economics of reprogramming robots far more than another incremental improvement in arm speed or payload.
BitcoinVersus.Tech
Advertisement
Editor’s Note
This report distinguishes Skild AI’s internal per-step benchmark results from complete-task production reliability. The featured cover is an original editorial illustration and is not duplicated in the article body.
We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb
BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment