TechForge

20th August 2026

The GEN-1.5 robot model from Generalist AI learns new physical tasks from one demo, without gradient updates.

The company describes GEN-1.5 as the first model it knows of to show one-shot and few-shot learning of physical skills at scale. That claim points at a specific trade-off: whether teaching a machine a new job still requires weeks of task-specific programming, or whether it can now be shown once and left to work out the rest.

Success rates on the tasks it tested are modest, and the tasks themselves are short and simple. Generalist AI nonetheless argues the pattern, learning from a handful of seconds of data with no training step, has not previously been demonstrated across such a range of physical tasks.

How physical prompting works

Generalist AI built GEN-1.5 as a large multimodal model that processes video input alongside other sensor, language, and proprioceptive data, and produces action trajectories at 100 Hz. It holds 30 seconds of memory in context.

A single demonstration, lasting between three and twelve seconds, can be inserted into that memory window as what the company calls a physical prompt. The model then attempts the task straight away, with no training step in between. Generalist AI says it added no architectural changes to promote in-context learning, no meta-learning loop, and no auxiliary objectives to encourage improvisation.

Physical prompts can come from a human wearing handheld grippers or from a rollout of the robot itself. Generalist AI offers two hypotheses for why the behaviour appears at all. One draws on research into language models, where repetitive and unevenly distributed patterns in training data have been linked to in-context learning. The other points to the naturally repetitive cycles found in physical work, which the model may have learned to detect and extend the way language models extend general sequences.

One-shot success rates remain modest

10 tasks made up the evaluation set for GEN-1.5. One-shot in-context prompting – run directly from the pretrained model with no gradient updates – achieved an average success rate of 59 percent, with a standard deviation of 10 percentage points.

Few-shot learning through ten gradient steps on five minutes of data – roughly 50 demonstrations – raised the average to 83 percent, with a standard deviation of nine percentage points.

The 10 tasks included retrieving money from a purse, twisting the lid off a glass jar, and folding paper. Other tasks required stacking two cups, sweeping trash with a brush, and opening a book cover. The remaining four involved brushing a cube into a bowl, flipping a phone upside down, unzipping a pencil pouch, and removing a vacuum pad.

Eight months of continuous pretraining

GEN-0, announced nine months before GEN-1.5, showed predictable scaling behaviour during pretraining. GEN-1 followed five months later and could be post-trained to task mastery at success rates above 99 percent, alongside initial signs of improvisation.

GEN-1.5’s pretraining has been running continuously for more than eight months, across three training phases, with next-action prediction error on a held-out validation set continuing to fall throughout that period.

Fewer gradient steps were needed to adapt GEN-1.5 to new tasks as its training progressed, Generalist AI reports. The requirement fell from hundreds of steps, to tens, and eventually to a single gradient step on one minute of data.

That progression led the team to test whether the model could learn without any training step, purely from context. According to the company, this had not been observed in a robot model before.

Chaining prompts and crossing the sim-to-real gap

Physical prompts can be combined. Generalist AI demonstrated placing two independently recorded demonstrations, unzipping a pencil pouch and retrieving money from it, into the model’s context window together.

GEN-1.5 chained the two into one continuous behaviour, producing repositioning and error-recovery motions that appeared in neither original recording. The company calls this physical prompt engineering and compares it to chaining instructions inside a language prompt, building longer tasks from a library of short, reusable examples.

In-context learning also crossed the boundary between simulation and the real world. A demonstration recorded entirely inside a simulator – despite none of GEN-1.5’s pretraining data containing simulated video or dynamics – worked as a physical prompt for a real robot performing the equivalent task. Generalist AI reports the resulting behaviour generalised to different hands and to new object positions and sizes in the physical scene.

In a separate test, a person demonstrated a task with their own hands in view of the robot’s cameras, and the robot reproduced the action afterward with its own hands.

Generalist AI also tested a second adaptation route: light fine-tuning through gradient descent. Previous robot models often required tens of thousands of gradient steps to learn a new task, sometimes more. GEN-1.5 adapted in 1–10 steps using 1–5 minutes of data (equivalent to roughly 10–50 demonstrations.)

Ten gradient steps changed the model’s weights on held-out tasks by less than 0.15 percent. In the single-step regime, trained on one minute of data, the model reached 66.5 percent success on a held-out task, and that figure improved with larger batch sizes and higher learning rates. The company says it did not tune this adaptation procedure or sweep hyperparameters specific to it.

Substituting tools without instruction

Fine-tuned models generalised beyond their training data in Generalist AI’s tests, extending to new embodiments, objects, environments, and manipulation strategies.

In one example, researchers fine-tuned GEN-1.5 on five minutes of human demonstrations showing a brush sweeping a block into a bowl. Handed a banana instead, the model used it as a substitute brush. 

Handed a dustpan, it departed further from the demonstrated strategy, using the tool to lift the block and tip it into the bowl. Generalist AI ran a nearest-neighbour language search across 1,891,392 scenes in its pretraining data and found no close match for a dustpan used that way, in either the fine-tuning set or the wider pretraining corpus, to the best of its knowledge.

Other fine-tuned behaviours followed a similar pattern. A model fine-tuned only to place a block in a bowl removed a sheet of paper covering the bowl before completing the task, sometimes replacing the paper afterward, despite no such obstacle appearing in its five minutes of training data or – as far as the company can tell – in pretraining.

When a Lego brick stuck to one hand’s fingertips, the model used its other hand to remove it. Models trained to twist a lid off a jar with one hand sometimes switched to using both, and models fine-tuned on a single block and bowl sometimes began sorting several blocks by colour. Generalist AI reports this improvisation grew more frequent as the number of fine-tuning gradient steps fell, a pattern it attributes to lightly-adapted models staying closer to their pretrained behaviour.

What one demo means for robot deployment

The comparison Generalist AI draws is to the older ambitions of industrial robotics. Teach-by-guiding methods trace back to the Unimate patented in 1954 and MIT’s Copy Demo in 1970, decades of work aimed at getting machines to copy a demonstrated motion.

Robots marketed as general-purpose machines have needed an expert programmer, a process the company says took months of specialised effort. Generalist AI argues that once pretraining passes a certain point, the cost of adapting a model to a new task becomes small enough that showing the robot once starts to substitute for that programming step.

Generalist AI is careful to flag the limits of the one-shot route. Skills learned purely in context remain more brittle than those reached through fine-tuning, according to the company, even though they can handle some unexpected variation and recover from mistakes mid-task.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: Cisco and Rockwell connect IT and OT for industrial AI

Banner for IoT Tech Expo

Want to learn more about the IoT from industry leaders? Check out IoT Tech Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including AI & Big Data Expo and the Cyber Security Expo. Click here for more information.

IoT News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

About the Author

Senior Editor

Ryan Daws is a senior editor at TechForge Media with over a decade of experience in weaving narratives and dissecting complex topics. His articles and interviews with industry leaders have earned him recognition as a key tech influencer from numerous organisations. Under his leadership, publications have been praised by analyst firms for their excellence and performance. Connect with him on X, Mastodon, Bluesky, Threads, and/or LinkedIn.

Related

17th August 2026

13th August 2026

11th August 2026

10th August 2026

Join our Community

Subscribe now to get all our premium content and latest tech news delivered straight to your inbox

Popular

6047 view(s)
2428 view(s)
2314 view(s)
1866 view(s)

Subscribe

All our premium content and latest tech news delivered straight to your inbox

This field is for validation purposes and should be left unchanged.