
Image: METAL
Summary
- GEN-1.5 performs new tasks with no gradient updates once a single 3-to-12-second demonstration is loaded into its 30-second context window
- Across 10 manipulation tasks, one-shot prompting hit a 59% average success rate, and fine-tuning with 10 steps on five minutes of task-specific data pushed that to 83%
- There are no public weights, API, or pricing yet — it's a research release accessible only through partnerships for now
A person sits in front of a robotic arm and shows it, just once, how to sweep blocks into a bowl. The moment that 12-second clip gets loaded into the robot's "memory window," the robot repeats the exact same motion — no training involved. That's GEN-1.5, the robot foundation model just released by robotics foundation-model startup Generalist AI.
One Demo, No Gradient Updates
GEN-1.5 is a multimodal model that takes in video, sensor data, language, and proprioception (the sense of one's own body position and force) all at once. It holds 30 seconds of memory while generating action trajectories at 100Hz. The core trick behind it is something Generalist AI calls "physical prompting." You drag and drop sensor data and an actual motion trajectory into that 30-second window, the rest of the space fills up with live observations, and the robot performs the task on the spot. No gradient updates that change the weights, no separate fine-tuning, no task-specific programming.
Generalist AI says it didn't deliberately design this capability. There was no architectural change aimed at encouraging in-context learning, no meta-learning loop, no auxiliary objective pushing the model toward improvisation. Instead, the company says the ability emerged naturally from more than eight months of continued pretraining on physical interaction data collected in homes, warehouses, and factories. The team drew a direct comparison to how one-shot prompting emerged in OpenAI's GPT-3 purely from the scale of pretraining, without any special design for it.
The Numbers Behind the Success Rate
Across 10 different manipulation tasks, success rates varied by training method as follows.
| Training Method | Data Scale | Success Rate |
|---|---|---|
| One-shot in-context prompting | 1 demo, 3–12 seconds | 59% (±10%) |
| Fine-tuning (10 gradient steps) | 5 minutes per task, ~50 demos | 83% (±9%) |
| Minimal adaptation (1 gradient step) | 1 minute of data | 66.5% (held-out tasks) |
The compute figures are what stand out here. Adapting a robot policy to a new task normally takes tens of thousands of gradient steps, but GEN-1.5 saw a major performance jump with just 10. And those 10 steps only changed less than 0.15% of the model's weight values. That suggests the model isn't building new representations from scratch — it's reorganizing knowledge it already has. Generalist AI describes this as "test-time learning with an extremely small amount of data."
Learning in Simulation, Working in Reality
Three transfer capabilities GEN-1.5 demonstrated are also worth noting. First, when two demonstrations recorded separately at different locations were loaded into memory together, the robot generated and stitched together posture adjustments, regrasping, and error-recovery motions that appeared in neither original demonstration. Second, even though the pretraining data contained no rendered video or simulated physics whatsoever, demonstrations recorded purely in simulation worked directly on the real robot. That means some tasks may no longer require collecting physical demonstrations at all. Third, the team confirmed human-to-robot imitation: when a person demonstrated an action with their own hand in front of the robot's camera, the robot reproduced it with its own hand.
Even more interesting generalization showed up after fine-tuning. After five minutes of training on sweeping blocks into a bowl, the robot picked up a banana instead of a brush and swept with it, and in another scenario chose a completely different contact strategy — scooping the blocks with a dustpan and moving them. It also cleared away paper covering the bowl on its own, and even though it was only demonstrated using one hand, it was seen alternating between both hands.
Still Research-Only, Available Through Partnerships
GEN-1.5 isn't something you can try out right now. There are no public weights, no API, no pricing, no self-serve access. Generalist AI is currently running the model only within its own robot fleet and data engine, and outside access requires a direct partnership. The company itself acknowledges that the tasks tested so far are simple, short-horizon actions. Still, it says this is, to the team's knowledge, the first time one-shot learning of physical skills has appeared at this scale.
Editor's Take
Most of the numbers we've seen so far in the robot foundation model race have been bragging rights about training data volume. When Dyna Robotics released DYNA-2 on August 11, it led with the fact that pretraining used 1 million hours of human video — roughly 170 years' worth. Google DeepMind's Gemini Robotics 2 competed on per-task success rates, like 76.3% for shelf-picking and 36% for light bulb insertion. GEN-1.5 picked a different axis entirely: instead of total data volume, it's making "how much can be learned from a single demonstration" the headline metric. That's essentially porting over the trend from language models — where GPT-3 could handle new tasks from just a few examples with no fine-tuning — directly into robotics.
From a practical standpoint, the value of this approach lies in adaptation cost. Getting a robot to handle a new task used to require tens of thousands of gradient steps and mountains of demonstration data. GEN-1.5 pushed success from 59% to 83% with just 10 steps. That means the time and labor needed to teach a robot a new task on a factory floor or in a warehouse could shrink dramatically — a real reduction in the barrier to deploying robots. That said, numbers like 59% and 83% still mean failures happen often. It's too early to put this straight into logistics or manufacturing environments; for now, it's best read as a research signal.
The thing to watch over the coming months is how fast other companies catch up on this "physical prompting" approach. If it's true that this capability simply emerges once pretraining scale gets big enough, other robotics startups with similar data pipelines are likely to report the same phenomenon before long.





Comments