
Image: METAL
Summary
- Sakana AI, working with the University of Tokyo, unveiled SAIL, a VLM-based method for generating robot trajectories; the work will be presented at IROS 2026.
- Without changing the weights of Gemini Robotics-ER 1.5, SAIL refines trajectories through MCTS, similar-demonstration retrieval and step-level feedback, raising the average success rate across six tasks from 25% to 73%.
- A physical LeRobot SO-101 arm succeeded in five of six trials, and a policy trained on 120 trajectories collected by the search cut execution time from 644.72 seconds to 72.306 seconds.
Sakana AI on September 27 (UTC) unveiled SAIL (Scaling In-Context Imitation Learning), a method for generating robot trajectories. A vision-language model (VLM) plans the full motion of a robot arm from only a few successful demonstrations, but instead of executing its first answer, it runs the plan in a simulator first and then revises it. Across six manipulation tasks, raising the budget of candidate trajectories from one to 45 lifted the average rate of finding a successful trajectory from 25% to 73%, the company said. The research, carried out with the University of Tokyo, will be presented at the robotics conference IROS 2026.
The model's weights were never changed during the work. SAIL used Google's Gemini Robotics-ER 1.5 as is, both as the policy that writes trajectories and as the evaluator that scores the results. In language models, test-time scaling, which spends more computation at inference to generate, review and revise candidates, has become an established way to raise accuracy. Sakana AI asked whether the same trade holds for physical action. In its blog post the company noted that GPT-6 Astra has handled several manipulation tasks on a physical robot by setting motion targets from camera images and robot state alone, building on the premise that foundation models already contain knowledge usable for robot control.
The starting point is the observation that a single generation is hard to trust. A model's prediction shifts depending on which demonstrations it receives as context, and it wobbles from sample to sample even under the same conditions. "A small error in a movement target can cause the entire task to fail," Sakana AI explained. Executing the first prediction without evaluation leaves no chance to correct that error before the robot moves.
SAIL tackles the problem with Monte Carlo tree search (MCTS). Each node in the tree is one complete trajectory, and each edge is a single revision of the previous trajectory. First, the system pulls the demonstration that looks most like the current scene from an archive of successful trajectories, using LPIPS distance, and places it in context. The generated trajectory runs in the simulator, and an evaluation VLM samples 50 evenly spaced frames from the execution video to estimate what percentage of each subtask, such as reaching, grasping or handing over, has been completed. Those scores are attached to the trajectory's waypoints and fed back to the policy model, which keeps the high-scoring segments and rewrites only the parts where progress stalled. Successful trajectories found during search are added to the archive and reused for other initial configurations.
Makoto Sato, a University of Tokyo researcher and the paper's first author, and his three co-authors wrote that the design "shifts the focus from static imitation to a dynamic paradigm where a robot can effectively think longer to resolve the ambiguities of a novel task." Sato did the work as an intern at Sakana AI, and the co-authors are Yusuke Iwasawa of the University of Tokyo and Yujin Tang and So Kuroki of Sakana AI.
The numbers rose with computation. Experiments ran in the ALOHA manipulation simulator on six tasks, handing over a banana, handing over a pen, placing a bowl on a rack, opening a drawer, closing a laptop and removing a marker's cap, with 20 initial configurations per task. With budgets of 6, 15, 30 and 45 candidates, the average success rate was 55%, 65%, 71% and 73%, respectively. Drawer opening climbed from 10% to 50%, laptop closing from 15% to 70%, marker cap removal from 5% to 45%, and banana handover reached 95%. The bowl task hit 100% with just six candidates and stayed there.
Search strategy made a difference. At the same 15-candidate budget, SAIL's average success rate of 65% beat breadth-first search, which draws 15 candidates at once without feedback, at 51%, and depth-first search, which revises one trajectory 15 times in a row, at 37%. On laptop closing, however, breadth-first search scored 60% to SAIL's 50%. The authors attributed this to the fine positional adjustment needed for the gripper to pass safely over the thin top edge of the screen, a geometric requirement that suits the more random breadth-first search.

Which demonstrations go in mattered more than how many. At a 15-candidate budget, SAIL with one retrieved similar-scene demonstration reached 65%, versus 45% with only the fixed original demonstration and 50% with random retrieval from the archive. Raising fixed demonstrations to three yielded only 49%, and three random ones 53%. The authors concluded in the paper that "performance is more sensitive to the relevance of in-context demonstrations than to their sheer quantity." In the feedback comparison, trajectory text only, rollout images only, or both together scored 45% to 48%, and a single final score scored 49%, all below the 65% of step-level feedback that attaches scores to each waypoint.
Physical validation used an SO-101 robot arm from the LeRobot family, built around the open-source robot learning library. The task was to pick up a blue block and drop it into a red bowl. GroundingDINO locates the objects from the color and depth data of an Intel RealSense D435i camera, SAM2 traces their outlines, and coordinates are aligned using an ArUco marker fixed to the robot base, so the same scene is rebuilt inside the simulator. Searching that digital twin with a 15-candidate budget and executing the result on the physical arm succeeded in five of six trials in which the block and bowl positions were shuffled by hand. The authors attributed the one failure to small errors in depth-based pose estimation and to real-world contact dynamics the simulator did not capture.
Speed was addressed through distillation. The MCTS search took an average of 644.72 seconds per trial. The researchers varied object placement in simulation, used SAIL to collect 120 successful trajectories, and trained an ACT (Action Chunking with Transformers) policy on them through behavior cloning. That policy also succeeded in five of six physical trials while cutting average execution time to 72.306 seconds. "Our MCTS framework can serve as an automated data collection engine to train fast, deployable robot policies," the authors wrote.

According to the paper and project page that METAL reviewed, the work first appeared on arXiv on March 9 and a revised version followed on September 19. Trajectories are executed open-loop, following the plan without visual feedback during motion, and physical evaluation was limited to one task and six trials, Sakana AI itself noted. As a next step, the researchers said they aim to combine the framework with photorealistic digital twins built with Gaussian Splatting to narrow the visual gap between simulation and reality. METAL has previously reported that Sakana AI unveiled PC-ALM, a learning method without backpropagation, and also covered the problems the Gemini Robotics family ran into on the real home robot Apollo, the model line SAIL builds on.
Seen through an AI engineer's lens, SAIL shifts the cost structure of robot learning. Until now, teaching a robot a new task has meant paying for data: people collecting demonstrations and retraining a policy. SAIL fills that slot with inference compute and puts a measured curve behind the claim that success rises as compute rises. The price is time. A search that takes nearly 11 minutes per attempt is hard to deploy as is, but a loop in which 120 successful trajectories found that way become training data for a 72-second policy looks like a factory that turns compute into data. The quality of that factory is ultimately tied to how closely the simulator resembles reality. Sakana AI pointed at the same bottleneck when it said, "We think there is more to learn about what existing models can do with this kind of feedback, and how far those improvements carry over to physical robots."





Comments