Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Researchers turned first-person human manipulation videos into 18,561 hours of training data for 15 robot types and tested whether it actually helps robot policies generalize
Training robots to manipulate objects well needs lots of varied demonstrations, but collecting them directly on real robots is slow and expensive. Ego2Robot is a pipeline that converts egocentric human videos into robot training data through action alignment, visual alignment, and quality filtering, producing 18,561 hours of data across 15 robot arm types. Mixing this synthetic data with real robot data during pretraining consistently improved success rates under unseen visual, layout, embodiment, and language conditions, with gains confirmed on a real robot as well.
METAL LAB explanatory visual
Ego2Robot pipeline: from human video to robot training data
Evidence statusMeasured results reported
- Input: egocentric manipulation videoAbout 1,940 hours of first-person human hand manipulation footage from four sources: ANT, EgoDex, ViTRA, EgoVerse
- Action alignment21 hand keypoints per frame are converted into smooth robot end-effector trajectories, with speed matched to robot teleoperation pace
- Visual alignmentSAM 3 segments and removes the human arm via inpainting; a robot base pose is solved via inverse kinematics and a rendered robot arm (one of 15 types) is composited into the scene
- Quality curationThree levels of filtering remove IK failures, self-collisions, statistical outliers, and semantically inconsistent clips flagged by a vision-language model
- Output: 18,561 hours of training dataParallel data generated for 15 robot morphologies, mixed with real robot data to pretrain a vision-language-action model
What they did
- The team collected about 1,940 hours of first-person human hand manipulation videos from four sources (ANT, EgoDex, ViTRA, EgoVerse), then built a pipeline with three stages: action alignment (converting 21 hand keypoints into robot end-effector trajectories), visual alignment (removing human arms via SAM 3 segmentation and inpainting, then compositing a rendered robot arm using inverse kinematics), and multi-level quality curation to filter bad frames and trajectories.
- They ran this pipeline separately for 15 different robot arm morphologies (Panda, UR5e, xArm7, and 12 others), turning each source video into 15 parallel training streams and yielding a total of 18,561 hours of synthesized robot data, described as the largest ego-to-robot dataset so far.
- To measure generalization precisely, the authors extended the RoboTwin2.0 benchmark so that four factors -- visual appearance (background, lighting, robot color), scene layout (table height, clutter, camera offset), embodiment (swapping in different robot arms), and task semantics (unseen objects, paraphrased instructions) -- could each be tested independently instead of all mixed together.
- Compared with pretraining on real robot data alone (about 6,565 hours from DROID, AgibotWorld, and InternData), mixing the synthesized data with real data at a 1:1 ratio raised success rate on RoboTwin's randomized setting to 53.5% (+2.6 points), and produced gains across background (+4), lighting (+8), robot color (+6), camera offset (+6), unseen objects (+11 at a 3:1 ratio), and paraphrased instructions (reaching 69% at 1:1).
- On a real ARX ACone robot across five tasks (putting fruit in a basket, placing blocks in a drawer, folding a towel, sweeping trash, inserting a screw), the model trained with both teleoperated demonstrations and the synthesized ego-derived data outperformed the robot-only model on every task, with the biggest gains on placing blocks (+14 points) and inserting a screw (+13 points).

| Pretraining | RoboTwin | Per-Dimension (RoboTwin) | EBench | ||||
|---|---|---|---|---|---|---|---|
| Clean | Rand | Visual | Scene | Embody | Task | Avg | |
| Robot-only | 62.2 | 50.9 | 61.4 | 52.9 | 23.8 | 46.2 | 39.6 |
| Ego2R+Robot (1:3) | 61.4 –0.8 | 51.0 +0.1 | 61.2 –0.2 | 52.5 –0.4 | 21.9 –1.9 | 49.5 +3.3 | 47.4 +7.8 |
| Ego2R+Robot (3:1) | 64.1 +1.9 | 49.2 –1.7 | 62.7 +1.3 | 54.3 +1.4 | 28.2 +4.4 | 51.6 +5.4 | 51.7 +12.1 |
| Ego2R+Robot (1:1) | 68.1 +5.9 | 53.5 +2.6 | 67.3 +5.9 | 56.9 +4.0 | 27.2 +3.4 | 54.1 +7.9 | 49.8 +10.2 |

| Pretraining | Visual | Scene | Embodiment | Task | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BG | Light | Color | Height | Clutter | Camera | ARX | UR5 | Franka | Obj | Lang | |
| Robot-only | 66.6 | 58.2 | 59.4 | 60.1 | 48.3 | 50.4 | 44.1 | 20.2 | 7.0 | 29.3 | 63.1 |
| Ego2R+Robot (1:3) | 65.0 –1.6 | 58.3 +0.1 | 60.3 +0.9 | 58.6 –1.5 | 49.3 +1.0 | 49.6 –0.8 | 43.7 –0.4 | 17.6 –2.6 | 4.5 –2.5 | 36.8 +7.5 | 62.2 –0.9 |
| Ego2R+Robot (3:1) | 65.5 –1.1 | 60.9 +2.7 | 61.8 +2.4 | 62.0 +1.9 | 49.2 +0.9 | 51.6 +1.2 | 47.6 +3.5 | 31.4 +11.2 | 5.6 –1.4 | 40.0 +10.7 | 63.1 +0.0 |
| Ego2R+Robot (1:1) | 70.3 +3.7 | 65.8 +7.6 | 65.8 +6.4 | 62.4 +2.3 | 52.0 +3.7 | 56.3 +5.9 | 51.2 +7.1 | 25.0 +4.8 | 5.3 –1.7 | 39.6 +10.3 | 68.5 +5.4 |
| Robot | DOF | Gripper (mm) | Reach (m) |
|---|---|---|---|
| Panda | 7 | 0–80 | 1.272 |
| Kinova Gen3 | 7 | 0–85 | 1.337 |
| IIWA | 7 | 0–85 | 1.411 |
| Sawyer | 7 | 14–79 | 1.420 |
| FR3 | 7 | 0–80 | 1.272 |
| xArm7 | 7 | 0–85 | 1.290 |
| UR5e | 6 | 0–85 | 1.236 |
| UR10e | 6 | 0–85 | 1.627 |
| Jaco | 6 | 0–125 | 1.200 |
| ViperX | 6 | 15–87 | 0.911 |
| WidowX | 6 | 11–55 | 0.787 |
| ARX-L5 | 6 | 0–88 | 0.855 |
| Piper | 6 | 0–70 | 0.883 |
| YAM | 6 | 4–75 | 0.866 |
| Aloha-Agilex | 6 | 7–102 | 0.853 |
| Task | Step 1 | Step 2 | Step 3 | Step 4 |
|---|---|---|---|---|
| Put Fruits | 33.3 | 33.3 | 33.3 | — |
| Put Blocks | 25 | 25 | 25 | 25 |
| Fold Towel | 50 | 50 | — | — |
| Sweep Trash | 25 | 25 | 25 | 25 |
| Insert Screw | 25 | 25 | 25 | 25 |

| Task | Pi0.5 | Robot-only | Ego2R (1:3) | Ego2R (3:1) | Ego2R (1:1) | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Clean | Rand | Clean | Rand | Clean | Rand | Clean | Rand | Clean | Rand | |
| adjust_bottle | 62.0 | 18.0 | 94.0 | 77.0 | 92.0 | 79.0 | 100.0 | 72.0 | 89.0 | 91.0 |
| beat_block_hammer | 48.0 | 16.0 | 63.0 | 27.0 | 64.0 | 56.0 | 57.0 | 18.0 | 65.0 | 71.0 |
| blocks_ranking_rgb | 74.0 | 16.0 | 56.0 | 60.0 | 64.0 | 63.0 | 49.0 | 43.0 | 78.0 | 60.0 |
| blocks_ranking_size | 32.0 | 0.0 | 30.0 | 22.0 | 46.0 | 21.0 | 43.0 | 7.0 | 52.0 | 16.0 |
| click_alarmclock | 44.0 | 56.0 | 96.0 | 93.0 | 100.0 | 87.0 | 100.0 | 85.0 | 100.0 | 100.0 |
| click_bell | 32.0 | 40.0 | 100.0 | 95.0 | 100.0 | 94.0 | 100.0 | 93.0 | 100.0 | 100.0 |
| dump_bin_bigbin | 72.0 | 70.0 | 54.0 | 73.0 | 72.0 | 69.0 | 81.0 | 76.0 | 86.0 | 72.0 |
| grab_roller | 94.0 | 46.0 | 98.0 | 71.0 | 68.0 | 48.0 | 91.0 | 64.0 | 89.0 | 63.0 |
| handover_block | 16.0 | 0.0 | 5.0 | 2.0 | 18.0 | 4.0 | 39.0 | 9.0 | 36.0 | 9.0 |
| handover_mic | 26.0 | 4.0 | 69.0 | 13.0 | 86.0 | 34.0 | 83.0 | 17.0 | 97.0 | 23.0 |
| hanging_mug | 10.0 | 4.0 | 14.0 | 16.0 | 14.0 | 9.0 | 12.0 | 9.0 | 10.0 | 10.0 |
| lift_pot | 8.0 | 2.0 | 93.0 | 28.0 | 84.0 | 14.0 | 95.0 | 44.0 | 93.0 | 44.0 |
| move_can_pot | 28.0 | 0.0 | 47.0 | 50.0 | 30.0 | 40.0 | 42.0 | 85.0 | 47.0 | 65.0 |
| move_pillbottle_pad | 60.0 | 44.0 | 63.0 | 60.0 | 68.0 | 69.0 | 41.0 | 67.0 | 74.0 | 82.0 |
| move_playingcard_away | 90.0 | 52.0 | 74.0 | 58.0 | 88.0 | 92.0 | 88.0 | 42.0 | 92.0 | 63.0 |
| move_stapler_pad | 22.0 | 6.0 | 23.0 | 19.0 | 34.0 | 21.0 | 30.0 | 13.0 | 43.0 | 31.0 |
| open_laptop | 68.0 | 20.0 | 77.0 | 61.0 | 74.0 | 56.0 | 64.0 | 46.0 | 75.0 | 58.0 |
| open_microwave | 26.0 | 8.0 | 77.0 | 59.0 | 42.0 | 29.0 | 37.0 | 32.0 | 50.0 | 25.0 |
| pick_diverse_bottles | 56.0 | 18.0 | 60.0 | 37.0 | 58.0 | 53.0 | 71.0 | 50.0 | 70.0 | 50.0 |
| pick_dual_bottles | 82.0 | 14.0 | 91.0 | 59.0 | 80.0 | 77.0 | 83.0 | 55.0 | 97.0 | 54.0 |
| place_a2b_left | 60.0 | 10.0 | 40.0 | 55.0 | 58.0 | 42.0 | 54.0 | 48.0 | 73.0 | 44.0 |
| place_a2b_right | 58.0 | 14.0 | 42.0 | 49.0 | 52.0 | 41.0 | 53.0 | 51.0 | 64.0 | 53.0 |
| place_bread_basket | 68.0 | 44.0 | 75.0 | 55.0 | 72.0 | 64.0 | 76.0 | 59.0 | 79.0 | 62.0 |
| place_bread_skillet | 86.0 | 46.0 | 76.0 | 41.0 | 76.0 | 50.0 | 80.0 | 61.0 | 87.0 | 53.0 |
| place_burger_fries | 90.0 | 78.0 | 96.0 | 83.0 | 98.0 | 81.0 | 97.0 | 88.0 | 96.0 | 88.0 |
| place_can_basket | 40.0 | 0.0 | 49.0 | 25.0 | 38.0 | 9.0 | 34.0 | 15.0 | 34.0 | 16.0 |
| place_cans_plasticbox | 94.0 | 74.0 | 90.0 | 68.0 | 54.0 | 68.0 | 97.0 | 59.0 | 70.0 | 59.0 |
| place_container_plate | 96.0 | 58.0 | 86.0 | 76.0 | 92.0 | 79.0 | 93.0 | 73.0 | 93.0 | 81.0 |
| place_dual_shoes | 54.0 | 12.0 | 50.0 | 30.0 | 34.0 | 27.0 | 35.0 | 17.0 | 35.0 | 19.0 |

Findings
- Mixing Ego2R-synthesized data with real robot data at a 1:1 ratio reached 53.5% success on RoboTwin's randomized setting, 2.6 points above robot-only pretraining, while still keeping 68.1% on the clean setting.
- At the 1:1 mixing ratio, per-factor visual improvements were background +4, lighting +8, robot color +6, and camera offset +6 compared to robot-only pretraining.
- Unseen-object generalization improved from 29% to 40% (+11 at a 3:1 ratio), and success on paraphrased instructions reached 69% at the 1:1 ratio.
- In cross-embodiment tests swapping in different robot arms, ARX improved from 44% to 51% and UR5 peaked at 31% (3:1 ratio), while Franka Panda stayed below 7%.
- On the real ARX ACone robot across five tasks, the model trained with both teleoperated demonstrations and the pipeline-converted egocentric data outperformed the robot-only model on every task, with the largest gains on placing blocks and inserting a screw.
Where it can be used
- Teams that lack large amounts of real demonstration data for a new robot arm or new task could consider converting existing human manipulation footage into supplementary pretraining data.
- Service robot settings where backgrounds, lighting, objects, and phrasing change frequently could consider using diverse human video as an additional training source to improve robustness.
- Projects supporting multiple robot arm models at once could reference this approach of rendering a single human video into several robot embodiments in parallel.

Limits and open work
- Hand motions are converted only into simple parallel-jaw gripper movements, so the approach does not directly extend to dexterous multi-fingered manipulation.
- Visual synthesis relies on inpainting and depth-aware compositing, which can introduce artifacts under heavy occlusion or difficult lighting conditions.
- Evaluation is confined to the task scope of RoboTwin2.0, so it remains untested whether the benefits hold for a broader range of tasks and robot configurations.
- For robots with a large kinematic gap from the training embodiments, such as Franka Panda, cross-embodiment success stayed below 7%, showing the gains are not uniform across all robot types.
Why it matters
Collecting real robot demonstrations is expensive and slow, so being able to automatically convert the enormous existing supply of first-person human manipulation videos into usable robot training data could substantially cut data-collection costs. The fact that this specifically improves robustness to unseen backgrounds, robot types, objects, and phrasing matters directly for deploying robots in real, unpredictable environments.
Terms in this paper
- Egocentric video · Video recorded from a person's own point of view, typically with a head- or body-mounted camera, showing their hands manipulating objects
- VLA model (vision-language-action model) · An AI model that takes camera images and language instructions as input and directly outputs the robot actions to perform
- Inverse kinematics (IK) · A method for computing the joint angles a robot arm needs so its end-effector reaches a desired position and orientation
- Out-of-distribution (OOD) generalization · The ability to keep performing well under backgrounds, objects, robot types, or phrasing that were not seen during training
- End-effector (EEF) · The part at the tip of a robot arm, such as a gripper, that actually makes contact with objects
Original abstract (English)
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that con
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI assistants would rather double-check facts than ask you a question, even when asking is the right call
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
Latest from METAL LAB
- Sakana AI Signs Deal With Japan's Defense Ministry for Intelligence Analysis AI Trial
- Hermes Agent builds its own skills the more you use it
- Is training AI on copyrighted books legal? Courts are still fighting it out
- Chinese gray market sells Anthropic Claude tokens at 10% of list price
- Even the Best AI Runaway Response Plan Among Five Major Labs Scores Only 3
Figures: Ye Wang et al., arXiv:2608.02580, arxiv-nonexclusive