ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
A capture rig records a person's first-person view, whole-body motion, hand motion, sound, and touch all at once to build training data for robots
ACE turns real homes into recording studios, capturing a person's egocentric video, multi-view third-person video, full-body and hand motion, object 6-DoF trajectories, audio, and tactile pressure all synchronized on the same timeline and spatial frame while they cook or tidy up. The resulting ACE-Data-0 dataset spans 150 hours, 200 task categories, 50 participants, 17M frames, and 75,000 interaction episodes. The team built a three-tier benchmark, testing touch prediction, body/hand pose recovery, and hand-object interaction estimation, evaluating more than 30 existing methods.
METAL LAB explanatory visual
How ACE captures and structures its data
Evidence statusMeasured results and planned work
- Two capture configurationsA table-scale rig (8 close-range cameras + 16 mocap cameras) for hand manipulation and a room-scale rig (8 wide-baseline cameras + 12 mocap cameras) for whole-body movement, run in parallel
- Synchronized multi-sensor recordingEgocentric headset video, multi-view exocentric video, full-body/hand motion capture, object 6-DoF trajectories, audio, and tactile glove pressure all aligned on one timeline and spatial frame
- Automatic annotationCamera calibration, pose projections, object mesh and trajectories, and text descriptions generated automatically from tracked ground-truth states
- Three-tier benchmark evaluationTactile signal prediction (signal level), body/hand pose recovery (component level), and ego/exo hand-object interaction estimation (interaction level) tested across 30+ methods
What they did
- Existing datasets were fragmented: large egocentric datasets like Ego4D and EPIC-Kitchens lack ground-truth body/object motion, while motion-captured datasets like BEHAVE, GRAB, and ARCTIC lack egocentric view, audio, or tactile signals.
- To fix this, the authors built two complementary setups: a table-scale rig (8 close-range cameras plus 16 motion-capture cameras) for fine hand manipulation, and a room-scale rig (8 wide-baseline cameras plus 12 motion-capture cameras) for whole-body movement across a furnished apartment.
- Participants wore a motion-capture suit, tactile gloves, and a four-fisheye-camera headset while every sensor stream was synchronized in hardware or software and registered into a shared spatial frame.
- The resulting ACE-Data-0 dataset comes with automatically derived annotations: camera calibration, body/hand pose projected onto every camera frame, per-object 3D mesh and 6-DoF pose, and text descriptions of events.
- Using this data, the researchers built a three-level benchmark - tactile signal prediction, body/hand pose estimation, and hand-object interaction estimation from ego and exo views - and evaluated more than 30 state-of-the-art methods.

| Dataset | Year | Hours | #Frames | #Subj | #Obj | #Tasks | Ego | #Exo | Body | Hand | Obj. 6D | Tactile | Audio | Sync | Setup | LH |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Egocentric video | ||||||||||||||||
| Ego4D [33] | 2022 | 3670 | – | 931 | – | Open | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | – | In-the-wild | ✓ |
| EPIC-KITCHENS-100 [19] | 2021 | 100 | 20M | 37 | – | Open | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | – | Kitchens | ✓ |
| HoloAssist [95] | 2023 | 166 | – | 222 | 16 | 20 | ✓ | ✗ | ✗ | (✓) | ✗ | ✗ | ✓ | – | Desktop | ✓ |
| Ego-Exo4D [34] | 2024 | 1286 | – | 740 | – | 8 Dom. | ✓ | 4–5 | (✓) | ✗ | ✗ | ✗ | ✓ | ✓ | In-the-wild | ✓ |
| EgoLife [100] | 2025 | 266 | – | 6 | – | Open | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | Shared house | ✓ |
| HD-EPIC [73] | 2025 | 41 | 4.46M | 9 | – | 69 | ✓ | ✗ | ✗ | ✗ | (✓) | ✗ | ✓ | – | Home kitchens | ✓ |
| Hand–object interaction | ||||||||||||||||
| ContactPose [7] | 2020 | – | 2.9M | 50 | 25 | 2 | ✗ | 3 | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | Table-top | ✗ |
| HO-3D [36] | 2020 | – | 78K | 10 | 10 | – | ✗ | 1–5 | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | Table-top | ✗ |
| GRAB [88] | 2020 | – | 1.6M | 10 | 51 | 4 | ✗ | ✗ | ✓ | ✓ | ✓ | (✓) | ✗ | – | Mocap lab | ✗ |
| DexYCB [15] | 2021 | – | 582K | 10 | 20 | 1 | ✗ | 8 | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | Table-top | ✗ |
| H2O [52] | 2021 | – | 571K | 4 | 8 | 36 | ✓ | 4 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| OakInk [101] | 2022 | – | 230K | 12 | 100 | 5 | ✗ | 4 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| HOI4D [61] | 2022 | 22 | 2.4M | 9 | 800 | 54 | ✓ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | – | Indoor rooms | ✗ |
| ARCTIC [22] | 2023 | – | 2.1M | 10 | 11 | 2 | ✓ | 8 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Mocap lab | ✗ |
| TACO [60] | 2024 | – | 5.2M | 14 | 196 | 151 | ✓ | 12 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| HOT3D [3] | 2024 | 13.9 | 3.7M | 19 | 33 | Open | ✓ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | – | Lab rooms | ✗ |
| OakInk2 [110] | 2024 | – | 4.0M | 9 | 75 | 150 | ✓ | 3 | (✓) | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | (✓) |
| GigaHands [26] | 2025 | 34 | 183M | 56 | 417 | Open | ✗ | 51 | ✗ | (✓) | (✓) | ✗ | ✗ | ✗ | Table-top | ✗ |
| Full-body HOI, human-scene interaction, and daily motion | ||||||||||||||||
| BEHAVE [5] | 2022 | – | 15K | 8 | 20 | – | ✗ | 4 | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | Lab rooms | ✗ |
| InterCap [41] | 2022 | – | 67K | 10 | 10 | – | ✗ | 6 | ✓ | (✓) | ✓ | ✗ | ✗ | ✓ | Lab room | ✗ |
| CHAIRS [44] | 2022 | 17.3 | – | 46 | 81 | 32 | ✗ | 4 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Lab room | – |
| EgoBody [113] | 2022 | – | 220K | 36 | – | 5 Cat. | ✓ | 3–5 | (✓) | (✓) | ✗ | ✗ | ✗ | ✓ | Indoor rooms | ✗ |
| Aria Digital Twin [69] | 2023 | 6.6 | – | – | 398 | Open | ✓ | ✗ | (✓) | ✗ | ✓ | ✗ | ✓ | – | Apartment, office | ✗ |
| OMOMO [54] | 2023 | 10 | – | 17 | 15 | – | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | – | Lab room | ✗ |
| HIMO [65] | 2024 | – | 4.1M | 34 | 53 | – | ✓ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | – | Mocap lab | ✗ |
| TRUMANS [45] | 2024 | 15 | 1.6M | 7 | 20 | – | ✗ | ✗ | ✓ | ✗ | (✓) | ✗ | ✗ | – | Scene mockups | ✗ |
| ParaHome [50] | 2024 | 8.1 | – | 38 | 22 | – | ✗ | 70 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Home room | ✓ |

| Device | Role | Qty | Resolution / rate | Notes |
|---|---|---|---|---|
| OptiTrack PrimeX 22 | Optical motion capture | 28 | 2048×1088, 60 Hz | IR tracking: 41 body markers, objects, ego rig |
| ZED One | Exocentric RGB capture | 8 | 1920×1080, 30 FPS | GMSL2 to a single Jetson Orin host; shared frame trigger |
| GoPro | Exocentric RGB capture | 8 | 1920×1080, 30 FPS | Rigidly mounted on stands; audio-triggered recording |
| ACE-Ego-Head-V02 Lite | Egocentric capture | 4 cameras | 4×1088×1280, 20 FPS | One headset: front/back fisheye pairs, IMU, 5 markers |
| Manus | Hand pose | 2 gloves | 60 Hz | Per-finger articulation, both hands |
| ACE-Sense-Glove Lite | Contact pressure | 2 gloves | – | Full-palm pressure map, both hands |
| Jetson Orin | Recording host | 1 | – | Ingests all ZED One streams; NatNet-coordinated |
| Motive host (PC) | Mocap host, sync reference | 2 | – | Renders the optical clock for cross-system sync |

| Method | Temp Acc. ↑ | C-IoU ↑ | V-IoU ↑ | CoP ↓ |
|---|---|---|---|---|
| PressureVision [32] | 0.0093 | 0.0007 | 0.0000 | 10.9807 |
| EgoPressureDiff [109] | 0.2912 | 0.0197 | 0.0025 | 8.5152 |
| TouchAnything [115] | 0.7095 | 0.1646 | 0.1357 | 6.5846 |

| Method | PA-MPJPE ↓ | PA-PVE ↓ | MPJPE ↓ | PVE ↓ | WA-MPJPE ↓ |
|---|---|---|---|---|---|
| — Single-view exocentric, per-frame — | |||||
| Multi-HMR [4] | 110.5 | – | 115.3 | – | – |
| Multi-HMR2 [24] | 85.5 | – | 92.8 | – | – |
| SAM-3D-Body [103] | 71.0 | – | 76.4 | – | – |
| PyMAF-X [111] | 94.7 | 177.5 | 99.9 | 177.0 | – |
| PARE [51] | 98.7 | 132.3 | 102.0 | 134.8 | – |
| CameraHMR [70] | 78.7 | 98.7 | 85.7 | 105.2 | – |
| OSX [59] | 59.2 | 73.7 | 62.5 | 76.6 | – |
| — Single-view exocentric, temporal — | |||||
| SMPLer-X [11] | 57.0 | 70.2 | 59.8 | 71.7 | – |
| SMPLest-X [107] | 55.7 | 65.8 | 58.8 | 67.6 | – |
| Humans-in-4D [31] | 77.1 | 94.4 | 79.6 | 96.4 | – |
| GVHMR [84] | 88.0 | 104.1 | 96.3 | 112.0 | 217.1 |
| WHAM [85] | 131.1 | 160.2 | 144.0 | 166.4 | 243.1 |
| EasyMoCap [21, 86] | 103.4 | 155.8 | 106.0 | 156.1 | 256.2 |
| — Single-view exocentric, scene-aware — | |||||
| Phy-SIC [68] | 64.3 | 78.9 | 71.3 | 84.6 | – |
| UniSH [55] | 80.0 | 111.2 | 82.8 | 114.0 | 196.3 |
| JOSH [62] | 64.0 | 89.2 | 69.9 | 95.0 | 245.1 |
| Human3R [16] | 60.1 | 75.4 | 63.5 | 80.6 | 180.2 |
| — Multi-view exocentric — | |||||
| MAMMA [18] | 70.1 | 86.8 | 74.8 | 89.6 | 230.8 |
| U-HMR [57] | 134.4 | 176.0 | 150.2 | 179.2 | – |
| HSfM [67] | 92.8 | 133.8 | 93.8 | 134.8 | – |
| — Egocentric — | |||||
| EgoEgo [53] | 159.6 | 245.3 | 163.4 | 263.9 | 306.2 |
| EgoAllo [106] | 131.7 | 196.8 | 147.9 | 220.3 | 252.2 |

| Method | PA-MPJPE ↓ | MPJPE ↓ | F@5 ↑ | F@15 ↑ | AUCJ ↑ | Traj. err. ↓ |
|---|---|---|---|---|---|---|
| WildHands [76] | 11.2 | 12.6 | 0.175 | 0.774 | 0.776 | – |
| Dyn-HaMR [108] | 18.9 | 21.1 | 0.042 | 0.413 | 0.624 | 98.2 |
| HaWoR [112] | 13.8 | 17.4 | 0.130 | 0.666 | 0.729 | 102.1 |

| Method | PA-MPJPE ↓ | MPJPE ↓ | F@5 ↑ | F@15 ↑ | AUCJ ↑ | Traj. err. ↓ |
|---|---|---|---|---|---|---|
| HORT [17] | 10.8 | 12.3 | 0.276 | 0.776 | 0.784 | – |
| HaMeR [72] | 9.6 | 10.4 | 0.281 | 0.848 | 0.812 | – |
| HaPTIC [105] | 10.0 | 10.7 | 0.247 | 0.838 | 0.804 | 63.0 |
| WiLoR [75] | 9.1 | 9.9 | 0.313 | 0.854 | 0.819 | – |
| OmniHands [58] | 10.7 | 11.5 | 0.211 | 0.814 | 0.791 | – |

Findings
- For egocentric hand-motion estimation, the per-frame method WildHands achieved the best articulation accuracy, while video-based methods Dyn-HaMR and HaWoR showed large world-frame trajectory errors of 98-102 mm, likely due to camera-motion estimation errors.
- For exocentric hand-motion estimation, five methods had similar articulation accuracy with PA-MPJPE ranging 9.1-10.8 mm; WiLoR performed best at 9.1 mm PA-MPJPE with F@5 of 0.313 and AUCJ of 0.819, and HaMeR followed closely at 9.6 mm.
- HaPTIC, using a fixed exocentric camera, achieved a trajectory error of 63 mm, substantially lower than the 98-102 mm from egocentric world-space methods, indicating a static camera provides a more stable reference frame.
- HORT, which jointly models the manipulated object, achieved PA-MPJPE of 10.8 mm on manipulation frames, comparable to other methods, showing joint object modeling neither clearly helps nor hurts articulation accuracy.
- Comparing egocentric versus exocentric views directly, exocentric methods outperformed egocentric ones in both pose and trajectory accuracy, attributed to the stability of a fixed camera coordinate frame versus head-motion-derived frames.

Where it can be used
- Serving as synchronized vision-motion-tactile training data for imitation learning or policy learning in household or humanoid robots
- Acting as a benchmark to test models that estimate hand motion or tactile signals from egocentric video alone
- Providing a foundation for research combining egocentric and exocentric views for hand-object interaction understanding

Limits and open work
- Data comes from only two home sites, limiting variation in layout, furnishing, and lighting.
- Ground truth is limited to instrumented entities: tracked objects must be pre-scanned and marker-equipped, so state changes of articulated mechanisms, fluids, or deformable materials are not annotated.
- The motion-capture suit, gloves, headset, and markers remain visible in recordings, potentially introducing dataset-specific visual cues.
- Combining egocentric and exocentric streams, feeding measured headset motion into egocentric methods, or supplying object pose as auxiliary input are proposed future experiments not yet performed.

Why it matters
Teaching robots or AI to manipulate objects like humans do requires data showing how vision, motion, sound, and touch unfold together at the same moment, which fragmented prior datasets could not provide. ACE offers a way to record all of this simultaneously in real homes, laying groundwork for imitation learning and world-model research in robotics.

Terms in this paper
- Egocentric / Exocentric video · First-person video from a wearable camera versus third-person video from cameras fixed in the environment
- 6-DoF pose · An object's position (three directions) and rotation (three angles) combined into six values describing its full 3D pose
- MANO · A standard 3D model representing the joints and shape of a human hand
- PA-MPJPE · A joint-position error metric computed after aligning scale, position, and rotation, mainly reflecting finger articulation accuracy
- Motion capture (OptiTrack) · A system that tracks markers attached to the body or objects using multiple cameras to record precise 3D motion

Original abstract (English)
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environme
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI assistants would rather double-check facts than ask you a question, even when asking is the right call
- SynFlow: A Multidimensional Diachronic Semantic Analysis ToolkitAn open-source tool that breaks down how a word's meaning changed, not just that it changed
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
Latest from METAL LAB
- Sakana AI Signs Deal With Japan's Defense Ministry for Intelligence Analysis AI Trial
- Hermes Agent builds its own skills the more you use it
- Is training AI on copyrighted books legal? Courts are still fighting it out
- Chinese gray market sells Anthropic Claude tokens at 10% of list price
- Even the Best AI Runaway Response Plan Among Five Major Labs Scores Only 3
Figures: Yukang Cao et al., arXiv:2607.28625, arxiv-nonexclusive