AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

arXiv:2607.286252026-07-29

A capture rig records a person's first-person view, whole-body motion, hand motion, sound, and touch all at once to build training data for robots

ACE turns real homes into recording studios, capturing a person's egocentric video, multi-view third-person video, full-body and hand motion, object 6-DoF trajectories, audio, and tactile pressure all synchronized on the same timeline and spatial frame while they cook or tidy up. The resulting ACE-Data-0 dataset spans 150 hours, 200 task categories, 50 participants, 17M frames, and 75,000 interaction episodes. The team built a three-tier benchmark, testing touch prediction, body/hand pose recovery, and hand-object interaction estimation, evaluating more than 30 existing methods.

METAL LAB explanatory visual

How ACE captures and structures its data

Evidence statusMeasured results and planned work

  1. Two capture configurationsA table-scale rig (8 close-range cameras + 16 mocap cameras) for hand manipulation and a room-scale rig (8 wide-baseline cameras + 12 mocap cameras) for whole-body movement, run in parallel
  2. Synchronized multi-sensor recordingEgocentric headset video, multi-view exocentric video, full-body/hand motion capture, object 6-DoF trajectories, audio, and tactile glove pressure all aligned on one timeline and spatial frame
  3. Automatic annotationCamera calibration, pose projections, object mesh and trajectories, and text descriptions generated automatically from tracked ground-truth states
  4. Three-tier benchmark evaluationTactile signal prediction (signal level), body/hand pose recovery (component level), and ego/exo hand-object interaction estimation (interaction level) tested across 30+ methods
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. Existing datasets were fragmented: large egocentric datasets like Ego4D and EPIC-Kitchens lack ground-truth body/object motion, while motion-captured datasets like BEHAVE, GRAB, and ARCTIC lack egocentric view, audio, or tactile signals.
  2. To fix this, the authors built two complementary setups: a table-scale rig (8 close-range cameras plus 16 motion-capture cameras) for fine hand manipulation, and a room-scale rig (8 wide-baseline cameras plus 12 motion-capture cameras) for whole-body movement across a furnished apartment.
  3. Participants wore a motion-capture suit, tactile gloves, and a four-fisheye-camera headset while every sensor stream was synchronized in hardware or software and registered into a shared spatial frame.
  4. The resulting ACE-Data-0 dataset comes with automatically derived annotations: camera calibration, body/hand pose projected onto every camera frame, per-object 3D mesh and 6-DoF pose, and text descriptions of events.
  5. Using this data, the researchers built a three-level benchmark - tactile signal prediction, body/hand pose estimation, and hand-object interaction estimation from ego and exo views - and evaluated more than 30 state-of-the-art methods.
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Table 1: Comparison of ACE-Data-0 with related datasets, grouped by category. ✓: provided with measured (mocap-grade) ground truth; (✓): provided but estimated, pseudo-labeled, or device-tracked; ✓: partially provided (limited coverage or a subset); ✗: capability absent; –: number not applicable or not reported; a ✓in the #Exo column denotes that exocentric video is provided without a reported camera count; Sync: synchronization across modality families (ego video, exo video, motion, audio, tactile); a ✓requires at least two families, chiefly ego-exo or video-motion, recorded simultaneously and aligned, while ✓denotes that the synchronization is conducted only among one type of viewpoint, and ‘–’ means that only one perspective is captured. LH: long-horizon (goal-directed activities of minutes or longer). Setup: capture environment. In the #Tasks column, dom., cat., scen., and skills denote domains, interaction categories, scenarios, and skills, following each dataset’s own task organization.
DatasetYearHours#Frames#Subj#Obj#TasksEgo#ExoBodyHandObj. 6DTactileAudioSyncSetupLH
Egocentric video
Ego4D [33]20223670931OpenIn-the-wild
EPIC-KITCHENS-100 [19]202110020M37OpenKitchens
HoloAssist [95]20231662221620(✓)Desktop
Ego-Exo4D [34]202412867408 Dom.4–5(✓)In-the-wild
EgoLife [100]20252666OpenShared house
HD-EPIC [73]2025414.46M969(✓)Home kitchens
Hand–object interaction
ContactPose [7]20202.9M502523Table-top
HO-3D [36]202078K10101–5Table-top
GRAB [88]20201.6M10514(✓)Mocap lab
DexYCB [15]2021582K102018Table-top
H2O [52]2021571K48364Table-top
OakInk [101]2022230K1210054Table-top
HOI4D [61]2022222.4M980054Indoor rooms
ARCTIC [22]20232.1M101128Mocap lab
TACO [60]20245.2M1419615112Table-top
HOT3D [3]202413.93.7M1933OpenLab rooms
OakInk2 [110]20244.0M9751503(✓)Table-top(✓)
GigaHands [26]202534183M56417Open51(✓)(✓)Table-top
Full-body HOI, human-scene interaction, and daily motion
BEHAVE [5]202215K8204Lab rooms
InterCap [41]202267K10106(✓)Lab room
CHAIRS [44]202217.34681324Lab room
EgoBody [113]2022220K365 Cat.3–5(✓)(✓)Indoor rooms
Aria Digital Twin [69]20236.6398Open(✓)Apartment, office
OMOMO [54]2023101715Lab room
HIMO [65]20244.1M3453Mocap lab
TRUMANS [45]2024151.6M720(✓)Scene mockups
ParaHome [50]20248.1382270Home room
(b) Home floor plan and camera setup of site II.
(b) Home floor plan and camera setup of site II.
Table 2: Capture hardware of ACE for synchronized multi-modal recording. Quantities are the totals for both configurations combined; a dash marks an entry that does not apply to that device.
DeviceRoleQtyResolution / rateNotes
OptiTrack PrimeX 22Optical motion capture282048×1088, 60 HzIR tracking: 41 body markers, objects, ego rig
ZED OneExocentric RGB capture81920×1080, 30 FPSGMSL2 to a single Jetson Orin host; shared frame trigger
GoProExocentric RGB capture81920×1080, 30 FPSRigidly mounted on stands; audio-triggered recording
ACE-Ego-Head-V02 LiteEgocentric capture4 cameras4×1088×1280, 20 FPSOne headset: front/back fisheye pairs, IMU, 5 markers
ManusHand pose2 gloves60 HzPer-finger articulation, both hands
ACE-Sense-Glove LiteContact pressure2 glovesFull-palm pressure map, both hands
Jetson OrinRecording host1Ingests all ZED One streams; NatNet-coordinated
Motive host (PC)Mocap host, sync reference2Renders the optical clock for cross-system sync
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Table 3: Tactile estimation from ego-view on ACE-Data-0.
MethodTemp Acc. ↑C-IoU ↑V-IoU ↑CoP ↓
PressureVision [32]0.00930.00070.000010.9807
EgoPressureDiff [109]0.29120.01970.00258.5152
TouchAnything [115]0.70950.16460.13576.5846
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Table 4: Human motion estimation on ACE-Data-0. All metrics in mm.
MethodPA-MPJPE ↓PA-PVE ↓MPJPE ↓PVE ↓WA-MPJPE ↓
— Single-view exocentric, per-frame —
Multi-HMR [4]110.5115.3
Multi-HMR2 [24]85.592.8
SAM-3D-Body [103]71.076.4
PyMAF-X [111]94.7177.599.9177.0
PARE [51]98.7132.3102.0134.8
CameraHMR [70]78.798.785.7105.2
OSX [59]59.273.762.576.6
— Single-view exocentric, temporal —
SMPLer-X [11]57.070.259.871.7
SMPLest-X [107]55.765.858.867.6
Humans-in-4D [31]77.194.479.696.4
GVHMR [84]88.0104.196.3112.0217.1
WHAM [85]131.1160.2144.0166.4243.1
EasyMoCap [21, 86]103.4155.8106.0156.1256.2
— Single-view exocentric, scene-aware —
Phy-SIC [68]64.378.971.384.6
UniSH [55]80.0111.282.8114.0196.3
JOSH [62]64.089.269.995.0245.1
Human3R [16]60.175.463.580.6180.2
— Multi-view exocentric —
MAMMA [18]70.186.874.889.6230.8
U-HMR [57]134.4176.0150.2179.2
HSfM [67]92.8133.893.8134.8
— Egocentric —
EgoEgo [53]159.6245.3163.4263.9306.2
EgoAllo [106]131.7196.8147.9220.3252.2
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Table 5: Hand motion estimation from ego-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
WildHands [76]11.212.60.1750.7740.776
Dyn-HaMR [108]18.921.10.0420.4130.62498.2
HaWoR [112]13.817.40.1300.6660.729102.1
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Table 6: Hand motion estimation from exo-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
HORT [17]10.812.30.2760.7760.784
HaMeR [72]9.610.40.2810.8480.812
HaPTIC [105]10.010.70.2470.8380.80463.0
WiLoR [75]9.19.90.3130.8540.819
OmniHands [58]10.711.50.2110.8140.791
Figure 7: Examples of different task categories captured in ACE-Data-0.
Figure 7: Examples of different task categories captured in ACE-Data-0.

Findings

  • For egocentric hand-motion estimation, the per-frame method WildHands achieved the best articulation accuracy, while video-based methods Dyn-HaMR and HaWoR showed large world-frame trajectory errors of 98-102 mm, likely due to camera-motion estimation errors.
  • For exocentric hand-motion estimation, five methods had similar articulation accuracy with PA-MPJPE ranging 9.1-10.8 mm; WiLoR performed best at 9.1 mm PA-MPJPE with F@5 of 0.313 and AUCJ of 0.819, and HaMeR followed closely at 9.6 mm.
  • HaPTIC, using a fixed exocentric camera, achieved a trajectory error of 63 mm, substantially lower than the 98-102 mm from egocentric world-space methods, indicating a static camera provides a more stable reference frame.
  • HORT, which jointly models the manipulated object, achieved PA-MPJPE of 10.8 mm on manipulation frames, comparable to other methods, showing joint object modeling neither clearly helps nor hurts articulation accuracy.
  • Comparing egocentric versus exocentric views directly, exocentric methods outperformed egocentric ones in both pose and trajectory accuracy, attributed to the stability of a fixed camera coordinate frame versus head-motion-derived frames.
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.

Where it can be used

  • Serving as synchronized vision-motion-tactile training data for imitation learning or policy learning in household or humanoid robots
  • Acting as a benchmark to test models that estimate hand motion or tactile signals from egocentric video alone
  • Providing a foundation for research combining egocentric and exocentric views for hand-object interaction understanding
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.

Limits and open work

  • Data comes from only two home sites, limiting variation in layout, furnishing, and lighting.
  • Ground truth is limited to instrumented entities: tracked objects must be pre-scanned and marker-equipped, so state changes of articulated mechanisms, fluids, or deformable materials are not annotated.
  • The motion-capture suit, gloves, headset, and markers remain visible in recordings, potentially introducing dataset-specific visual cues.
  • Combining egocentric and exocentric streams, feeding measured headset motion into egocentric methods, or supplying object pose as auxiliary input are proposed future experiments not yet performed.
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

Why it matters

Teaching robots or AI to manipulate objects like humans do requires data showing how vision, motion, sound, and touch unfold together at the same moment, which fragmented prior datasets could not provide. ACE offers a way to record all of this simultaneously in real homes, laying groundwork for imitation learning and world-model research in robotics.

Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

Terms in this paper

  • Egocentric / Exocentric video · First-person video from a wearable camera versus third-person video from cameras fixed in the environment
  • 6-DoF pose · An object's position (three directions) and rotation (three angles) combined into six values describing its full 3D pose
  • MANO · A standard 3D model representing the joints and shape of a human hand
  • PA-MPJPE · A joint-position error metric computed after aligning scale, position, and rotation, mainly reflecting finger articulation accuracy
  • Motion capture (OptiTrack) · A system that tracks markers attached to the body or objects using multiple cameras to record precise 3D motion
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.

Original abstract (English)

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environme

Authors · Yukang Cao

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Yukang Cao et al., arXiv:2607.28625, arxiv-nonexclusive