工作日早上 7 点读 AI,周日早上 8 点读周报订阅邮件

METAL LAB

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

arXiv:2607.286252026-07-29

一套设备同时记录一个人做家务时的第一人称视角、全身动作、手部动作、声音和触觉,用来做机器人训练数据

ACE把真实的家变成录制棚,在同一时间轴和空间坐标下,同步记录一个人做饭或整理房间时的第一人称视频、多视角第三人称视频、全身和手部动作、物体六自由度轨迹、声音和触觉压力。由此收集到的ACE-Data-0数据集包含150小时录像、200种任务、50名参与者、1700万帧和7.5万个交互片段。研究团队据此建立了三层基准测试,评估了30多种现有方法在触觉预测、身体/手部姿态恢复、手物交互估计上的表现。

METAL LAB 解读图

ACE数据采集与结构

证据状态实测结果与计划中的工作并存

  1. 两套互补采集配置桌面级(8台近距相机+16台动捕相机)用于手部操作,房间级(8台广基线相机+12台动捕相机)用于全身移动,两者并行运作
  2. 多传感器同步录制第一人称头戴视频、多视角第三人称视频、全身及手部动捕、物体六自由度轨迹、音频、触觉手套压力数据全部对齐到同一时间轴和空间坐标系
  3. 自动标注生成根据追踪到的真实物理状态自动生成相机标定、姿态投影、物体网格与轨迹、事件文字描述
  4. 三层基准评测依次评测触觉信号预测(信号层)、身体/手部姿态恢复(组件层)、第一/第三人称手物交互估计(交互层),覆盖30余种方法
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 现有数据集存在割裂问题:Ego4D、EPIC-Kitchens等大规模第一人称数据集缺少身体和物体的真实动作标注,而BEHAVE、GRAB、ARCTIC等动捕数据集又缺少第一人称视角、声音或触觉信号。
  2. 为此团队搭建了两套互补装置:桌面级配置(8台近距离相机加16台动捕相机)用于精细手部操作,房间级配置(8台广基线相机加12台动捕相机)用于覆盖整间公寓的全身移动。
  3. 参与者穿戴动捕服、触觉手套,并佩戴带有四个鱼眼镜头的头戴设备,所有传感器数据流通过硬件或软件方式同步,并统一注册到同一空间坐标系。
  4. 最终得到的ACE-Data-0数据集附带自动生成的丰富标注,包括相机标定、投射到每个摄像机画面上的身体和手部姿态、每个物体的三维网格与六自由度姿态,以及事件文字描述。
  5. 研究团队据此建立了三层基准:触觉信号预测、人体/手部姿态估计、以及从第一人称和第三人称视频中估计手物交互动作,并评估了30多种现有先进方法。
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Table 1: Comparison of ACE-Data-0 with related datasets, grouped by category. ✓: provided with measured (mocap-grade) ground truth; (✓): provided but estimated, pseudo-labeled, or device-tracked; ✓: partially provided (limited coverage or a subset); ✗: capability absent; –: number not applicable or not reported; a ✓in the #Exo column denotes that exocentric video is provided without a reported camera count; Sync: synchronization across modality families (ego video, exo video, motion, audio, tactile); a ✓requires at least two families, chiefly ego-exo or video-motion, recorded simultaneously and aligned, while ✓denotes that the synchronization is conducted only among one type of viewpoint, and ‘–’ means that only one perspective is captured. LH: long-horizon (goal-directed activities of minutes or longer). Setup: capture environment. In the #Tasks column, dom., cat., scen., and skills denote domains, interaction categories, scenarios, and skills, following each dataset’s own task organization.
DatasetYearHours#Frames#Subj#Obj#TasksEgo#ExoBodyHandObj. 6DTactileAudioSyncSetupLH
Egocentric video
Ego4D [33]20223670931OpenIn-the-wild
EPIC-KITCHENS-100 [19]202110020M37OpenKitchens
HoloAssist [95]20231662221620(✓)Desktop
Ego-Exo4D [34]202412867408 Dom.4–5(✓)In-the-wild
EgoLife [100]20252666OpenShared house
HD-EPIC [73]2025414.46M969(✓)Home kitchens
Hand–object interaction
ContactPose [7]20202.9M502523Table-top
HO-3D [36]202078K10101–5Table-top
GRAB [88]20201.6M10514(✓)Mocap lab
DexYCB [15]2021582K102018Table-top
H2O [52]2021571K48364Table-top
OakInk [101]2022230K1210054Table-top
HOI4D [61]2022222.4M980054Indoor rooms
ARCTIC [22]20232.1M101128Mocap lab
TACO [60]20245.2M1419615112Table-top
HOT3D [3]202413.93.7M1933OpenLab rooms
OakInk2 [110]20244.0M9751503(✓)Table-top(✓)
GigaHands [26]202534183M56417Open51(✓)(✓)Table-top
Full-body HOI, human-scene interaction, and daily motion
BEHAVE [5]202215K8204Lab rooms
InterCap [41]202267K10106(✓)Lab room
CHAIRS [44]202217.34681324Lab room
EgoBody [113]2022220K365 Cat.3–5(✓)(✓)Indoor rooms
Aria Digital Twin [69]20236.6398Open(✓)Apartment, office
OMOMO [54]2023101715Lab room
HIMO [65]20244.1M3453Mocap lab
TRUMANS [45]2024151.6M720(✓)Scene mockups
ParaHome [50]20248.1382270Home room
(b) Home floor plan and camera setup of site II.
(b) Home floor plan and camera setup of site II.
Table 2: Capture hardware of ACE for synchronized multi-modal recording. Quantities are the totals for both configurations combined; a dash marks an entry that does not apply to that device.
DeviceRoleQtyResolution / rateNotes
OptiTrack PrimeX 22Optical motion capture282048×1088, 60 HzIR tracking: 41 body markers, objects, ego rig
ZED OneExocentric RGB capture81920×1080, 30 FPSGMSL2 to a single Jetson Orin host; shared frame trigger
GoProExocentric RGB capture81920×1080, 30 FPSRigidly mounted on stands; audio-triggered recording
ACE-Ego-Head-V02 LiteEgocentric capture4 cameras4×1088×1280, 20 FPSOne headset: front/back fisheye pairs, IMU, 5 markers
ManusHand pose2 gloves60 HzPer-finger articulation, both hands
ACE-Sense-Glove LiteContact pressure2 glovesFull-palm pressure map, both hands
Jetson OrinRecording host1Ingests all ZED One streams; NatNet-coordinated
Motive host (PC)Mocap host, sync reference2Renders the optical clock for cross-system sync
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Table 3: Tactile estimation from ego-view on ACE-Data-0.
MethodTemp Acc. ↑C-IoU ↑V-IoU ↑CoP ↓
PressureVision [32]0.00930.00070.000010.9807
EgoPressureDiff [109]0.29120.01970.00258.5152
TouchAnything [115]0.70950.16460.13576.5846
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Table 4: Human motion estimation on ACE-Data-0. All metrics in mm.
MethodPA-MPJPE ↓PA-PVE ↓MPJPE ↓PVE ↓WA-MPJPE ↓
— Single-view exocentric, per-frame —
Multi-HMR [4]110.5115.3
Multi-HMR2 [24]85.592.8
SAM-3D-Body [103]71.076.4
PyMAF-X [111]94.7177.599.9177.0
PARE [51]98.7132.3102.0134.8
CameraHMR [70]78.798.785.7105.2
OSX [59]59.273.762.576.6
— Single-view exocentric, temporal —
SMPLer-X [11]57.070.259.871.7
SMPLest-X [107]55.765.858.867.6
Humans-in-4D [31]77.194.479.696.4
GVHMR [84]88.0104.196.3112.0217.1
WHAM [85]131.1160.2144.0166.4243.1
EasyMoCap [21, 86]103.4155.8106.0156.1256.2
— Single-view exocentric, scene-aware —
Phy-SIC [68]64.378.971.384.6
UniSH [55]80.0111.282.8114.0196.3
JOSH [62]64.089.269.995.0245.1
Human3R [16]60.175.463.580.6180.2
— Multi-view exocentric —
MAMMA [18]70.186.874.889.6230.8
U-HMR [57]134.4176.0150.2179.2
HSfM [67]92.8133.893.8134.8
— Egocentric —
EgoEgo [53]159.6245.3163.4263.9306.2
EgoAllo [106]131.7196.8147.9220.3252.2
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Table 5: Hand motion estimation from ego-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
WildHands [76]11.212.60.1750.7740.776
Dyn-HaMR [108]18.921.10.0420.4130.62498.2
HaWoR [112]13.817.40.1300.6660.729102.1
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Table 6: Hand motion estimation from exo-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
HORT [17]10.812.30.2760.7760.784
HaMeR [72]9.610.40.2810.8480.812
HaPTIC [105]10.010.70.2470.8380.80463.0
WiLoR [75]9.19.90.3130.8540.819
OmniHands [58]10.711.50.2110.8140.791
Figure 7: Examples of different task categories captured in ACE-Data-0.
Figure 7: Examples of different task categories captured in ACE-Data-0.

研究结果

  • 在第一人称视频的手部动作估计中,逐帧方法WildHands的关节姿态精度最高,而视频类方法Dyn-HaMR和HaWoR的世界坐标轨迹误差高达98-102毫米,原因可能是相机运动估计误差。
  • 在第三人称视频的手部动作估计中,五种方法的关节姿态精度相近,PA-MPJPE在9.1到10.8毫米之间;WiLoR表现最好,PA-MPJPE为9.1毫米,F@5为0.313,AUCJ为0.819,HaMeR紧随其后为9.6毫米。
  • 使用固定第三人称相机的HaPTIC轨迹误差仅为63毫米,远低于第一人称世界坐标方法的98-102毫米,说明固定视角能提供更稳定的参考坐标系。
  • 同时重建被操作物体的HORT在操作帧上的PA-MPJPE为10.8毫米,与其他方法相近,说明联合建模物体既没有明显提升也没有明显损害手部姿态精度。
  • 将第一人称和第三人称方法直接比较后发现,第三人称方法在姿态和轨迹精度上都优于第一人称方法,这归因于固定相机坐标系比由头部运动推算出的坐标系更稳定。
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.

可应用场景

  • 作为同步的视觉-动作-触觉训练数据,用于家庭或人形机器人的模仿学习和策略学习
  • 作为基准测试,检验仅凭第一人称视频估计手部动作或触觉信号的模型表现
  • 为结合第一人称和第三人称视角的手物交互理解研究提供基础数据
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.

局限与待验证事项

  • 目前数据仅来自两处住宅场地,户型、家具和光照的多样性有限。
  • 真实标注仅限于经过预先扫描并安装标记点的物体,门等关节机构、液体或可变形材料的状态变化未被标注。
  • 动捕服、触觉手套、头戴设备和标记点在画面中始终可见,可能带来数据集特有的视觉痕迹。
  • 结合第一人称与第三人称数据流、向第一人称方法提供实测头戴设备运动信息、或将物体姿态作为辅助输入等实验尚未开展,仅作为未来研究方向提出。
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

为什么重要

要让机器人或AI学会像人一样用手操作物体,需要能同时展示视觉、动作、声音和触觉如何在同一时刻交织的数据,而以往割裂的数据集难以满足这一需求。ACE提供了一种在真实家庭环境中同步精确记录这些信号的方法,为模仿学习和世界模型等机器人研究打下基础。

Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

本文术语

  • 第一人称/第三人称视频(egocentric/exocentric) · 佩戴在身上的相机拍到的第一人称画面,与固定在环境中的相机拍到的第三人称画面
  • 六自由度姿态(6-DoF pose) · 用三个方向的位置和三个方向的旋转角度共六个数值来描述物体的完整三维姿态
  • MANO · 一种表示手部关节和形状的标准三维手部模型
  • PA-MPJPE · 在分别对齐尺度、位置和旋转后计算的关节位置误差,主要反映手指姿态的准确度
  • 动作捕捉(OptiTrack) · 通过多台相机追踪身体或物体上贴附的标记点,记录精确三维动作的系统
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.

论文原文摘要(英文)

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environme

作者 · Yukang Cao

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Yukang Cao et al., arXiv:2607.28625, arxiv-nonexclusive