월~금 오전 7시, 일요일 오전 8시 — AI 뉴스와 용어를 보내드립니다메일로 받아보기

METAL LAB

집안일을 하는 사람의 눈, 몸, 손, 소리, 촉감을 동시에 녹화하는 장비로 로봇 학습용 데이터를 만들었다

arXiv:2607.286252026-07-29

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

집안일을 하는 사람의 눈, 몸, 손, 소리, 촉감을 동시에 녹화하는 장비로 로봇 학습용 데이터를 만들었다

ACE는 실제 가정집을 촬영 스튜디오로 바꿔, 사람이 요리하고 정리하는 모습을 1인칭 영상, 3인칭 다중 시점 영상, 전신·손 모션, 물체 위치, 소리, 촉각 압력까지 모두 같은 시간·공간 기준으로 동시에 기록하는 시스템이다. 이렇게 모은 ACE-Data-0은 150시간, 200가지 작업, 50명 참가자, 17M프레임, 7만5000개 상호작용 장면으로 구성된 대규모 데이터셋이다. 연구팀은 이 데이터로 촉각 예측, 몸·손 동작 복원, 손-물체 상호작용 추정이라는 세 단계 벤치마크를 만들어 30여 개 기존 방법을 평가했다.

METAL LAB 해설 도표

ACE 데이터 수집 구조

증거 상태측정 결과와 예정된 검증이 함께 있음

  1. 두 가지 촬영 설치손 조작용 탁상 규모(근접 카메라 8대+모션캡처 16대)와 전신 이동용 방 규모(광각 카메라 8대+모션캡처 12대)를 동시에 운영
  2. 다중 센서 동시 녹화1인칭 헤드셋 영상, 3인칭 다중 시점 영상, 전신·손 모션캡처, 물체 6자유도 궤적, 오디오, 촉각 압력 장갑을 같은 시간·공간 기준으로 기록
  3. 자동 주석 생성추적된 실제 물리 상태로부터 카메라 캘리브레이션, 자세 투영, 물체 메시·궤적, 텍스트 설명을 자동 생성
  4. 3단계 벤치마크 평가촉각 신호 예측(신호), 전신·손 자세 복원(구성요소), 1인칭·3인칭 손-물체 상호작용 추정(상호작용) 순으로 30여 개 방법을 테스트
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 기존 데이터셋들은 1인칭 영상만 있거나(Ego4D, EPIC-Kitchens 등), 정밀한 몸·물체 동작은 있지만 1인칭 시점이나 소리·촉각이 빠져 있는 등(BEHAVE, GRAB, ARCTIC 등) 정보가 조각나 있었다는 문제를 지적했다.
  2. 이를 해결하기 위해 탁상 규모(정밀한 손 조작용, 8대 근접 카메라와 16대 모션캡처 카메라)와 방 규모(전신 이동용, 8대 광각 카메라와 12대 모션캡처 카메라) 두 가지 설치를 만들고, 참가자는 모션캡처 슈트와 촉각 장갑, 4개 어안렌즈가 달린 헤드셋을 착용해 촬영했다.
  3. 모든 센서 신호는 하드웨어 또는 소프트웨어로 시간 동기화되고 공통 좌표계에 정렬되어, 같은 순간의 영상·소리·동작·촉각이 서로 어긋나지 않게 기록된다.
  4. 수집된 ACE-Data-0에는 카메라 캘리브레이션, 전신·손 자세와 각 카메라 화면에의 투영, 물체의 3D 메시와 6자유도 위치, 텍스트 설명 등 풍부한 주석이 자동으로 붙어 있다.
  5. 이 데이터로 촉각 신호 예측, 사람 전신·손 동작 추정, 1인칭·3인칭 영상에서의 손-물체 상호작용 추정이라는 세 단계 벤치마크를 만들어 30여 개 최신 방법을 평가했다.
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Figure 2: Overview of the Ambient Capture Engine. We design two complementary home scenes and capture at two spatial scales, with the equipment installed accordingly.
Table 1: Comparison of ACE-Data-0 with related datasets, grouped by category. ✓: provided with measured (mocap-grade) ground truth; (✓): provided but estimated, pseudo-labeled, or device-tracked; ✓: partially provided (limited coverage or a subset); ✗: capability absent; –: number not applicable or not reported; a ✓in the #Exo column denotes that exocentric video is provided without a reported camera count; Sync: synchronization across modality families (ego video, exo video, motion, audio, tactile); a ✓requires at least two families, chiefly ego-exo or video-motion, recorded simultaneously and aligned, while ✓denotes that the synchronization is conducted only among one type of viewpoint, and ‘–’ means that only one perspective is captured. LH: long-horizon (goal-directed activities of minutes or longer). Setup: capture environment. In the #Tasks column, dom., cat., scen., and skills denote domains, interaction categories, scenarios, and skills, following each dataset’s own task organization.
DatasetYearHours#Frames#Subj#Obj#TasksEgo#ExoBodyHandObj. 6DTactileAudioSyncSetupLH
Egocentric video
Ego4D [33]20223670931OpenIn-the-wild
EPIC-KITCHENS-100 [19]202110020M37OpenKitchens
HoloAssist [95]20231662221620(✓)Desktop
Ego-Exo4D [34]202412867408 Dom.4–5(✓)In-the-wild
EgoLife [100]20252666OpenShared house
HD-EPIC [73]2025414.46M969(✓)Home kitchens
Hand–object interaction
ContactPose [7]20202.9M502523Table-top
HO-3D [36]202078K10101–5Table-top
GRAB [88]20201.6M10514(✓)Mocap lab
DexYCB [15]2021582K102018Table-top
H2O [52]2021571K48364Table-top
OakInk [101]2022230K1210054Table-top
HOI4D [61]2022222.4M980054Indoor rooms
ARCTIC [22]20232.1M101128Mocap lab
TACO [60]20245.2M1419615112Table-top
HOT3D [3]202413.93.7M1933OpenLab rooms
OakInk2 [110]20244.0M9751503(✓)Table-top(✓)
GigaHands [26]202534183M56417Open51(✓)(✓)Table-top
Full-body HOI, human-scene interaction, and daily motion
BEHAVE [5]202215K8204Lab rooms
InterCap [41]202267K10106(✓)Lab room
CHAIRS [44]202217.34681324Lab room
EgoBody [113]2022220K365 Cat.3–5(✓)(✓)Indoor rooms
Aria Digital Twin [69]20236.6398Open(✓)Apartment, office
OMOMO [54]2023101715Lab room
HIMO [65]20244.1M3453Mocap lab
TRUMANS [45]2024151.6M720(✓)Scene mockups
ParaHome [50]20248.1382270Home room
(b) Home floor plan and camera setup of site II.
(b) Home floor plan and camera setup of site II.
Table 2: Capture hardware of ACE for synchronized multi-modal recording. Quantities are the totals for both configurations combined; a dash marks an entry that does not apply to that device.
DeviceRoleQtyResolution / rateNotes
OptiTrack PrimeX 22Optical motion capture282048×1088, 60 HzIR tracking: 41 body markers, objects, ego rig
ZED OneExocentric RGB capture81920×1080, 30 FPSGMSL2 to a single Jetson Orin host; shared frame trigger
GoProExocentric RGB capture81920×1080, 30 FPSRigidly mounted on stands; audio-triggered recording
ACE-Ego-Head-V02 LiteEgocentric capture4 cameras4×1088×1280, 20 FPSOne headset: front/back fisheye pairs, IMU, 5 markers
ManusHand pose2 gloves60 HzPer-finger articulation, both hands
ACE-Sense-Glove LiteContact pressure2 glovesFull-palm pressure map, both hands
Jetson OrinRecording host1Ingests all ZED One streams; NatNet-coordinated
Motive host (PC)Mocap host, sync reference2Renders the optical clock for cross-system sync
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Figure 3: Overall workflow of ACE. Each participant wears the motion-capture suit before the multi-modal recording. Synchronization is conducted after data export to align the timeline of each modality. Finally, we annotate the collected data with rich and high-quality labels derived from the captured ground truth.
Table 3: Tactile estimation from ego-view on ACE-Data-0.
MethodTemp Acc. ↑C-IoU ↑V-IoU ↑CoP ↓
PressureVision [32]0.00930.00070.000010.9807
EgoPressureDiff [109]0.29120.01970.00258.5152
TouchAnything [115]0.70950.16460.13576.5846
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Figure 4: Synchronized modalities captured by ACE. ACE is capable of capturing synchronized multi-modal data, including egocentric and exocentric videos, audio, and human and object motions, etc.
Table 4: Human motion estimation on ACE-Data-0. All metrics in mm.
MethodPA-MPJPE ↓PA-PVE ↓MPJPE ↓PVE ↓WA-MPJPE ↓
— Single-view exocentric, per-frame —
Multi-HMR [4]110.5115.3
Multi-HMR2 [24]85.592.8
SAM-3D-Body [103]71.076.4
PyMAF-X [111]94.7177.599.9177.0
PARE [51]98.7132.3102.0134.8
CameraHMR [70]78.798.785.7105.2
OSX [59]59.273.762.576.6
— Single-view exocentric, temporal —
SMPLer-X [11]57.070.259.871.7
SMPLest-X [107]55.765.858.867.6
Humans-in-4D [31]77.194.479.696.4
GVHMR [84]88.0104.196.3112.0217.1
WHAM [85]131.1160.2144.0166.4243.1
EasyMoCap [21, 86]103.4155.8106.0156.1256.2
— Single-view exocentric, scene-aware —
Phy-SIC [68]64.378.971.384.6
UniSH [55]80.0111.282.8114.0196.3
JOSH [62]64.089.269.995.0245.1
Human3R [16]60.175.463.580.6180.2
— Multi-view exocentric —
MAMMA [18]70.186.874.889.6230.8
U-HMR [57]134.4176.0150.2179.2
HSfM [67]92.8133.893.8134.8
— Egocentric —
EgoEgo [53]159.6245.3163.4263.9306.2
EgoAllo [106]131.7196.8147.9220.3252.2
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Figure 5: Calibrated cameras in the room-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture; the green cameras denote the ZED One for capturing exocentric RGB videos; the orange camera is one of the egocentric cameras in ACE-Ego-Head.
Table 5: Hand motion estimation from ego-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
WildHands [76]11.212.60.1750.7740.776
Dyn-HaMR [108]18.921.10.0420.4130.62498.2
HaWoR [112]13.817.40.1300.6660.729102.1
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Figure 6: Calibrated cameras in the table-scale configuration. The blue cameras represent the OptiTrack PrimeX 22 for motion capture.
Table 6: Hand motion estimation from exo-view on ACE-Data-0.
MethodPA-MPJPE ↓MPJPE ↓F@5 ↑F@15 ↑AUCJ ↑Traj. err. ↓
HORT [17]10.812.30.2760.7760.784
HaMeR [72]9.610.40.2810.8480.812
HaPTIC [105]10.010.70.2470.8380.80463.0
WiLoR [75]9.19.90.3130.8540.819
OmniHands [58]10.711.50.2110.8140.791
Figure 7: Examples of different task categories captured in ACE-Data-0.
Figure 7: Examples of different task categories captured in ACE-Data-0.

실제로 확인된 결과

  • 1인칭 영상 기반 손 동작 추정에서 프레임별 방법(WildHands)이 관절 자세 정확도는 가장 좋았지만, 비디오 기반 방법(Dyn-HaMR, HaWoR)은 카메라 움직임 추정 오차 때문에 전역 궤적 오차가 98~102mm로 크게 나타났다.
  • 3인칭 영상 기반 손 동작 추정에서는 다섯 방법의 관절 자세 정확도가 PA-MPJPE 9.1~10.8mm로 비슷했고, WiLoR가 9.1mm로 가장 좋았으며 HaMeR이 9.6mm로 근접했다.
  • 고정된 3인칭 카메라를 이용한 HaPTIC은 궤적 오차가 63mm로, 1인칭 방법들의 98~102mm보다 훨씬 낮아, 고정 시점이 궤적 추정에 유리함을 보였다.
  • 물체를 함께 복원하는 HORT는 조작 프레임에서 PA-MPJPE 10.8mm를 기록해 다른 방법들과 비슷한 수준이었고, 물체 정보를 함께 쓰는 것이 뚜렷한 이득이나 손해를 주지는 않았다.
  • 1인칭과 3인칭을 비교했을 때 3인칭 방법들이 자세와 궤적 모두에서 더 나은 결과를 보였는데, 이는 고정 카메라가 안정적인 기준 좌표계를 제공하기 때문으로 분석됐다.
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.
Figure 8: Overview statistics of ACE-Data-0. ACE-Data-0 covers three types of interaction tasks: atomic HOI tasks, chains of HOI tasks, and HSI tasks.

어디에 쓸 수 있나

  • 가정용 로봇이나 휴머노이드의 모방학습·정책학습에 쓸 수 있는 동기화된 시각-동작-촉각 데이터로 활용
  • 1인칭 영상만으로 손 동작이나 촉각을 추정하는 모델의 성능을 검증하는 벤치마크로 사용
  • 1인칭·3인칭 영상을 함께 쓰는 손-물체 상호작용 인식 연구의 출발점으로 활용
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.
Figure 9: Examples of data annotation for captured objects. ACE-Data-0 provides rich annotations for the captured objects, including name tags, 6-DoF poses, bounding boxes, and motion trails.

한계와 남은 검증

  • 현재 두 곳의 가정집에서만 촬영되어 실내 구조, 가구, 조명의 다양성이 제한적이다.
  • 추적 대상은 사전에 3D 스캔하고 마커를 부착한 물체로 한정되어, 문 여닫힘 같은 관절 기구나 액체, 변형되는 재질의 상태 변화는 기록되지 않는다.
  • 모션캡처 슈트, 촉각 장갑, 헤드셋, 마커가 영상에 그대로 보여 데이터셋 특유의 시각적 흔적이 남을 수 있다.
  • 1인칭과 3인칭 스트림을 함께 활용하거나 헤드셋의 측정된 움직임 정보, 물체 자세를 보조 입력으로 주는 등의 실험은 아직 수행되지 않고 향후 연구로 제안됐다.
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 10: Examples of data annotation for the captured human body and hands. ACE-Data-0 provides ground-truth poses for both the human body and hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

왜 중요한가

로봇이나 AI가 사람처럼 손으로 물건을 다루는 법을 배우려면 시각, 동작, 소리, 촉감이 같은 순간에 어떻게 얽혀 있는지 보여주는 데이터가 필요한데, 기존 자료는 이런 정보가 흩어져 있어 학습에 한계가 있었다. ACE는 실제 가정에서 이 모든 신호를 동시에 정확히 기록하는 방법을 제시해, 모방학습이나 세계모델 같은 로봇 학습 연구에 쓸 수 있는 토대를 마련했다.

Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.
Figure 11: Examples of data annotation for captured hands. ACE-Data-0 provides ground-truth poses for human hands. We also project the poses to both egocentric and exocentric videos for better embodied AI learning with visual inputs.

이 논문의 용어

  • 1인칭(egocentric)/3인칭(exocentric) 영상 · 사람이 직접 착용한 카메라로 찍은 시야(1인칭)와 외부에 고정된 카메라로 찍은 시야(3인칭)
  • 6자유도(6-DoF) 포즈 · 물체의 3차원 위치(앞뒤·좌우·상하)와 회전(세 방향의 기울기)을 합쳐 6개 값으로 나타낸 자세 정보
  • MANO · 손의 관절과 형태를 3D로 표현하는 표준 손 모델
  • PA-MPJPE · 크기·위치·회전을 각각 맞춘 뒤 관절 위치 오차를 재는 지표, 손가락 자세 정확도를 주로 측정
  • 모션캡처(OptiTrack) · 몸이나 물체에 붙인 마커를 여러 카메라로 추적해 정밀한 3D 동작을 기록하는 장비
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.
Figure 12: Examples of data annotation for tactile signals. ACE-Data-0 provides ground-truth full-hand grasp pressure values, resolving interaction events that remain ambiguous under visual occlusion.

저자 · Yukang Cao

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Yukang Cao et al., arXiv:2607.28625, arxiv-nonexclusive