工作日早上 7 点读 AI,周日早上 8 点读周报订阅邮件

METAL LAB

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

arXiv:2608.082852026-08-07

一套成本不到200美元的头戴设备采集了550小时第一人称立体视频,开源发布用于机器人学习研究

Ego-OSCAR是一款完全由市售零件和3D打印部件组装、单件成本低于200美元的头戴式立体摄像头加IMU采集设备。作者在印度招募25名贡献者,历时约六个月,在40多个室内环境中采集了每台相机约550小时的同步立体视频与惯性数据,并附带自由文本动作字幕和逐帧3D手部重建后发布,而非原始传感器数据。该项目的目标不是追求Project Aria等研究级设备的精度,而是找到一种最廉价、可靠、可大规模众包采集第一人称数据的方案。

METAL LAB 解读图

Ego-OSCAR 采集到数据集的流程

证据状态已报告实测结果

  1. 头戴硬件全局快门立体摄像头、6轴IMU、单板计算机与ESP32微控制器组成,单件成本低于200美元
  2. 时间同步ESP32将摄像头曝光信号与IMU数据合并,并用LED闪光标定帧号,残余偏差降至700微秒
  3. 实地部署25名贡献者在印度40多个室内环境中历时六个月完成1,462个会话,每台相机采集约550小时
  4. 质量筛选看门狗机制、批次校验和手部可见性筛选,使可用会话比例达到96%
  5. 标注层209,315条自由文本动作字幕与基于WiLoR的逐帧3D手部重建覆盖全部语料
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 设备将硬件同步的全局快门立体摄像头、6轴IMU、嵌入式Linux单板计算机和实时微控制器组合在一起,完整物料清单成本约200美元(合19,100印度卢比)一台。
  2. ESP32微控制器读取摄像头的曝光起始信号并与IMU数据合并,再用LED闪光标定帧号,将视觉与惯性数据间的残余时间偏差降至700微秒。
  3. 在13台已部署设备上,矫正后的平均逐像素对极误差为0.4像素,单相机重投影误差低于0.03像素,SGBM与RAFT-Stereo无需特殊调参即可在全视场内生成密集视差图。
  4. 六个月的实地部署中,96%的采集会话最终产生可用数据;手部检测器(WiLoR)在全部帧中检测到手部的比例为94%;立体惯性视觉里程计(VINS-Fusion)在20段留出序列中有12段生成稳定轨迹,而采用主动立体深度的Intel RealSense在同等条件下为15/20。
  5. 发布的Ego-OSCAR-550h数据集包含1,462个采集会话、25名贡献者、40多个环境,附有209,315条自由文本动作片段标注和逐帧3D手部重建,标注呈现长尾分布,排名前20的动作表达仅占全部实例的1.5%。
Figure 1: Open-source egocentric capture system.
Figure 1: Open-source egocentric capture system.
Table 1: Bill of Materials.
ComponentINRFunction
Dexcin USB stereo camera6,300Stereo capture, global shutter
Radxa Rock 5C (2 GB)6,500SBC: capture, encode, store
Heatsink (Radxa)800Thermal management
256 GB SD card2,500OS + ∼16–18 hr recording
USB cable, A–C 90°300Camera ↔ SBC
ICM-20948 6-axis IMU800Inertial sensing
Seeed Xiao ESP32-S3650UX + watchdog MCU
Misc. electronics300LEDs, buzzer, button, wiring
3D-printed shells + screws300Enclosure
Visor / cap300Head mount
USB-C PD cable, 90°150SBC ↔ power
10,000 mAh power bank500Power
Total19,100∼USD 200
(b) Mechanical Outline
(b) Mechanical Outline
Table 2: Comparison with existing egocentric datasets. Figures are as reported by each dataset’s own publication. “Cam-h” denotes camera-hours. Ego-Exo4D hours combine egocentric and exocentric video.
DatasetHoursWearersCameraSync. IMUDense action labelsOpen HW
Ego4D [7]3,670931Mono consumer, rolling shutter (stereo in a subset)PartialTimestamped narrationsNo
Ego-Exo4D [8]1,286740Aria: mono RGB + 2 mono SLAMYesNarrations + expert commentaryNo
EPIC-K.-100 [2]10037Mono head-mountedNo90K segments, closed taxonomy (97 verbs / 300 nouns)No
Nymeria [15]300264Aria + body mocapYes301.5K narration sentencesNo
Ours550/cam (1,100 cam-h)25Calibrated RGB stereo, global shutter, per-session calib.Yes 120 Hz209,315 segments, open vocabulary, ≈100% coverageYes ∼USD 200
(c) Device in action
(c) Device in action
Table 3: Dataset overview: measured dimensions and what each supports.
DimensionEvidenceRelevance
Video scale≈550 h per camera (≈1,100 stereo cam-h)Large calibrated stereo RGB corpus
Capture geometrySynchronized pair + per-session calibrationMetric binocular depth cues, not just RGB
Delivery formatMP4 (H.264, 1280×720, 30 fps)Direct ingestion into video pipelines
Label densityMedian 94 segments/sessionDense temporal supervision, not clip-level tags
Action vocabulary460 observed verbsCoverage of manipulation primitives
Object vocabulary32,630 object phrasesWide object, material and tool coverage
Effective breadth57,104 verb–object combinationsCompositional, resistant to rare-label inflation
Temporal structure192,509 ordered task transitionsSequence structure for world models
Contributors25 user IDs / 13 devicesVariation in behavior, routine and execution style
Figure 2: System Overview: Hardware architecture.
Figure 2: System Overview: Hardware architecture.
Table 4: Most frequent observed task expressions and their share of all labeled segments.
Task expressionOccurrencesShare
idle / no manipulation6590.31%
cut sewing thread with scissors2690.13%
close refrigerator door2380.11%
peel garlic clove2120.10%
open refrigerator door2120.10%
turn on kitchen faucet1390.07%
pick up iron from side table1340.06%
turn off kitchen faucet1240.06%
adjust stove burner knob1220.06%
rinse small metal cup under running water1190.06%
adjust stove control knob1150.05%
roll dough on rolling board with rolling pin1090.05%
Figure 3: Data utility results.
Figure 3: Data utility results.
Table 5: Activity-domain composition of the Ego-OSCAR-550h dataset (keyword-derived, approximate).
Activity domainLabeled hoursPrimary in
Cooking and food preparation187 h664 sessions
Dishwashing and kitchen cleanup90 h258 sessions
Textile and craft (sewing, tailoring, flowers)54 h214 sessions
Laundry and clothing care45 h145 sessions
Organizing and storage39 h87 sessions
Cleaning and housekeeping29 h87 sessions
Generic manipulation and transitions106 h7 sessions
(b) Stereo depth maps
(b) Stereo depth maps
Table 6: Per-contributor coverage (percentiles across the 25 contributors).
Contributor-level diversity measure25thMedian75th
Labeled action segments5,1836,25812,075
Distinct task expressions4,0305,6369,115
Distinct action verbs90152178
Figure 4: Task diversity in the Ego-OSCAR-550h dataset.
Figure 4: Task diversity in the Ego-OSCAR-550h dataset.
Table 7: Sequence and composition signals in the dataset.
SignalEvidenceWhy it matters
Verb–object composition57,104 unique combinationsSystematic generalization across skills and objects
Broadly recombined verbs132 verbs with 25+ objects; 66 with 100+Reuse of manipulation primitives across object types
Ordered task transitions192,509 unique transitionsTemporal structure for sequence learning
Per-session sequence richnessMedian 86 transitions/session95.8% of sessions contain 10+ distinct transitions
Cross-contributor support19.4% of instances seen across 2+ contributorsReduces reliance on a single execution style

研究结果

  • 在13台部署设备上测得,矫正后的平均对极误差为0.4像素,单相机重投影误差低于0.03像素,验证了立体几何的可用性。
  • 经Kalibr相机-IMU偏移测试验证,校正后的视觉-惯性残余时间偏差为700微秒。
  • 立体惯性里程计(VINS-Fusion)在20段留出序列中有12段生成稳定轨迹(相比之下采用主动立体的Intel RealSense为15/20,但后者具有不公平的深度优势),手部检测器在全语料中的检测率为94%。
  • 六个月部署期内96%的会话产生了端到端可用数据,最终获得1,462个会话、25名贡献者、40多个环境,每台相机约550小时(约1,100立体相机小时)的数据。
  • 发布数据集中含有209,315条自由文本动作片段标注和覆盖全语料的逐帧3D手部重建,排名前20的动作表达仅占全部实例的1.5%,呈现明显长尾分布。

可应用场景

  • 为机器人视觉-语言-动作模型或世界模型的预训练采集大规模第一人称视觉-动作数据
  • 基于开源硬件设计自行搭建设备,众包采集现有数据集覆盖不足的特定领域第一人称数据(如家庭手工艺)
  • 利用已标定的立体图像和手部重建标注开展深度估计或手部姿态研究,无需自行搭建采集硬件
  • 为无法获取Project Aria等封闭研究平台的团队提供自建采集流水线的参考

局限与待验证事项

  • 论文并未证明基于该数据训练的机器人策略优于基于现有数据集训练的策略,这一验证被列为未来工作。
  • 缺乏运动捕捉或测量级的相机轨迹真值,因此12/20的里程计数据只反映收敛率,不代表轨迹精度。
  • 语料库地域和人群较为集中:25名贡献者均在印度,共享13台设备,活动内容偏向家庭日常。
  • 消费级IMU是姿态误差的主要来源,虽可在同一总线上更换更高级型号,但尚未对此进行基准测试。
  • 设备仅用于采集,无法在会话过程中实时剔除不良数据,且仍存在头带压痛点、长时间佩戴设备下垂、外壳无防水等耐用性问题。

为什么重要

面向机器人的视觉-语言-动作模型需要大规模、多样化的多模态数据,但遥操作难以规模化、仿真存在仿真到现实的落差,而现有第一人称数据集大多依赖未标定的消费级相机或无法自由复制的封闭研究级硬件。通过开源一套低成本、可自行组装的采集设备及其软件和经过验证的数据集,这项工作降低了任何团队自行大规模采集第一人称数据的门槛。

本文术语

  • 第一人称视频(Egocentric video) · 由头戴摄像头拍摄、呈现佩戴者自身视角的影像
  • 全局快门 · 传感器所有像素同时曝光的方式,可避免快速运动下的图像畸变
  • IMU(惯性测量单元) · 测量加速度和旋转角速度以感知设备运动的传感器
  • 立体标定 · 测定左右两台相机之间的几何关系和镜头畸变,以便准确计算深度的校准过程
  • 视觉惯性里程计(VIO) · 结合图像与IMU数据来估计设备移动轨迹的技术

论文原文摘要(英文)

We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced

作者 · Gunjan Paul

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Gunjan Paul et al., arXiv:2608.08285, CC BY 4.0