Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
一套成本不到200美元的头戴设备采集了550小时第一人称立体视频,开源发布用于机器人学习研究
Ego-OSCAR是一款完全由市售零件和3D打印部件组装、单件成本低于200美元的头戴式立体摄像头加IMU采集设备。作者在印度招募25名贡献者,历时约六个月,在40多个室内环境中采集了每台相机约550小时的同步立体视频与惯性数据,并附带自由文本动作字幕和逐帧3D手部重建后发布,而非原始传感器数据。该项目的目标不是追求Project Aria等研究级设备的精度,而是找到一种最廉价、可靠、可大规模众包采集第一人称数据的方案。
METAL LAB 解读图
Ego-OSCAR 采集到数据集的流程
证据状态已报告实测结果
- 头戴硬件全局快门立体摄像头、6轴IMU、单板计算机与ESP32微控制器组成,单件成本低于200美元
- 时间同步ESP32将摄像头曝光信号与IMU数据合并,并用LED闪光标定帧号,残余偏差降至700微秒
- 实地部署25名贡献者在印度40多个室内环境中历时六个月完成1,462个会话,每台相机采集约550小时
- 质量筛选看门狗机制、批次校验和手部可见性筛选,使可用会话比例达到96%
- 标注层209,315条自由文本动作字幕与基于WiLoR的逐帧3D手部重建覆盖全部语料
他们做了什么
- 设备将硬件同步的全局快门立体摄像头、6轴IMU、嵌入式Linux单板计算机和实时微控制器组合在一起,完整物料清单成本约200美元(合19,100印度卢比)一台。
- ESP32微控制器读取摄像头的曝光起始信号并与IMU数据合并,再用LED闪光标定帧号,将视觉与惯性数据间的残余时间偏差降至700微秒。
- 在13台已部署设备上,矫正后的平均逐像素对极误差为0.4像素,单相机重投影误差低于0.03像素,SGBM与RAFT-Stereo无需特殊调参即可在全视场内生成密集视差图。
- 六个月的实地部署中,96%的采集会话最终产生可用数据;手部检测器(WiLoR)在全部帧中检测到手部的比例为94%;立体惯性视觉里程计(VINS-Fusion)在20段留出序列中有12段生成稳定轨迹,而采用主动立体深度的Intel RealSense在同等条件下为15/20。
- 发布的Ego-OSCAR-550h数据集包含1,462个采集会话、25名贡献者、40多个环境,附有209,315条自由文本动作片段标注和逐帧3D手部重建,标注呈现长尾分布,排名前20的动作表达仅占全部实例的1.5%。

| Component | INR | Function |
|---|---|---|
| Dexcin USB stereo camera | 6,300 | Stereo capture, global shutter |
| Radxa Rock 5C (2 GB) | 6,500 | SBC: capture, encode, store |
| Heatsink (Radxa) | 800 | Thermal management |
| 256 GB SD card | 2,500 | OS + ∼16–18 hr recording |
| USB cable, A–C 90° | 300 | Camera ↔ SBC |
| ICM-20948 6-axis IMU | 800 | Inertial sensing |
| Seeed Xiao ESP32-S3 | 650 | UX + watchdog MCU |
| Misc. electronics | 300 | LEDs, buzzer, button, wiring |
| 3D-printed shells + screws | 300 | Enclosure |
| Visor / cap | 300 | Head mount |
| USB-C PD cable, 90° | 150 | SBC ↔ power |
| 10,000 mAh power bank | 500 | Power |
| Total | 19,100 | ∼USD 200 |

| Dataset | Hours | Wearers | Camera | Sync. IMU | Dense action labels | Open HW |
|---|---|---|---|---|---|---|
| Ego4D [7] | 3,670 | 931 | Mono consumer, rolling shutter (stereo in a subset) | Partial | Timestamped narrations | No |
| Ego-Exo4D [8] | 1,286 | 740 | Aria: mono RGB + 2 mono SLAM | Yes | Narrations + expert commentary | No |
| EPIC-K.-100 [2] | 100 | 37 | Mono head-mounted | No | 90K segments, closed taxonomy (97 verbs / 300 nouns) | No |
| Nymeria [15] | 300 | 264 | Aria + body mocap | Yes | 301.5K narration sentences | No |
| Ours | 550/cam (1,100 cam-h) | 25 | Calibrated RGB stereo, global shutter, per-session calib. | Yes 120 Hz | 209,315 segments, open vocabulary, ≈100% coverage | Yes ∼USD 200 |

| Dimension | Evidence | Relevance |
|---|---|---|
| Video scale | ≈550 h per camera (≈1,100 stereo cam-h) | Large calibrated stereo RGB corpus |
| Capture geometry | Synchronized pair + per-session calibration | Metric binocular depth cues, not just RGB |
| Delivery format | MP4 (H.264, 1280×720, 30 fps) | Direct ingestion into video pipelines |
| Label density | Median 94 segments/session | Dense temporal supervision, not clip-level tags |
| Action vocabulary | 460 observed verbs | Coverage of manipulation primitives |
| Object vocabulary | 32,630 object phrases | Wide object, material and tool coverage |
| Effective breadth | 57,104 verb–object combinations | Compositional, resistant to rare-label inflation |
| Temporal structure | 192,509 ordered task transitions | Sequence structure for world models |
| Contributors | 25 user IDs / 13 devices | Variation in behavior, routine and execution style |

| Task expression | Occurrences | Share |
|---|---|---|
| idle / no manipulation | 659 | 0.31% |
| cut sewing thread with scissors | 269 | 0.13% |
| close refrigerator door | 238 | 0.11% |
| peel garlic clove | 212 | 0.10% |
| open refrigerator door | 212 | 0.10% |
| turn on kitchen faucet | 139 | 0.07% |
| pick up iron from side table | 134 | 0.06% |
| turn off kitchen faucet | 124 | 0.06% |
| adjust stove burner knob | 122 | 0.06% |
| rinse small metal cup under running water | 119 | 0.06% |
| adjust stove control knob | 115 | 0.05% |
| roll dough on rolling board with rolling pin | 109 | 0.05% |

| Activity domain | Labeled hours | Primary in |
|---|---|---|
| Cooking and food preparation | 187 h | 664 sessions |
| Dishwashing and kitchen cleanup | 90 h | 258 sessions |
| Textile and craft (sewing, tailoring, flowers) | 54 h | 214 sessions |
| Laundry and clothing care | 45 h | 145 sessions |
| Organizing and storage | 39 h | 87 sessions |
| Cleaning and housekeeping | 29 h | 87 sessions |
| Generic manipulation and transitions | 106 h | 7 sessions |

| Contributor-level diversity measure | 25th | Median | 75th |
|---|---|---|---|
| Labeled action segments | 5,183 | 6,258 | 12,075 |
| Distinct task expressions | 4,030 | 5,636 | 9,115 |
| Distinct action verbs | 90 | 152 | 178 |

| Signal | Evidence | Why it matters |
|---|---|---|
| Verb–object composition | 57,104 unique combinations | Systematic generalization across skills and objects |
| Broadly recombined verbs | 132 verbs with 25+ objects; 66 with 100+ | Reuse of manipulation primitives across object types |
| Ordered task transitions | 192,509 unique transitions | Temporal structure for sequence learning |
| Per-session sequence richness | Median 86 transitions/session | 95.8% of sessions contain 10+ distinct transitions |
| Cross-contributor support | 19.4% of instances seen across 2+ contributors | Reduces reliance on a single execution style |
研究结果
- 在13台部署设备上测得,矫正后的平均对极误差为0.4像素,单相机重投影误差低于0.03像素,验证了立体几何的可用性。
- 经Kalibr相机-IMU偏移测试验证,校正后的视觉-惯性残余时间偏差为700微秒。
- 立体惯性里程计(VINS-Fusion)在20段留出序列中有12段生成稳定轨迹(相比之下采用主动立体的Intel RealSense为15/20,但后者具有不公平的深度优势),手部检测器在全语料中的检测率为94%。
- 六个月部署期内96%的会话产生了端到端可用数据,最终获得1,462个会话、25名贡献者、40多个环境,每台相机约550小时(约1,100立体相机小时)的数据。
- 发布数据集中含有209,315条自由文本动作片段标注和覆盖全语料的逐帧3D手部重建,排名前20的动作表达仅占全部实例的1.5%,呈现明显长尾分布。
可应用场景
- 为机器人视觉-语言-动作模型或世界模型的预训练采集大规模第一人称视觉-动作数据
- 基于开源硬件设计自行搭建设备,众包采集现有数据集覆盖不足的特定领域第一人称数据(如家庭手工艺)
- 利用已标定的立体图像和手部重建标注开展深度估计或手部姿态研究,无需自行搭建采集硬件
- 为无法获取Project Aria等封闭研究平台的团队提供自建采集流水线的参考
局限与待验证事项
- 论文并未证明基于该数据训练的机器人策略优于基于现有数据集训练的策略,这一验证被列为未来工作。
- 缺乏运动捕捉或测量级的相机轨迹真值,因此12/20的里程计数据只反映收敛率,不代表轨迹精度。
- 语料库地域和人群较为集中:25名贡献者均在印度,共享13台设备,活动内容偏向家庭日常。
- 消费级IMU是姿态误差的主要来源,虽可在同一总线上更换更高级型号,但尚未对此进行基准测试。
- 设备仅用于采集,无法在会话过程中实时剔除不良数据,且仍存在头带压痛点、长时间佩戴设备下垂、外壳无防水等耐用性问题。
为什么重要
面向机器人的视觉-语言-动作模型需要大规模、多样化的多模态数据,但遥操作难以规模化、仿真存在仿真到现实的落差,而现有第一人称数据集大多依赖未标定的消费级相机或无法自由复制的封闭研究级硬件。通过开源一套低成本、可自行组装的采集设备及其软件和经过验证的数据集,这项工作降低了任何团队自行大规模采集第一人称数据的门槛。
本文术语
- 第一人称视频(Egocentric video) · 由头戴摄像头拍摄、呈现佩戴者自身视角的影像
- 全局快门 · 传感器所有像素同时曝光的方式,可避免快速运动下的图像畸变
- IMU(惯性测量单元) · 测量加速度和旋转角速度以感知设备运动的传感器
- 立体标定 · 测定左右两台相机之间的几何关系和镜头畸变,以便准确计算深度的校准过程
- 视觉惯性里程计(VIO) · 结合图像与IMU数据来估计设备移动轨迹的技术
论文原文摘要(英文)
We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit一款把单词意义变化拆解到语法细节的开源分析工具
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)用AI总结股市新闻发现:简单的摘要方法反而比时髦的检索增强技术更靠谱
METAL LAB 最新报道
图片来源: Gunjan Paul et al., arXiv:2608.08285, CC BY 4.0