One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Dyna Robotics unveils robot model trained on 1 million hours of human video

DYNA-2, a robot foundation model pretrained on 170 years' worth of human behavior video, emerges

로봇팔이 도마에서 오이를 썰어 통에 담아놓은 모습

이미지: X — 뉴스 앰프

Summary

  • Dyna Robotics has unveiled its robot foundation model DYNA-2
  • It was pretrained on more than 1 million hours of human video, equivalent to about 170 years' worth
  • The company emphasized breakthroughs beyond the dataset's size itself, but specifics have not been confirmed
Video from the source
모델명
DYNA-2
개발사
Dyna Robotics
사전학습 데이터
인간 영상 100만 시간 이상
환산 규모
약 170년 치 연속 깨어있는 경험 분량
공개 시점
2026년 8월 11일(현지시간) 소셜미디어를 통해 알려짐

A robot that consumed 1 million hours of human video

Robotics startup Dyna Robotics has unveiled its new robot foundation model, DYNA-2. The key point is the scale of the training data. It was pretrained on more than 1 million hours of human video, reportedly equivalent to the amount of experience a person would accumulate by staying awake without sleep for 170 years. However, the company reportedly stated that the sheer size of this massive dataset is not itself the real breakthrough. What exactly was demonstrated is not clearly confirmed in the source.

Robots learning by watching human hands

A robot foundation model refers to a general-purpose model in which a single neural network is trained to broadly perform a wide range of object manipulation tasks. The problem is data. Demonstration data of robots actually grasping and moving objects is slow and costly to collect, since humans must manually operate robot arms to create it one instance at a time. In contrast, video of humans picking up objects, assembling things, and tidying up already exists in vast quantities across the internet, including on YouTube. Recently, the robotics industry has been racing to use this human video to first build broad "behavioral intuition," then fine-tune it with a small amount of actual robot data.

This trend has become clear over the past few weeks. On July 30, Google DeepMind unveiled Gemini Robotics 2, which controls everything from a robot's legs to its five fingers with a single model, and disclosed success rates of 76.3% for picking from shelves, 45.7% for picking from the floor, and 36% for screwing in a lightbulb. In early August, NVIDIA also released Cosmos 3, an open-weight model combining vision reasoning, world generation, and action prediction, along with Alpamayo 2 Super, a 34-billion-parameter model dedicated to autonomous driving. DYNA-2 is another entrant in this race, distinguished by pushing reliance on human video to an extreme degree.

So what actually changes

The long-standing bottleneck in robot learning has been a shortage of data generated by robots themselves. As more efforts emerge to bypass this bottleneck by drawing on massive amounts of human video, a trend is solidifying in which extensive pretraining becomes possible without even a single robot arm. However, human hand movements and the physical structure of robot arms differ. The real question is how accurately intuition learned from video can be translated into actual gripper movements, and this announcement alone is too early to answer that.