τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
로봇이 애매한 상황에서는 답을 바로 내지 않고 여러 수를 미리 상상해본 뒤 결정한다
τ0-VLA는 청소, 요리, 밀크티 만들기처럼 몇 분에서 12분까지 걸리는 긴 로봇 작업에서 다음에 할 하위 작업을 정할 때, 어려운 순간에만 여러 후보를 상상하고 결과 이미지를 예측해 점수를 매긴 뒤 최종 선택을 하는 계층형 시스템이다. 40,115시간 분량의 실제 로봇 데이터로 학습됐고, 실물 로봇 실험에서 추가 연산을 쓸수록 다음 행동 예측 정확도와 실제 작업 성공률이 함께 올라갔다. 저자는 Xiaowei Cai이며 논문은 arXiv 2608.16885에 공개됐다.
무엇을 했나
- 기존 계층형 로봇 AI는 다음에 할 일을 한 번의 계산만으로 정해서, 어려운 결정에 더 신경 쓸 방법이 없었다
- τ0-VLA는 확신이 낮을 때만 여러 후보 하위 작업을 만들고, 각 후보가 실행됐을 때의 최종 화면을 예측하는 세계모델과 그 결과를 채점하는 가치모델로 빔서치를 수행한 뒤 최종 결정을 내린다
- 선택된 하위 작업은 40차원 통일 동작 공간을 쓰는 하위 실행 모델이 실제 로봇 팔다리 움직임으로 바꿔 여러 로봇 몸체에서 수행한다
- 청소, 재료 준비, 볶음요리, 밀크티 제작, 빨래 수거, 책 정리 등 실물 로봇 과제에서 연산을 더 쓸수록 다음 작업 예측 정확도와 최종 성공률이 함께 향상됐다
- 학습 데이터가 없던 낯선 책 배열 상황에서도 같은 경향이 유지돼 분포가 달라져도 방법이 견고함을 보였다

| Method | Clean Room | Prepare Ingredients | Tomato and Egg Stir Fry | Make Milk Tea | Avg. | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | |
| GR00T N1.7 [26] | 0/10 | 59.80% | 1/10 | 68.57% | 0/10 | 24.32% | 0/10 | 28.46% | 2.50% | 45.29% |
| LingBot-VLA [40] | 0/10 | 66.60% | 0/10 | 35.00% | 0/10 | 12.27% | 0/10 | 63.85% | 0.00% | 44.43% |
| π0.5 [2] | 4/10 | 86.20% | 2/10 | 73.93% | 0/10 | 49.77% | 3/10 | 82.31% | 22.50% | 73.05% |
| τ0-VLA | 4/10 | 92.80% | 2/10 | 66.43% | 0/10 | 65.00% | 5/10 | 96.15% | 27.50% | 80.10% |
| τ0-VLA (Hierarchical System, Plan Once) | 5/10 | 94.80% | 4/10 | 82.86% | 4/10 | 81.82% | 5/10 | 91.92% | 45.00% | 87.85% |

| Method | Collect Laundry | Tidy Makeup Table | ||||||
|---|---|---|---|---|---|---|---|---|
| T-shirt | Cotton Pad | Eyelash Curler | Makeup Puff | |||||
| SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | |
| GR00T N1.7 [26] | 4/10 | 76.00% | 10/10 | 87.50% | 8/10 | 77.50% | 7/10 | 52.50% |
| LingBot-VLA [40] | 2/10 | 35.00% | 9/10 | 67.50% | 3/10 | 22.50% | 3/10 | 33.75% |
| π0.5 [2] | 9/10 | 88.00% | 9/10 | 85.00% | 8/10 | 85.00% | 7/10 | 73.75% |
| τ0-VLA | 10/10 | 97.00% | 10/10 | 95.00% | 9/10 | 92.50% | 10/10 | 95.00% |
| Method | Make Milk Tea | Book Organization | Clean Room | |||
|---|---|---|---|---|---|---|
| SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | SR ↑ | Progress ↑ | |
| Plan Once | 5/10 | 91.92% | 6/10 | 66.67% | 5/10 | 94.80% |
| TTC | 7/10 | 95.38% | 9/10 | 93.33% | 7/10 | 97.60% |
| Coordinates | Dimensions | State representation |
|---|---|---|
| Left EEF position | 1–3 | Cartesian position in meters |
| Left EEF orientation | 4–9 | Rot6D(𝐑L) |
| Right EEF position | 10–12 | Cartesian position in meters |
| Right EEF orientation | 13–18 | Rot6D(𝐑R) |
| Left gripper | 19 | native opening coordinate |
| Right gripper | 20 | native opening coordinate |
| Waist | 21–22 | two native coordinates |
| Planar base velocity | 23–24 | two native coordinates |
| Left arm joints | 25–32 | q1L,…,q8L in radians |
| Right arm joints | 33–40 | q1R,…,q8R in radians |
| Task | Maximum duration |
|---|---|
| Clean Room | 20 min |
| Prepare Ingredients | 20 min |
| Tomato and Egg Stir Fry | 20 min |
| Make Milk Tea | 10 min |
| Book Organization | 5 min |
| Collect Laundry | 5 min |
| Tidy Makeup Table (each group) | 5 min |
| Family | Sampling position | Input → target memory | Target subtask | Deployment failure countered | Mix |
|---|---|---|---|---|---|
| within-subtask | anywhere in seg. n | ℳn→ℳn | seg. n | — (aligned, normal progression) | 58% |
| transition | tail of seg. n | ℳn→ℳn+1 | seg. n+1 | starting a new subtask after completion | 15% |
| catch-up | head of seg. n | ℳn−1→ℳn | seg. n | memory lag (behind the visual state) | 10% |
| rollback | late in seg. n | ℳn+1…n+3→ℳn | retry seg. n | memory run-ahead (over-optimistic) | 12% |
| error-think | annotated failure frame | ℳn→ type-dependent | recovery step | unnoticed execution failure | 5% |
왜 중요한가
긴 작업을 하는 로봇은 잘못된 하위 결정을 내리면 아무리 손동작이 정확해도 실패하는데, 이 연구는 어려운 순간에만 더 오래 생각하게 만드는 방법을 실물 로봇으로 검증했다. 이는 언어모델의 테스트 타임 연산 확장 아이디어를 로봇 제어에 옮긴 사례로, 앞으로 범용 로봇 시스템 설계에 참고가 될 수 있다.
이 논문의 용어
- VLA (Vision-Language-Action) 모델 · 카메라로 본 장면과 언어 지시를 받아 로봇 동작을 출력하는 인공지능 모델
- 테스트 타임 연산(Test-Time Computation) · 모델을 다시 학습시키지 않고, 실제 사용 시점에 더 많은 계산을 들여 답의 질을 높이는 방법
- 세계모델(World Model) · 어떤 행동을 하면 환경이 어떻게 바뀔지 미리 예측하는 모델
- 빔서치(Beam Search) · 여러 가능한 다음 수 중 점수가 높은 몇 개만 남기고 계속 확장해 나가는 탐색 방법
- 실행 메모리(Execution Memory) · 지금까지 로봇이 어디까지 작업을 끝냈는지 요약해 기억해두는 정보
논문 원문 초록 (영문)
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
arXiv에서 원문 보기최신 논문
- Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention트랜스포머의 모든 어텐션 헤드가 같은 길이의 문맥을 볼 필요는 없다는 실험
- Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models새 취약점이 계속 나와도 잊지 않고 하나의 모델로 합치는 스마트컨트랙트 취약점 탐지 AI
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
- A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries환자가 증상을 다 말하지 않아 애매한 질문에, AI가 먼저 되물어서 답을 맞히는 방법
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI 비서가 뭘 기억할지 결정할 때, '물어봐야 할 순간'에 되레 세상에 확인하고 넘어간다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
METAL LAB 최신 기사
그림 출처: Xiaowei Cai et al., arXiv:2608.16885, arxiv-nonexclusive