매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

arXiv:2608.168852026-08-16

로봇이 애매한 상황에서는 답을 바로 내지 않고 여러 수를 미리 상상해본 뒤 결정한다

τ0-VLA는 청소, 요리, 밀크티 만들기처럼 몇 분에서 12분까지 걸리는 긴 로봇 작업에서 다음에 할 하위 작업을 정할 때, 어려운 순간에만 여러 후보를 상상하고 결과 이미지를 예측해 점수를 매긴 뒤 최종 선택을 하는 계층형 시스템이다. 40,115시간 분량의 실제 로봇 데이터로 학습됐고, 실물 로봇 실험에서 추가 연산을 쓸수록 다음 행동 예측 정확도와 실제 작업 성공률이 함께 올라갔다. 저자는 Xiaowei Cai이며 논문은 arXiv 2608.16885에 공개됐다.

무엇을 했나

  1. 기존 계층형 로봇 AI는 다음에 할 일을 한 번의 계산만으로 정해서, 어려운 결정에 더 신경 쓸 방법이 없었다
  2. τ0-VLA는 확신이 낮을 때만 여러 후보 하위 작업을 만들고, 각 후보가 실행됐을 때의 최종 화면을 예측하는 세계모델과 그 결과를 채점하는 가치모델로 빔서치를 수행한 뒤 최종 결정을 내린다
  3. 선택된 하위 작업은 40차원 통일 동작 공간을 쓰는 하위 실행 모델이 실제 로봇 팔다리 움직임으로 바꿔 여러 로봇 몸체에서 수행한다
  4. 청소, 재료 준비, 볶음요리, 밀크티 제작, 빨래 수거, 책 정리 등 실물 로봇 과제에서 연산을 더 쓸수록 다음 작업 예측 정확도와 최종 성공률이 함께 향상됐다
  5. 학습 데이터가 없던 낯선 책 배열 상황에서도 같은 경향이 유지돼 분포가 달라져도 방법이 견고함을 보였다
Fig. 2: The hierarchical τ0-VLA architecture. (a) At a high-level inference step t, the proposal model P conditions on the latest multi-view observation ot, task instruction ℓ, carried execution memory ℳt−1, and previously generated subtask zt−1⋆. It produces observation-aligned memory ℳt and a direct proposal ztdir. (b) The low-level policy conditions on the generated subtask zt⋆, multi-view observation ot, proprioceptive state 𝐬t, and textual control metadata η. A vision-language backbone and Mixture-of-Transformers (MoT) action expert generate the action chunk 𝐚t:t+H−1 through conditional flow matching from a noisy action chunk. (c) On the TTC route, the proposal model is invoked N times for each retained branch to generate N candidates. Given the branch’s head-camera image and a candidate, the world model predicts the terminal head-camera image, and the value model assigns a candidate-quality score conditioned on the task instruction, candidate, and predicted image. Beam search globally retains the top-B branches by cumulative score and recursively expands them to depth D. The figure illustrates the root expansion with N=3 and B=2, where local and cumulative scores coincide. Deeper expansion is omitted for clarity. The reflective model then conditions on h¯t and the final branch summaries 𝒞t to generate the final subtask zt⋆. This output may coincide with a retained proposal but is not restricted to the retained set and is passed to the low-level policy for execution.
Fig. 2: The hierarchical τ0-VLA architecture. (a) At a high-level inference step t, the proposal model P conditions on the latest multi-view observation ot, task instruction ℓ, carried execution memory ℳt−1, and previously generated subtask zt−1⋆. It produces observation-aligned memory ℳt and a direct proposal ztdir. (b) The low-level policy conditions on the generated subtask zt⋆, multi-view observation ot, proprioceptive state 𝐬t, and textual control metadata η. A vision-language backbone and Mixture-of-Transformers (MoT) action expert generate the action chunk 𝐚t:t+H−1 through conditional flow matching from a noisy action chunk. (c) On the TTC route, the proposal model is invoked N times for each retained branch to generate N candidates. Given the branch’s head-camera image and a candidate, the world model predicts the terminal head-camera image, and the value model assigns a candidate-quality score conditioned on the task instruction, candidate, and predicted image. Beam search globally retains the top-B branches by cumulative score and recursively expands them to depth D. The figure illustrates the root expansion with N=3 and B=2, where local and cumulative scores coincide. Deeper expansion is omitted for clarity. The reflective model then conditions on h¯t and the final branch summaries 𝒞t to generate the final subtask zt⋆. This output may coincide with a retained proposal but is not restricted to the retained set and is passed to the low-level policy for execution.
TABLE I: Long-horizon task performance. Each method–task setting uses 10 independently collected physical-robot trials. SR reports successful trials as x/10, and Progress is the normalized milestone-completion score. Avg. is the unweighted mean of the four task-level rates. The first four rows use direct execution. The final row uses the Hierarchical System with Plan Once and no beam search.
MethodClean RoomPrepare IngredientsTomato and Egg Stir FryMake Milk TeaAvg.
SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑
GR00T N1.7 [26]0/1059.80%1/1068.57%0/1024.32%0/1028.46%2.50%45.29%
LingBot-VLA [40]0/1066.60%0/1035.00%0/1012.27%0/1063.85%0.00%44.43%
π0.5 [2]4/1086.20%2/1073.93%0/1049.77%3/1082.31%22.50%73.05%
τ0-VLA4/1092.80%2/1066.43%0/1065.00%5/1096.15%27.50%80.10%
τ0-VLA (Hierarchical System, Plan Once)5/1094.80%4/1082.86%4/1081.82%5/1091.92%45.00%87.85%
Fig. 3: Representative physical-robot evaluation tasks. (a) Clean Room requires collecting two dirty garments, placing them in a laundry basket, hanging a handbag, handing a blanket to a person, and disposing of table trash. (b) Prepare Ingredients requires retrieving a tomato and an egg, then cracking and stirring the egg while returning the tools. (c) Tomato and Egg Stir Fry requires chopping and transferring the tomatoes, cooking and seasoning the ingredients, plating the dish, and returning the cookware and utensils. (d) Make Milk Tea requires adding toppings, pouring milk and tea, sealing the cup, and inserting a straw. (e) Collect Laundry requires transferring a T-shirt from the bedside table to a laundry basket. (f) Tidy Makeup Table comprises three separately scored instruction-following groups that require different object selections and action sequences from matched visual states.
Fig. 3: Representative physical-robot evaluation tasks. (a) Clean Room requires collecting two dirty garments, placing them in a laundry basket, hanging a handbag, handing a blanket to a person, and disposing of table trash. (b) Prepare Ingredients requires retrieving a tomato and an egg, then cracking and stirring the egg while returning the tools. (c) Tomato and Egg Stir Fry requires chopping and transferring the tomatoes, cooking and seasoning the ingredients, plating the dish, and returning the cookware and utensils. (d) Make Milk Tea requires adding toppings, pouring milk and tea, sealing the cup, and inserting a straw. (e) Collect Laundry requires transferring a T-shirt from the bedside table to a laundry basket. (f) Tidy Makeup Table comprises three separately scored instruction-following groups that require different object selections and action sequences from matched visual states.
TABLE II: Direct-execution performance across embodiments. Tidy Makeup Table comprises three independently scored instruction-following groups. All methods execute the full task instruction without a high-level policy. Each task or task group is evaluated over 10 trials. SR reports successful trials as x/10, and Progress is the normalized milestone-completion score.
MethodCollect LaundryTidy Makeup Table
T-shirtCotton PadEyelash CurlerMakeup Puff
SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑
GR00T N1.7 [26]4/1076.00%10/1087.50%8/1077.50%7/1052.50%
LingBot-VLA [40]2/1035.00%9/1067.50%3/1022.50%3/1033.75%
π0.5 [2]9/1088.00%9/1085.00%8/1085.00%7/1073.75%
τ0-VLA10/1097.00%10/1095.00%9/1092.50%10/1095.00%
Fig. 4: Next-subtask prediction accuracy under different high-level inference methods. We compare Plan Once, Best-of-N, and TTC (Ours) across four evaluation settings: Make Milk Tea, Book Organization (In-Domain), Book Organization (OOD), and Clean Room.
Fig. 4: Next-subtask prediction accuracy under different high-level inference methods. We compare Plan Once, Best-of-N, and TTC (Ours) across four evaluation settings: Make Milk Tea, Book Organization (In-Domain), Book Organization (OOD), and Clean Room.
TABLE III: Closed-loop physical-robot performance with test-time computation. Each entry uses 10 independently collected trials. Book Organization uses shuffled initial arrangements and is reported without the in-domain and OOD split used in the open-loop evaluation. SR denotes task success rate, and Progress is the normalized milestone-completion score.
MethodMake Milk TeaBook OrganizationClean Room
SR ↑Progress ↑SR ↑Progress ↑SR ↑Progress ↑
Plan Once5/1091.92%6/1066.67%5/1094.80%
TTC7/1095.38%9/1093.33%7/1097.60%
Fig. 5: Relationship between computational cost and subtask-prediction accuracy. The accuracy increases with additional computation for the Make Milk Tea and Book Organization tasks, shown in panels (a) and (b), respectively. Each point represents an experimental result obtained at a different computational cost. The orange dashed curves show saturation fits to these results, and the gray dashed lines indicate the Plan Once baselines.
Fig. 5: Relationship between computational cost and subtask-prediction accuracy. The accuracy increases with additional computation for the Make Milk Tea and Book Organization tasks, shown in panels (a) and (b), respectively. Each point represents an experimental result obtained at a different computational cost. The orange dashed curves show saturation fits to these results, and the gray dashed lines indicate the Plan Once baselines.
TABLE IV: Canonical 40-D state and action layout. Dimensions are one-indexed. For a rotation matrix 𝐑=[𝐫1,𝐫2,𝐫3], we use Rot6D⁡(𝐑)=[𝐫1⊤,𝐫2⊤]⊤.
CoordinatesDimensionsState representation
Left EEF position1–3Cartesian position in meters
Left EEF orientation4–9Rot6D⁡(𝐑L)
Right EEF position10–12Cartesian position in meters
Right EEF orientation13–18Rot6D⁡(𝐑R)
Left gripper19native opening coordinate
Right gripper20native opening coordinate
Waist21–22two native coordinates
Planar base velocity23–24two native coordinates
Left arm joints25–32q1L,…,q8L in radians
Right arm joints33–40q1R,…,q8R in radians
TABLE V: Maximum duration of each physical-robot trial.
TaskMaximum duration
Clean Room20 min
Prepare Ingredients20 min
Tomato and Egg Stir Fry20 min
Make Milk Tea10 min
Book Organization5 min
Collect Laundry5 min
Tidy Makeup Table (each group)5 min
TABLE VI: High-level instance families, synthesized by perturbing only the input memory while reading the corrected target from the demonstration. ℳn is the memory upon entering segment n. A single unified <think>/<memory>/<subtask> format instantiates all of them at zero extra annotation.
FamilySampling positionInput → target memoryTarget subtaskDeployment failure counteredMix
within-subtaskanywhere in seg. nℳn→ℳnseg. n— (aligned, normal progression)58%
transitiontail of seg. nℳn→ℳn+1seg. n+1starting a new subtask after completion15%
catch-uphead of seg. nℳn−1→ℳnseg. nmemory lag (behind the visual state)10%
rollbacklate in seg. nℳn+1​…​n+3→ℳnretry seg. nmemory run-ahead (over-optimistic)12%
error-thinkannotated failure frameℳn→ type-dependentrecovery stepunnoticed execution failure5%

왜 중요한가

긴 작업을 하는 로봇은 잘못된 하위 결정을 내리면 아무리 손동작이 정확해도 실패하는데, 이 연구는 어려운 순간에만 더 오래 생각하게 만드는 방법을 실물 로봇으로 검증했다. 이는 언어모델의 테스트 타임 연산 확장 아이디어를 로봇 제어에 옮긴 사례로, 앞으로 범용 로봇 시스템 설계에 참고가 될 수 있다.

이 논문의 용어

  • VLA (Vision-Language-Action) 모델 · 카메라로 본 장면과 언어 지시를 받아 로봇 동작을 출력하는 인공지능 모델
  • 테스트 타임 연산(Test-Time Computation) · 모델을 다시 학습시키지 않고, 실제 사용 시점에 더 많은 계산을 들여 답의 질을 높이는 방법
  • 세계모델(World Model) · 어떤 행동을 하면 환경이 어떻게 바뀔지 미리 예측하는 모델
  • 빔서치(Beam Search) · 여러 가능한 다음 수 중 점수가 높은 몇 개만 남기고 계속 확장해 나가는 탐색 방법
  • 실행 메모리(Execution Memory) · 지금까지 로봇이 어디까지 작업을 끝냈는지 요약해 기억해두는 정보

논문 원문 초록 (영문)

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

저자 · Xiaowei Cai

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Xiaowei Cai et al., arXiv:2608.16885, arxiv-nonexclusive