METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

NVIDIA proposes robot policy trained on video instead of vision-language models

Cosmos 3-based World Action Model suggested as a replacement for VLA robot policy architecture

NVIDIA proposes robot policy trained on video instead of vision-language models

Summary

  • NVIDIA introduced the concept of a World Action Model (WAM), built on video world models rather than vision-language models (VLMs)
  • The open model Cosmos 3, a Mixture-of-Transformers architecture trained on hundreds of millions of image, video, and action samples, was presented as the foundation for WAM
  • WAM is said to learn physical dynamics first, strengthening zero-shot transfer performance to new tasks, robots, and environments

NVIDIA has proposed the World Action Model (WAM) as a new foundation for robot policies. Existing vision-language-action (VLA) models work by adding an action module to a pretrained vision-language model (VLM); while strong at describing scenes, they were noted to be weak at modeling dynamics — that is, predicting how a scene will change.

NVIDIA proposes robot policy trained on video instead of vision-language models
이미지: NVIDIA Dev Blog

WAM replaces this language backbone with a video world model, pretraining on physical changes such as how a cup deforms when grasped or how a towel folds. NVIDIA's research paper, "World Action Models are Zero-shot Policies," explains that jointly predicting video and action gives the model a capability that VLA struggles to achieve: the ability to learn from diverse data.

NVIDIA proposes robot policy trained on video instead of vision-language models
이미지: NVIDIA Dev Blog

Cosmos 3, an open model, was presented as the foundation for this approach. Cosmos 3 uses a Mixture-of-Transformers architecture and was trained on a large-scale multimodal dataset containing hundreds of millions of image, video, and action samples. NVIDIA stated that Cosmos 3 supports deployment across multiple tiers, from serving on high-performance workstations to real-time on-device inference using Jetson hardware, and significantly reduces the amount of task-specific data needed to adapt to different robot forms.

Comments