AI GlossaryㅂTechnical words in the news
Vision-Language-Action Model
An AI architecture that lets a robot understand what it sees and what it's told, then directly generate the movement commands to act on it
In plain words
A Vision-Language-Action Model (VLA) is a single, connected system that lets a robot understand a scene captured by its camera together with an instruction given in speech or text, and then directly compute how its limbs should move.
Think of a tour guide. The guide looks at the surroundings (vision), understands what the visitor says (language), and then actually walks and leads the way (action). A VLA works the same way: it looks at a cup through a camera, receives the instruction "pick up the cup," and produces the exact degree by which each arm joint should move, all in one step.
These models are usually built by taking an existing AI that has already been trained to understand images and text, and adding a new module on top that handles physical movement. However, this approach has been criticized for being strong at describing a scene in words while being weak at predicting physical changes, such as how an object gets pressed or deformed when grasped. Because of this, there are recent attempts to replace the language-based pretraining with models that learn physical phenomena from video instead.
How it shows up in the news
The article explains that "existing Vision-Language-Action (VLA) models work by adding an action module on top of a pretrained vision-language model." A common misconception here is thinking VLA is the only way to control a robot. In reality, it's just one of several approaches, and in this article NVIDIA proposed an alternative that learns physical phenomena from video instead of language.
See also
Stories using this term
- NVIDIA proposes robot policy trained on video instead of vision-language modelsAI · 2026.08.09
- Google Unveils Gemini Robotics 2 for Full-Body Robot ControlAI · 2026.08.04
- NVIDIA Unveils 34B-Parameter Reasoning Model for Autonomous DrivingAI · 2026.08.09
- Microsoft Unveils Video-Capable Vision-Language Model Mage-VLAI · 2026.08.10
- NVIDIA unveils 'Ising Calibration 1.5' VLM for automated quantum computer calibrationAI · 2026.08.09
- French-specialized small AI 'Luth-2' outperforms models three times its sizeAI · 2026.08.11
