METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅂTechnical words in the news

video-action model

An AI model trained to predict not just the video it generates, but also the movements and actions happening within that video.

In plain words

A video-action model doesn't just stop at drawing a plausible-looking video on screen — it also predicts how the objects or robots in that scene should move next. If a typical video-generation AI is a painter stringing together pretty pictures in sequence, a video-action model is more like a choreographer who draws the picture while simultaneously working out the action plan, such as 'this hand moves in this direction next.'

This kind of model matters when training machines that actually need to move their bodies, like robots or self-driving cars. If the ability to predict how the world in a video will change and the ability to predict what actions are needed to produce that change live inside the same model, a robot can infer actions just from watching video, without needing to be taught separate action data item by item.

Recently, multimodal models that generate images, video, and sound all at once have started going further, training action prediction within the same framework as well. Instead of generating video, sound, and movement separately, a single architecture is trained to understand all three at the same time.

How it shows up in the news

The article describes Black Forest Labs' FLUX 3 as having "trained image, video, audio, and action prediction together within a single architecture." This 'action prediction' is the core concept behind the video-action model. However, it's worth not confusing this with what was actually released on August 4 — only the text-and-image-to-video generation feature — while the robotics collaboration model that makes use of action prediction was introduced separately, only through a technical blog post. In other words, video-action model elements are built into the design of the larger FLUX 3 model, but the FLUX 3 video feature you can currently call via API does not itself control robots.

See also

Stories using this term

Browse every entry