METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅊTechnical words in the news

Reference-to-Video

A video generation method that assigns different roles to reference images—like a person's face or a background—so their distinct features are preserved in the output.

In plain words

Reference-to-Video is a way of feeding an AI video generator reference images, such as photos of a face or a background, while telling it in advance exactly what part of the final video each image is responsible for. Think of it like a film shoot: a director hands the actor's headshot to one crew member with instructions to reference only the face, and hands a set photo to another crew member with instructions to reference only the background.

In the past, if you gave an AI one or two reference images, it would guess on its own how to blend the overall mood and look. Reference-to-Video is different because you can explicitly split up the roles—'this photo controls only the face,' 'that photo controls only the outfit,' 'this other photo controls only the background.' That makes it possible to generate a complex scene with multiple characters and a separate setting all in one shot, without the different reference images bleeding into each other.

How it shows up in the news

In the article, Pika used Alibaba's video generation model WAN 3.0 to create a 30-second period-drama scene, assigning one photo of a female character's face, one photo of a male character's face, and one background photo to separate roles using Reference-to-Video (R2V) mode. A common point of confusion: R2V isn't a proprietary technology name from one specific company—it refers to the general feature of feeding in reference images split by role. Pika didn't build WAN 3.0; it's the service that makes the model available through its platform.

Try it yourself

If your video generation tool supports multiple reference images, try locking down each image's role in your prompt like this:

Image 1 controls only Character A's face and outfit. Image 2 controls only Character B's face and outfit. Image 3 controls only the background and atmosphere.

Spelling out the roles this way makes it far more likely that the characters and background stay separate and each faithfully reflects its assigned reference image.

See also

Stories using this term

Browse every entry