AI GlossaryㅊTechnical words in the news
Reference-to-Video
A video generation method that assigns different roles to reference images—like a person's face or a background—so their distinct features are preserved in the output.
In plain words
Reference-to-Video is a way of feeding an AI video generator reference images, such as photos of a face or a background, while telling it in advance exactly what part of the final video each image is responsible for. Think of it like a film shoot: a director hands the actor's headshot to one crew member with instructions to reference only the face, and hands a set photo to another crew member with instructions to reference only the background.
In the past, if you gave an AI one or two reference images, it would guess on its own how to blend the overall mood and look. Reference-to-Video is different because you can explicitly split up the roles—'this photo controls only the face,' 'that photo controls only the outfit,' 'this other photo controls only the background.' That makes it possible to generate a complex scene with multiple characters and a separate setting all in one shot, without the different reference images bleeding into each other.
How it shows up in the news
In the article, Pika used Alibaba's video generation model WAN 3.0 to create a 30-second period-drama scene, assigning one photo of a female character's face, one photo of a male character's face, and one background photo to separate roles using Reference-to-Video (R2V) mode. A common point of confusion: R2V isn't a proprietary technology name from one specific company—it refers to the general feature of feeding in reference images split by role. Pika didn't build WAN 3.0; it's the service that makes the model available through its platform.
Try it yourself
If your video generation tool supports multiple reference images, try locking down each image's role in your prompt like this:
Image 1 controls only Character A's face and outfit. Image 2 controls only Character B's face and outfit. Image 3 controls only the background and atmosphere.
Spelling out the roles this way makes it far more likely that the characters and background stay separate and each faithfully reflects its assigned reference image.
See also
Stories using this term
- Alibaba Opens Wan3.0 Public Beta, Native 30-Second Video GenerationAI · 2026.08.09
- Alibaba unveils Wan-Animate-2, adds real-time streaming generationAI · 2026.08.11
- Runway's Ruby converts AI video to broadcast HDR standardsCreative · 2026.08.25
- MiniMax H3 video generation now outpaces playback timeCreative · 2026.09.02
- Alibaba's Wan Adds 'Agent Skills' for Sharing Video WorkflowsAI · 2026.08.14
- Luma unveils 'Luma Scenes,' letting creators approve shots before final renderAI · 2026.08.12
