
이미지: @pika_labs (X) 영상 갈무리 · METAL LAB 편집
Summary
- Pika revealed that it used the Pika API Club to feed WAN 3.0 multiple reference images and a highly detailed prompt, producing a 30-second period drama video.
- The prompt assigned distinct roles to each of the three reference images — the female character's face, the male character's face, and the background environment — and called for a single shot with synchronized dialogue.
- WAN 3.0 is a video generation model released in public beta on August 6 by Alibaba's Wan team, and Pika offers access to it through its own API Club.
- 모델
- WAN 3.0 (개발: 알리바바 Wan)
- 제공 경로
- Pika API Club
- 생성 모드
- R2V(Reference-to-Video), 16:9, 30초 단일 숏, 영어 대사 동기화
- 레퍼런스 수
- 이미지 3장 (인물 2명 얼굴·의상 + 배경 환경)
- 장면 설정
- 19세기 초 영국 전원 로맨스(period drama)
- WAN 3.0 공개일
- 2026년 8월 6일(UTC), 퍼블릭 베타
Two faces, one background — reference images split by role
Pika posted a 30-second period drama clip on its account, made using Alibaba's video generation model WAN 3.0. What stands out in the demo is how the three reference images were each assigned a different role. The prompt included one photo capturing the face and costume of the female character "Miss Hartley," another photo capturing the face and costume of the male character "Elias," and a third photo of a background scene showing a meadow and a manor house — each designated as a separate control target. Pika isn't the developer of this model; it's the party that makes WAN 3.0 available through the Pika API Club it runs.
WAN 3.0 is a video model built by Alibaba
WAN 3.0 is a model released in public beta on August 6 (UTC) by the Wan team, Alibaba's video generation model lineup. It touts native generation of 30-second clips in one pass along with "reality grade" rendering, and features an Omni-Reference capability that accepts not just text, images, audio, and video as references but also documents, spreadsheets, slides, and web pages. Pika's demo showcases the R2V (Reference-to-Video) mode within that reference feature set — the one that lets users separate images by character and by environment.
The full prompt
The image Pika posted included the complete prompt used. It laid out the mode, the scope of control for each character, and even the emotional arc of the scene in detail, so here it is in full.
The mode was specified first: "MODE: R2V, one coherent 30-second 16:9 live-action sequence with synchronized English dialogue."
Reference roles were spelled out image by image. It stated, "Image 1 controls only Miss Hartley's adult face, natural skin tone, dark low updo, proportions and pale full-length period dress," clearly defining the scope for the first character, and then split off the second character's scope with "Image 2 controls only Elias's adult face, dark wavy hair, proportions and refined coatless wardrobe: tailored waistcoat, neckcloth, fitted trousers and low leather shoes." The background image was assigned with "Image 3 controls only the environment layout and atmosphere: high meadow, large mature tree at screen-right, distant manor at screen-left, rolling countryside and late-afternoon light."
There was also an instruction meant to prevent confusion: "Ignore white character-sheet backgrounds, duplicate views and reference poses. Show one woman and one man only."
The scene's emotional arc was spelled out too. The line "The scene's emotional turn is that an intelligent, self-possessed woman realizes a wealthy young gentleman is quietly promising to remain" pins down the core of the story, while the cinematography was set with "Fine 35mm grain, gentle halation, soft highlights, mild lens bloom and natural skin. Warm late-afternoon side-backlight, muted greens and cool." The original text cuts off mid-sentence at this point.

How to try it on Pika API Club
Pika said it made this video through the Pika API Club. Pika API Club is a service Pika operates that lets users plug in external models like WAN 3.0 via API. The tweet also included a link to the API Club.
This approach of assigning separate images to characters and to the background is worth keeping in mind for videos where character consistency matters, such as advertisements or short-form dramas. For instance, if you're making a brand video that needs to keep the same character's face consistent across multiple scenes, or a web drama trailer that needs the same actor shot against different backgrounds, the role-splitting structure from this prompt could be applied directly.
Editor's take
Lately the conversation around video generation models has shifted away from resolution or length and toward how finely a model can parse separate references. What WAN 3.0 demonstrated here is the ability to process two faces and one background without letting them bleed into each other, which looks like an attempt to solve the "character consistency" problem — something Runway and the Midjourney lineage have wrestled with for a long time — through a different approach: assigning explicit roles to each reference image. Expanding the types of reference input, as Omni-Reference does, and explicitly assigning roles to each reference image, as seen here, are two separate axes of progress, and the fact that Alibaba's Wan team is pushing on both at once is notable.
Anyone who's actually put reference-based video models like this to practical use tends to hit the same wall. The moment more than one character appears, faces start blending, or the model ends up copying a pose from the background image's subject. The fact that this prompt bothered to include the line "ignore white character-sheet backgrounds, duplicate views and reference poses" reads like it reflects a practitioner's hard-won experience trying to prevent exactly that kind of contamination.
Teams in Korea working with video content might want to get in the habit of spelling out the control scope for each reference directly in the prompt. Structuring it as "Image 1 controls only the face, Image 2 controls only the outfit, Image 3 controls only the background" gives you a framework you can reuse even as the underlying model changes. That said, demanding synchronized dialogue within a single 30-second shot is still something only a handful of API providers attempt, so it's safer to validate this in storyboarding or previs first rather than deploying it directly into commercial production.
Given how Pika API Club keeps rolling out external models like WAN 3.0, it seems likely that within weeks we'll see prompt examples controlling three or more characters at once, or demos that extend dialogue beyond English.



Comments