One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Alibaba unveils Wan-Animate-2, adds real-time streaming generation

Upgrade to open-source character animation model includes multi-character support and text-driven camera control

이미지: METAL LAB 생성

Summary

  • The official Wan account unveiled Wan-Animate-2, the successor to its open-source character animation model, on August 11 (UTC)
  • The company listed four upgrades: high-fidelity character animation, multi-character support, text-directed camera viewpoints, and real-time streaming generation
  • Weights, demos, and documentation were released together, making it available for direct download and use
Video from the source
모델명
Wan-Animate-2
발표 주체
Wan 공식 계정(@Alibaba_Wan)
발표 시점
2026년 8월 11일(UTC)
성격
오픈소스 캐릭터 애니메이션 모델의 대규모 업그레이드
핵심 기능
고충실도 캐릭터 애니메이션, 다중 캐릭터 애니메이션
제어 방식
텍스트로 카메라 시점 지정
생성 방식
실시간 스트리밍 생성
공개 자료
가중치(weights), 데모, 문서

It's not just one person performing — multiple people can act simultaneously in the same scene. Camera angles are directed through text prompts. And you don't have to wait for the entire output video to finish rendering. These are the claims behind Wan-Animate-2, unveiled by Alibaba's Wan video model series through its official account on August 11 (UTC).

What Was Added

Wan presented four upgrades. The announcement itself is brief — just a list of items, with no benchmark scores or parameter counts included in the post. However, the company noted that weights, demos, and documentation were released together — meaning this isn't just a paper preview, but a release available for download and use right now.

FeatureOriginal TermWhat It Means
High FidelityHigh Fidelity Character AnimationKeeps a character's face, clothing, and details consistent frame to frame
Multi-CharacterMulti-Character AnimationAnimates two or more people simultaneously in one scene
Camera ControlText-Controlled Camera ViewpointDirects viewpoint and angle through text prompts
Real-TimeReal-Time Streaming GenerationOutput streams out rather than requiring the user to wait for generation to complete

Why Character Animation Is Hard

What the Wan-Animate series does differs from generating a new video from a single line of text. It takes a still image of a character and a reference video of an actual person moving, then transfers that motion and expression onto the character. This corresponds to the retargeting work animators used to do by hand.

The difficulty lies in consistency. A face can subtly turn into a different person a few frames later, clothing patterns can shift slightly, and fingers can become distorted. Difficulty jumps again when a second character is added — the model starts confusing whose arm belongs to whom, and boundaries break down the moment two people overlap. This is why "multi-character" was called out as a separate item.

Real-time streaming is a different kind of challenge. Video generation typically means submitting a request and waiting anywhere from tens of seconds to several minutes. If output can stream as it's generated, it opens up uses like live-stream avatars or real-time character performances. However, this announcement doesn't specify the conditions for "real-time" — such as frame rate or resolution.

A Different Axis From Wan3.0

Though part of the same Wan family, the two models play different roles. Wan3.0, released as a public beta on August 6 (UTC), natively generates 30-second videos and features Omni-Reference, which accepts text, images, audio, video, and even documents and slides as reference input. It's a broad, general-purpose generator.

Wan-Animate-2 sits at the opposite end — a specialized tool for making a fixed character perform a fixed set of actions. This is the kind of work that demands more creative control.

The Strategy Behind Open Weights

The competition in video generation is splitting into two camps. One side offers access only through APIs and web services — a closed approach. The other releases weights so that anyone can run and modify the model on their own hardware — an open approach. Chinese labs have aggressively pushed the latter, and Wan is one of the leading examples of that trend.

Once weights are open, fine-tuning and workflow tools follow. Experiments — training the model further for a specific character, plugging it into a local pipeline, or chaining it with other models — pour out of the community within days. Even if closed models lead on quality, this is why open models often reach real production workflows first.

What Actually Changes

For creators making webtoons, virtual characters, or short ads, this opens room to move away from stitching together single-character clips to assemble a scene. If a shot of two characters facing each other in conversation can be made in one pass, the editing process itself shrinks. The ability to direct angles through text also lets creators try out different viewpoints without drawing storyboards.

Still, the gap between a short announcement and actual quality is always filled in by the community. Whether Wan-Animate-2 holds up will be answered within days by people actually running the released weights. What's confirmed for now is just one thing — tools for animating characters are moving from rendering you wait for, to generation that streams out as it happens.