
Summary
- Alibaba's Wan account released Wan 3.0 in production, a clip from the eighth live session, on September 17.
- Filmmaker Keith Zhang named 30-second single-shot generation and uploading five videos to omni-reference as the most surprising features.
- He proposed generating long clips and cutting them to the script in post-production instead of building scenes clip by clip.
Alibaba's official account for Wan, its video generation model, released a clip from the eighth episode of its live sessions on September 17. The 93-second video is titled Wan 3.0 in production, and it takes on a single question: where does Wan3.0 fit into an actual film production? The person answering is Keith Zhang, the filmmaker behind Soulscape, joining from the United States, and the question comes from Johnny Mai of Alibaba Cloud. The post framed the clip as the point where Wan3.0 starts fitting into real production.
Length was the first thing Zhang pointed to. "It can do one shot 30 seconds, which is definitely crazy," he said in the video. In the past the model managed 5 seconds, then 15, and now a single 30-second shot comes straight out of the model. The post itself lays it out in the same order: 5-second clips a year ago, then 15 seconds, and now a single 30-second shot direct from the model.
The second was references. Mentioning director-level control, he said he was surprised that the omni-reference interface accepts up to five videos. "Five videos, not five images," he stressed. METAL reported at the Wan3.0 public beta launch that omni-reference goes beyond text, images, audio and video to take documents, spreadsheets, slides and webpages, and this clip confirms one piece of that, the number of video references, from someone who has actually used it.
The August 6 announcement card listed three items: native 30-second video generation, reality-grade rendering, and an omni-reference that reaches past text, images, audio and video to documents, spreadsheets, slides and webpages. That post carried the slogan simple input, smart creation, and said it was taking applications for the public beta. This live session revisits the first and last of those items, length and references, through a filmmaker's hands-on experience.
What changes in the workflow is that clip-by-clip work goes away. According to Zhang, there is no longer a need to generate each clip separately; you generate long clips and then do post-production around the script and the story. The cards in the video draw the same flow: don't generate clip by clip, generate long clips, do post-production based on script and story, and use a lot more references. The post summed it up as "No more stitching together endless short clips."
From an engineer's point of view, both numbers point at the same bottleneck. When you stitch 5-second clips together, the most time-consuming part is not generation but keeping continuity between one clip and the next, and if references can only be images, motion and camera work have to be described in words every time. A 30-second single shot cuts down the number of seams in the first place, and five video references are a channel for handing over motion as footage rather than description. Having both features together is what makes script-level work possible, and that is the backbone of the production use Zhang described.
In an earlier clip posted by the same account a day before, on September 16, Zhang talked about how far the model has moved. Asked by Mai what he noticed in the latest update and what surprised him, he set last year's Wan 2.1 and 2.2 beside this year's Wan3.0 and answered that it is "definitely more than a generational leap, like what happened within a year." The items he named were multi-reference, consistency, creative control and resolution, and he said that as the models move toward the next generation they become friendlier to creators, filmmakers and storytellers. That post wrote, "The tools are finally getting out of the storyteller's way."
Both posts METAL reviewed end with a DingTalk sign-up form link for people who want to bring Wan3.0 to their team. The case videos double as an enterprise sales channel. The September 17 clip carries Wan3.0 and the names of both speakers in the top-left corner of the screen; the September 16 clip runs 65 seconds and the September 17 clip 93 seconds. At the time of checking, the September 16 clip had a little over 9,000 views, and the September 17 clip, only hours old, was around 1,800. Set against the August 6 public beta announcement, which passed 20 million views, this episode is not a broad announcement but a demonstration aimed at people who already use the model.
The Alibaba Wan account is building a series out of voices from the production floor. METAL reported that a commercial director in London finished an ad with Wan3.0 without color grading, and now a filmmaker in the United States has placed the same model inside a long-form production workflow. Ads and films make different demands: the earlier case was about the look, and this one is about length and references.
The 30-second single shot and the five video references were specifications on the August announcement card, and this clip records those specifications coming back out of the mouth of someone who has actually worked with them. Whether they stay numbers on a spec sheet or harden into a production workflow will depend on how quickly testimony like this piles up.





Comments