
이미지: X — 모델·오픈소스
Summary
- Wan, Alibaba's video generation model family, announced via its official account that it has released Wan3.0 in public beta.
- It presented native generation of 30-second-long videos and "reality-grade" rendering as its core features.
- It highlighted Omni-Reference, which expands reference inputs from text, image, voice, and video to documents, spreadsheets, slides, and web pages.
- 모델명
- Wan3.0
- 공개 상태
- 퍼블릭 베타(Public Beta)
- 발표 주체
- Wan 공식 계정(@Alibaba_Wan)
- 발표 시점
- 2026년 8월 6일(UTC)
- 영상 길이
- 네이티브 30초 영상 생성
- 렌더링
- Reality-Grade Rendering으로 표기
- 참조 입력
- 텍스트·이미지·음성·영상 + 문서·스프레드시트·슬라이드·웹페이지
- 슬로건
- "Simple Input. Smart Creation." (Wan 공식 계정)
Alibaba's Wan Enters Public Beta with 3.0
Wan, Alibaba's video generation model family, has released its next version, Wan3.0, in public beta. According to a post from its official account, the new version was introduced with the tagline "Simple Input. Smart Creation." and was opened immediately as a beta accessible to general users. The post was published on August 6, 2026 (UTC).

Native 30-Second Generation and Rendering Quality
The first item in the announced feature list is native generation of 30-second-long videos. This was presented to mean generating a full 30 seconds at once rather than stitching together multiple clips, though the announcement did not include implementation details or specifications such as resolution or frame rate. The second item, "Reality-Grade Rendering," was used to describe near-photorealistic rendering quality, but no benchmark figures or comparison points were disclosed to support the claim.
Omni-Reference: Expanded Reference Inputs
The most notable item is Omni-Reference. It was explained as accepting not only the existing text, image, voice, and video references but also documents, spreadsheets, slides, and web pages as reference inputs. This reads as a direction toward directly incorporating tabular data, presentations, or web documents as source material for video production, but since the original post was cut off midway, it is unconfirmed whether the listed input types end there.
| Item | Wan3.0 Announcement Details |
|---|---|
| Release Form | Public beta |
| Video Length | Native 30 seconds |
| Rendering | Reality-Grade Rendering |
| Reference Inputs | Text, image, voice, video, documents, spreadsheets, slides, web pages |
| Disclosed Metrics | None |
What Remains Unconfirmed
The announcement made no mention of whether model weights would be released, supported languages, access channels, or commercial usage terms. Regardless of whatever distribution approach the Wan family has taken in the past, the currently available material offers no basis for assuming it applies directly to Wan3.0. Since this information is based on a single post, specifications may become clearer once a technical report or official documentation follows.
The competitive landscape for video generation models and recent trends in multimodal model releases can be followed further in metallab.ai's video generation coverage.
Points to Watch
The axis of competition in video generation is shifting toward length, realism, and controllability. The three features Wan3.0 has put forward each correspond to one of these axes. In particular, the structure of accepting documents and spreadsheets as references can be seen as designed to target production workflows—such as advertising, reports, or educational content—where source material already exists. However, actual quality and generation stability can only be judged once results accumulate from beta users.


