Alibaba has opened beta testing for Wan3.0, an artificial intelligence model capable of generating videos up to 30 seconds long from inputs that include text, images, video, audio, webpages, PDFs, and PowerPoint presentations.
The maximum output is twice the 15-second limit of its predecessor, Wan2.7-Video. Alibaba said the longer duration is intended to accommodate continuous shots, more complex camera movements, and extended narratives.
Wan3.0 also includes an intelligent duration function that recommends a video length based on the user’s prompt, as well as tools for extending existing clips.
“Mainstream AI video generators typically produce clips lasting only a few seconds to 15 seconds, which is the maximum clip duration of Wan2.7-Video, the preceding Wan video generation model,” Alibaba said.
The model can process several types of reference material simultaneously, allowing users to convert text-heavy documents and webpages into video content.
Alibaba said Wan3.0 was designed to reduce visual drift and distortion, recurring problems in AI-generated video.
The model seeks to maintain consistency in characters, products, objects, spatial arrangements, visual styles, voices, and software interfaces across a clip.
It can also generate multilingual speech and synchronize facial expressions with audio, according to the company.
Alibaba is positioning Wan3.0 for film production, short-form dramas, social media, marketing, and educational content.
The company said developers could also use it to produce simulation videos for training autonomous-driving and robotics systems.
Users may apply to test the beta version through Alibaba Cloud’s Model Studio and Qwen Cloud platforms.
Alibaba introduced the Wan series of image and video generation models in July 2023.


