WAN 3.0 AI Video Model - 30-Second Video Generation
AI VIDEO MODELS
WAN 3.0: native 30-second AI video generation
WAN 3.0 is the newest model in Alibaba's WAN AI video family, and the first that holds a shot for a full 30 seconds without cutting away. Feed it text, images, audio, video - or a document, or a web page URL - and it reads all of it as reference. 1080p, adaptive ratio, smart duration.
30 seconds, in one continuous take. The more useful part is that you no longer have to pick the number: describe the action and smart duration reads the pacing you implied, then sets the length to match. Because the clip runs unbroken, light, motion and pacing hold all the way through instead of resetting at a splice. Need longer? Extend any finished clip and keep building from a take that already works.
Point it at a document, or a web page
The 4 references you already use are all here - text, image, audio, video. WAN 3.0 adds two more. Drop in a document (.doc, .pdf, .ppt, .xls) and it reads the thing. Or just give it a URL: a product page, an article, a research paper, your own site. The spec sheet becomes the ad film. The landing page becomes the launch video. Long, multi-part instructions get parsed properly too, so a detailed request survives all the way to the final frame.
What you can create with WAN 3.0
Hand over the deck, the PDF or the spec sheet you already wrote - or just the product page URL - and let the model build from it instead of starting at a blank prompt box.
Put words on screen and WAN 3.0 renders them legibly and accurately, as text you can actually read - which counts for most in busy, information-heavy scenes, where there is the most to get wrong. Detail sits closer to real footage throughout. References hold at pixel level, so characters, objects, scenes, styles and audio all stay themselves across the sequence, and motion and emotion carry more range. Pin the first and last frame, and adaptive ratio shapes everything in between.
WAN 3.0 in Picsart: Find the right model for your video
You can find WAN 3.0 in Picsart's AI Playground, where a single prompt runs against 150+ other models at once and you see the difference before committing. It generates at 1080p, at roughly one to two seconds of render per second of finished video, with every input type in the same panel. WAN 2.7 sits alongside it: 15 seconds in fixed blocks of 5, 10 or 15, up to 5 reference images, and it still holds the 4K advantage. Pick 3.0 for length, input range and readable text. Pick 2.7 when resolution matters most.
Explore more models like WAN 3.0
Compare WAN 3.0 with other video models for motion, ads, and social clips.
WAN 3.0 FAQ
WAN 3.0 is the latest model in Alibaba's WAN AI video family. It generates a single continuous clip of up to 30 seconds at 1080p, accepts text, image, audio, video, document and web page URL references, and supports start and end frame control with adaptive aspect ratio.
Up to 30 seconds in one continuous generation, double the 15-second ceiling of WAN 2.7. Smart duration control suggests a length based on the action in your prompt, and the extend function lets you build further from a finished clip.
Text, images, audio and video, plus two new types: documents (.doc, .pdf, .ppt, .xls) and web page URLs. Point the model at a product page, article or spec sheet and it uses that content as reference material for the video.
3 things: 30 continuous seconds instead of 15, document and web page inputs on top of the existing four reference types, and pixel-level consistency across characters, objects, scenes, styles and audio. Text rendering in information-dense scenes is also significantly stronger.
WAN 3.0 is available in Picsart's AI Playground, where you can compare it against 130+ other AI models using a single prompt. All input types, duration control and frame settings are available from the same panel.
No. You describe what you want in plain language and add whatever references you have - a photo, a clip, a document, a link. No technical video editing experience is required.
Yes. Videos generated through Picsart's tools powered by WAN 3.0 can be used for marketing, social media, brand content and other commercial applications, subject to Picsart's terms of use.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.