Wan 3.0 Prime: the full Wan 3.0 toolset, up to 7x faster
Wan 3.0 Prime is the speed-tuned member of Alibaba's Wan 3.0 AI video family: the complete Wan 3.0 toolset - native 30-second single-take video, text, image, audio, and video references, start and end frames, adaptive ratio, smart duration, optional synced audio, and up to 1080p output - running at up to 7x the generation speed. Same range, far less waiting. Available now in Picsart's AI Playground.
Prime keeps everything Wan 3.0 can do and runs it at up to 7x the generation speed - so a 30-second single take, smart duration, and every reference type come back in a fraction of the wait. Nothing is stripped out to get there: describe the action and smart duration reads the pacing you implied, then sets the length to match, up to a full 30 seconds in one continuous take. Because the clip runs unbroken, light, motion, and pacing hold all the way through instead of resetting at a splice - and you can extend any finished clip to keep building.
EVERY INPUT IN ONE PANEL
Text, image, video, and audio - all as reference
Wan 3.0 Prime takes the full range of references in a single panel: a text prompt, images (including start and end frames), reference video, and reference audio. Pin the first and last frame and adaptive ratio shapes everything in between, while optional synced audio is generated right alongside the visuals. A Deep Thinking mode helps it parse long, multi-part instructions, so a detailed request survives all the way to the final frame - now at Prime speed.
Speed doesn't cost quality. Put words on screen and Wan 3.0 Prime renders them legibly and accurately, as text you can actually read - which counts for most in busy, information-heavy scenes. Detail sits close to real footage throughout. References hold at pixel level, so characters, objects, scenes, styles and audio all stay themselves across the sequence, and motion and emotion carry real range. Pin the first and last frame, and adaptive ratio shapes everything in between.
INSIDE PICSART
Find the right Wan model for your video
Wan 3.0 Prime lives in Picsart's AI Playground, where a single prompt runs against 150+ other models at once so you see the difference before committing. It's the full Wan 3.0 toolset at up to 7x the speed - pick Prime when turnaround matters most. Prefer maximum resolution? Wan 2.7 still holds the 4K advantage. And you can reach Wan 3.0 Prime whichever way you work: on the web, in the desktop app, or built into your own projects via CLI, MCP, REST API, and SDK - with no third-party API keys or separate subscriptions to manage.
What you can create with Wan 3.0 Prime
Bring a text prompt, an image, a reference clip, or an audio track - or pin start and end frames - and let Wan 3.0 Prime build from what you already have, at up to 7x the speed.
Explore more models like Wan 3.0 Prime
Compare Wan 3.0 Prime with other video models for motion, ads, and social clips.
Wan 3.0 Prime FAQ
Wan 3.0 Prime is the speed-tuned member of Alibaba's Wan 3.0 AI video family. It delivers the complete Wan 3.0 toolset - native 30-second single-take video at up to 1080p, text/image/video/audio references, start and end frame control, adaptive ratio, smart duration, and optional synced audio - at up to 7x the generation speed.
Wan 3.0 Prime runs the same Wan 3.0 feature set at up to 7x the speed. Nothing is removed to get there - you get the full toolset with far less waiting. Choose Prime when turnaround matters most, and standard Wan 3.0 when you don't need the extra speed.
Up to 30 seconds in one continuous generation. Smart duration control suggests a length based on the action in your prompt, and the extend function lets you build further from a finished clip.
Text prompts, images (including start and end frames), reference video, and reference audio. Wan 3.0 Prime can also generate optional synced audio alongside the video, and includes a Deep Thinking mode for parsing detailed, multi-part prompts.
Wan 3.0 Prime generates at 480P, 720P, and 1080P, with aspect ratios including 16:9, 9:16, 1:1, 4:3, 3:4, and adaptive.
Wan 3.0 Prime is available in Picsart's AI Playground. You can also use it on the web, in the Picsart desktop app, and build it into your own projects via CLI, MCP, REST API, and SDK.
Yes. Videos generated through Picsart's tools powered by Wan 3.0 Prime can be used for marketing, social media, brand content and other commercial applications, subject to Picsart's terms of use.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.