WAN 2.7 comes from Alibaba's WAN AI video model family. It supports text-to-video, image-to-video, first-to-last frame control, up to 5 reference images, 4K resolution, and clips of 5, 10, or 15 seconds - all with roughly 30% cleaner output than its predecessor.
WAN 2.7 is an AI video generation model from Alibaba's WAN model family. Picsart made WAN 2.7 its default AI video model, running it across mobile native, mobile web, and desktop web. It supports every modality - text-to-video, image-to-video, first-to-last frame generation, up to 5 reference images, and audio synchronization — delivering noticeably cleaner output with sharper edges, more believable skin tones, and more grounded color balance.
WAN 2.7 capabilities
WAN 2.7 generates video at up to 4K resolution (4096×2160) in clips of 5, 10, or 15 seconds across all standard aspect ratios. It accepts up to 5 reference images as visual anchors - keeping characters, products, environments, and brand assets consistent throughout the video. Camera control responds to plain-language direction (pan, dolly, zoom), and audio sync supports ambient sound, dialogue-matched lip-sync, and background music in a single run.
What you can create with WAN 2.7
Create video scenes with multiple characters where every face and outfit stays consistent across the entire clip, powered by up to 5 reference images as visual anchors.
How Picsart uses WAN 2.7
Picsart runs WAN 2.7 as its default AI video model across mobile native, mobile web, and desktop web. Five reference images, first-to-last frame control, 4K output, natural-language camera direction, and tighter audio sync are all set as defaults across every video tool - no extra configuration needed.
WAN 2.7 is integrated into Picsart's AI Video Generator and AI Playground, where you can compare it with 130+ other AI models using a single prompt.
Why creators choose WAN 2.7
Creators choose WAN 2.7 for its multi-reference image support - up to 5 visual anchors that no other WAN model or competitor currently matches. Combined with 4K resolution, physically realistic motion, plain-language camera control, and integrated audio sync, it delivers the most versatile AI video generation workflow available in Picsart.
Explore more models like WAN 2.7
Compare WAN 2.7 with other video models for motion, ads, and social clips.
WAN 2.7 FAQ
WAN 2.7 is an AI video generation model from Alibaba's WAN model family. It supports text-to-video, image-to-video, first-to-last frame control, up to 5 reference images, 4K resolution (4096×2160), and clips of 5, 10, or 15 seconds with integrated audio synchronization.
Picsart runs WAN 2.7 as its default AI video model across mobile native, mobile web, and desktop web. All of its capabilities — reference images, frame control, 4K output, camera direction, and audio sync — are built directly into Picsart's AI Video Generator.
WAN 2.7 supports up to 5 reference images as visual anchors, which no other WAN model and no competitor currently matches. This keeps characters, products, and environments consistent throughout the video. It also delivers roughly 30% cleaner output than its predecessor, with sharper edges and more believable color balance.
No. WAN 2.7 works behind the scenes within Picsart's AI Video Generator. You describe what you want using text prompts and reference images — no technical video editing experience is required.
Picsart offers a free tier that includes up to 5 seconds of 720p video generation. Access to higher resolutions, longer clips, and additional features depends on your subscription plan.
WAN 2.7 generates video at up to 4K resolution (4096×2160) in all standard aspect ratios. Clip lengths are available at 5, 10, or 15 seconds.
Yes. Videos generated through Picsart's tools powered by WAN 2.7 can be used for marketing, social media, brand content, and other commercial applications, subject to Picsart's terms of use.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.