Next-gen cinematic video with optional audio and reference image. Up to 4K.Reference inputAudio4KCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Fast 4K generation with accurate text and search-grounded accuracy.4KFast generationImage generationSee model
Top-tier 4K images with precise multilingual text rendering.4KPro qualityImage generationSee model
Adaptable generation across varied visual styles up to 4K.4KImage generationSee model
Speedy 3K output with negative prompt and dual-image input support.Reference inputFast generationImage generationSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.AudioPro qualityCinematicSee model
Flexible generation across creative styles using V3 Omni architecture, with optional 4K output.4KCinematicVideo generationSee model
O1-architecture video generation with 5 or 10 second output.CinematicVideo generationSee model
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Lip-synced talking head from a portrait and speech audio.AudioCinematicVideo generationSee model
Cinematic visuals with up to 4K resolution and 10 reference images.Reference input4KCinematicSee model
O1-architecture image generation with multi-reference support.Reference inputCinematicImage generationSee model
Apply curated Kling visual effects to photos — single or dual-image scenes.CinematicPhotorealVideo generationSee model
Text-to-audio clips of 3–10 seconds from a prompt description.AudioCinematicMusic generationSee model
Extract or generate a matching audio track from an uploaded video.AudioCinematicMusic generationSee model
4K output with audio and aspect ratio control — production-ready v2.3.Audio4KPro qualitySee model
Fast 2.3 with long video support — up to 20s at 1080p with aspect ratio control.1080pFast generationVideo generationSee model
Generate video driven by an audio track — 2-20s, optional image for first frame.AudioVideo generationSee model
Seamlessly extend an existing video forward or backward — up to 20s.Video generationSee model
Retake video with new direction — replace audio, video, or both.AudioVideo generationSee model
Product showcase from still images with gentle camera motion.Video generationSee model
Image-driven video with layered ambient atmosphere and optional audio.AudioVideo generationSee model
Quick ambient video from images with optional audio overlay.AudioFast generationVideo generationSee model
Straightforward text/image-to-video at 720p with broad style coverage.
Animate a portrait with realistic body movement driven by audio.AudioPhotorealVideo generationSee model
Denoise, color-correct and super-resolve existing footage up to 8K, with frame-rate conversion.Video generationSee model
Turn a still photo into polished video with automated composition.PhotorealVideo generationSee model
Infographic-friendly generation with readable text and cfg control.Image generationSee model
Stylized 768p animation with strong character expression and emotion.Video generationSee model
1080p output focused on detailed scenes and polished short-form content.1080pPro qualityVideo generationSee model
Quick 768p previews with expressive characters for rapid experimentation.Fast generationVideo generationSee model
Fast 1080p output for short, polished clips with varied styles.1080pFast generationPro qualitySee model
MiniMax H3 2K video from text, start/last frame, or image/video/audio references.Reference inputAudioVideo generationSee model
Wan 2.7 R2V — generate video from reference images/video with style direction.Reference inputCinematicVideo generationSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Wan 3.0 Prime — the same all-in-one model as Wan 3.0, up to 7x faster.Fast generationCinematicVideo generationSee model
Smooth video with a dreamy, polished aesthetic — up to 4K resolution.4KCinematicVideo generationSee model
Quick image-to-video with smooth, stylized motion — up to 4K.Image to video4KFast generationCinematicSee model
Reframe a video to a new aspect ratio using Luma Ray 2.CinematicVideo generationSee model
Reframe a video to a new aspect ratio using Luma Flash 2.Fast generationCinematicVideo generationSee model
Luma UNI-1 — agentic image generation and editing with up to 9 reference images.Reference inputCinematicImage generationSee model
Luma UNI-1 Max — higher-quality UNI-1 variant with the same multi-reference editing controls.Reference inputCinematicImage generationSee model
Luma Ray 3.2 — high-fidelity video generation with start/end frames, HDR, and looping (early access).CinematicVideo generationSee model
Edit a prior video from a prompt using Luma Ray 3.2 — preservation-vs-reimagination presets (early access).CinematicVideo generationSee model
Reframe a video to a new aspect ratio using Luma Ray 3.2 (early access).CinematicVideo generationSee model
Latest cinematic video with audio, multi-reference input, and mp4/mov output. Up to 30s.Reference inputAudioCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Stitch up to 10 clips into one continuous, extended video.Video generationSee model
Lightweight cinematic video with audio, reference images, and start/end frame control.Reference inputAudioCinematicSee model
Lightweight video edit — modify scenes with reference images.Video editingReference inputVideo generationSee model
Stitch up to 3 clips into one continuous, extended video.Video generationSee model
Quickly stitch up to 3 clips into one continuous video.Fast generationVideo generationSee model
Lightweight: stitch up to 3 clips into one continuous video.Video generationSee model
Top-tier single-image generation with up to 10 reference images and 2K detail.Reference inputPro qualityImage generationSee model
Reliable all-purpose generation with readable text overlay.Image generationSee model
Detailed 4K renders with clean in-image text and dual-image input.Reference input4KImage generationSee model
Synthesize natural speech in 20 languages — pick a named voice or clone one from a reference audio.Reference inputAudioMusic generationSee model
Synthesize natural English or Chinese speech — pick a named voice or clone one from a reference audio.Reference inputAudioMusic generationSee model
Fastest generation pipeline — 720p with audio in seconds, up to 15s.AudioFast generationVideo generationSee model
Next-gen Grok video — faster, higher fidelity, up to 15s with audio.AudioFast generationVideo generationSee model
Restyle or remix an existing video with a new prompt direction.Video editingVideo generationSee model
Extend an existing video forward with a new prompt — up to 10 seconds.Video generationSee model
Rapid image creation with wide aspect-ratio selection and image input.Reference inputImage generationSee model
Higher-fidelity Grok Imagine variant for production-grade images.Image generationSee model
Latest Grok Imagine generation — sharper detail with a low/medium quality tier.Image generationSee model
Expressive text-to-speech from xAI Grok with multilingual support.Music generationSee model
4K video with built-in audio — voices, music, and effects match every scene.Audio4KVideo generationSee model
Quick 4K video with synchronized audio for rapid iteration.Audio4KFast generationSee model
Lightweight video with built-in audio — fast and affordable, 720p/1080pAudio1080pFast generationSee model
Generate speaking avatar videos from preset characters with natural lip-sync.CinematicVideo generationSee model
Photorealistic motion at 1080p with nuanced camera and lighting.1080pCinematicPhotorealSee model
Next-gen video restyling with keyframe image guidance for precise motion and style control.CinematicVideo generationSee model
Generate a still image from up to 3 reference images with consistent identity.Reference inputCinematicImage generationSee model
Sharp images up to 4K with fine-tuned color accuracy and detail.4KPro qualityImage generationSee model
Maximum detail for intricate compositions and demanding scenes.Image generationSee model
Edit and compose from up to 4 reference images with context awareness.Reference inputImage generationSee model
Single-image context-aware editing and generation — fast.Fast generationPro qualityImage generationSee model
Text-to-video with synchronized audio, plus image-to-video (animate up to 10 images) and video continuation.Text to videoImage to videoAudioSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Lightweight Nano Banana 2 variant for faster, high-volume image generation.Fast generationImage generationSee model
Quick, lightweight image creation for high-volume workflows.Fast generationImage generationSee model
Google Gemini native text-to-speech with expressive multilingual voices.Fast generationMusic generationSee model
Premium Gemini TTS with richer expressiveness and multi-speaker support.Pro qualityMusic generationSee model
Google Gemini multimodal video — text, image, or video as input.Fast generationVideo generationSee model
Gemini Omni with frame interpolation, video extension, reference-guided generation, and up to 4K output.Reference input4KFast generationSee model
Most capable GPT Image tier — premium edits and campaign-grade output, with longer generation times.Image generationSee model
Fast GPT Image tier — everyday generation at roughly half the latency of GPT Image 2.Fast generationImage generationSee model
Strong text-in-image and infographic rendering with multi-image input.Reference inputImage generationSee model
Latest voice engine with expanded tone and pacing control.Music generationSee model
Stable multilingual speech across 29+ languages with natural rhythm.Music generationSee model
Create custom sound effects from a text description — up to 30 seconds.AudioMusic generationSee model
Generate music with vocals or instrumental from a text prompt.AudioMusic generationSee model
Swap your voice to a different speaker while keeping timing and emotion.Music generationSee model
Voice swap across 29 languages — preserves emotion and cadence.Music generationSee model
Isolate vocals and remove background noise from an audio file.AudioMusic generationSee model
Dub audio or video across languages with automatic voice matching.AudioMusic generationSee model
Design a new voice from a text description using v3 engine.Music generationSee model
Design a new voice from a text description with multilingual support.Music generationSee model
Generate voice previews from a description to audition before committing.Music generationSee model
Animate any photo into a speaking avatar with natural lip-sync.PhotorealVideo generationSee model
Generate a speaking avatar video from a stock avatar and text script.Video generationSee model
Text-to-music with vocals or instrumentals from a style prompt and lyrics prompt.AudioMusic generationSee model
Text-to-music with vocals or instrumentals from a style prompt and optional lyrics, with configurable audio encoding.AudioMusic generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Ideogram's latest model — class-leading text rendering at up to ~3K resolution.Image generationSee model
Tiered Ideogram text-to-image — pick a speed/quality tier from very-low (fastest) to high (max quality).Fast generationImage generationSee model
Best-in-class text placement for logos, posters, and graphic design.Image generationSee model
Maintain a consistent character across scenes using a single reference photo.Reference inputPhotorealImage generationSee model
Fast music clips from text and image prompts using Google Lyria 3.AudioFast generationMusic generationSee model
Extended music generation up to 184s with vocals, powered by Google Lyria 3 Pro.AudioPro qualityMusic generationSee model
Full-length song generation with vocals from text and image prompts, powered by Google Lyria 3.5.AudioMusic generationSee model
Premium Qwen 2 (2026-04-22) with highest quality output.Pro qualityImage generationSee model
Qwen-Image 3.0 Pro (GA) — flagship text-to-image and image editing with prompt-rewrite modes and thinking mode.Pro qualityImage generationSee model
Next-generation raster output with refined detail and 10K-character prompts.Image generationSee model
Pro-tier V4.1 with enhanced quality and detail for premium output.Pro qualityImage generationSee model
V4.1 tuned for utility output — icons, logos, and functional design assets.Image generationSee model
Pro-tier V4.1 utility — premium quality for icons, logos, and design assets.Pro qualityImage generationSee model
Dedicated SVG vector output using V4.1 with clean lines.Vector outputImage generationSee model
Pro-tier V4.1 SVG vector output with enhanced detail.Pro qualityVector outputImage generationSee model
V4.1 utility tuned for SVG vector output — icons, logos, design assets.Vector outputImage generationSee model
Pro-tier V4.1 utility SVG vector output for premium design assets.Pro qualityVector outputImage generationSee model
Raster and vector output with clean text placement and 10K-character prompts.Vector outputImage generationSee model
SVG vector, illustration, and photo styles with readable in-image text.Vector outputPhotorealImage generationSee model
Pro-quality raster and vector output with enhanced detail and 10K-character prompts.Pro qualityVector outputImage generationSee model
Dedicated SVG vector output with clean lines and 10K-character prompts.Vector outputImage generationSee model
Pro-quality SVG vector output with enhanced detail and 10K-character prompts.Pro qualityVector outputImage generationSee model
Style-focused raster output with 10K-character prompts.Image generationSee model
Style-focused SVG vector output with 10K-character prompts.Vector outputImage generationSee model
Pro-quality style-focused raster output with enhanced detail and 10K-character prompts.Pro qualityImage generationSee model
Pro-quality style-focused SVG vector output with enhanced detail and 10K-character prompts.Pro qualityVector outputImage generationSee model
Dedicated SVG vector output with substyle options and negative prompts.Vector outputImage generationSee model
Convert raster images to clean SVG vector format.Vector outputImage generationSee model
AI-enhanced upscaling that adds creative detail to enlarged images.Image generationSee model
Clean upscaling that preserves sharp edges and fine detail.Image generationSee model
Replace the background of an image using a text prompt.Image generationSee model
Explore creative image ideas from a text prompt using Recraft V4.Image generationSee model
Find images visually similar to a previously explored Recraft image.Image generationSee model
Image upscaling and enhancement with Topaz AI — Standard, Hi-Fi, CGI, Recovery and Wonder models.Image generationSee model
Video upscaling and enhancement with Topaz AI — Proteus, Artemis, Nyx, Gaia and Starlight models.Video generationSee model
Swap the background of a photo using a text prompt for the new scene.PhotorealImage generationSee model
Remove the background from any image with precision, leaving a clean cutout.Image generationSee model
AI image enhancement with upscale and face enhancement.Image generationSee model
General-purpose image editing for swaps, fixes, style changes, and creative edits.Image generationSee model
Apply virtual makeup to portraits — lipstick, eye looks, blush, and full styled looks.Image generationSee model
Fast Flux 2 Klein 4B — up to 3 optional reference images.Reference inputFast generationImage generationSee model
Fast text-to-image generation powered by SANA-Sprint.Fast generationImage generationSee model
Apply curated Picsart effect presets to a photo — multi-step Magic Flow pipelines, one tap.PhotorealImage generationSee model
Animate a photo with curated Picsart video presets — multi-step Magic Flow pipelines, one tap.PhotorealVideo generationSee model
Happy Horse 1.0 — up to 15s at 1080P with optional first-frame guidance.Text to video1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Edit video — style transfer or object replacement, with up to 5 references.Video editingReference inputVideo generationSee model
Happy Horse 1.1 — up to 15s at 1080P with optional first-frame guidance.Text to video1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from a text prompt with PixVerse V6.Video generationSee model
Animate a source image into video with PixVerse V6.Video generationSee model
Fuse up to 7 reference images (subjects/backgrounds) into a new video scene with PixVerse V6.Reference inputVideo generationSee model
Generate video from a text prompt with PixVerse C1.Video generationSee model
Animate a source image into video with PixVerse C1.Video generationSee model
Fuse up to 7 reference images (subjects/backgrounds) into a new video scene with PixVerse C1.Reference inputVideo generationSee model
Generate natural speech from text with Async AI’s Flash voice engine.Fast generationMusic generationSee model
Meta's agentic image model — plans with reasoning, web and image search before rendering.Image generationSee model
Next-gen GPT image model with arbitrary output dimensions and multi-image input.Reference inputImage generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Wan 3.0 all-in-one — text, image/video/audio references, and start/end frames with adaptive ratio, intelligent duration, and audio.Reference inputAudioCinematicSee model
Upscale videos toward 4K (1.5x–3x) in precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K.4KVideo generationSee model
Edit videos with a text instruction — change objects, styles or scenes while preserving motion, timing and audio. Source clips up to 15 seconds; output at 24 fps, up to 720p.Video editingAudioVideo generationSee model
Top-tier MiniMax H3 Max video from text, a start/end frame, or reference images, videos, and audio — refer to references in the prompt as Image 1, Video 1, Audio 1, in input order. Up to 15s at 1080p.Reference inputAudio1080pSee model
Faster-than-realtime MiniMax H3 Max Turbo video from text or a start/end frame, with prompt expansion. Up to 15s at 1080p.1080pFast generationVideo generationSee model
MiniMax H3 Max video from a start frame with scripted camera motion — 2-12 keyframes set the camera angle, height, and distance over time. Leave the prompt blank to freeze the scene and move only the camera. Up to 15s at 1080p.1080pVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model
Generate video from up to 9 reference images — refer to them in the prompt as `[Image 1]`, `[Image 2]`, … in the same order they appear in the input list.Reference inputVideo generationSee model