Gemini Omni 1.1 Flash AI Video Model: Up to 4K, Frame Interpolation | Picsart
GEMINI OMNI 1.1 FLASH
Gemini Omni 1.1 Flash: Google's multimodal AI video model
Gemini Omni 1.1 Flash is Google's multimodal AI video model: generate video from a text prompt, images, or existing footage, pin the start and end frame and let it interpolate the motion between, guide results with up to 5 reference images and 3 reference videos, extend any clip, and output up to 4K. Available now in Picsart's AI Playground.
Turn text, images, and video into one Gemini AI video
Gemini Omni 1.1 Flash is genuinely multimodal: start from a text prompt, a single image, or existing footage, and combine them in one generation. Feed up to 5 reference images to lock characters, products, or style, and up to 3 reference videos to carry motion and look across shots - so every clip stays on-brand from the first frame to the last.
START & END FRAME CONTROL
Set the first and last frame, interpolate everything between
Pin a start frame and an end frame and Gemini Omni 1.1 Flash fills in the motion between them, so a transition lands exactly where you want it. It's frame-level control most text-to-video models don't give you - ideal for logo reveals, product turns, and seamless loops - with output from 360p all the way up to 4K, in 16:9 or 9:16.
UP TO 4K OUTPUT
Google video generation, sharp enough for up to 4K
Every Gemini Omni 1.1 Flash clip renders up to native 4K, in 16:9 for landscape or 9:16 for Shorts, Reels, and TikTok - no upscaling or reformatting. Generate up to 10 seconds per clip, then extend the result when a scene needs more room to breathe.
INSIDE PICSART
Find the right Google video model in Picsart
Gemini Omni 1.1 Flash lives in Picsart's AI Playground, where a single prompt runs against 150+ other models so you can compare before committing. Pick Gemini Omni 1.1 Flash for reference-guided, frame-controlled video; choose Veo 3.1 or Veo 3.1 Fast when you need synchronized audio in the same render. And you can reach it whichever way you work - on the web, in the Picsart desktop app, or built into your own projects via CLI, MCP, REST API, and SDK, with no third-party API keys or separate subscriptions to manage.
What you can create with Gemini Omni 1.1 Flash
Bring up to 5 reference images and 3 reference videos to keep characters, products, and style consistent across every generated clip.
Explore more models like Gemini Omni 1.1 Flash
Compare Gemini Omni 1.1 Flash with other video models for motion, ads, and social clips.
Gemini Omni 1.1 Flash FAQ
Gemini Omni 1.1 Flash is Google's multimodal AI video model. It generates video from text, images, or video, supports start- and end-frame interpolation, up to 5 reference images and 3 reference videos, clip extension, and output up to 4K.
Gemini Omni 1.1 Flash focuses on multimodal, reference-guided generation with start/end frame interpolation and clip extension. Veo 3.1 focuses on cinematic 4K with synchronized native audio. Choose Gemini Omni for reference and frame control; choose Veo 3.1 when you need audio in the same render.
Each generation produces up to 10 seconds of video (from about 3 seconds). You can bring a source video of up to 30 seconds as input, and extend or continue clips to build a longer sequence.
A text prompt, images (including start and end frames), and video. You can add up to 5 reference images and up to 3 reference videos to keep characters, products, and style consistent.
Gemini Omni 1.1 Flash outputs at 360p, 720p, 1080p, and 4K, in 16:9 and 9:16 aspect ratios.
It's available in Picsart's AI Playground. You can also use it on the web, in the Picsart desktop app, and build it into your own projects via CLI, MCP, REST API, and SDK.
Yes. Videos generated through Picsart's tools powered by Gemini Omni 1.1 Flash can be used for marketing, social media, brand content, and other commercial applications, subject to Picsart's terms of use.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.