Flux 3 generates the picture and the sound in the same pass, which quietly removes an entire step from video production. No separate audio generation, no manual sync, no hunting for a library track that almost fits. Black Forest Labs built its first fully multimodal model, and the practical consequence is that a 20-second clip arrives finished: frames, dialogue, effects and ambience, from one prompt. Here is what the model does, where it breaks from the usual stack, and how to run it inside Picsart.

What Flux 3 actually is

Flux 3 is Black Forest Labs’ first fully multimodal model, unifying image, video and audio in a single architecture. The makers of FLUX previously shipped image models, with everything else bolted on downstream. This one generates a video clip of up to 20 seconds and its synchronized soundtrack together, from the same request.

It also takes more than text going in. Flux 3 accepts text, audio, video and up to 10 image references at once, which is what makes editing, remixing and multi-shot storyboarding possible rather than just prompt-and-hope generation. Output holds cinematic realism and physical coherence, so objects behave like objects across the length of a clip.

Every generation starts from one of three places. Text to video begins with nothing but the prompt. Image to video builds the clip around images you pin. Video continuation picks up an existing clip and carries it forward. The rest of the request looks the same either way, so moving between them is a change of starting point rather than a change of tool.

Mikayel Vardanyan, Picsart’s COO and Co-Founder, put it this way when the model launched: “Having motion and audio grounded in the same model lets us treat each clip as a self-contained shot and handle narrative and timing at the pipeline level. Our users expect access to the latest and greatest models, and Flux 3 raises the bar for what they can create with AI.”

That pipeline point is the practical one. A clip that arrives complete can be treated as a unit and placed, rather than assembled from parts that each need checking against the others.

 

The numbers worth knowing

Durations run 5, 10, 15 and 20 seconds, or auto if you would rather the model fit the length to the content. Everything outputs at 24 fps. Resolution is your choice of HD or FHD, so you can trade pixels for speed while an idea is still moving. Aspect ratios cover eight options, including 21:9, 4:3, 1:1 and 9:16, which is unusually wide for a video model.

Audio is on by default and can be switched off for a silent clip. Continuing an existing clip accepts a start video of up to 15 seconds. There is also a draft mode that returns a fast preview at a fraction of a full render, and sending the good one back reproduces that exact generation at full quality rather than a fresh interpretation.

Three things that change in practice

Sound arrives with the frames

Native audio is generated in the same pass as the video, covering atmospheric sound, physical interactions and lip-synced dialogue in multiple languages. That matters more than it sounds. A clip generated with separate audio always has a sync problem waiting somewhere, and fixing it costs an editing session per asset. When picture and sound come from one generation, they match from the first frame because they were never separate.

Twenty seconds is a whole beat

Most top models cap out around 15 seconds, which is enough for a moment but not for a scene. Twenty seconds fits a full narrative beat: a character enters, interacts with something, delivers a line. That is the difference between generating footage you then have to assemble and generating something that already works on its own.

Ten references, and a place on the timeline

Text, audio, video and up to 10 image references can go in together. Pure text-to-video models cannot hold a character, a product and a set consistent across several shots, because there is nothing to hold them to. Reference stacking is what turns a single generation into a sequence that looks deliberate.

Images can also be pinned in time rather than just handed over. One becomes the opening frame, a pair fixes the start and end with the motion filled in between, and several pinned to timestamps act as an ordered storyboard. That is closer to blocking a shot than describing one.

Where Flux 3 breaks from the usual stack

The normal workflow for a finished clip involves at least three tools: one model for video, another for voice or sound design, then an editor to marry them. Flux 3 collapses that into one request, and the knock-on effect is turnaround rather than quality. A concept that took a day to assemble becomes something you can iterate on several times in an afternoon.

What it does not do is replace a specialist. Tight brand-locked design work and long sustained simulation are still jobs for dedicated tools or dedicated models. Capabilities are also rolling out in stages, so some editing and reference features arrive after the initial release. Flux 3 is a preview model, and the gaps will move.

What to make with it

Social spots with actual dialogue

A 20-second clip with lip-synced speech covers a full ad beat without a voiceover session. Write the line into the prompt, name the language, and keep the sentence short enough to breathe.

Product motion with sound design

Physical interactions generate their own audio, so a bottle opening, a zip closing or a shoe hitting pavement comes with the noise it should make. That detail is what separates a product clip that reads as real from one that reads as rendered.

Multi-shot storyboards

Reference stacking plus multi-shot storyboarding holds a character across several generations, which makes a sequence of clips feel like one piece rather than a collection of tests.

Localized variants

Multilingual dialogue means one scene can be regenerated per market with the line swapped and the lip-sync matched, instead of commissioning a dub or reshooting.

Stylized and animated looks

Style range runs well past cinematic realism into animation, motion design and stylized looks, so the model suits brand work that deliberately avoids looking like stock footage. On-screen typography holds its shape through motion too, which covers titles, signage and lower-thirds.

How to use Flux 3 in Picsart AI Playground

Picsart AI Playground is where single generations and model comparison belong. It is the fastest way to find out whether Flux 3 suits a brief before committing a workflow to it.

  1. Open AI Playground and select Flux 3 from the model list, alongside the other 150+ models available from the same prompt bar.
  2. Write the prompt as a scene with sound. Describe what happens, then describe what it sounds like. Audio is part of the generation, so leaving it unspecified means the model decides for you.
  3. Add references if consistency matters. Up to 10 images, plus audio or video, when a character, product or set has to survive across shots.
  4. Draft first, then commit. A draft preview costs a fraction of a full render, so explore freely and pay full price only for the take you keep.
  5. Iterate on one variable. Change the light, or the line, or the camera move, but not all three, or you will not know which change helped.

How to use Flux 3 in Picsart Flow

Picsart Flow is the surface for repeatable work. Where AI Playground answers whether a model works for an idea, Flow answers how to do it 40 times without redoing the setup.

  1. Start with your input nodes. Flow takes text prompts, uploaded images and audio as starting points, which maps neatly onto what Flux 3 accepts as references.
  2. Add Flux 3 as the generation step in the chain, so the same prompt structure and the same references apply on every run.
  3. Chain what comes after generation. Resizing per placement, style passes and variant branching all become part of the workflow rather than manual follow-up.
  4. Branch from one base. Flow lets several concepts run from the same starting node, which is how one approved scene becomes a set of market or format variants.
  5. Run it once, then reuse it. The setup cost is paid once. Every later campaign is a new input into a workflow that already works.

Which surface for which job

Use AI Playground when the question is creative: what does this model do with my idea, and how does it compare against the alternatives. Iteration is fast and nothing is committed.

Move to Flow when the answer is settled and the work becomes production: many assets from one approved approach, consistent treatment on each, and no appetite for rebuilding the setup every time. Most teams end up using both, with Playground as the place ideas get tested and Flow as the place approved ideas get manufactured.

Prompting a model that also makes sound

Flux 3 rewards prompts written like a brief for a crew rather than a search query.

  • Describe the audio explicitly. Name the dialogue, the ambience and the key physical sounds. Silence is rarely what you want, and an unspecified soundtrack is a guess.
  • Keep spoken lines short. Lip-sync holds better on a sentence with breathing room than on a dense paragraph, and short lines survive localization more cleanly.
  • Give the full 20 seconds instructions. Prompts that describe only the opening tend to drift in the final third, because nothing told the model what happens there.
  • Name one camera move. Locked framing, a slow push-in or a follow. Directorial control works when it is specific rather than cinematic-sounding.
  • Order your references. When several images go in, be explicit about which is the character, which is the product and which sets the look.

Get answers to common questions

Video of up to 20 seconds with native synchronized audio, from a single generation. Because picture and sound are produced in one pass, they are matched from the first frame.

Start generating

The interesting thing about Flux 3 is not any single spec, it is that the output needs less work after it arrives. Sound already fits the picture, the clip is long enough to say something, and the references keep it consistent enough to use twice.

Open Picsart AI Playground to run your first prompt, generate a finished clip in the AI video generator, or see the rest of the lineup in the model picker.