LTX 2.5 Pro: cinematic video with native audio, in one pass
LTX 2.5 Pro is the quality-optimized tier of the LTX 2.5 video family from Lightricks. One prompt in, a finished shot out - with synchronized native audio generated in the same pass, polished 1080p detail, and precise camera motion control. Describe your scene, pick a camera move, and LTX 2.5 renders picture and sound together.
LTX 2.5 Pro generates synchronized native audio in the same pass as the video - no separate scoring or dubbing step. Dialogue, ambience, and effects land in time with the motion.
Native audio, on by default
Speech, music, and sound effects are generated together with the footage and synced to the action, so a finished clip arrives with its soundtrack already in place.
Camera motion you direct
Choose static, dolly in and out, dolly left and right, jib up and down, or a focus shift to set exactly how the shot moves - no post-production rig required.
Production-ready 1080p
Render up to 1080p at 24, 25, or 50 fps, tuned for per-frame quality so the take you deliver looks finished, not drafted.
WHAT IS LTX 2.5 PRO
A single-pass model built for the final take
LTX 2.5 Pro is the quality-optimized member of the LTX 2.5 lineup - a text-to-video model that produces synchronized sound and picture in one generation. Where a fast tier is tuned for quick drafts and longer runtimes, Pro concentrates on per-frame quality: cleaner detail, steadier motion, and audio written to match the scene. Feed it a prompt of up to 5,000 characters and it returns a finished 6, 8, or 10-second clip at 720p or 1080p, in 16:9 or 9:16. It belongs to the same LTX family as LTX 2.3 Audio to Video, LTX 2.3 Reframe, and LTX 2.3 Outpaint, sharing the lineup's audio-native approach to generation.
DIRECT EVERY MOVE
Camera moves you choose, not chance
Great coverage is about motion. LTX 2.5 Pro gives you explicit camera motion control - keep the frame static, push with a dolly in or out, track with a dolly left or right, lift and drop with a jib, or pull focus with a focus shift. Set the movement per shot and the model animates the camera the way you asked, so your footage cuts together with intent instead of guesswork.
FINISHED, NOT DRAFTED
1080p detail with sound that fits the scene
Pro is tuned for the deliverable. Render at up to 1080p and 24, 25, or 50 fps for smooth, broadcast-friendly motion, while native audio generation lays down dialogue, ambience, and effects that track the picture. The result is a clip you can hand off - lighting that reads naturally, detail that holds, and a soundtrack already in sync.
INSIDE PICSART
How LTX 2.5 Pro works inside Picsart
LTX 2.5 Pro is available in Picsart's AI Playground, where you can generate with it directly and compare its output against 150+ other AI models from a single prompt - no setup or model configuration required. And you can reach LTX 2.5 whichever way you work: on the web, in the desktop app, or built straight into your own projects via CLI, MCP, REST API, and SDK.
PRO OR FAST
When to reach for the Pro tier
Choose LTX 2.5 Pro when the shot is the deliverable and quality per frame matters most - hero product moments, narrative beats, and polished social spots that need clean 1080p and synced sound in up to 10 seconds. Reach for the faster LTX 2.5 Fast tier when you want to iterate quickly or push longer runtimes and higher ceilings. Same LTX 2.5 family, two speeds: Pro for the finished take, Fast for the rapid draft.
What you can create with LTX 2.5 Pro
Turn a concept into a polished product moment with clean 1080p detail, a directed camera move, and native audio that sells the shot.
Tutorials and guides for LTX 2.5
Explore more models like LTX 2.5 Pro
Compare LTX 2.5 Pro with other video models for cinematic motion, native audio, and social clips.
LTX 2.5 Pro AI model FAQ
LTX 2.5 Pro is the quality-optimized tier of Lightricks' LTX 2.5 video model. It generates synchronized native audio in a single pass with the video, renders at 720p or 1080p in 16:9 or 9:16, and supports precise camera motion control - built for finished, production-ready clips.
LTX 2.5 Pro renders at 720p or 1080p (1080p by default) in 16:9 or 9:16, at 24, 25, or 50 fps (25 by default). Clips can be 6, 8, or 10 seconds long, with 6 seconds as the default.
Yes. Native audio generation is on by default - LTX 2.5 Pro produces dialogue, music, and sound effects in the same pass as the video, synced to the motion, so no separate scoring or dubbing step is needed.
Yes. LTX 2.5 Pro offers camera motion control with options for none, static, dolly in, dolly out, dolly left, dolly right, jib up, jib down, and focus shift, so you can direct exactly how each shot moves.
Yes. LTX 2.5 Pro accepts an optional start frame and end frame image, so you can guide how a shot begins and ends while it generates the motion and audio in between.
You can use LTX 2.5 Pro whichever way you work: on the web, in the Picsart desktop app, or built straight into your own projects via CLI, MCP, REST API, and SDK. It is also in the AI Playground, where you can compare it against 150+ other AI models from a single prompt.
LTX 2.5 Pro is tuned for quality per frame - cleaner 1080p detail and synced native audio for finished takes up to 10 seconds. The faster LTX 2.5 tier trades some of that polish for quicker iteration and longer runtimes. Choose Pro for the deliverable and Fast for rapid drafts.
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.
Map body movement from a video clip onto a portrait photo — V3 quality.CinematicPhotorealVideo generationSee model
Transfer body movement from a reference video onto a portrait photo.Reference inputCinematicPhotorealSee model
Fast cinematic video with audio, reference images, and start/end frame control.Reference inputAudioFast generationCinematicSee model
Edit video — replace subjects, add or remove objects, restyle scenes with reference images.Video editingReference inputVideo generationSee model
Fast video edit — modify scenes with reference images.Video editingReference inputFast generationSee model
Wan 2.7 T2V — up to 15s at 1080p with audio input and prompt enhancement.Text to videoAudio1080pCinematicSee model
Long-form video up to 15s with native audio and start/end frame control.AudioCinematicVideo generationSee model
Faster V3 variant — long-form video up to 15s with native audio, start/end frame control, and 720p/1080p output.Audio1080pFast generationCinematicSee model
Mature pipeline with audio and pro-tier rendering.