One visual direction
Joint training across images, video and audio keeps the character, palette and lighting connected as a still becomes a moving scene.
FLUX 3 by Black Forest Labs generates AI video with native audio and multilingual dialogue. Turn a text-to-video prompt, image or keyframes into a scene with one visual direction.
Joint training across images, video and audio keeps the character, palette and lighting connected as a still becomes a moving scene.
Give a line of dialogue, a reaction and a camera move room to unfold, with up to 20 seconds of video and audio per generation.
Chain clips into longer multi-shot scenes. Keep subject, lighting and camera direction specific in each prompt so the shots connect.
Set the character, palette and framing in a reference image, or start directly with a text-to-video prompt.
Describe the action and camera movement. A first frame, end frame or keyframes can guide the shot.
Include dialogue, language and ambient sound in the prompt. Native audio is generated with the picture.
| Dimension | FLUX 2 | FLUX 3 |
|---|---|---|
| Modalities | Images — text-to-image and image editing in one model. | Image, video, audio and action prediction in one model. |
| Video | Not a video model. | Up to 20 seconds with audio in a single generation. |
| Audio | None. | Native audio generation with multilingual dialogue. |
| Architecture | Latent flow matching, coupling a 24B vision-language model with a rectified flow transformer. | Built on Self-Flow, jointly trained across images, video and audio. |
| Availability | Available on YouArt today, with create and edit modes. | Available on YouArt today. |
YouArt offers FLUX 3 Video at 720p or 1080p with synchronized audio. For image generation and editing, use FLUX 2 or Flux Kontext.
Turn a campaign frame into a twenty-second spot and create language versions with a consistent product, palette and message.
Animate product stills with sound while preserving the shape, colour and texture shoppers need to recognize.
Previsualize a scene with dialogue, camera movement and sound cues already in place, so the team can judge its pacing.
Keep a spokesperson consistent across training modules and generate additional language versions from the same direction.
One credit balance for FLUX 3 Video, FLUX 2, Flux Kontext and more.
For hobbyists and explorers
For creators and pro users
For power users and teams
For teams and studios
FLUX 3 is Black Forest Labs' multimodal model, announced on 23 July 2026. It spans image, video, audio and action prediction in a single model that was trained jointly across those modalities, rather than assembled from separate systems. As with any early access release, availability and supported modes can evolve; YouArt provides FLUX 3 Video for hands-on generation today.
Yes — with FLUX 2 and Flux Kontext, both available on YouArt today. FLUX 2 handles text-to-image and reference-guided editing in create and edit modes, and Flux Kontext does instruction-based editing from a single reference image with mask-based inpainting. You do not need to manage separate model weights to use these workflows in YouArt.
FLUX 2 is an image model — text-to-image plus single- and multi-reference editing in one checkpoint. FLUX 3 widens that to text to video, video, audio and action prediction in a single jointly trained model, and generates up to 20 seconds of video with audio in one pass.
Yes — up to 20 seconds of video with audio in a single generation, including multilingual dialogue, plus continuation from an input video and its audio. Longer pieces chain clips into multi-shot sequences rather than stretching one generation.
Up to 20 seconds in a single generation. Longer sequences chain individual clips into multi-shot scenes rather than extending one generation further.
Pricing works on a credit system across every model on YouArt. FLUX 3 Video is billed by duration at 55 credits per second, and you can see plans and per-model pricing on the pricing page. For API planning, compare render duration, resolution, prompt iteration, and production workflow requirements before you estimate a per-project cost.
In Black Forest Labs' own preference testing — preliminary results from a midtraining checkpoint, on 10-second 720p clips with audio — FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against the strongest competitors. Kling and FLUX 3 may suit different video-generation needs, so compare the output quality, prompt control, motion, audio, and available API workflow that match your project. Independent testing is still limited.
Yes. FLUX 3 runs on YouArt as FLUX 3 Video: 720p or 1080p video with synchronized native audio, from a prompt alone, a first frame to animate, a first and last frame to fill the motion between them, or up to 10 keyframes pinned along the timeline. FLUX 2 and Flux Kontext remain the FLUX models for images.
FLUX, FLUX.2 and FLUX 3 are trademarks of Black Forest Labs Inc. YouArt is not affiliated with or endorsed by Black Forest Labs.
Every handoff between tools is a place where your look drifts. FLUX 3 collapses those handoffs into a single model, so the thing you approved in the first frame is still there in the last. For a creative team, that means fewer transfers between models, fewer prompt rewrites, and a more consistent result from image to motion to sound.
FLUX 3 covers image generation and editing, text to video, video, and sound in a single model rather than a chain of specialists — so what you establish in a still carries into the motion and the mix. The model's multimodal understanding connects the visual direction and the audio direction before the generation begins.
Up to 20 seconds of video with audio in a single pass — long enough to carry a whole beat. Longer pieces chain clips into multi-shot sequences, so you build the story through intentional shots rather than stretching one generation past its useful length.
Native audio arrives with the video — multilingual dialogue included — instead of being layered on afterwards and nudged into sync. That makes language versions a creative decision at the prompt stage, not a separate audio-replacement workflow.
Image synthesis and editing sit in the same model as the video, so the still you refine and the frame you animate come from one system instead of two that only roughly agree. For users working through early access or model-preview workflows, this can reduce the need to move assets between incomplete tools while the final capability set evolves.
One jointly trained model. A consistent look from still to motion to sound
Up to 20 seconds in one generation. A complete moment, not a fragment
Native audio with the generation. Dialogue already sitting in the picture
Multilingual dialogue. One shoot, several language cuts
Teams tired of stitching an image model to a video model to an audio tool, who want one consistent system end to end.
A scene already made with FLUX 2 on YouArt, pushed into motion with dialogue — that handoff is the one FLUX 3 is designed to remove. Use text to video prompts when you are starting from an idea, or use a still and keyframes when framing and visual weight are already defined.
Single quick stills where FLUX 2 already gets you there faster.
One jointly trained model means the character, palette and lighting you set in a still are still there once the shot is moving.
Long enough for a line of dialogue and a reaction, instead of a fragment you have to cut around.
Native audio generation removes the alignment pass that normally sits between a finished render and a finished scene.
Multilingual dialogue means a second language version is a generation, not a reshoot.
Creative directors hold one visual idea across a still, a moving shot and a soundtrack, instead of re-establishing it in every tool.
Performance marketers spin variants and language cuts from the same look, without a reshoot for each one.
Filmmakers previsualize scenes with dialogue in place, so pacing can be judged rather than imagined.
Social creators go from an idea to a short piece with sound already in it, without assembling a toolchain to get there.