Up to 20 seconds in one generation
FLUX 3 Video can generate diverse clips with synchronized audio for up to 20 seconds. The official preliminary evaluations used 10-second, 720p outputs with audio.
FLUX 3 · EARLY ACCESS · JULY 2026
One model for image, video, audio & action.
Explore the modelBlack Forest Labs introduced FLUX 3 on July 23, 2026 as a unified multimodal foundation model trained jointly across images, video, and audio. Its first Early Access release focuses on video with native audio, while image, action-prediction, and open-weight FLUX 3 Dev releases are planned in phases.
FLUX 3 VIDEO
FLUX 3 Video can generate diverse clips with synchronized audio for up to 20 seconds. The official preliminary evaluations used 10-second, 720p outputs with audio.
Start from a prompt, animate a first frame, use images as visual references, transform a source clip, continue video with audio, or define key moments.
Sound is generated with the picture. FLUX 3 is designed to connect physical events with effects, ambience, speech, and lip movement across languages.
ONE MULTIMODAL BACKBONE
FLUX 3 scales BFL’s self-supervised flow-matching approach across image, video, and audio so generation and representation learning happen in one architecture.
BFL says FLUX 3 Image will support varied styles, aspect ratios, resolutions, complex prompts, and multilingual typography after a separate Early Access phase.
FLUX-mimic adapts the video backbone for robot action generation. BFL and mimic robotics report that it is already being tested on industrial tasks with Audi.
FLUX 3 is Black Forest Labs' first unified multimodal foundation model for image, video, audio, and action prediction. Instead of treating sound or motion as add-ons, it learns how the world looks, moves, sounds, and responds inside one architecture.
Direct subject, action, scene, camera, style, speech, and sound from one prompt.
Animate a start frame or carry characters and visual language into a new context.
Generate dialogue, ambience, and physical sound effects with the picture.
Adapt the same learned world representation to physical robot actions.
Prompt anatomy
A useful FLUX 3 video prompt reads like a compact shot brief. Keep each instruction concrete and connect sound to the event that causes it.
Build a FLUX 3 promptExample
A ceramic robot repairs a greenhouse at dawn, carefully tightening a brass valve. Slow dolly-in, 50mm lens, soft mist and warm backlight. It says, “Pressure stable.” Metal clicks, leaves rustle, and the valve releases a short hiss in sync.
Release status
| Capability | Status | Access |
|---|---|---|
| FLUX 3 Video | Early Access | Application; API and private weights for selected users |
| FLUX 3 Image | Coming next | A separate Early Access phase is planned |
| FLUX 3 Action | Partner preview | Selected research and commercial partners |
| FLUX 3 Dev | Planned | Open-weight multimodal backbone; no public date yet |
Frequently asked questions
FLUX 3 is a multimodal foundation model from Black Forest Labs. It is trained jointly across images, video, and audio and is designed to support image generation, video with native audio, and action prediction through one shared backbone.
FLUX 3 Video is available only through a limited Early Access program. The official model page still says Coming Soon, and FLUX 3 Image, FLUX 3 Action, and FLUX 3 Dev are following separate phased releases.
Black Forest Labs says FLUX 3 can generate video with native audio up to 20 seconds long in a single generation. Its preliminary launch evaluations used 10-second, 720p clips with audio.
Yes. FLUX 3 Video generates audio together with video, including sound effects, ambience, and multilingual dialogue. This is native joint generation rather than a separate audio track added afterward.
Yes. The announced workflows include text to video, image to video from a starting frame or visual reference, video to video, video-and-audio continuation, and keyframe to video.
Not yet. Black Forest Labs has announced FLUX 3 Dev as a future open-weight multimodal backbone, but there is currently no official FLUX 3 checkpoint, model card, license, or ComfyUI workflow available for public download.
Public API pricing has not been announced. The official application currently describes Early Access as free, with priority based on use case and fit. Treat prices shown by unrelated third-party sites as their own service pricing, not official BFL pricing.
FLUX 2 is primarily an image generation and editing family. FLUX 3 expands the foundation into jointly trained image, video, audio, and action prediction, with video serving as the main signal for learning motion and physical dynamics.
No. flux3-video.app is an independent guide and prompt-planning studio. It is not affiliated with, endorsed by, or operated by Black Forest Labs. FLUX and Black Forest Labs are trademarks of their respective owner.
FREE PROMPT WORKBENCH
Structure a FLUX 3-ready brief now: subject, action, camera, style, dialogue, sound, references, duration, and aspect ratio. The workbench is a planning preview while the official model remains in Early Access.