FLUX 3 · EARLY ACCESS · JULY 2026

FLUX 3 Video

One model for image, video, audio & action.

Explore the model

FLUX 3Real-world model

Black Forest Labs introduced FLUX 3 on July 23, 2026 as a unified multimodal foundation model trained jointly across images, video, and audio. Its first Early Access release focuses on video with native audio, while image, action-prediction, and open-weight FLUX 3 Dev releases are planned in phases.

Up to 20 secondsNative audioVisual referencesEarly Access

FLUX 3 VIDEO

Video and sound. Generated together

20smaximum announced length
01

Up to 20 seconds in one generation

FLUX 3 Video can generate diverse clips with synchronized audio for up to 20 seconds. The official preliminary evaluations used 10-second, 720p outputs with audio.

05announced video workflows
02

Text, image, video, and keyframe control

Start from a prompt, animate a first frame, use images as visual references, transform a source clip, continue video with audio, or define key moments.

A/Vjoint generation
03

Native audio and multilingual dialogue

Sound is generated with the picture. FLUX 3 is designed to connect physical events with effects, ambience, speech, and lip movement across languages.

ONE MULTIMODAL BACKBONE

Beyond video. Toward world models

01shared backbone
01

Jointly trained with Self-Flow

FLUX 3 scales BFL’s self-supervised flow-matching approach across image, video, and audio so generation and representation learning happen in one architecture.

SOONimage early access
02

Image synthesis and editing are next

BFL says FLUX 3 Image will support varied styles, aspect ratios, resolutions, complex prompts, and multilingual typography after a separate Early Access phase.

<80msreported backbone latency
03

The same foundation predicts actions

FLUX-mimic adapts the video backbone for robot action generation. BFL and mimic robotics report that it is already being tested on industrial tasks with Audi.

Updated July 24, 2026Independent launch tracker

What is FLUX 3 Video?

FLUX 3 is Black Forest Labs' first unified multimodal foundation model for image, video, audio, and action prediction. Instead of treating sound or motion as add-ons, it learns how the world looks, moves, sounds, and responds inside one architecture.

Text to video

Direct subject, action, scene, camera, style, speech, and sound from one prompt.

Image & video references

Animate a start frame or carry characters and visual language into a new context.

Native audio

Generate dialogue, ambience, and physical sound effects with the picture.

Action prediction

Adapt the same learned world representation to physical robot actions.

Prompt anatomy

Write for picture, motion, and sound.

A useful FLUX 3 video prompt reads like a compact shot brief. Keep each instruction concrete and connect sound to the event that causes it.

Build a FLUX 3 prompt
  1. 01Subject & setting
  2. 02Action & timing
  3. 03Camera & framing
  4. 04Light & visual style
  5. 05Dialogue in quotes
  6. 06Ambience & sound effects

Example

A ceramic robot repairs a greenhouse at dawn, carefully tightening a brass valve. Slow dolly-in, 50mm lens, soft mist and warm backlight. It says, “Pressure stable.” Metal clicks, leaves rustle, and the valve releases a short hiss in sync.

Release status

What is actually available?

Phased rollout; details may change
CapabilityStatusAccess
FLUX 3 VideoEarly AccessApplication; API and private weights for selected users
FLUX 3 ImageComing nextA separate Early Access phase is planned
FLUX 3 ActionPartner previewSelected research and commercial partners
FLUX 3 DevPlannedOpen-weight multimodal backbone; no public date yet

Frequently asked questions

FLUX 3 Video FAQ

What is FLUX 3?+

FLUX 3 is a multimodal foundation model from Black Forest Labs. It is trained jointly across images, video, and audio and is designed to support image generation, video with native audio, and action prediction through one shared backbone.

Is FLUX 3 available now?+

FLUX 3 Video is available only through a limited Early Access program. The official model page still says Coming Soon, and FLUX 3 Image, FLUX 3 Action, and FLUX 3 Dev are following separate phased releases.

How long can FLUX 3 videos be?+

Black Forest Labs says FLUX 3 can generate video with native audio up to 20 seconds long in a single generation. Its preliminary launch evaluations used 10-second, 720p clips with audio.

Does FLUX 3 generate audio?+

Yes. FLUX 3 Video generates audio together with video, including sound effects, ambience, and multilingual dialogue. This is native joint generation rather than a separate audio track added afterward.

Can FLUX 3 use image or video references?+

Yes. The announced workflows include text to video, image to video from a starting frame or visual reference, video to video, video-and-audio continuation, and keyframe to video.

Can I download FLUX 3 for ComfyUI?+

Not yet. Black Forest Labs has announced FLUX 3 Dev as a future open-weight multimodal backbone, but there is currently no official FLUX 3 checkpoint, model card, license, or ComfyUI workflow available for public download.

How much does FLUX 3 cost?+

Public API pricing has not been announced. The official application currently describes Early Access as free, with priority based on use case and fit. Treat prices shown by unrelated third-party sites as their own service pricing, not official BFL pricing.

What is the difference between FLUX 2 and FLUX 3?+

FLUX 2 is primarily an image generation and editing family. FLUX 3 expands the foundation into jointly trained image, video, audio, and action prediction, with video serving as the main signal for learning motion and physical dynamics.

Is flux3-video.app an official Black Forest Labs website?+

No. flux3-video.app is an independent guide and prompt-planning studio. It is not affiliated with, endorsed by, or operated by Black Forest Labs. FLUX and Black Forest Labs are trademarks of their respective owner.

FREE PROMPT WORKBENCH

Direct the world.Before access opens

Structure a FLUX 3-ready brief now: subject, action, camera, style, dialogue, sound, references, duration, and aspect ratio. The workbench is a planning preview while the official model remains in Early Access.