Models & Research

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

· July 23, 2026
Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

What it does

Black Forest Labs has launched Flux 3, a new multimodal foundation model that processes images, video, and audio simultaneously. For the first time, it can generate videos up to 20 seconds long that include native, synced audio. This blends visual and sound generation in a single model, pushing beyond typical video generation tools that usually mock or omit sound.

Why it matters

Combining native audio with video generation changes the game for content creation and interactive AI applications. Producing short video clips with synchronized sound natively opens doors for automated marketing, storytelling, virtual assistants, and even robotics where audio-visual coherence is crucial. It pressures existing video AI players who lack integrated sound generation to catch up or risk losing ground in markets demanding richer outputs.

Who it is for

Flux 3 targets AI developers and companies building multimodal applications that benefit from coordinated audio-visual content. Its use extends to robotics, where synchronized sensory data matters for task execution, per BFL’s early trials. Content creators, marketers, and interactive experience designers will find practical value in a tool that can produce fully synchronized video and audio in one step.

The catch

So far, independent benchmarks are unavailable. Black Forest Labs claims Flux 3 edges out market leader Seedance 2.0 based on internal tests, but outside verification is needed to confirm actual performance gains. The maximum clip length is currently capped at 20 seconds, which limits use cases requiring longer video streams. The underlying complexity of handling three data modalities in one model might also affect compute cost and scalability.

What to watch next

Look for independent benchmarks of Flux 3 to validate Black Forest Labs’ performance claims against Seedance and others. Monitoring how quickly the 20-second limit extends will show if Flux 3 can scale to longer video outputs. Watch how BFL leverages this tech in robotics, since applying native audio/video generation there signals a push beyond entertainment to real-world automation. Also, note competition dynamics as other labs race to improve multimodal synthesis with sound.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.