Flux 3 adds native audio to multimodal AI video

Black Forest Labs has released Flux 3, a multimodal foundation model trained on images, video, and audio together. It can generate videos up to 20 seconds long with native audio, while related work on Flux-mimic is already being tested at Audi.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

Flux 3 is mainly a product launch, but its world-model framing and future robotics ambitions mildly point toward more powerful autonomous systems.

Flux 3 adds native audio to multimodal AI video

Black Forest Labs has released Flux 3, a multimodal foundation model built to learn from images, videos, and audio at the same time. The German AI company is positioning the model as a step toward systems that can understand media, generate it, and support future work in robotics.

The most immediate change is video generation with native audio. Flux 3 can create clips up to 20 seconds long, and Black Forest Labs says the model can connect what appears on screen with the sounds that physical events should produce.

A model trained across images, video, and audio

Black Forest Labs describes Flux 3 as part of its push toward "real-world visual intelligence." The company defines that goal as models that can "perceive, predict, and act across physical and digital environments." In practical terms, that means Flux 3 is not only a video generator. It is designed around the idea that different forms of media provide different signals about the same world.

The company’s argument is straightforward. Images can represent spatial structure. Video can show how that structure changes over time. Audio can reveal relationships between events and the sounds they create, especially when motion and sound are linked.

Training on all three together gives the model a broader context than training on each type of data separately. If one modality leaves information unclear, another can help fill the gap. Black Forest Labs presents this as an important part of building so-called world models.

What Flux 3 can generate

Flux 3 supports text-to-video, image-to-video, video-to-video, and keyframe-based transitions. It also supports multilingual dialogue and agent-driven links between clips, which can be used to build longer multi-shot sequences from individual clips.

The native-audio capability is central to this release. Earlier video generation systems often treat sound as a separate step or omit it entirely. Flux 3, as described by Black Forest Labs, generates video and audio together for clips up to 20 seconds long.

The company says the model performs especially well in two areas: human facial expressions and matching sounds to physical events. Both are difficult problems for video models because they require more than isolated image quality. Facial expression needs temporal consistency, while sound matching needs the model to understand when visual action and audio should align.

Those features matter because video generation is increasingly judged not just by how realistic a single frame looks, but by whether the whole sequence remains coherent. A clip can look strong visually while still failing if motion, expression, dialogue, or sound feel disconnected.

Early comparisons show strong results, with limits

Black Forest Labs reported early evaluations using 10-second clips at 720p. In those tests, Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent.

The reported margins were smaller against several other systems. Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 and Gemini Omni Flash at 52 percent each.

Those results are notable, but the company also says they are preliminary. No independent tests are available yet. That matters because preference tests can depend on prompts, evaluation methods, and the clips selected for comparison.

Still, the comparison group shows where Black Forest Labs wants Flux 3 to compete. The source notes that matching leading systems such as Seedance and Gemini Omni Flash would place Flux 3 among the top video models.

Image generation and robotics are also part of the plan

Flux 3 is not limited to video output. Black Forest Labs expects the model to improve image generation as well, especially for complex prompts and accurate text rendering in multiple languages. The company plans to release Flux 3 Image in early access within the next few weeks.

The model also connects to robotics work. Black Forest Labs says Flux 3 can predict actions based on its understanding of the world. The company worked with Mimic Robotics to develop Flux-mimic, a video-action model now being tested on production tasks at Audi.

That detail shows why the company emphasizes perception and action rather than only media generation. A system that learns from visual and audio signals may be useful for generating clips, but the same underlying approach can also support models that need to predict what should happen next in a physical task.

Flux 3 is based on Self-Flow, Black Forest Labs’ approach for teaching one model to generate and understand content at the same time. A multimodal transformer uses dedicated components to convert images, video, and audio into a shared internal representation, then turn that representation back into outputs.

The same architecture includes a component for actions, which provides the basis for robotics applications. Black Forest Labs says this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model’s understanding of the physical world.

A staged rollout with open-weight access planned

Black Forest Labs is rolling out Flux 3 in phases. Flux 3 Video is already available, while Flux 3 Image is expected to follow in the coming weeks. Action prediction will initially be available through select partners.

The company also plans to release open-weight access to the multimodal backbone under the name "Flux 3 Dev." Longer term, Black Forest Labs says it is working on next-generation models that combine perception, action, and language prediction in a single model.

For now, Flux 3’s importance is in how it combines several AI video priorities at once: native audio, multimodal learning, longer clip generation, and a path from media synthesis toward action prediction. The open question is how the model performs once independent testing becomes available.