Black Forest Labs Unveils FLUX.3 Multimodal Model with Native Synchronized Audio and Video Generation for Up to 20 Seconds

Technology24.Jul.2026 02:073 min read

German startup Black Forest Labs has unveiled FLUX.3, a multimodal foundation model built on its Self-Flow architecture, integrating image, video, audio, and motion encoders and decoders. The model is the company's first to support up to 20 seconds of native synchronized audio and video generation while enabling a wide range of video generation and editing tasks, demonstrating strong overall competitiveness.

Black Forest Labs Unveils FLUX.3 Multimodal Model with Native Synchronized Audio and Video Generation for Up to 20 Seconds

On July 23, German AI startup Black Forest Labs introduced Flux3, a new multimodal foundation model designed to handle both understanding and generation across physical and digital environments. Built on the company’s Self-Flow architecture, the model combines dedicated encoders and decoders for images, video, audio, and motion, signaling a broader ambition beyond conventional media generation.

One of Flux3’s most notable capabilities is native synchronized audio-and-video generation. According to the release, the model can produce clips of up to 20 seconds in a single pass, with sound generated alongside the visuals rather than added later as a separate layer. That makes Flux3 the company’s first model to support native audio generation and gives it a more complete multimedia output pipeline.

In practical terms, this means the system is not limited to creating moving images alone. It can generate video content together with matching audio, which helps make the results feel more cohesive and production-ready.

Broader task coverage beyond text-to-video

Black Forest Labs is positioning Flux3 as a flexible multimodal system rather than a model built for one narrow workflow. In addition to synchronized audio-video generation, it supports a range of creation modes, including:

  • Text-to-video

  • Image-to-video

  • Video-to-video

  • Keyframe-based transition generation

  • Multilingual dialogue generation

This wider task support reflects the company’s larger goal: reducing the separation between different media types and enabling a single model to work across multiple forms of input and output.

Early testing suggests strong competitiveness

In early tests conducted at 720p resolution with 10-second clips, Flux3 reportedly showed competitive performance against several well-known models in the market. Based on the figures shared by the company, Flux3 achieved a 93% win rate against Luma Ray3.2 and a 77% win rate against Runway Gen-4.5.

Black Forest Labs also said the model held a slight edge in comparisons with other leading systems, including Seedance 2.0 and Gemini Omni Flash. While these results come from early-stage testing, they suggest the company sees Flux3 as a serious entrant in the increasingly crowded multimodal generation space.

Expansion into robotics and industrial use

The release of Flux3 is not limited to general-purpose content generation. Black Forest Labs has also worked with Mimic Robotics on Flux-mimic, a video-action model aimed at robotics applications. The model is already being tested in production tasks at an Audi factory, pointing to possible uses beyond digital content and into real-world industrial environments.

This extension is especially notable because it connects media generation with motion understanding and action modeling, an area that is increasingly important for AI systems intended to interact with the physical world.

Phased rollout of the product family

Black Forest Labs is introducing the Flux3 lineup in stages. Flux3Video is available now, while Flux3Image and the open-weight version, Flux3Dev, are expected to follow soon.

This staggered launch approach suggests the company is building a broader ecosystem around the model family rather than treating Flux3 as a single standalone release.

A step toward AI “world models”

From a strategic perspective, Flux3 appears to be aimed at more than improving isolated generation tasks. Black Forest Labs is framing the model as part of a push toward AI systems that can perceive, interpret, and generate across images, video, audio, and motion in a more unified way.

By trying to bridge those modalities, the company is aligning Flux3 with the broader idea of world models—AI systems that do not just create content, but also develop stronger representations of how real and digital environments work. As capabilities in perception, understanding, and action continue to advance, Flux3 is being positioned as an important step in that direction.