FLUX 3 is a multimodal video AI from Black Forest Labs that generates videos up to 20 seconds long with audio generated alongside and synchronized to the video. The model is trained on images, video, and audio within one shared system, i.e., multimodality, across image, video, and audio data.
In head-to-head evaluations FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons, while it was preferred over Gemini Omni and Seedance in 52% of evaluations across those comparative judgments overall.
FLUX 3 is trained on images, video, and audio within one shared system, i.e., multimodality. The multimodal training incorporates image, video, and audio data into a single unified system rather than separate modality-specific pipelines. The model generates videos with audio produced alongside and synchronized to the video.
FLUX-mimic is a lightweight decoder that translates the model’s motion predictions into robot motions, and the article states that it converts FLUX 3’s predicted movements—drawn from multimodal training on images, video, and audio—into physical actions by robots. Audi is testing FLUX-mimic for tasks like fitting flexible door seals, examples given in the article to illustrate its use in soft-body manipulation operations.
Mimic co-founder Stephan-Daniel Gravert said, “Audi represents the kind of manufacturing partner we built FLUX-mimic for.”
Audi’s Christoph Schneider said the robots now “solve complex soft-body manipulation work” that older machines couldn’t touch.
Black Forest Labs was founded in August 2024 by researchers who helped build the original Stable Diffusion models at Stability AI. The founding group is described in the article as drawn from researchers with direct contributions to those earlier image-generation models. Robin Rombach is co-founder and CEO of Black Forest Labs. The article identifies the company’s origins in that August 2024 founding and its personnel’s prior work on Stable Diffusion.
The article reports that the Flux Dev and Schnell models were claimed by AI artists to hold the title of “best open source image generator.” Those models were reported to have beaten Stability’s Stable Diffusion 3.5 expectations. The article presents these assessments as claims from AI artists rather than formal benchmark figures. The mention of Flux Dev and Schnell appears alongside the company’s other model work in the article.
FLUX 3 is a multimodal video AI from Black Forest Labs that generates videos up to 20 seconds long with audio produced alongside and synchronized to the video, and the system is trained on images, video, and audio in a single shared architecture.
In head-to-head evaluations FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons, over Luma Ray 3.2 in 93% of comparisons, and over Gemini Omni and Seedance in 52% of evaluations, and the full system reportedly reacts in about 101 milliseconds.
The company’s ecosystem includes FLUX-mimic, a lightweight decoder that translates the model’s motion predictions into robot motions, and it is being tested with Audi on soft-body manipulation tasks such as fitting flexible door seals.


