Black Forest Labs Launches FLUX 3: Multimodal Video, Audio, and Robotics AI Model
What happened: German AI startup Black Forest Labs (BFL) released FLUX 3, a multimodal foundation model trained on images, video, audio, and robotic action prediction.
What happened: German AI startup Black Forest Labs (BFL) released FLUX 3, a multimodal foundation model trained on images, video, audio, and robotic action prediction. The model can generate up to 20 seconds of synchronized video and audio, supports multilingual dialogue, and can be fine-tuned for robotic manipulation tasks with as little as 30 minutes of data. BFL is piloting FLUX-mimic, its video-action decoder, with Audi to automate soft-body assembly tasks. The company claims FLUX 3 outperforms Runway Gen-4.5 and Luma Ray 3.2 in human-preference tests, though these results are self-reported. BFL recently raised a $300 million Series B at a $3.25 billion valuation.
Why it matters: FLUX 3’s early-access launch marks a step forward in integrating generative AI with real-world robotics, a domain where most competitors remain focused on image or text modalities. The Audi partnership demonstrates immediate industrial application, though independent third-party evaluations of FLUX 3’s performance are not yet available. With open weights for developers planned later in 2026, BFL’s approach could broaden access to advanced multimodal AI, though for now, access remains gated.
Source: Decrypt