Black Forest Labs Launches FLUX 3 Multimodal AI Model
Black Forest Labs has launched FLUX 3 in early access, a unified multimodal foundation model capable of generating up to 20-second videos with native audio in a single generation. The model jointly trains on video, image, and audio data, expanding the capabilities of the previous FLUX.1 and FLUX.2 series. By natively combining image, video, and audio generation into a single architecture, FLUX 3 advances unified multimodal AI and opens new possibilities, including robotics behavior prediction through a partnership with Mimic Robotics. The model is built on the Self-Flow learning framework and supports diverse tasks like text-to-video, image-to-video, and multilingual dialogue. In human evaluations of 10-second 720p videos with audio, FLUX 3 achieved win rates of 69% against Grok Imagine Video and 52% against Gemini Omni Flash.
## BACKGROUND
Black Forest Labs is a prominent AI research lab known for its state-of-the-art FLUX image generation models. Traditional generative AI systems often use separate models for generating video and audio, which can lead to synchronization issues, whereas unified multimodal models aim to generate both simultaneously for better coherence.