Black Forest Labs Announces Flux 3 Multimodal Model
Black Forest Labs has announced Flux 3, a unified multimodal model capable of generating video, audio, and images, as well as predicting actions. The company plans to release an open-weight version of the model, called "FLUX 3 Dev," in the coming weeks and months. The release of Flux 3 represents a significant step forward in unified multimodal AI, combining content creation across multiple mediums with action prediction in a single architecture. Providing open-weight access via FLUX 3 Dev will allow developers and researchers to run, customize, and build upon this advanced model locally. Flux 3 jointly learns from images, videos, and audio within a single architecture, rather than processing these modalities in isolation. While the model promises capabilities like 20-second video generation, early promotional materials have faced scrutiny for using jumpcuts rather than continuous video generation.
## BACKGROUND
Black Forest Labs previously gained prominence for its FLUX.1 suite of text-to-image models, which set new standards in image detail and prompt adherence. Open-weight models differ from fully closed APIs by publicly releasing their core neural network weights, enabling users to download and run them on their own hardware.