Black Forest Labs Releases FLUX 3 Action, a 7B Robot Control Model
Black Forest Labs has released FLUX 3 Action, an open-weights 7-billion parameter world action model designed to convert visual intelligence into direct action control. The model allows developers to generate executable control signals for robots, simulators, and games based on multimodal input. Open-sourcing a 7B foundation model specifically built for action control significantly lowers the barrier to entry for embodied AI research and development. By bridging multimodal video understanding with physical action execution, it accelerates progress toward versatile autonomous agents in robotics. FLUX 3 Action operates as a world action model within BFL's unified FLUX 3 architecture, which integrates image, video, audio, and action processing. It translates environmental context and visual representations directly into control commands for physical hardware or simulation platforms.
## BACKGROUND
In robotics, vision-language-action (VLA) and world action models are multimodal foundation architectures designed to connect visual perception directly with executable physical actions. Black Forest Labs (BFL) is prominent for its FLUX generative AI models, and the FLUX 3 family expands these capabilities into a single model capable of processing visual, auditory, and robotic action control domains simultaneously.