Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction
Black Forest Labs has launched FLUX 3, a multimodal foundation model capable of processing images, video, and audio within a unified architecture. This release marks the first time the company has integrated video, audio, and robot action prediction into a single set of weights. By consolidating these capabilities into one model, the developers aim to improve how AI systems interpret and interact with diverse sensory inputs. This approach signals a shift toward more integrated foundation models for complex, cross-modal tasks.
Covered by 1 source
- MMarkTechPost↗Michal Sutter3d ago