FLUX 3: Multimodal Video, Image & Audio
Black Forest Labs released FLUX 3, a multimodal foundation model trained jointly on video, images, and audio, now available in Early Access.
- FLUX 3 Video generates clips with native audio up to 20 seconds in a single generation.
- In early evaluations, FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%.
- FLUX 3 Image early access opens in the following weeks; APIs, private weight access, and open-weight FLUX 3 Dev are planned.
- With mimic robotics, Black Forest Labs developed FLUX-mimic, a video-action model tested on production tasks at Audi.
Content creators and robotics developers can now build on a single model that generates video, audio, and images and predicts actions, instead of chaining separate models.
Sources
Read this as text
Back to the AI news