# FLUX 3: Multimodal Video, Image & Audio

Black Forest Labs released FLUX 3, a multimodal foundation model trained jointly on video, images, and audio, now available in Early Access.

- FLUX 3 Video generates clips with native audio up to 20 seconds in a single generation.
- In early evaluations, FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%.
- FLUX 3 Image early access opens in the following weeks; APIs, private weight access, and open-weight FLUX 3 Dev are planned.
- With mimic robotics, Black Forest Labs developed FLUX-mimic, a video-action model tested on production tasks at Audi.

## Why it matters

Content creators and robotics developers can now build on a single model that generates video, audio, and images and predicts actions, instead of chaining separate models.

## Sources

- [Black Forest Labs: FLUX 3: Multimodal Video, Image & Audio](https://bfl.ai/blog/flux-3)

---

Summarized by dstilled on 2026-07-23. https://dstilled.ai/story/3bbe86d0-4923-4c75-bdf6-32a66d2839da
