FLUX 3: Multimodal Video, Image & Audio

Black Forest Labs released FLUX 3, a multimodal foundation model trained jointly on video, images, and audio, now available in Early Access.

Content creators and robotics developers can now build on a single model that generates video, audio, and images and predicts actions, instead of chaining separate models.

Sources

Read this as text

Back to the AI news