Black Forest Labs launched Flux 3 on July 23, a multimodal frontier model that jointly handles image generation, video with native audio, and robotic action prediction from a single trained backbone. The Freiburg-based lab says the architecture learns one representation of the visual world rather than stitching together separate models for each modality.
Flux 3 ships in four product tiers: Flux 3 Video with optional synchronized audio, Flux 3 Image for text-to-image generation, Flux 3 Action for robotics and computer-use tasks, and an upcoming open-weight Flux 3 Dev release. Video and Action tiers enter a gated early access program immediately, with image generation rolling out in the coming weeks.
The launch marks Black Forest Labs’ first public video model and its most ambitious release since the original Flux image models took the open-source community by storm. The company has not yet published pricing, production SLAs, or detailed benchmarks, and the absence of downloadable weights at launch marks a departure from its usual open-weight strategy.