|

Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

Black Forest Labs (BFL) has launched FLUX 3, a multimodal basis mannequin that learns from photos, movies and audio inside a single structure. It can be the primary FLUX mannequin to ship video, audio and motion prediction from one set of weights.

The Black Forest Labs (BFL) analysis staff argues that no single modality offers a whole description of the world. Images seize spatial construction at one immediate. Video restores time and exposes bodily dynamics. Audio reveals causal relationships between mechanical occasions and sound. Each is handled as a lossy projection of the identical underlying actuality.

Training on all of them directly means the modalities constrain one another. The sound has to match the influence. The movement has to obey the mass. The analysis staff calls FLUX 3 its first mannequin constructed solely on that precept.

The technique beneath: Self-Flow

FLUX 3 builds on Self-Flow, BFL’s technique for aligning multimodal era and understanding in a single structure. Self-Flow combines the stream matching goal with a self-supervised function reconstruction goal. The reference implementation on GitHub is Apache-2.0 and makes use of SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token masks ratio and self-distillation from an EMA trainer at layer 20 to a pupil at layer 8.

That launched checkpoint is an ImageInternet 256×256 analysis mannequin, not FLUX 3. BFL states that it ‘considerably scaled up compute and knowledge assets’ on the identical strategy to coach FLUX 3 throughout video, photos and audio concurrently. Self-Flow itself was launched in March 2026, so it isn’t new to this launch. What is new is the size.

What FLUX 3 Video does

FLUX 3 Video generates clips as much as 20 seconds lengthy in a single era, with native audio. The supported modes cowl text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for managed transitions, and generative video-audio continuation from enter video and audio.

BFL additionally lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and robust typography era with animated designs. The BFL staff reviews specific energy in human facial expressions and in associating sounds with bodily occasions.

Performance

BFL staff revealed preliminary human choice outcomes. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was most well-liked over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. Against Grok Imagine Video the determine is as much as 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result’s 52%, near a coin flip.

Interactive Explorer


bfl@flux-3:~/real-world-models
Early Access




enter
Sources: bfl.ai/weblog/flux-3 · bfl.ai/weblog/flux-3-mimic · mimicrobotics.com · figures dated 23 Jul 2026
Built by Marktechpost

Similar Posts