MODEL
Flux 3
modeltopic-notemultimodalvideo-gen
Overview
Flux 3 is Black Forest Labs‘s multimodal foundation model trained jointly on image, video, and audio — BFL’s first native-audio video release and the first native-audio frontier video release from a European lab. Supports text/image/video-to-video generation, keyframe stitching, and video with native audio up to 20 seconds; paired with a Flux-Mimic robotics action model in limited early-access.
Timeline
- 2026-07-24-AI-Digest — Flux 3 released as Black Forest Labs‘s multimodal foundation model — text/image/video-to-video generation, keyframe stitching, video with native audio up to 20 seconds. Paired Flux-Mimic robotics action model in limited early-access with unnamed research and commercial robotics partners. BFL’s internal evals claim a 93% win-rate vs. Luma Ray 3.2 and ~52% parity vs. Seedance and Gemini Omni Flash — vendor-reported and not yet independently benchmarked. Narrow read: first-of-a-kind for BFL (native audio in generated video) and first-of-a-kind for European labs. Limited early access via API to “initial partners” — no public pricing disclosed. The Decoder explicitly flags the vendor claims as awaiting independent validation. Structural read the corpus carries: Flux 3 moves BFL from “leading open-weight image lab” to “multimodal frontier candidate” — Stability AI attempted the same trajectory and stalled. The Flux-Mimic robotics arm is the cross-vertical bet — a multimodal image/video/audio foundation model with a paired action model targeting robotics is the multi-vertical play that made Google’s Gemini strategy load-bearing. Whether the vendor-reported wins hold under independent evaluation is the load-bearing question; the ambition shape is unambiguous. 60-day watch: whether an independent benchmark corroborates the 93% Luma Ray win-rate; whether BFL names a first robotics customer for Flux-Mimic.
Key Developments
- First European Frontier Multimodal With Native-Audio 20-Second Video (July 24, 2026): Multimodal foundation model jointly trained on image/video/audio; native-audio generated video up to 20 seconds; paired Flux-Mimic action model in limited robotics early-access. BFL internal evals: 93% win-rate vs. Luma Ray 3.2, ~52% parity vs. Seedance and Gemini Omni Flash — vendor-reported, awaiting independent validation. Limited early access via API to initial partners; no public pricing. Corpus framing: ambition shape is unambiguous but vendor claims need independent evaluation before the read consolidates.
- 2026-07-25-AI-Digest — FLUX 3 Action ships as Black Forest Labs‘s first robotics-oriented model, built on the Flux 3 unified multimodal architecture (image / video / audio) shipped this week. Bet: cross-modal grounding on a single architecture beats specialist stacks for the cause-and-effect reasoning robotics needs. Narrow read: variant of the Flux 3 stack already covered in 2026-07-24-AI-Digest, not a separate architecture — the robotics-fine-tuned surface on the same underlying multimodal foundation. Structural read the corpus carries: the temptation is to fold this into a “European frontier labs pivot to physical AI” thesis — the evidence doesn’t support the plural. Mistral, Aleph Alpha, and Silo remain LLM/multimodal-focused; BFL is a one-lab move, not a coalition rotation. What is real: the frontier image-model labs (BFL specifically) can amortise their multimodal training investment across a second downstream market. Read as one lab’s option value on a second market, not as a continent-wide strategic re-alignment. Track FLUX 3 Action under this family note rather than spawning a variant-specific topic note — same underlying architecture, robotics-fine-tuned surface.
Related
See also: Black Forest Labs, Luma Ray 3.2, Gemini Omni Flash, MOC - Major Companies.