- Black Forest Labs launched FLUX 3 on July 23, 2026: its first frontier model jointly trained across image, video, audio, and robotic action in one architecture.
- FLUX 3 Video generates 20-second clips with native audio, and Black Forest Labs claims 77-93% preference over rivals in early, self-reported testing.
- The bigger signal: FLUX-mimic, built on FLUX 3 by robotics company mimic, already runs production tasks at Audi, a real deployment, not a demo reel.
- The rollout stays staged: video and action first, image and API access later in 2026, now the default pattern for frontier launches.
Black Forest Labs launched FLUX 3 on July 23, 2026, and the announcement reads like two different products stapled together. One half is a text-to-video generator with native audio. The other half is a robotics controller running on a car manufacturer's assembly line. Both halves come from the same set of weights.
That is the actual news here, more than any single spec. FLUX 3 is Black Forest Labs' first multimodal frontier model: image, video, audio, and robotic action, jointly trained inside one architecture rather than bolted together as separate systems. FLUX 3 Video and FLUX 3 Action are live now in gated early access. FLUX 3 Image follows in the coming weeks, and API access, private model weights, and an open-weight "FLUX 3 Dev" version are planned later this year, according to Black Forest Labs' own launch post.
The most interesting detail in the launch is not the demo reel. It is that the first production customer named in the announcement is Audi, running FLUX 3's action model on a real factory floor.
What Black Forest Labs Actually Shipped
FLUX 3 Video generates clips up to 20 seconds long, with audio (dialogue, sound effects, ambient noise) generated natively alongside the visuals rather than added in a separate pass, per VentureBeat's coverage. It supports text-to-video, image-to-video, and video-to-video generation, plus video continuation, keyframe-controlled transitions, multilingual dialogue, on-screen typography, and clip-chaining into longer sequences.
None of that is available to the general public yet. FLUX 3 Video and FLUX 3 Action opened in a gated early access program on launch day. Anyone can apply, but Black Forest Labs approves each request individually.
- Now: FLUX 3 Video, in gated early access
- Now, partners only: FLUX 3 Action, including the FLUX-mimic robotics variant built with mimic robotics
- Coming weeks: FLUX 3 Image, in its own early access phase
- Later in 2026: FLUX 3 Video API access and private model weights
- Later in 2026: FLUX 3 Dev, an open-weight version of the multimodal backbone
That staged sequencing is deliberate, and it matters for how teams should read every frontier launch from here, which the later section on rollout norms picks up.

Black Forest Labs published its own comparison data alongside the launch: in blind pairwise human evaluations of 10-second text-to-video clips, FLUX 3 was preferred 77% of the time over Runway Gen-4.5, 93% over Luma Ray 3.2, 69% over Grok Imagine Video, 60% over Kling v3 Pro, and 52% over Seedance 2.0 and Gemini Omni Flash. Black Forest Labs describes these as preliminary, self-run comparisons rather than an independently audited benchmark, and the company has not published its rater count, prompt set, or selection methodology. Treat the numbers as a company's own claim about its own model until an outside lab reproduces them, not as settled fact.
Why Train Video, Audio, and Robot Action Together
Every major video model to date, including Sora, Veo, and Kling, has been trained primarily on video and learned to render motion convincingly within that single modality. FLUX 3's pitch is different: train image, video, audio, and physical action from the same underlying representation, on the theory that each modality is a different sensor reading of the same reality.
Black Forest Labs frames the reasoning directly on its own blog: each modality is "a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process." Train them together, the company argues, and the model has to reconcile constraints that a single-modality model never has to face: "the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past."
That is a concrete, testable claim, not marketing language. A video-only model can render a hand knocking over a glass and still get the physics wrong: an object that clips through a surface, a shadow that doesn't track the light, a sound that arrives a frame late. A model trained jointly on video and physical action data has to predict what actually happens next in a way that a robot can act on, which is a much harder bar than looking plausible for six seconds.
That is also the real difference between a video generator and what NVIDIA and others now call a world model: one renders a plausible-looking scene, the other has to represent how a scene behaves well enough that a downstream system, a robot arm or a video editor doing a scene extension, can rely on it. NVIDIA made a similar case when it released Cosmos 3, a release that named Black Forest Labs as a founding member of its Cosmos Coalition months before FLUX 3 shipped. That detail alone suggests Black Forest Labs was building toward this convergence well before July's launch.
The First Customer Is a Car Factory, Not a Marketing Team
Here is the detail that separates FLUX 3 from a typical model-launch news cycle: Black Forest Labs did not lead with a filmmaker or a marketing team. It led with a robotics deployment already running in production.
FLUX-mimic, built by robotics company mimic on the FLUX 3 backbone, is running at Audi today. Christoph Schneider of Audi Production Lab put it plainly:
In partnership with mimic, Audi has been testing and deploying FLUX-mimic. We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics. This can have a major impact in assisting our employees, increasing efficiency, and expanding flexible automation across production and logistics operations.

The specific tasks matter because they explain why conventional automation skipped them for so long: kitting parts into structured trays, inserting electronic control units into tight-fitting fixtures, assembling components, and handling soft, flexible materials like seals and cables, according to Black Forest Labs' mimic launch post. Rigid, pre-programmed robot arms are good at repeating the exact same motion on the exact same rigid part. They are bad at the "close enough" judgment a hand makes when a cable is bent slightly differently than the previous unit.

The efficiency claim behind FLUX-mimic is just as notable as the deployment itself. mimic robotics says the model can learn a new manipulation task from about 30 minutes of robot demonstration data, versus the 30-plus hours that prior approaches typically required, cutting deployment timelines from months to weeks. mimic co-founder and Chief Product Officer Stephan-Daniel Gravert said Audi "represents the kind of manufacturing partner we built FLUX-mimic for," pointing to production environments that "demand automation that is flexible enough to handle unstructured tasks, reliable enough for continuous operation, and can be integrated without months of engineering overhead."
mimic robotics itself is a two-year-old company, founded in 2024 out of ETH Zurich, with a 50-plus person team that includes alumni of Google DeepMind and Tesla's Optimus program, and offices in Zurich and San Francisco.
This Is Part of a Bigger Money Wave
FLUX 3's Audi deployment is not an isolated bet. It lands in the middle of a genuine venture-capital rush into models that connect vision, video, and physical action.

Three of the biggest rounds of 2026 all went to companies working some version of this problem. Yann LeCun's AMI raised a $1.03 billion seed round in March, the largest seed round ever for a European startup, to build what LeCun calls world models rather than language models. Fei-Fei Li's World Labs closed a $1 billion round in February, with $200 million of it from Autodesk. And Skild AI raised $1.4 billion in January at a valuation over $14 billion, more than triple what it was valued at seven months earlier, to build what it calls an "omni-bodied" foundation model that can control different robot hardware without retraining from scratch.

Zoom out further and the pattern holds at the industry level, not just in a handful of marquee rounds. Physical AI and robotics startups raised $27.6 billion across 1,009 deals in 2025, according to PitchBook data cited in an industry funding analysis, more than double the $13.7 billion raised in 2024. Within that total, robot foundation models specifically pulled in more than $2.2 billion, defense and security robotics took $8.0 billion, and autonomous drones took $6.2 billion.

Black Forest Labs' own numbers fit that pattern. The company raised a $300 million Series B in December 2025 at a valuation north of $3 billion, backed by investors including a16z, Nvidia, and Salesforce Ventures, before FLUX 3 had even shipped. The money was already betting on exactly the kind of unified visual-and-physical model FLUX 3 turned out to be.
Staged, Gated Launches Are the New Normal for Frontier Models
FLUX 3's rollout, video and action first, image weeks later, API and open weights sometime after that, is not a one-off. It is close to the default pattern for frontier model launches now.

OpenAI's Sora 2 launched invite-only on September 30, 2025, priority access for ChatGPT subscribers, a reported "one invite unlocks four more" referral system, and a staged regional rollout that started in the US and Canada, according to TechCrunch's launch coverage. That gating did not save it. OpenAI shut down the Sora web and app experiences on April 26, 2026, with the API following in September, per The Decoder. The shutdown was a hard lesson in what AI video's unit economics actually demand once the hype cycle ends.
The lesson is not that gated rollouts are a red flag. Almost every serious model launches gated now; compute cost and safety review make an instant, ungated release irresponsible at frontier scale. The lesson is that a gated rollout by itself proves nothing about whether a model earns a durable customer. Sora had enormous consumer hype and an invite system and still could not turn that into a business. FLUX 3, so far, has a named industrial customer running real production tasks on day one. That is a meaningfully different, and harder to fake, signal than a waitlist.
What This Means If You're Evaluating AI Video Tools Today
For a team choosing an AI video generator or comparing options for text-to-video production, FLUX 3 itself is not something most teams will touch directly this year. It's gated, it's aimed first at partners and researchers, and the image variant that would matter most for everyday marketing content is not out yet.
What does matter is the direction. Every serious ai video generation platform is converging on the idea that a model gets more useful, not less, once it has to hold up under physical scrutiny: object permanence across a scene, plausible motion, sound that lines up with what's happening on screen. Those are exactly the properties that make longer clips, cleaner scene extensions, and multilingual dialogue actually usable in production, not just impressive in a launch demo.
That is also the argument for staying model-agnostic rather than betting on one video engine: an agent picks the storyboard and visual direction for a scene, then generates against whichever underlying model fits that shot best, rather than locking a team into one vendor's roadmap. As frontier models get better at holding physical consistency, that ceiling rises for anyone building on top of them; the workflow of prompt, script, storyboard, video doesn't have to change just because the engine underneath does. ngram builds on that logic directly: Premium and Ultimate plans already unlock every advanced model in the product today, on the bet that the model layer will keep moving and users shouldn't have to re-shop every time it does. Try ngram if you want a workflow that stays useful as the models underneath it keep changing.
What FLUX 3 Doesn't Mean Yet
It's worth being precise about what this launch does not prove. Black Forest Labs' preference numbers are self-reported and unaudited. FLUX-mimic is running specific tasks at one manufacturer, not a general-purpose robot brain. FLUX 3 Image, the variant most relevant to everyday marketing and content teams, has not shipped. And a model that reasons well about physics still needs a human to decide what a video should say, who it's for, and how long it should run.
The three questions worth tracking as FLUX 3 and its rivals mature: can the model hold physical consistency across a full scene rather than one impressive clip, can it take structured direction like a storyboard or reference image rather than just a text prompt, and can it revise one part of a scene without regenerating the whole thing. Those, not clip length, are the questions that will separate a genuinely useful production tool from a longer demo reel.
Frequently Asked Questions
What is Black Forest Labs?
Black Forest Labs is a German AI company founded in August 2024 by Robin Rombach and other researchers who previously built the latent diffusion technology behind Stable Diffusion. It makes the FLUX family of image and, as of FLUX 3, video, audio, and action-prediction models, and raised a $300 million Series B in December 2025 at a valuation above $3 billion.
What is FLUX 3?
FLUX 3 is Black Forest Labs' first multimodal frontier model, jointly trained across image, video, audio, and robotic action prediction in one architecture, rather than as separate single-modality systems stitched together.
Is FLUX 3 available to the public?
Not yet, fully. FLUX 3 Video and FLUX 3 Action opened in gated early access on July 23, 2026, with applications open to anyone but approval decided by Black Forest Labs. FLUX 3 Image is expected in the following weeks, and API access, private weights, and an open-weight FLUX 3 Dev version are planned later in 2026.
What is FLUX-mimic?
FLUX-mimic is a video-action model built by robotics company mimic on top of the FLUX 3 backbone. It is running in Audi production facilities today, handling soft-body manipulation tasks like kitting parts and inserting components that conventional rigid automation struggles to do cost-effectively.
Why did Black Forest Labs train FLUX 3 on robot action data, not just video?
The company's stated reasoning is that video, audio, and physical action are different sensor readings of the same underlying reality, and training a model to reconcile all of them (sound matching impact, motion obeying mass) produces stronger physical consistency than training on video alone. That consistency is also what a robot needs to act on a prediction rather than just render it.
How does FLUX 3 compare to Runway, Kling, or Luma's video models?
Black Forest Labs published its own preliminary preference data claiming FLUX 3 beat Runway Gen-4.5, Luma Ray 3.2, Grok Imagine Video, Kling v3 Pro, and Seedance 2.0 and Gemini Omni Flash in blind human comparisons of 10-second clips. These are the company's own unaudited results, not an independent benchmark, so treat them as a starting claim rather than a settled ranking.
You just read it. Now watch it.
ngram turns this post into a short explainer video: scenes, voiceover, and motion graphics included.






