- MiniMax H3 (Hailuo 3.0) launched July 31, 2026: native-2K, 4 to 15-second video with synchronized stereo audio generated in the same pass as the picture, live now under model ID MiniMax-H3 and in the Hailuo AI app.
- It debuts at #2 on the Artificial Analysis video arena leaderboard (Elo 1238) behind Gemini Omni Flash (1245), and ahead of Veo 3.1 and Kling 3.0 Omni.
- It's the third frontier video model to ship native audio generation in about six months, after Kling 3.0 Omni (February 2026) and Gemini Omni Flash (June 2, 2026). The gap between the last two launches was 59 days, and native audio has gone from novelty to expected default.
- Native audio collapses the render-and-mix steps of a video pipeline, not the planning and editability steps. Multi-speaker consistency, brand voice, and structure before generation still need dedicated tooling.
On July 31, 2026, MiniMax released MiniMax H3, also sold under the consumer name Hailuo 3.0. It's an omni-modal video model that generates native-2K video with synchronized stereo dialogue, ambience, and sound effects, all produced in the same pass that renders the picture. That's the headline. The more interesting story is what happens to a video production pipeline once a model stops treating audio as a separate problem to solve afterward.
H3 is live now in the MiniMax platform API under the model ID MiniMax-H3, and in the consumer Hailuo AI app, per MiniMax's own release notes. It ranks #2 on the Artificial Analysis video arena leaderboard, MiniMax is making aggressive pricing claims against mainstream competitors, and open weights are due "in the coming days." None of that is really the news either. The news is that native-audio AI video generation just went from novelty to the third default in six months: MiniMax H3 follows Kling 3.0 Omni and Gemini Omni Flash in shipping synchronized audio as a standard part of what a frontier video model produces, not an add-on you request separately.
This post covers what H3 actually shipped, where it lands competitively, how fast native audio is spreading across the model layer, and, more usefully, what genuinely disappears from a production pipeline when the base model handles audio, and what still doesn't.
What MiniMax H3 Actually Shipped
MiniMax built H3 around four pieces of new architecture, according to the company's technical writeup: Contextual Omni Representation, a captioning pipeline that compresses roughly 100,000 tokens of source material down to about 4,000 tokens of inference, using language as the bridge between mixed inputs and video output; H3-VAE, a rebuilt tokenizer that delivers a 4x gain in effective sequence length and makes native 2K output computationally viable; the H3-Omni Transformer, which separates understanding and generation workloads and lifts training throughput by nearly 30%; and In-Context Regeneration, where the base model regenerates its own low-resolution output in-context instead of relying on a separate super-resolution module for the 2K upscale.
What that architecture translates to for a working video generator:
- Native 2K output (1440px short edge) at 24fps, 4 to 15 seconds per clip, across seven aspect ratios including 21:9, 16:9, 9:16, and 1:1.
- Synchronized stereo audio in the same generation pass. Dialogue, ambience, room tone, and sound effects are jointly modeled rather than generated as separate tracks and mixed afterward.
- An omni-reference system that accepts up to 9 reference images, 3 reference video clips, and 3 reference audio clips per generation (12 files max), plus audio-to-audio reference and instruction-based editing of existing clips.
On pricing, MiniMax claims H3's per-second cost at 2K is less than a third of mainstream competing 2K models, and at 768p, less than half of mainstream 720p pricing. Those are the vendor's own numbers, not an independent benchmark, and they're upper bounds ('less than'), so treat them as directional rather than exact.

Open weights are planned for public release "in the coming days," subject to legal and compliance review, MiniMax says. That would put a native-audio, 2K generation model into open-weight circulation faster than most of its closed competitors have moved.
Where H3 Lands on the Video Arena Leaderboard
H3 entered the Artificial Analysis video arena leaderboard on July 31 at #2 with an Elo score of 1238, in blind human comparisons, just behind Gemini Omni Flash at 1245. That's a tight gap between the top two models, and both sit well ahead of the rest of the field: Dreamina's Seedance 2.0 (720p) follows at 1223, then a real drop to Veo 3.1 at 1095 and Kling 3.0 Omni (1080p) at 1093.

The leaderboard position matters less than what all five of those models have in common: every one of them now generates audio natively, not as an optional bolt-on. Compare that to Veo 3.1, which was itself a notable step forward for native audio-in-video less than a year ago. The bar for what a frontier video model ships by default has moved fast.
Native Audio Is Becoming the Default, Not the Differentiator
Line up the release dates and a pattern shows up. Kling 3.0 shipped native audio, 4K, and 15-second generation in February 2026, with a Kling 3.0 Omni and Turbo variant following on June 17. Gemini Omni Flash rolled out to YouTube Shorts and the Gemini app on June 2, 2026, bringing conversational, multi-input, native-audio generation to a platform with roughly 2.7 billion monthly users. MiniMax H3 followed on July 31. Three separate model families, three separate labs, the same capability, inside a six-month window.
The two most precisely dated launches in that sequence, Gemini Omni Flash and MiniMax H3, are 59 days apart. That's the tightest gap between major native-audio flagship releases so far this year, and it's a useful proxy for how fast this specific capability is spreading across the model layer, even accounting for the fact that Kling's exact February date isn't publicly pinned down to the day.

Methodology note: the timeline above is ngram's own count of frontier text-to-video models that ship synchronized audio generation as a default capability, built from each vendor's release announcement (MiniMax, Google) and third-party coverage of Kling's February 2026 launch, cross-checked in early August 2026. It's a small sample by design, three data points, because this is still an emerging capability. That's exactly why the cadence is worth watching.
What Actually Collapses in a Video Production Pipeline

The pre-native-audio workflow for a talking video was, roughly: generate or shoot the visual, write a voiceover script, generate the voice with a text-to-speech model, run a lip-sync pass to match mouth movement to the new audio, then mix dialogue, ambience, and any sound effects into a final track. That's four or five separate render steps, each with its own queue, its own review pass, and its own chance to drift out of sync with the others.
When a model like H3, Gemini Omni Flash, or Kling 3.0 Omni generates dialogue, ambience, and effects jointly with the picture, several of those steps genuinely disappear for the scenarios they cover. There's no separate lip-sync pass because the mouth movement and the audio were never separate outputs to begin with. There's no manual level-setting between a dialogue track and a music bed because the model reasoned about them together. For a single short clip, that's a real reduction in render steps and review cycles, not just marketing language.
What Native Audio Still Doesn't Solve

Collapsing a step is not the same as removing the problem it existed to solve. Four gaps show up quickly once you try to use native audio for anything beyond a single short clip:
- Fine-grained control. If one line of dialogue comes out wrong, the usual fix is regenerating the whole clip, not editing a single audio track in isolation, because the audio and video were never separate layers.
- Multi-speaker consistency across scenes. Kling 3.0 Omni's own documentation highlights speaker mapping within a single clip as a new feature, which tells you it's a hard problem: keeping the same two voices, tones, and turn-taking consistent across many separate scenes and clips in one project is a different and harder task than getting one clip right.
- Brand voice. A native-audio model generates a plausible voice for the scene. It doesn't know your team's approved brand voice, cloned or otherwise, unless you feed it as a reference, and doing that consistently across a multi-scene project is its own workflow problem.
- Structure before generation. A script, a storyboard, and a scene-by-scene plan still have to exist before any model call happens. Native audio changes what comes out of one generation step. It doesn't decide what the video should say or how it should be organized.
Why Production Tools Still Orchestrate Across Multiple Models
This is why tools built for real video production keep composing across specialized providers instead of betting the whole workflow on one model's built-in voice or one model's built-in video. ngram's own voiceover stack, for example, already routes across multiple TTS providers, ElevenLabs and MiniMax TTS, with OpenAI TTS configured as an additional fallback, rather than depending on a single vendor's voice engine. That's not a workaround for a missing capability. It's the same reasoning this whole piece is about: no single provider stays the best fit for every language, every voice style, and every price point at once, so a production tool needs the option to route around any one of them.

That's a 25% compound annual growth rate for the standalone AI voiceover market, according to SkyQuest's market research, the exact layer that native-audio video generation is now folding into the base model. Both things are true at once: base models are absorbing simple, single-clip voiceover generation, and demand for dedicated, controllable, multi-provider voice infrastructure keeps growing, because production tools need consistency and redundancy that a single model's built-in voice output doesn't guarantee on its own.
The AI Video Market Is Growing Into This Shift
The broader AI video generation market was worth an estimated $716.8 million in 2025 and $847 million in 2026, on track to reach roughly $3.35 billion by 2034, an 18.8% compound annual growth rate, per Fortune Business Insights. Text-to-video already accounts for 46.25% of that market globally, and marketing and advertising is the single largest use case at 33.88%.

A market growing that fast is exactly the environment where you'd expect capability races like this one. Resolution and duration were the first fronts, then came reference-image consistency and instruction-based editing. Native audio is the current front, and based on the last six months, it won't be the last one.
What This Means for AI Video Production Timelines
For a single short clip, a social cutdown, a quick reaction video, native audio is a genuine speed win. What used to be a script-to-voice-to-lip-sync-to-mix chain is now one generation call, and that shortens time-to-first-cut meaningfully.
For a multi-scene, branded production, the picture is more mixed. The planning work upstream (deciding what the video says, to whom, and in what order) still has to happen before any model is called, native audio or not. And the control work downstream (fixing one line without a full regenerate, keeping the same brand voice across ten scenes, handling multilingual dubbing) still benefits from dedicated, swappable tooling rather than a single model's all-in-one output.
That's a reasonable way to think about where a model like H3 fits into a broader production tool: an omni-modal model with native audio is exactly the kind of capability a tool like ngram could draw on for a scene's video-and-ambient-audio layer, the same way it already draws on multiple TTS and image providers today. But the parts of a production pipeline that stay their own specialized steps, dialogue consistency across speakers, brand voice, and multilingual dubbing with lip-sync, don't go away just because one model call produces a synced clip. That's the real shift here: not that one model replaces a whole pipeline, but that the pipeline's shape keeps changing every few months as each layer gets absorbed or specialized in turn.
Frequently Asked Questions
What is MiniMax H3 (Hailuo 3.0)?
MiniMax H3, also marketed as Hailuo 3.0, is an omni-modal AI video model released July 31, 2026. It generates native-2K video up to 15 seconds long with synchronized stereo dialogue, ambience, and sound effects produced in the same generation pass as the picture. It's live in the MiniMax platform API (model ID MiniMax-H3) and the consumer Hailuo AI app.
How is MiniMax H3 different from Hailuo 2.3?
Hailuo 2.3 topped out at 1080p and roughly 10 seconds, generated from a prompt or a single image, with no native audio. H3 moves to native 2K and up to 15 seconds, adds an omni-reference system that accepts up to 9 images, 3 video clips, and 3 audio clips per generation, and generates synchronized audio and voice transfer in a single pass.
Where does MiniMax H3 rank on the Artificial Analysis leaderboard?
H3 debuted at #2 on the Artificial Analysis video arena leaderboard with an Elo score of 1238, just behind Gemini Omni Flash at 1245, and ahead of Seedance 2.0 (720p), Veo 3.1, and Kling 3.0 Omni.
Is MiniMax H3 cheaper than other AI video models?
MiniMax claims H3's per-second price at 2K is less than a third of mainstream competing 2K models, and at 768p, less than half of mainstream 720p pricing. Those are the vendor's own upper-bound figures rather than an independent benchmark, so treat the exact multiple as directional.
Will MiniMax H3's weights be open source?
MiniMax says it plans to release H3's model weights publicly "in the coming days," subject to legal and compliance review. No firm date has been confirmed as of this writing.
What is native audio in AI video generation?
Native audio means a video model generates dialogue, ambience, and sound effects jointly with the picture, in the same inference pass, rather than requiring a separate text-to-speech step and a separate lip-sync pass afterward. Gemini Omni Flash, Kling 3.0 Omni, and MiniMax H3 all ship this capability as of mid-2026.
Does native audio replace the need for separate voiceover and dubbing tools?
For a single short clip, largely yes. For a multi-scene branded video, not entirely. Native audio doesn't currently offer fine-grained editing of one line without a full regenerate, doesn't guarantee a consistent brand voice across many scenes unless it's fed as a reference every time, and doesn't handle multilingual dubbing with lip-sync as its own dedicated workflow the way purpose-built video translation tools do.
How is MiniMax H3 different from Kling 3.0 Omni and Gemini Omni Flash?
All three generate native-2K-class video with synchronized audio as a default. Kling 3.0 Omni is known for native lip-synced dialogue across five languages and multi-speaker mapping within a scene. Gemini Omni Flash is distributed through YouTube Shorts and the Gemini app, reaching a huge existing user base for free. MiniMax H3 differentiates on price claims, a large omni-reference input system, and a planned open-weights release, an option the other two don't currently offer.
You just read it. Now watch it.
ngram turns this post into a short explainer video: scenes, voiceover, and motion graphics included.






