- Generative AI video models have no universal 2026 winner: Gemini Omni Flash leads picture-only text-to-video at 1324 Elo and $6.00/min, while seven cheaper SKUs also sit on the price-quality frontier.
- Only 2 of the top 15 models have open weights; MiniMax H3 leads that segment at 1301 Elo.
- Audio changes the field: the With Audio board has 33 rows, the No Audio board has 80, and only 30 model names overlap.
- Official documentation left 51 of 192 audited capability cells unknown across 16 prominent SKUs, especially provenance and watermarking.
A leaderboard gives you a rank. A procurement decision needs a boundary. We analyzed 80 current generative AI video models across picture preference, list price, open-weight status, lab origin, audio-pool membership, and 12 capability fields. The result is less tidy than a top-ten list: Gemini Omni Flash leads the picture-only board, eight SKUs sit on the price-quality frontier, and the most important spec-sheet questions are often unanswered.
The primary dataset is the Artificial Analysis Text-to-Video leaderboard filtered to Current, Global, and No Audio. We captured the live board on 1 September 2026 at 1:32 PM Pacific time. Elo, prices, ranks, sample counts, and confidence intervals can move after that timestamp.
How to read the AI video model leaderboard
For generative AI video models, this report uses the No Audio board as its primary view because that setting isolates visual preference. Voters compare two clips generated from the same prompt and choose the one they prefer. Artificial Analysis aggregates those pairwise votes with a Bradley-Terry model, rescales the result to an Elo-like score, and recomputes ratings hourly. The benchmark keeps modality pools separate, so text-to-video, image-to-video, With Audio, and No Audio scores are not interchangeable.
The benchmark's default target is a 10-second, 16:9 clip at 1080p or the nearest supported setting. Price is the provider's published cost to generate one minute at those default settings. That makes $/min useful for comparison, but it is not a complete invoice. Retry rates, failed generations, queue priority, input processing, upscaling, and enterprise discounts sit outside the number.
Elo also needs its confidence interval. A model at 1222 with a +/-7 interval cannot be cleanly separated from a model at 1219 with a +/-8 interval. Rank order may look precise even when the underlying ranges overlap. Conversely, a gap whose intervals do not touch deserves more weight than a one-point table lead. The Artificial Analysis methodology documents those choices and states that providers do not pay for placement or favorable outcomes.
| Rank | Model | Elo (95% CI) | $/min | Open |
|---|---|---|---|---|
| 1 | Gemini Omni Flash | 1324 (1315-1333) | $6.00 | No |
| 2 | MiniMax H3 | 1301 (1291-1311) | $7.80 | Yes |
| 3 | HappyHorse-1.0 | 1282 (1274-1290) | $13.20 | No |
| 4 | Dreamina Seedance 2.0 720p | 1267 (1260-1274) | $9.07 | No |
| 5 | HappyHorse-1.1 | 1262 (1254-1270) | $9.90 | No |
| 6 | Wan2.7-260612 | 1243 (1235-1251) | $9.00 | No |
| 7 | Kling 3.0 1080p (Pro) | 1238 (1231-1245) | $13.44 | No |
| 8 | Kling 3.0 Omni 1080p (Pro) | 1227 (1219-1235) | $13.44 | No |
| 9 | grok-imagine-video | 1222 (1215-1229) | $4.20 | No |
| 10 | Bach-1.0 Preview | 1219 (1211-1227) | $3.00 | No |
| 11 | Kling 3.0 Omni 720p (Standard) | 1214 (1206-1222) | $10.08 | No |
| 12 | LTX-2.5 Fast | 1213 (1202-1224) | $7.80 | Yes |
| 13 | Wan 2.7 | 1213 (1204-1222) | $9.00 | No |
| 14 | Kling 3.0 720p (Standard) | 1212 (1205-1219) | $10.08 | No |
| 15 | Runway Gen-4.5 | 1212 (1205-1219) | No API price | No |
The table is a dated picture-quality snapshot. Wan 3.0 and MiniMax H3 Max do not appear because those SKUs were in the separate With Audio pool. Runway Gen-4.5 appears at rank 15 with no public API price on the captured board, so it is excluded from price comparisons.
Gemini leads picture quality, but the market has no one-model answer
Gemini Omni Flash ranks first at 1324 Elo and $6.00 per generated minute. MiniMax H3 follows at 1301 Elo and $7.80 per minute. HappyHorse-1.0 ranks third at 1282 Elo and $13.20 per minute. The 95% confidence ranges are 1315 to 1333 for Gemini, 1291 to 1311 for H3, and 1274 to 1290 for HappyHorse. None of those three intervals overlap.

The top line is unusually practical. Gemini is both the highest-rated model and the cheapest of ranks one through five. H3 costs $1.80 more per minute while scoring 23 Elo lower. HappyHorse-1.0 costs $7.20 more per minute while scoring 42 Elo lower. On these two measured axes, both are dominated by Gemini.
That does not make Gemini the universal choice. The board measures average visual preference under a standard generation setup. It does not directly score editability, prompt adherence by business domain, data handling, regional availability, self-hosting, time to first frame, or how often a team accepts the first result. A production team should treat 1324 as strong evidence about picture preference, then test the workflow constraints the arena does not cover.
The gap below the top three also matters. grok-imagine-video sits at 1222 Elo and $4.20 per minute. Bach-1.0 Preview is nearly tied at 1219 Elo and $3.00. Their confidence intervals overlap. A buyer paying for Gemini instead of grok spends an extra $1.80 per generated minute, or $0.30 per 10-second clip, for a 102-point Elo advantage on this snapshot. Against Bach, the premium is $3.00 per minute, or $0.50 per 10-second clip, for 105 Elo.
The price-quality frontier contains eight models, not eight winners
A model sits on the price-quality Pareto frontier when no other priced model is at least as good in Elo and no more expensive, with one of those comparisons strictly better. The rule removes models that cost more and score lower. It does not declare every surviving model equally good.

The eight frontier SKUs are Gemini Omni Flash (1324, $6.00), grok-imagine-video (1222, $4.20), Bach-1.0 Preview (1219, $3.00), Hailuo 02 Standard (1169, $2.80), Hailuo 2.3 (1169, $2.80), LTX-2.3 Fast (1120, $2.40), Seedance 1.0 Mini (1079, $2.22), and Agnes-Video-V2.0 (1050, $0.30). The two Hailuo rows share the same price and Elo, so neither strictly dominates the other.
The endpoints explain why 'frontier' is not another word for 'best.' Gemini has the maximum Elo. Agnes has the minimum price. Agnes is 274 Elo behind Gemini and costs one twentieth as much. Those are different offers. Putting both on a short list makes sense only when the short list spans sharply different budgets.
The middle of the curve is more useful for routine buying. grok saves $1.80 per minute versus Gemini with a large quality gap. Bach saves another $1.20 per minute versus grok while their integer Elo scores differ by only three and their intervals overlap. Hailuo cuts another $0.20. Each move trades some measured preference for a lower list price. A team can set an acceptable visual floor, then choose the cheapest point above it.
Several famous models fall off the frontier. MiniMax H3 ranks second but costs more than Gemini. HappyHorse-1.0 ranks third but costs more than Gemini. Kling 3.0 1080p Pro scores 1238 at $13.44 per minute, while Gemini scores higher at less than half the price. Veo 2 sits at 1114 and $30.00 per minute, five times Gemini's list price and 210 Elo lower. Brand familiarity does not rescue a dominated row.
Open, closed, US, and China are different competitive sets
Global rank hides four different procurement questions. A team choosing a hosted API compares all priced models. A team that must run weights in its own environment has a much smaller field. A buyer with regional processing or vendor-policy constraints may need an origin-specific shortlist. Recompute the frontier inside the allowed set instead of treating the global order as portable.

Closed models hold most of the board
Closed models account for 64 of 80 Current rows and 57 of 73 priced rows. Gemini leads that set. The closed-only frontier has seven SKUs because the open LTX-2.3 Fast drops out. For a buyer who wants a managed API and has no self-hosting requirement, the closed set offers more price points and most of the upper leaderboard.
Open weights have one clear quality leader
Open-weight models account for 16 of 80 rows, but only two reach the top 15: MiniMax H3 at rank two and LTX-2.5 Fast at rank 12. H3 leads LTX-2.5 Fast by 88 Elo at the same $7.80 benchmark price. The open-only Pareto frontier contains Krea Realtime at 965 and $1.50, LTX-2.3 Fast at 1120 and $2.40, and H3 at 1301 and $7.80.
The label still needs a license check. MiniMax publishes H3 weights under its Community License, and the local 2K regeneration module was not open at the research cutoff. LTX-2.5 also uses a community license with revenue conditions. 'Open weights' describes model access; it does not automatically mean an OSI-approved license or unrestricted commercial deployment.
China-origin labs occupy ten of the top 15 positions
Using the originating lab's documented headquarters or corporate domicile, China-origin labs account for 41 of 80 rows and 10 of the top 15. US-origin labs account for 8 of 80 and two of the top 15: Gemini Omni Flash and grok-imagine-video. LTX is the remaining documented Other row in the top 15; Bach-1.0 Preview and Runway Gen-4.5 remain unclassified because their origins were not established from retained primary sources.
The geography result is a leaderboard slice, not a market census. One lab can contribute several resolutions, modes, or generations. Kling alone places four SKUs in the top 15. The count says China-origin labs have broad representation in this picture benchmark; it does not say they hold two thirds of customers, revenue, or generated minutes.
Audio changes who is evaluated and how votes are framed
Artificial Analysis maintains separate With Audio and No Audio Elo pools. The With Audio snapshot contained 33 Current rows. The No Audio snapshot contained 80. Thirty model names appeared in both, three appeared only in With Audio, and 50 appeared only in No Audio. Pool membership changes enough that a ranking copied without its audio setting can describe a different contest.

Wan 3.0, MiniMax H3 Max post-trained by fal, and Agnes-Video-2.5 were With Audio only. Bach-1.0 Preview, Runway Gen-4.5, Veo 3, Ray 3, and Pika 2.5 were among the No Audio-only rows. H3 Max is a closed hosted SKU and should not be collapsed into the open MiniMax H3 row. Agnes-Video-2.5 is also distinct from Agnes-Video-V2.0.
Rank-order movement among shared names can be described, but Elo movement cannot. grok moves from rank 21 in With Audio to rank nine in No Audio. LTX-2.5 Fast moves from 23 to 12. HappyHorse-1.0 moves from eight to three. Those rank changes suggest that audio affects preference and competition, yet they do not measure an audio penalty because each Elo scale is fit independently.
We found no published percentage for how often business or social-video teams keep a model's native soundtrack. Some models generate dialogue, music, and effects by default; other APIs offer separate controls. Many production workflows mute, replace, or finish audio later. Treat every keep-rate estimate as unknown until a provider or independent study publishes usage data.
That separation is practical in ngram's AI video generation workflow: the visual model is one layer, while script, voiceover, captions, music, and editing remain separate production choices. A team can evaluate picture quality without pretending that the first generated soundtrack must ship.
The spec sheet still has 51 important unknowns
We audited 16 prominent SKUs against 12 fields drawn from official model pages, API documentation, license files, safety materials, and provenance claims. The fields cover text-to-video, image-to-video, video editing, native audio, stated duration, stated resolution, open weights, public API, lab origin, C2PA, invisible watermarking, and visible watermarking. Of 192 cells, 51 remained unknown.

Provenance produced most of the missing data. Eleven of 16 C2PA cells were unknown, 12 invisible-watermark cells were unknown, and 15 visible-watermark cells were unknown. An unknown does not prove that a model lacks the feature. It means we found no verified public claim at the model or hosting layer reviewed.
That distinction prevents several common mistakes. An arena row named '720p' does not establish the model's maximum resolution. A separate soundtrack product does not prove native audio in the named video SKU. A host's C2PA overlay does not prove that self-hosted model weights emit Content Credentials. A launch promise to release weights does not establish that every module has shipped.
C2PA and invisible watermarking also solve different problems. The C2PA Content Credentials standard provides cryptographically verifiable provenance metadata that can record origin and edits. Google describes SynthID as an invisible watermark embedded in generated media. Metadata can be removed; watermark signals can survive some transformations. A procurement review should ask which layer creates the signal, whether export or re-encoding preserves it, and which detector can verify it.
The highest unknown counts belonged to Seedance 2.0, Pika 2.5, SkyReels V4, LTX-2.5, and Grok Imagine in the exact SKUs audited. Sora's historical row had no unknown cells because OpenAI documented the old product extensively, but the Sora web and app experiences ended in April 2026 and its API was scheduled to end later in September. Documentation completeness is not the same thing as current availability.
A buying checklist that survives the next leaderboard update
Buying generative AI video models starts with constraints, then uses Elo and price inside the allowed set. Five questions cover most real decisions:
- Which pool matches the deliverable? Choose text-to-video or image-to-video first, then decide whether native audio belongs in the evaluation.
- What is the acceptable visual floor? Use confidence intervals and sample counts, not integer rank alone.
- What is the real accepted-clip cost? Multiply list generation cost by retries and add any upscale, audio, or editing charges.
- Which deployment constraints remove candidates? Check region, data handling, API access, self-hosting, license conditions, and latency.
- What evidence must travel with the output? Verify C2PA, watermarking, visible labels, and what survives the final export path.
Then run a small acceptance test on your own prompts. Track first-pass acceptance, retries per approved clip, hands-on correction time, and final cost. A model that costs $3.00 per minute but needs four generations per accepted clip can be more expensive than a $6.00 model that succeeds on the first try. The public leaderboard gives a prior; your workflow test gives the purchasing answer.
Five predictions this report can be wrong about in 2027
A dated report should say what would falsify it. We will revisit these five checks against the same Current, Global, text-to-video, No Audio view:
- Gemini keeps the top picture rank. The prediction fails if another Current row outranks Gemini Omni Flash, or if Gemini leaves the board.
- MiniMax H3 remains dominated by Gemini. The prediction fails if H3 is no longer both more expensive and lower in Elo, or if the two 95% confidence intervals overlap.
- Wan 3.0 stays absent from No Audio Current. The prediction fails when a Current No Audio row named Wan 3.0 appears.
- China-origin labs retain at least 10 of the top 15. The prediction fails if the count falls to nine or fewer under the same lab-origin definition.
- Open weights remain a small part of the top 15. The prediction fails when four or more top-15 rows carry the benchmark's open-weights flag.
These bets can break because of a model release, a price change, a new set of votes, or a taxonomy correction. That is the point. A forecast that cannot be checked is decoration; a forecast tied to a frozen snapshot can become evidence.
Methodology, limitations, and disclosure
We captured Artificial Analysis's Text-to-Video leaderboard on 1 September 2026 at 20:32:10 UTC, with filters set to Current, Global, and No Audio. The snapshot contained 80 rows, 73 with numeric public API prices and seven marked unavailable or unpriced. We retained the live rank order, model name, creator, Elo, 95% confidence interval, samples, release label, price, and open-weight flag.
We calculated the price-quality frontier after removing null prices. Model A is dominated when another priced model B has Elo greater than or equal to A and price less than or equal to A, with at least one strict inequality. Equal price and equal Elo do not create domination, which is why both Hailuo rows remain. Clip costs assume linear conversion from the benchmark's per-minute price.
We classified open versus closed from the leaderboard flag, then checked prominent open models against their published licenses. 'Open weights' is the report's term; it does not imply an OSI license. For origin, we used the headquarters or corporate domicile of the originating lab, supported by official company, legal, or investor pages. We kept unresolved labs in Other-unverified instead of inferring geography from names, language, or API hosts.
The audio comparison uses a separate With Audio snapshot from earlier on 1 September 2026. We normalized only a trailing 'Open Weights' label to compare membership. We did not subtract Elo, compare prices as an audio surcharge, or combine ranks across pools. The capability audit is another separate dataset: 16 selected SKUs by 12 fields, with primary documentation favored and silence recorded as unknown.
Limits: preference arenas measure aggregate voter choice, not every production requirement. Prices and ratings move. Multiple rows can represent the same lab or model family at different modes. Public documentation can lag a product. The capability sample is purposive, not exhaustive. Origin counts describe rows, not revenue or usage. This report is analysis, not legal advice.
ngram publishes this report and sells business video software. ngram is not a row on the Artificial Analysis model leaderboard and did not supply Elo or price data. Readers should discount our workflow interpretation accordingly. The raw leaderboard snapshot, derived tables, source ledger, and calculation notes were retained in the report workspace so every number can be reproduced.
Frequently asked questions
What is the best generative AI video model in 2026?
Gemini Omni Flash is the highest-rated picture-only text-to-video model in this 1 September 2026 snapshot, at 1324 Elo and $6.00 per generated minute. It also costs less than ranks two through five. The best choice for a specific team can change when self-hosting, region, audio, API availability, or provenance requirements remove Gemini from the allowed set.
What does Elo mean for AI video models?
Artificial Analysis derives video Elo from pairwise human preferences: voters see two clips made from the same prompt and choose one. A Bradley-Terry model aggregates those votes into a relative score. Elo measures preference inside one modality and audio setting. It is not a percentage, a pass rate, or a score that can be subtracted across separate leaderboards.
Which AI video models offer the best price-quality tradeoff?
Eight SKUs were non-dominated among 73 priced rows: Gemini Omni Flash, grok-imagine-video, Bach-1.0 Preview, Hailuo 02 Standard, Hailuo 2.3, LTX-2.3 Fast, Seedance 1.0 Mini, and Agnes-Video-V2.0. Gemini maximizes Elo; Agnes minimizes price. A buyer should choose the cheapest frontier point that clears the team's visual-quality floor.
Are open-weight AI video models competitive with closed models?
MiniMax H3 is competitive on picture quality: it ranks second globally at 1301 Elo. The next open model, LTX-2.5 Fast, is 88 points lower at the same benchmark price. Open weights make up 16 of 80 rows and two of the top 15. License and module availability still need separate checks before self-hosting.
Does native audio change AI video model rankings?
Audio changes both membership and preference context. The With Audio board had 33 Current rows; the No Audio board had 80, with 30 shared names. Wan 3.0 and H3 Max appeared only in With Audio, while Bach and Runway Gen-4.5 appeared only in No Audio. Because each pool has its own Elo fit, the scores cannot be subtracted.
How often should an AI video model report be updated?
Recheck the live board before any high-volume purchase and publish a dated refresh after a major launch, price change, or pool redesign. Artificial Analysis recomputes ratings hourly. A quarterly report can show direction, but a procurement decision should use a fresh export plus an internal acceptance test on the buyer's prompts.
Need to compare models inside a complete production workflow? Browse ngram's AI video model directory for business-focused model pages, then test the finalists against your own prompts.
You just read it. Now watch it.
ngram turns this post into a short explainer video: scenes, voiceover, and motion graphics included.






