Back to Industry news
Industry news

Eleven v4: What Changes for Video Voiceovers in 2026

ElevenLabs launched Eleven v4 and Turbo on September 28. We examine language coverage, voice continuity, latency claims and the cost of revisions.

Eleven v4: What Changes for Video Voiceovers in 2026
12 min read•Updated at September 29, 2026
Written and edited by
Rishikesh Ranjan
Rishikesh Ranjan
all thing growth @ ngram.com

ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026. The voice-model release combines more expressive delivery, improved speaker continuity and broader language coverage. For teams making product videos, the useful question is how those changes affect a narration that has to survive edits and translation. (Official announcement)

A launch script rarely stays finished. Someone changes the feature name, the French version needs a shorter sentence, or the last scene gets a different call to action. Each change creates another chance for the narrator's tone to drift. We think that is the most interesting promise of this release: more control over a performance as the video changes.

The low-latency variant addresses a different job, starting speech during a live exchange. Below, we separate those workloads, inspect the launch claims and calculate what a small localization campaign could consume. The calculations are planning examples, not results from a production benchmark.

What Eleven v4 changes, and what still needs testing

ElevenLabs describes Eleven v4 as a new architecture for more expressive speech, with natural-language direction and inline audio tags. It also claims stronger speaker preservation across regenerated lines and dialogue. Those are supplier claims about model behavior; the announcement does not establish an acceptance rate for your scripts or voices. (Release details)

Consider a product walkthrough that opens with an enthusiastic launch announcement and ends with a calm explanation of setup. The writing alone may not communicate that shift. Direction attached to each passage can tell the model how to perform it. The editor still needs to listen for exaggerated excitement, misplaced pauses or a delivery that makes a routine instruction sound urgent.

A second issue is voice migration. A brand may have approved a synthetic narrator months ago. A more faithful clone can sound different from the version that stakeholders approved, even when both come from the same reference recording. The model's dedicated guide explicitly warns that v4 may sound substantially different from v3 and recommends comparing your own content. (Eleven v4 guide)

That makes migration a creative decision as well as a technical one. Preserve a small set of approved reference clips. Compare the new generation with those clips in the final music mix, then ask whether the voice still sounds like the same narrator. A clean solo sample can conceal a change that becomes distracting across a whole campaign.

Eleven v4 language coverage expands the evaluation surface

The current model overview lists 90+ languages for both v4 variants, compared with 70+ for v3, 32 for Flash v2.5 and 29 for Multilingual v2. These are coverage counts, not equal-quality guarantees across languages. (Model overview)

Eleven v4 documents 90+ languages, compared with 70+ for v3, 32 for Flash v2.5 and 29 for Multilingual v2
Published language counts as checked September 29, 2026. v3 and v4 bars plot the lower-bound figures 70 and 90; the supplier documents them as 70+ and 90+. Coverage does not establish equal quality in every language. Official source. View full-size image
Published language counts as checked September 29, 2026. v3 and v4 bars plot the lower-bound figures 70 and 90; the supplier documents them as 70+ and 90+. Coverage does not establish equal quality in every language.
ModelLanguages (v3/v4 lower bounds)
Multilingual v229
Flash v2.532
Eleven v370+
Eleven v490+

The plus signs matter. The chart uses the published lower-bound figures for v3 and v4 rather than pretending those are exact totals. Nor does a larger count tell us how a chosen voice handles a local accent, an acronym or a sentence that switches languages halfway through.

The v4 guide describes another behavior worth checking: when generated speech uses a different language from the reference voice, the model aims for a native accent in the target language. A narrator's original accent is preserved when the languages match. If your brand deliberately uses a recognizable accent across markets, migration can change that identity. (Accent behavior)

For video localization, listen to complete scenes rather than isolated greetings. Put a product name next to a date, a price and a sentence with the intended tone. Have a fluent reviewer judge pronunciation and meaning alongside delivery. A smooth reading of the wrong translation is still a failed asset.

ngram already supports changing a video's language after generation and re-voicing its narration through chat. Its translation workflow illustrates why model coverage belongs inside a reviewable video project. This release does not establish that ngram has integrated Eleven v4 or that the model's full language list is available in that workflow.

A bigger request limit changes batching, not finished duration

ElevenLabs currently documents a 10,000-character ceiling for v4, twice v3's 5,000. Multilingual v2 is also listed at 10,000, while Flash v2.5 allows 40,000. These are input ceilings for a request, not promised runtimes or recommended scene lengths. (Character limits)

Per-request character ceilings: Eleven v3 5,000; Eleven v4 10,000; Multilingual v2 10,000; Flash v2.5 40,000
Documented input ceilings, checked September 29, 2026. A higher character limit permits a larger request; it does not guarantee a particular speaking time or fewer revisions. Official source. View full-size image
Documented input ceilings, checked September 29, 2026. A higher character limit permits a larger request; it does not guarantee a particular speaking time or fewer revisions.
ModelCharacters per request
Eleven v35,000
Eleven v410,000
Multilingual v210,000
Flash v2.540,000

A higher ceiling can reduce how often a long script must be split. It does not make a single giant request the best editing unit. If the closing sentence changes, regenerating an entire presentation can waste audio and force another complete review. Scene-sized generations retain a useful boundary for revisions.

The character count also says little about the final speaking time. Punctuation, language and pauses change duration. In a product walkthrough, the narration must land alongside the cursor movement or the appearance of an interface. If a translation runs longer, adjust the script or scene timing before deciding the voice is unusable.

Request stitching and improved continuity are useful promises here, but they need to be tested at the join. Listen across the transition between two separately generated scenes. A narrator who sounds consistent within a request can still change loudness or energy at the boundary.

The 100 ms and 150 ms figures measure different things

The model overview reports roughly 100 ms median inference latency for Eleven v4 Turbo. The announcement separately reports roughly 150 ms median time to first audible speech over WebSocket streaming, with measured network latency removed. Neither number describes a listener's complete wait in your application. (Model latency; announcement methodology)

Eleven v4 Turbo: about 100 ms median inference and about 150 ms median time to first audible speech; different metrics, not end-to-end application latency
Supplier-reported medians. The approximately 150 ms speech-start figure uses WebSocket streaming with network latency measured and removed. The model overview excludes application and network latency for inference. Neither measures a complete voice-agent exchange. Official source. View full-size image
Supplier-reported medians. The approximately 150 ms speech-start figure uses WebSocket streaming with network latency measured and removed. The model overview excludes application and network latency for inference. Neither measures a complete voice-agent exchange.
MetricSupplier-reported valueMeasurement caveat
Median inference~100 msApplication and network excluded
Median time to first audible speech~150 msWebSocket; measured network removed

In a conversational system, a user may first wait for speech recognition, turn detection and a language-model response. Then there is speech generation, transport and playback buffering. A quicker voice model changes one part of that sequence. Measure the whole exchange on the devices and connections your audience uses, including interruptions and longer waits.

For an offline video, the relevant outcome is usually an approved voiceover that fits the scenes. Time to first audio can improve the preview experience, but the model still has to produce the rest of the narration. Visual generation, captions, mixing, review and export remain separate work. A 150 ms speech-start claim cannot support a claim that an entire video finishes in 150 ms.

Choose Turbo when the first spoken response is the constraint. For a prerecorded training clip or product launch video, compare the final performance and revision behavior before prioritizing a streaming latency headline. The two variants invite different evaluation criteria even when they share a model family.

Our planning scenario: 60 clips become 180 voiceovers

Here is a deliberately small campaign: 60 short clips, each delivered in three total languages, with 900 characters of narration per version. Assume every language version has the same character count. That assumption keeps the arithmetic reproducible; real translations will differ.

The first pass requires 180 voiceovers and 162,000 generated characters: 60 × 3 × 900. A full second pass doubles that to 324,000 characters. If only a quarter of versions need a complete reread, the total is 202,500. These figures describe an invented workload, not customer usage or a measured v4 retry rate.

Assumed localization workload: 162,000 characters with no rereads, 202,500 with 25% rereads, 243,000 with 50%, and 324,000 with 100%
Original scenario, not an observed benchmark. Assume 60 clips, three total languages and 900 characters per version. Totals include one full-version reread for the stated share of versions, using 162,000 × (1 + reread share). View full-size image
Original scenario, not an observed benchmark. Assume 60 clips, three total languages and 900 characters per version. Totals include one full-version reread for the stated share of versions, using 162,000 × (1 + reread share).
Full-version reread assumptionTotal generated characters
No rereads162,000
25% reread202,500
50% reread243,000
100% reread324,000

This calculation makes one operational point: the number of approved deliverables stays at 180 while generation volume changes with revisions. Track both. A campaign dashboard that counts exports alone can conceal the cost of a difficult voice or a script that keeps changing after generation.

For a script change affecting one scene, retaining the approved narration elsewhere can reduce the amount you regenerate. ngram's AI voiceover workflow supports re-voicing affected scenes after script edits while retaining visuals. That is a current workflow capability, independent of any claim about adopting the new model.

Methodology: original planning analysis prepared September 29, 2026 from the verified launch brief and official documentation. There were no audio trials, surveyed customers or observed production jobs. Generated characters equal clips × total languages × characters per version × (1 + assumed full-version reread share). The scenario excludes partial-scene retries, failed requests and translation-length differences.

Launch pricing is cheap enough to make review the next question

As checked September 29, ElevenLabs' API rate card displays launch rates of $0.022 per 1,000 characters for v4 and $0.011 for v4 Turbo, with a promotion advertised until October 12. The displayed regular rates are $0.08 and $0.04 respectively. Treat these as dated API usage rates, not the price of a finished video or a universal subscription quote. (API pricing)

Using the campaign assumptions above, v4's first pass is $3.56 at the promotional rate versus $12.96 at the displayed regular rate, rounded to cents. A complete second pass is $7.13 versus $25.92. Turbo's first pass would be $1.78 or $6.48 using its corresponding rates. These are calculated usage equivalents; plan minimums, taxes and other production costs are excluded.

Calculated v4 usage cost ranges from $3.56 to $7.13 at launch rates, versus $12.96 to $25.92 at displayed regular rates as rereads increase
Calculated v4 usage equivalents for the stated workload. Promotional rate: $0.022 per 1,000 characters, advertised until October 12; displayed regular rate: $0.08. Rounded to cents. Not a full production budget or account bill; minimums, taxes and other costs excluded. Official source. View full-size image
Calculated v4 usage equivalents for the stated workload. Promotional rate: $0.022 per 1,000 characters, advertised until October 12; displayed regular rate: $0.08. Rounded to cents. Not a full production budget or account bill; minimums, taxes and other costs excluded.
Full-version reread assumptionPromotional USDDisplayed regular USD
No rereads$3.56$12.96
25% reread$4.46$16.20
50% reread$5.35$19.44
100% reread$7.13$25.92

Budget for Eleven v4 with the regular rate and treat a temporary discount as a temporary benefit. The calculation is intentionally per character, because converting directly to minutes assumes a speaking pace and language mix. Audio with long pauses and audio with rapid narration can have similar character counts and quite different durations.

At this workload, a fluent reviewer finding a mistranslated feature name may be more valuable than another small reduction in raw generation spend. That is an inference from the scenario, not a measured return on investment. The earlier analysis of falling AI audio prices offers context for the supplier economics; v4 adds a separate question about directing and maintaining a performance.

Preference tests are evidence of taste, not campaign ROI

ElevenLabs reports about 75% listener preference in blind head-to-head tests against selected competing models. Its footnote says graders heard the same line from each model, judged expressiveness and naturalness, and counted ties as half. The announcement does not disclose a sample size or uncertainty interval for that figure. (Preference-test methodology)

That gives a reason to audition the model. It does not tell us whether a technical explainer will need fewer retries, whether a legal disclaimer will sound appropriate or whether a localized ad will convert better. Those outcomes require their own evidence. Our article makes no claim of independent audio testing or measured business improvement.

Use two separate checks during an audition. First, compare candidate performances without showing reviewers which model made each one. Then inspect operational failures with the model names visible: pronunciation errors, missing words, excessive pauses and continuity across replacements. A voice can win a preference vote and still create extra editing work on a particular script.

A migration checklist for the next production job

Keep the pilot small enough to hear every output. Start with one existing video whose script and pronunciation have already been approved. Preserve its original voiceover, then generate a candidate version and a replacement for one sentence. Review the replacement in context before translating the whole project.

Suggested audition workflow: preserve reference audio, generate the same scene, replace one line, review one translation, approve the finished mix
Suggested editorial evaluation process, not a measured result. Keep the previous approved audio available throughout the audition. View full-size image

Write down the performance instructions that worked. Keep the exact script, voice identifier and generation settings beside the approved audio. If the supplier changes model behavior later, those records give you a repeatable regression check. The v4 guide says the model is under active development and recommends periodic re-testing. (Ongoing model changes)

Voiceover acceptance checks: pronunciation, speaker continuity, target-language meaning, scene timing, authorized voice use and the final mix
Suggested acceptance checklist. Compare the complete output with the approved source material before publishing. View full-size image

Treat a cloned voice as an authorized production asset. Confirm the speaker's permission and the scope of the intended use before recording or uploading samples. ElevenLabs' voice documentation describes verification for Professional Voice Clones; that mechanism does not replace your own record of what the speaker approved. (Voice documentation)

When a presenter is visible, review lip movement along with the new narration. Speech synthesis and video translation with lip sync solve related but distinct parts of the finished video. Also inspect captions and any on-screen numbers. A correct audio track can still disagree with the visual explanation.

The acceptance decision for Eleven v4 should come from the video you plan to publish. Listen on a phone, at normal playback speed, with the music included. Approve that output, retain the previous version and only then expand to the rest of the campaign.

Frequently asked questions

When did Eleven v4 launch?

ElevenLabs announced the model family on September 28, 2026. The official release introduced both the expressive v4 model and the low-latency Turbo variant. (Announcement)

Is Eleven v4 Turbo the faster choice for every video?

Turbo is designed for realtime speech. Its speech-start latency does not measure complete narration generation, scene production or export. For a prerecorded video, compare the approved final performance and the cost of revisions.

Does 90+ languages mean every voice is equally good everywhere?

No. Coverage describes availability, while fluency, pronunciation and suitability still need review. Test the exact voice and script in each target market rather than extrapolating from an English demo. (Language documentation)

Will an existing cloned voice sound unchanged after migration?

The v4 guide warns that clones can sound different from their v3 versions because the new model captures the source voice more faithfully. Keep approved reference audio and compare the replacement before updating a recurring narrator. (Migration guidance)

Are the launch API rates permanent?

The rate card advertises a promotion until October 12, as checked September 29, 2026. Use the displayed regular rate when planning work after the promotion, and recheck your account's actual billing terms. (Current API rate card)

What should a video team test first?

Use an approved scene, a sentence replacement and one target-language version. Judge pronunciation, speaker identity and timing in the final mix. Expand only after the replacement fits the existing video without creating more editing work.

Related articles

The AI Video Disclosure Era Starts Today: NY Law, EU AI Act, and What $9.1B in Ad Spend Must Change
Industry news12 min read

The AI Video Disclosure Era Starts Today: NY Law, EU AI Act, and What $9.1B in Ad Spend Must Change

New York's Synthetic Performer Disclosure Law is live as of June 9, 2026, and EU AI Act Article 50 enforcement arrives August 2. Here's what both laws actually require, who is exposed, and a practical compliance checklist for the next 54 days.

Industry NewsAI Video
Rishikesh Ranjan
Rishikesh Ranjan
Growth Lead
Jun 9, 2026
50+ AI Video Statistics for 2026: The Data Behind Video's Biggest Shift
Industry news20 min read

50+ AI Video Statistics for 2026: The Data Behind Video's Biggest Shift

The most comprehensive collection of AI video statistics for 2026 - covering market size, adoption rates, production cost shifts, viewer behavior, and GTM impact. Every data point sourced and cross-referenced.

ngramAI Video
Anish Muppalaneni
Anish Muppalaneni
Co-founder & CEO
Aug 26, 2026
Anatomy of a Modern AI Agent: Muse vs Grok Bot in 2026
Industry news21 min read

Anatomy of a Modern AI Agent: Muse vs Grok Bot in 2026

A technical teardown of modern AI agent architecture through Meta's Muse and Grok Bot: agent loops, persistent computers, memory, credentials, egress, approvals, and the public evidence gaps that still matter.

Industry NewsAI Agents
Devadutta Ghat
Devadutta Ghat
Co-founder & CTO
Sep 21, 2026
Avataar's Varya and the Collapsing Cost of AI Video Generation
Industry news11 min read

Avataar's Varya and the Collapsing Cost of AI Video Generation

Avataar launched Varya, an India-built video model distilled from Wan 2.2 that generates video at about $0.005 per second. Here is what the launch says about collapsing AI video generation costs.

Industry NewsAI Video
Rishikesh Ranjan
Rishikesh Ranjan
Growth Lead
Jun 12, 2026
Black Forest Labs' FLUX 3: One AI Model for Video, Audio, and Robots
Industry news11 min read

Black Forest Labs' FLUX 3: One AI Model for Video, Audio, and Robots

Black Forest Labs launched FLUX 3, its first multimodal frontier model unifying video, audio, and robotic action. Here's why the first production customer is a car factory, not a marketing team.

Industry NewsAI Video
Rishikesh Ranjan
Rishikesh Ranjan
Growth Lead
Jul 27, 2026
California AI Advertising Law: What SB 1050 Changes in 2027
Industry news12 min read

California AI Advertising Law: What SB 1050 Changes in 2027

California SB 1050 makes synthetic-performer disclosure a production requirement for audio and video ads, with a January 1, 2027 effective date and a court-order removal process for advertising media.

Industry NewsAI Video
Rishikesh Ranjan
Rishikesh Ranjan
Growth Lead
Sep 22, 2026

Ready to create your first video?

Join thousands of product teams using AI to create professional videos in minutes.