- ElevenLabs launched v4 and v4 Turbo on September 28, 2026; v4 targets produced audio while Turbo targets realtime speech.
- The 100 ms inference and 150 ms speech-start claims use different measurements and exclude latency your application can add.
- More language coverage still requires voice, pronunciation and translation checks for each market.
- Our assumed 60-clip, three-language workload consumes 162,000 characters before rereads. Launch API pricing is temporary; budget with regular rates.
ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026. The voice-model release combines more expressive delivery, improved speaker continuity and broader language coverage. For teams making product videos, the useful question is how those changes affect a narration that has to survive edits and translation. (Official announcement)
A launch script rarely stays finished. Someone changes the feature name, the French version needs a shorter sentence, or the last scene gets a different call to action. Each change creates another chance for the narrator's tone to drift. We think that is the most interesting promise of this release: more control over a performance as the video changes.
The low-latency variant addresses a different job, starting speech during a live exchange. Below, we separate those workloads, inspect the launch claims and calculate what a small localization campaign could consume. The calculations are planning examples, not results from a production benchmark.
What Eleven v4 changes, and what still needs testing
ElevenLabs describes Eleven v4 as a new architecture for more expressive speech, with natural-language direction and inline audio tags. It also claims stronger speaker preservation across regenerated lines and dialogue. Those are supplier claims about model behavior; the announcement does not establish an acceptance rate for your scripts or voices. (Release details)
Consider a product walkthrough that opens with an enthusiastic launch announcement and ends with a calm explanation of setup. The writing alone may not communicate that shift. Direction attached to each passage can tell the model how to perform it. The editor still needs to listen for exaggerated excitement, misplaced pauses or a delivery that makes a routine instruction sound urgent.
A second issue is voice migration. A brand may have approved a synthetic narrator months ago. A more faithful clone can sound different from the version that stakeholders approved, even when both come from the same reference recording. The model's dedicated guide explicitly warns that v4 may sound substantially different from v3 and recommends comparing your own content. (Eleven v4 guide)
That makes migration a creative decision as well as a technical one. Preserve a small set of approved reference clips. Compare the new generation with those clips in the final music mix, then ask whether the voice still sounds like the same narrator. A clean solo sample can conceal a change that becomes distracting across a whole campaign.
Eleven v4 language coverage expands the evaluation surface
The current model overview lists 90+ languages for both v4 variants, compared with 70+ for v3, 32 for Flash v2.5 and 29 for Multilingual v2. These are coverage counts, not equal-quality guarantees across languages. (Model overview)

The plus signs matter. The chart uses the published lower-bound figures for v3 and v4 rather than pretending those are exact totals. Nor does a larger count tell us how a chosen voice handles a local accent, an acronym or a sentence that switches languages halfway through.
The v4 guide describes another behavior worth checking: when generated speech uses a different language from the reference voice, the model aims for a native accent in the target language. A narrator's original accent is preserved when the languages match. If your brand deliberately uses a recognizable accent across markets, migration can change that identity. (Accent behavior)
For video localization, listen to complete scenes rather than isolated greetings. Put a product name next to a date, a price and a sentence with the intended tone. Have a fluent reviewer judge pronunciation and meaning alongside delivery. A smooth reading of the wrong translation is still a failed asset.
ngram already supports changing a video's language after generation and re-voicing its narration through chat. Its translation workflow illustrates why model coverage belongs inside a reviewable video project. This release does not establish that ngram has integrated Eleven v4 or that the model's full language list is available in that workflow.
A bigger request limit changes batching, not finished duration
ElevenLabs currently documents a 10,000-character ceiling for v4, twice v3's 5,000. Multilingual v2 is also listed at 10,000, while Flash v2.5 allows 40,000. These are input ceilings for a request, not promised runtimes or recommended scene lengths. (Character limits)

A higher ceiling can reduce how often a long script must be split. It does not make a single giant request the best editing unit. If the closing sentence changes, regenerating an entire presentation can waste audio and force another complete review. Scene-sized generations retain a useful boundary for revisions.
The character count also says little about the final speaking time. Punctuation, language and pauses change duration. In a product walkthrough, the narration must land alongside the cursor movement or the appearance of an interface. If a translation runs longer, adjust the script or scene timing before deciding the voice is unusable.
Request stitching and improved continuity are useful promises here, but they need to be tested at the join. Listen across the transition between two separately generated scenes. A narrator who sounds consistent within a request can still change loudness or energy at the boundary.
The 100 ms and 150 ms figures measure different things
The model overview reports roughly 100 ms median inference latency for Eleven v4 Turbo. The announcement separately reports roughly 150 ms median time to first audible speech over WebSocket streaming, with measured network latency removed. Neither number describes a listener's complete wait in your application. (Model latency; announcement methodology)

In a conversational system, a user may first wait for speech recognition, turn detection and a language-model response. Then there is speech generation, transport and playback buffering. A quicker voice model changes one part of that sequence. Measure the whole exchange on the devices and connections your audience uses, including interruptions and longer waits.
For an offline video, the relevant outcome is usually an approved voiceover that fits the scenes. Time to first audio can improve the preview experience, but the model still has to produce the rest of the narration. Visual generation, captions, mixing, review and export remain separate work. A 150 ms speech-start claim cannot support a claim that an entire video finishes in 150 ms.
Choose Turbo when the first spoken response is the constraint. For a prerecorded training clip or product launch video, compare the final performance and revision behavior before prioritizing a streaming latency headline. The two variants invite different evaluation criteria even when they share a model family.
Our planning scenario: 60 clips become 180 voiceovers
Here is a deliberately small campaign: 60 short clips, each delivered in three total languages, with 900 characters of narration per version. Assume every language version has the same character count. That assumption keeps the arithmetic reproducible; real translations will differ.
The first pass requires 180 voiceovers and 162,000 generated characters: 60 × 3 × 900. A full second pass doubles that to 324,000 characters. If only a quarter of versions need a complete reread, the total is 202,500. These figures describe an invented workload, not customer usage or a measured v4 retry rate.

This calculation makes one operational point: the number of approved deliverables stays at 180 while generation volume changes with revisions. Track both. A campaign dashboard that counts exports alone can conceal the cost of a difficult voice or a script that keeps changing after generation.
For a script change affecting one scene, retaining the approved narration elsewhere can reduce the amount you regenerate. ngram's AI voiceover workflow supports re-voicing affected scenes after script edits while retaining visuals. That is a current workflow capability, independent of any claim about adopting the new model.
Methodology: original planning analysis prepared September 29, 2026 from the verified launch brief and official documentation. There were no audio trials, surveyed customers or observed production jobs. Generated characters equal clips × total languages × characters per version × (1 + assumed full-version reread share). The scenario excludes partial-scene retries, failed requests and translation-length differences.
Launch pricing is cheap enough to make review the next question
As checked September 29, ElevenLabs' API rate card displays launch rates of $0.022 per 1,000 characters for v4 and $0.011 for v4 Turbo, with a promotion advertised until October 12. The displayed regular rates are $0.08 and $0.04 respectively. Treat these as dated API usage rates, not the price of a finished video or a universal subscription quote. (API pricing)
Using the campaign assumptions above, v4's first pass is $3.56 at the promotional rate versus $12.96 at the displayed regular rate, rounded to cents. A complete second pass is $7.13 versus $25.92. Turbo's first pass would be $1.78 or $6.48 using its corresponding rates. These are calculated usage equivalents; plan minimums, taxes and other production costs are excluded.

Budget for Eleven v4 with the regular rate and treat a temporary discount as a temporary benefit. The calculation is intentionally per character, because converting directly to minutes assumes a speaking pace and language mix. Audio with long pauses and audio with rapid narration can have similar character counts and quite different durations.
At this workload, a fluent reviewer finding a mistranslated feature name may be more valuable than another small reduction in raw generation spend. That is an inference from the scenario, not a measured return on investment. The earlier analysis of falling AI audio prices offers context for the supplier economics; v4 adds a separate question about directing and maintaining a performance.
Preference tests are evidence of taste, not campaign ROI
ElevenLabs reports about 75% listener preference in blind head-to-head tests against selected competing models. Its footnote says graders heard the same line from each model, judged expressiveness and naturalness, and counted ties as half. The announcement does not disclose a sample size or uncertainty interval for that figure. (Preference-test methodology)
That gives a reason to audition the model. It does not tell us whether a technical explainer will need fewer retries, whether a legal disclaimer will sound appropriate or whether a localized ad will convert better. Those outcomes require their own evidence. Our article makes no claim of independent audio testing or measured business improvement.
Use two separate checks during an audition. First, compare candidate performances without showing reviewers which model made each one. Then inspect operational failures with the model names visible: pronunciation errors, missing words, excessive pauses and continuity across replacements. A voice can win a preference vote and still create extra editing work on a particular script.
A migration checklist for the next production job
Keep the pilot small enough to hear every output. Start with one existing video whose script and pronunciation have already been approved. Preserve its original voiceover, then generate a candidate version and a replacement for one sentence. Review the replacement in context before translating the whole project.

Write down the performance instructions that worked. Keep the exact script, voice identifier and generation settings beside the approved audio. If the supplier changes model behavior later, those records give you a repeatable regression check. The v4 guide says the model is under active development and recommends periodic re-testing. (Ongoing model changes)

Treat a cloned voice as an authorized production asset. Confirm the speaker's permission and the scope of the intended use before recording or uploading samples. ElevenLabs' voice documentation describes verification for Professional Voice Clones; that mechanism does not replace your own record of what the speaker approved. (Voice documentation)
When a presenter is visible, review lip movement along with the new narration. Speech synthesis and video translation with lip sync solve related but distinct parts of the finished video. Also inspect captions and any on-screen numbers. A correct audio track can still disagree with the visual explanation.
The acceptance decision for Eleven v4 should come from the video you plan to publish. Listen on a phone, at normal playback speed, with the music included. Approve that output, retain the previous version and only then expand to the rest of the campaign.
Frequently asked questions
When did Eleven v4 launch?
ElevenLabs announced the model family on September 28, 2026. The official release introduced both the expressive v4 model and the low-latency Turbo variant. (Announcement)
Is Eleven v4 Turbo the faster choice for every video?
Turbo is designed for realtime speech. Its speech-start latency does not measure complete narration generation, scene production or export. For a prerecorded video, compare the approved final performance and the cost of revisions.
Does 90+ languages mean every voice is equally good everywhere?
No. Coverage describes availability, while fluency, pronunciation and suitability still need review. Test the exact voice and script in each target market rather than extrapolating from an English demo. (Language documentation)
Will an existing cloned voice sound unchanged after migration?
The v4 guide warns that clones can sound different from their v3 versions because the new model captures the source voice more faithfully. Keep approved reference audio and compare the replacement before updating a recurring narrator. (Migration guidance)
Are the launch API rates permanent?
The rate card advertises a promotion until October 12, as checked September 29, 2026. Use the displayed regular rate when planning work after the promotion, and recheck your account's actual billing terms. (Current API rate card)
What should a video team test first?
Use an approved scene, a sentence replacement and one target-language version. Judge pronunciation, speaker identity and timing in the final mix. Expand only after the replacement fits the existing video without creating more editing work.
You just read it. Now watch it.
ngram turns this post into a short explainer video: scenes, voiceover, and motion graphics included.






