ElevenLabs Eleven v4: Expressive TTS Architecture and Voice-Agent Latency

ElevenLabs Eleven v4: Expressive TTS Architecture and Voice-Agent Latency

Eleven v4: Expressive TTS Architecture and Voice-Agent Latency

Most voice agents do not feel slow because the text-to-speech engine is slow. They feel slow because three or four stages, each individually respectable, stack up into a pause that a human listener reads as hesitation. On September 28, 2026, ElevenLabs shipped Eleven v4 and a lower-latency sibling, Eleven v4 Turbo, and the launch is a useful excuse to take that budget apart. The headline claims are 90-plus languages, instant voice clones from about ten seconds of audio, inline tags that direct emotion and sound, and a Turbo variant with roughly 100 ms median inference latency.

Those numbers matter now because voice is moving from demo to deployment in support desks, healthcare intake and telephony. The buyer question has shifted from “does it sound human?” to “does it sound human at 3 a.m. on a noisy phone line, in Portuguese, without a two-second gap?” This post explains what has been disclosed about the model, how modern expressive text-to-speech works, and how to budget end-to-end latency with real arithmetic.

What this covers: the v4 lineage and what changed, the reference architecture of an expressive TTS system and a voice agent around it, the sourced benchmark and pricing facts (with vendor claims separated from independent ones), a worked latency budget, failure modes, a decision matrix against peers, and an FAQ.

Context and Background

Text-to-speech (TTS) spent two decades as a concatenative and then parametric technology: stitch recorded units together, or predict acoustic features with an HMM, and hand them to a vocoder. Neural TTS changed the quality ceiling around 2016 with autoregressive waveform models, and by the early 2020s the dominant recipe was a two-stage pipeline: a text-to-acoustic model that predicts a mel spectrogram, and a neural vocoder that converts it to audio. Since then, the field has split. One branch keeps the acoustic-model-plus-vocoder split with diffusion or flow matching. The other treats audio as a discrete token sequence, uses a neural codec to compress the waveform into tokens, and lets a large language-model-style decoder predict them.

ElevenLabs has iterated through that landscape quickly. Multilingual v2 became the workhorse for narration. Flash v2.5 traded expressiveness for speed, and the current documentation lists it at about 75 ms latency with 32 languages. Eleven v3 pushed expressiveness and audio tags, at the cost of latency that suited produced content more than live calls, and its conversational variant is listed at about 280 ms. Version 4 is the attempt to stop making customers pick between the two, which is why it ships as a pair: one quality-first model and one real-time model. For the wider model landscape around it, see our explainer on multimodal AI architecture and how audio fuses with vision and language.

The competitive field is crowded. ElevenLabs’ own comparisons name Cartesia Sonic 3.6, Inworld TTS-2 and Google’s Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS, and press coverage adds Deepgram, Fish Audio, WellSaid Labs and OpenAI to the list of rivals. If you follow the Gemini family, our piece on Gemini 3.8 Flash architecture, pricing and benchmarks covers the model whose TTS variants appear in that comparison. The authoritative source for what ElevenLabs itself says is the Eleven v4 launch post, and its model reference lives in the ElevenLabs models documentation.

A note on method. Everything below tagged “documented” comes from ElevenLabs’ own pages as of September 30, 2026. Anything marked “vendor claim” is measured by ElevenLabs, and “independent” means a third party such as Artificial Analysis. Where nothing is published, for example the parameter count or training data, I say so rather than guess. ElevenLabs has not, in the material I could verify, disclosed the v4 architecture beyond describing it as new and context-aware.

Eleven v4 and Turbo: What Was Actually Announced

Eleven v4 is ElevenLabs’ quality-first text-to-speech model, documented under the model ID eleven_v4, with 90-plus languages and a 10,000-character limit per generation (roughly ten minutes of audio per the docs). Eleven v4 Turbo, eleven_v4_turbo, is the real-time sibling with about 100 ms median inference latency and about 150 ms median time to first speech.

Both launched on September 28, 2026 and are available in ElevenAgents (the voice-agent platform), ElevenCreative (the studio tools) and ElevenAPI. Coverage from TechCrunch frames the language jump as going from 70 languages in v3 to about 90, with named quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

What is new versus v3

The shift is less about a single big feature than about four connected changes. First, context: the model reads the script with scene-level awareness, so an angry line after a calm one is rendered as an escalation rather than as an isolated sentence. Second, control: inline audio tags such as [laughs], [whispers] or [door slams] steer delivery and ambience, and ElevenLabs says SSML break tags are disabled, so pauses are shaped through natural-language tags instead. Third, identity: instant voice clones need about ten seconds of audio, and Professional Voice Clones return after their absence in v3. Fourth, stability: a speaker-stability feature is meant to keep a voice from drifting when you regenerate a line repeatedly, and request stitching preserves pacing across long-form output.

Two details deserve a caveat. The launch material does not explicitly state whether Turbo supports the full audio-tag vocabulary, and one summary of it says bidirectional streaming is a Turbo capability. If you plan to rely on either, verify in your own account before committing a design.

Numbers you can cite, and their provenance

Claim Value Provenance
Release date September 28, 2026 Documented (vendor blog)
Languages 90+ (v3: 70+) Documented
Turbo median inference latency about 100 ms Vendor claim, excludes app and network latency
Turbo median time to first speech about 150 ms Vendor claim, WebSocket streaming, network removed
Character limit, v4 10,000 per generation Documented
Instant Voice Clone input about 10 seconds Documented
Ranking first on Artificial Analysis Provider Voice Arena, reported at 1319 Elo on Sept 28, 2026 Independent, as reported by a third-party write-up
Listener preference about 75 percent in blind tests (range 65 to 81 percent by matchup) Vendor claim

The single most important footnote is the one attached to latency. The 100 ms and 150 ms figures are for the TTS stage alone, and the sources I checked say neither includes speech recognition or language-model time. Keep that in mind for the budget later in this post.

How Expressive TTS Like Eleven v4 Works: Reference Architecture

An expressive TTS model turns text plus control cues into audio in four steps: it analyses the text and tags for context, plans prosody and speaker identity, generates a compact acoustic representation, and decodes that into a waveform. ElevenLabs has not published the v4 internals, so what follows is the general architecture that current expressive systems share, not a description of proprietary weights.

Eleven v4 expressive TTS reference architecture from tagged text to waveform

Figure 1: Generic reference architecture for an expressive TTS system such as Eleven v4. The internals of v4 are not published; this shows the common structure of the category.

Figure 1 reads left to right. Tagged text and a speaker reference enter a context encoder. A prosody and style planner combines that with the speaker embedding. A generative acoustic model produces codec tokens or latent frames, and a decoder turns them into audio that streams back in chunks.

Stage 1: reading the script, not the sentence

Older TTS engines processed one sentence at a time, which is why they sounded fine in isolation and flat in dialogue. Expressive models instead condition on a wider window: preceding lines, speaker labels, and the tags. That is what “scene-level context” means in practice. If line five is a reply to an insult in line four, the model has the information needed to lower its pitch and tighten its rhythm without anyone writing prosody markup by hand.

This also explains the shift away from SSML. Markup languages such as SSML describe surface features like pitch, rate and break duration. Natural-language tags describe intent, such as [said angrily], and leave the mapping to acoustics to the model. The trade is control for expressiveness: you lose the ability to say “exactly 250 ms here” and gain the ability to say “sound like you are hiding bad news”.

Stage 2: speaker identity as a conditioning vector

A voice is not a stored recording; it is a point in a learned space. Given a short reference clip, an encoder produces a speaker embedding, and the generator is conditioned on that embedding at every step. Zero-shot or “instant” cloning is exactly this: no fine-tuning, just a forward pass over roughly ten seconds of audio to obtain the embedding. Professional cloning, by contrast, typically involves more audio and adaptation of the model to the speaker, which is why it holds up better on unusual timbres and long-form work.

The ten-second figure is a floor, not a promise. A clean, expressive ten seconds in a quiet room yields a better clone than thirty seconds of a compressed phone recording. Clone quality is dominated by recording quality and by how representative the sample is of the delivery you want, because the embedding captures the emotion in the sample along with the identity.

Stage 3: generating codec tokens or latents

The generator is where architectures diverge. In the autoregressive camp, a transformer decoder predicts discrete audio-codec tokens one step at a time, exactly as a language model predicts words. Its strength is expressiveness and long-range coherence, since the same machinery that keeps a paragraph consistent keeps a voice consistent. Its weakness is that generation is sequential, so time to first audio depends on how many tokens must exist before the first chunk can be decoded.

In the non-autoregressive camp, a diffusion or flow-matching model refines a whole block of acoustic latents in a small number of steps. It parallelises well and can be distilled to very few steps, which suits low-latency serving, but streaming requires care because the block must be at least partly complete before decoding. Real production systems mix ideas: chunked generation with look-ahead, speculative decoding, or distillation of a slow expressive teacher into a fast student. The existence of a quality model and a Turbo model in the same release is consistent with that last pattern, though ElevenLabs has not said how Turbo is derived.

Stage 4: codec decoding and streaming

The final stage converts tokens or latents into a waveform, typically at 16 to 48 kHz depending on the output format. For telephony, ElevenLabs lists µ-law among the output formats, alongside MP3, WAV and PCM, which matters because phone lines are 8 kHz and the fidelity you pay for in the model is largely discarded at the carrier. A practical consequence: on a phone channel, expressiveness cues carried by prosody survive better than timbre subtleties carried by high frequencies.

Why a Turbo model is not just a smaller model

It is tempting to assume Turbo is a pruned copy. Latency in TTS is affected by at least five things, and parameter count is only one of them: how much text must be seen before the first audio chunk, how many generation steps run per chunk, the codec frame rate, the decoding hardware and batching, and the protocol overhead. A model can be large and still low-latency if it starts speaking after a few words and generates faster than real time. Conversely, a small model that waits for a full sentence will lose to it.

That is why the vendor quotes two numbers. Inference latency (about 100 ms median) is the compute time. Time to first speech (about 150 ms median over WebSocket streaming, network removed) includes the first-chunk logic. The 50 ms gap between them is a reminder that the streaming protocol is part of the model’s performance envelope, not an afterthought.

The Voice-Agent Pipeline Around the Model

A modern voice agent is a cascade of at least four components, and TTS is usually the smallest contributor to delay. Understanding where the time goes is more valuable than any single benchmark.

Voice agent pipeline with Eleven v4 Turbo showing speech recognition, LLM, TTS and network stages

Figure 2: Cascaded voice-agent pipeline. Audio in, transcription, turn detection, language model, TTS with Eleven v4 Turbo, audio out. Each hop adds latency.

The four stages and their honest costs

Speech recognition (ASR, automatic speech recognition) runs on streaming audio and emits partial transcripts. The cost that matters is not recognition speed but endpointing: deciding that the caller has finished. A conservative endpoint waits a few hundred milliseconds of silence, and that wait is often the largest single term in the budget. Turn-taking models try to shorten it by predicting end of turn from prosody and content rather than from silence alone.

The language model is the second heavy term. What matters is time to first token, not tokens per second, because TTS can begin as soon as the first clause is available. Launch coverage notes that audio generation can start as the LLM starts producing its answer, which is only possible if the pipeline streams text into TTS incrementally rather than waiting for a complete reply.

TTS is the third term. With Turbo, the vendor’s own figure is roughly 150 ms to first audio on a clean link. The fourth term is transport: WebRTC or WebSocket hops between the caller’s device, your edge, the ASR vendor, the LLM provider and the TTS vendor. Each cross-region hop can cost tens of milliseconds, and telephony adds carrier latency that no model can remove.

A worked latency budget (illustrative)

The numbers below are illustrative planning values, not measurements. Only the TTS row comes from the vendor claim above; the others are typical ranges you should replace with your own traces.

Stage Optimistic (ms) Typical (ms) Conservative (ms)
Endpointing and turn detection 150 300 600
ASR final segment 50 100 200
LLM time to first useful clause 250 450 900
TTS time to first audio (Turbo claim) 150 150 250
Network and telephony 60 120 250
Total silence before reply 660 1,120 2,200

Human conversational gaps average a few hundred milliseconds across languages, so even the optimistic column is above natural but tolerable, and the conservative column feels broken. The lesson is that shaving 50 ms from TTS moves the total by less than the variance in LLM time to first token. If you can only optimise one thing, measure p95 rather than median, because callers remember the bad turns.

Co-optimised stacks and why vendors bundle

ElevenLabs describes Turbo as part of a co-optimised stack with its own transcription and turn-taking models inside ElevenAgents. The engineering logic is sound: when one vendor owns ASR, turn detection and TTS, it can overlap them. For example, it can begin speculative synthesis of a likely opening while the caller is still finishing, and cancel it if the caller keeps talking. That kind of overlap is hard to achieve across three vendors’ APIs. The counterargument is lock-in and the loss of best-of-breed choice, which we return to in the recommendations.

Deeper Analysis: Expressiveness, Cloning, Languages and Streaming

Audio tags as a control interface

Audio tags are best understood as a small, model-interpreted markup layer. Documented examples include [laughs], [whispers], [door slams], and the launch post also cites richer forms like [said angrily in French accent], [light rain] and [phone buzzing]. Tags can reportedly stack and appear sequentially, so a line might carry an emotional direction and an ambient cue together.

Because the mapping is learned, tags are probabilistic. The same tag on two generations will not be identical, and interactions between tags and the surrounding text can surprise you. The practical technique is to treat tags like prompts: keep them short, put them where the change should occur, test them on your actual voices, and lock the winning generation once you have it rather than regenerating at runtime. For live agents, that means restricting yourself to a small, tested tag vocabulary, since you cannot audition output in real time.

Here is a minimal request against the REST endpoint. The model ID is from the documentation; check the current API reference for parameter names before shipping, since this is a sketch.

import requests

API_KEY = "YOUR_KEY"
VOICE_ID = "YOUR_VOICE_ID"

resp = requests.post(
    f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}",
    headers={"xi-api-key": API_KEY, "Content-Type": "application/json"},
    json={
        "model_id": "eleven_v4",
        "text": "[whispers] We found the fault. [exhales] It was the sensor.",
        "output_format": "mp3_44100_128",
    },
    timeout=30,
)
resp.raise_for_status()
open("out.mp3", "wb").write(resp.content)

Streaming input for agents

For real-time use, the WebSocket stream-input endpoint lets you send text chunks as the LLM produces them and receive base64 audio back with optional alignment data. The documentation describes parameters such as auto_mode, which reduces latency by disabling chunk scheduling and buffers and is recommended only when you send full sentences, and inactivity_timeout, which defaults to 20 seconds with a maximum of 180. It also warns that WebSockets are more complex than plain HTTP and add buffering overhead when you already hold the complete text. The default model on that endpoint in the docs I read was eleven_multilingual_v2, so set model_id explicitly to a v4 model or you may benchmark the wrong thing.

import asyncio, json, websockets

URL = ("wss://api.elevenlabs.io/v1/text-to-speech/VOICE_ID/"
       "stream-input?model_id=eleven_v4_turbo")

async def speak(chunks):
    async with websockets.connect(URL) as ws:
        await ws.send(json.dumps({"text": " ", "xi_api_key": "YOUR_KEY"}))
        for c in chunks:
            await ws.send(json.dumps({"text": c}))
        await ws.send(json.dumps({"text": ""}))
        async for msg in ws:
            data = json.loads(msg)
            if data.get("audio"):
                pass  # base64 audio: decode and feed the playout buffer
            if data.get("isFinal"):
                break

asyncio.run(speak(["Hello, this is ", "your service agent. ", "How can I help?"]))

Eleven v4 audio-tag and text streaming sequence between LLM, agent runtime and TTS

Figure 3: Streaming sequence for a Turbo voice agent. Text chunks flow from the LLM through the agent runtime to the TTS socket, and audio chunks return while the LLM is still generating.

Figure 3 shows the property that matters: the audio path and the text path overlap. The first audio chunk goes out while later text is still being generated, which is how the total silence in the budget table stays near a second rather than the sum of full-reply generation times.

Voice cloning: capability and responsibility

Instant clones from ten seconds lower the cost of personalisation dramatically, and they lower the cost of misuse by the same amount. The launch material states that clones require verified owner consent, and Professional Voice Clones go through a stricter verification path. For enterprise deployment, treat consent records, voice ownership and retention as compliance artefacts, not settings. ElevenLabs lists SOC 2 Type II, ISO 27001, PCI DSS Level 1, GDPR compliance, HIPAA-eligible workflows and a Zero Retention Mode for enterprise, per launch coverage; confirm which apply to your plan and region in contract, not in a blog post.

Multilingual reach and mid-call switching

Ninety-plus languages is a coverage claim, not a uniform-quality claim. The vendor singles out Japanese, Brazilian Portuguese, Mandarin and Cantonese as improved, which implicitly acknowledges that quality varies by language. It also says agents can switch languages mid-call while keeping voice consistency, and that Professional Voice Clones work across all supported languages. That cross-lingual identity is technically plausible when the speaker embedding is language-independent, but accent leakage is the usual failure: a cloned English speaker rendered in Tamil may sound like an English speaker reading Tamil. Native-speaker review of your top three languages is not optional.

Pronunciation is the other multilingual lever. The improved International Phonetic Alphabet (IPA) support and pronunciation dictionaries exist for brand names, drug names and place names, which are exactly the tokens where a text normaliser guesses wrong. In regulated domains, an incorrectly voiced drug name is a safety issue, so maintain a reviewed pronunciation dictionary and version it like code.

Benchmarks, Pricing and Access

What the benchmark evidence does and does not show

Two kinds of evidence support the “best TTS” claim. The independent one is a ranking: Eleven v4 was reported first on the Artificial Analysis Provider Voice Arena at 1319 Elo as of September 28, 2026. An arena ranking aggregates blind human preference votes, so it measures what listeners prefer on the prompts and voices the arena uses. It is a strong signal for narration-like quality and a weak one for your telephony, accented-speaker, jargon-heavy use case.

The vendor evidence is a set of blind head-to-head tests in which ElevenLabs reports about 75 percent listener preference overall, with individual matchups spanning 65 to 81 percent. The comparison set includes Cartesia Sonic 3.6, Inworld TTS-2, Google’s Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS. These tests were run by the vendor, the exact prompts are not something I could verify, and vendor-run preference tests reliably favour the sponsor. Treat 75 percent as “probably ahead, by an unknown margin”.

Latency comparisons follow the same rule. The product page reports Cartesia Sonic 3.6 as about 112 ms slower and OpenAI’s GPT-4o mini TTS as about 664 ms slower than v4 Turbo, and one news write-up quotes 262 ms and 814 ms for those rivals against Turbo’s 150 ms. Those are consistent with each other (150 plus 112 is 262, and 150 plus 664 is 814), which is reassuring on arithmetic but says nothing about the measurement conditions. They are vendor-measured, likely from one region, and rivals update their models frequently. Run your own test from your own infrastructure.

Benchmarks you should run yourself

A defensible evaluation for a voice agent has five parts:

  • Time to first audio byte from the client, measured at p50, p95 and p99 over at least a few hundred requests, from the region where your agent runs.
  • Word error rate of round-trip audio: synthesise a set of your real strings (names, SKUs, drug names), transcribe them with an independent ASR, and count errors. This detects mispronunciation at scale.
  • MOS-style blind listening with five to ten native speakers per language, using your own scripts and your target voices.
  • Stability: regenerate the same line thirty times and look for drift in loudness, speed and identity.
  • Interruption behaviour: how quickly audio stops when the caller barges in, and whether partial output is discarded cleanly.

Eleven v4 versus Turbo model family latency and language coverage comparison

Figure 4: The ElevenLabs model family as documented on September 30, 2026: languages and stated latency per model. Latency values are vendor figures and measured under different conditions.

Figure 4 summarises the documented family. Flash v2.5 lists about 75 ms and 32 languages, Multilingual v2 lists 29 languages, v3 lists 70-plus, v3 Conversational about 280 ms, and the v4 pair lists 90-plus with Turbo at about 100 ms inference. Read the chart as a positioning map, not a like-for-like benchmark: Flash’s 75 ms and Turbo’s 100 ms may be defined differently, and only a controlled test can settle which is faster for you.

Pricing: what is documented and what is not

The pricing page I could read lists v4 at $0.08 per 1,000 characters and v4 Turbo at $0.04 per 1,000 characters at list price, against $0.08 for v3 and Multilingual v2, $0.04 for v3 Conversational, and $0.04 for Flash and Turbo-class models. It also shows a 72 percent launch discount running until October 12, which would bring v4 to $0.022 and Turbo to $0.011 per 1,000 characters. A third-party analysis reads the same figures. Since prices and promotions change, confirm on the live ElevenLabs API pricing page before you model a budget.

Converting characters to minutes is the trap. The page equates roughly 1,000 characters with about a minute of audio, and the model documentation gives a limit of 10,000 characters as roughly ten minutes, so the rule of thumb is consistent. But real speech rates vary by language, punctuation and tags. Spoken English commonly runs near 900 to 1,000 characters per minute, while languages written in logographic scripts pack far more meaning per character, so per-minute cost can be quite different in Japanese or Mandarin. Measure characters per audio minute on your own corpus.

Here is a worked example using list prices, assuming the 1,000 characters per minute rule and TTS only:

Scenario Audio minutes per month Turbo at $0.04 per minute v4 at $0.08 per minute
Pilot 1,000 about $40 about $80
Mid-size support desk 50,000 about $2,000 about $4,000
Large contact centre 500,000 about $20,000 about $40,000

The 1,000-minute Turbo case matches the roughly $40 per month figure in third-party coverage, which is arithmetic rather than a quoted invoice. Remember that an agent speaks only part of a call, so audio minutes are lower than call minutes, while the ASR, LLM and telephony bills sit on top. In many deployments the LLM and telephony costs match or exceed TTS, so a model that is twice as expensive per character may still be the right choice if it reduces retries and escalations.

The consumer-facing plans are documented as a free tier of 10,000 credits per month, described as roughly ten minutes of audio, paid plans starting at $6 per month, and custom enterprise pricing. ElevenLabs states both v4 models are included on all plans. Whether API character prices and plan credits convert one-to-one is something to check in your dashboard.

Access and deployment

Eleven v4 is a hosted service. There are no open weights, and I found no indication of self-hosting, so the deployment decision is which ElevenLabs surface you use: ElevenAgents for a managed voice-agent platform, ElevenCreative for production tools, or ElevenAPI for REST and streaming with TypeScript and Python SDKs. Data residency, retention and zero-retention modes are contractual and plan-dependent. If you need on-premises or air-gapped speech synthesis, this launch does not change that answer, and you would look at open-weight TTS models instead.

How It Compares: Decision Matrix

The market is fluid, and I am not ranking rivals on numbers I cannot verify. The matrix below is a qualitative guide built on the documented positioning of each option, with the caveat that vendors ship monthly.

Use case Eleven v4 Eleven v4 Turbo Flash v2.5 Google Gemini TTS Cartesia Sonic
Audiobook, dubbing, game dialogue Best fit: context and tags Acceptable Weak on emotion Strong if already on Google Cloud Possible
Phone agent, English Too slow to start Strong candidate Strong on speed, less expressive Candidate Strong candidate
Phone agent, 30+ languages Batch only Strongest documented coverage Only 32 languages Depends on locale Check coverage
Regulated IVR with fixed prompts Pre-render and cache Pre-render and cache Pre-render and cache Pre-render and cache Pre-render and cache
Self-hosted or air-gapped Not possible Not possible Not possible Not possible Not possible

The row that usually surprises teams is the fixed-prompt IVR. If your greetings and menus are static, synthesise them once with the highest-quality model, store the audio, and use the real-time model only for dynamic sentences. That turns latency and cost concerns into a storage problem, and it lets you human-review every prompt.

For comparison with another recent omnimodal release that includes speech generation, see our analysis of Qwen3.8 Omni Flash and its omnimodal architecture and pricing, which takes the opposite architectural bet of speaking from inside a single multimodal model rather than through a dedicated TTS stage.

Trade-offs, Gotchas, and What Goes Wrong

Latency is a budget, not a feature. The 150 ms figure is TTS-only, network removed. Put ASR endpointing, an LLM and a phone network in front and you land closer to a second of silence unless you engineer every stage. A team that swaps TTS vendors to fix “slow” without profiling usually finds the problem was endpointing.

Tags are probabilistic. [laughs] does not guarantee a laugh, and tag behaviour can differ between the quality model and Turbo. The failure mode is an agent that laughs at a caller describing a bereavement. Restrict live agents to a tested tag set and write guardrails so the LLM cannot emit arbitrary tags from user-supplied text, which would otherwise be an injection channel into your voice output.

No SSML means no hard timing guarantees. Because break tags are disabled, exact pause lengths, mid-word emphasis and spoken-out numeric formats are less deterministic. For phone numbers, dates and account numbers, normalise the text yourself, for example writing digits in the grouping you want spoken, and verify by round-trip ASR.

Consistency drifts across regenerations and long sessions. The vendor targets this with speaker stability and request stitching, but you should still test long calls. Reliability on long calls is a stated selling point, not a guarantee.

Language quality is uneven. Ninety languages does not mean ninety equal experiences. Low-resource languages, code-switching (Hinglish, Spanglish) and rare proper nouns are where errors concentrate.

Cloning creates legal and reputational exposure. Ten seconds of audio is a very low bar for impersonation. Voice-likeness law is changing across jurisdictions, and consent verification is a control, not a legal shield. Keep signed consent for every cloned voice and log which voice produced which call.

Vendor concentration. A bundled ASR, turn-taking and TTS stack from one supplier is efficient but creates a single point of failure. If the provider has an outage, both your listening and your speaking stop. Design a fallback path, even a plain lower-quality TTS, so callers hear something rather than dead air.

Preview and pricing volatility. A 72 percent launch discount ending October 12 means cost projections made this week will change. Model both list and discounted cases, and do not sign a volume commitment on promotional numbers.

Practical Recommendations

Start by deciding whether the use is produced audio or live conversation, because the answer chooses the model. Produced content (narration, dubbing, e-learning, game dialogue) belongs on Eleven v4, where context and tags matter and latency does not. Live conversation belongs on Turbo, with v4 reserved for pre-rendered lines.

Second, instrument before you optimise. Log timestamps at every stage boundary: end of caller speech, final transcript, first LLM token, first TTS byte, first audio played. Without that, latency arguments are opinion. Third, stream text into TTS clause by clause, with auto_mode considered only when you send full sentences, and cancel synthesis promptly on barge-in.

Fourth, treat pronunciation and tags as code: keep a reviewed dictionary, a tested tag vocabulary, and regression audio you re-run when the vendor updates a model. Fifth, budget both list and promotional prices and measure real characters per audio minute in your languages. Sixth, if you clone voices, capture consent and set retention rules before launch.

Checklist before go-live:

  • [ ] p95 end-to-end silence measured on production-like network, target under 1.5 seconds
  • [ ] Round-trip ASR check on names, numbers and domain terms passes
  • [ ] Native-speaker sign-off for each launch language
  • [ ] Tag allow-list enforced, user text cannot inject tags
  • [ ] Fallback TTS path and outage runbook tested
  • [ ] Consent records stored for every cloned voice
  • [ ] Cost model covers list price, promo expiry and ASR plus LLM plus telephony

For teams building agents that also act on a desktop or in apps, voice is only the interface layer; our overview of Claude computer use and LLM desktop agent architecture covers the action layer behind it.

Frequently Asked Questions

What is Eleven v4?

Eleven v4 is ElevenLabs’ text-to-speech model released September 28, 2026, positioned as its most expressive yet. It supports 90-plus languages, inline audio tags for emotion and sound, instant voice cloning from about ten seconds of audio, and up to 10,000 characters per generation. It ships alongside Eleven v4 Turbo, a lower-latency variant for real-time voice agents. Both are available through ElevenAgents, ElevenCreative and the ElevenAPI, and the model IDs are eleven_v4 and eleven_v4_turbo.

How fast is Eleven v4 Turbo?

ElevenLabs reports a median inference latency of about 100 ms and a median time to first speech of about 150 ms over WebSocket streaming, with network latency removed. These are vendor figures for the TTS stage only. They exclude speech recognition and language-model time, so a full voice agent will still have roughly a second of total silence unless every stage is tuned. Measure p95 from your own region and network before committing.

How much does Eleven v4 cost?

At list price, the pricing page I checked shows $0.08 per 1,000 characters for v4 and $0.04 for v4 Turbo. A 72 percent launch discount runs until October 12, bringing them to about $0.022 and $0.011. With roughly 1,000 characters per audio minute, list cost is about $0.08 or $0.04 per minute, but the ratio varies by language. Plans start at $6 per month, with a free tier of 10,000 credits. Always confirm live pricing.

Does Eleven v4 support SSML?

No. ElevenLabs states that SSML break tags are disabled in v4, and pauses and delivery are controlled through natural-language audio tags and punctuation instead. The trade-off is that you lose exact millisecond timing control but gain intent-level direction such as whispering or hesitating. For numbers, dates and codes, normalise the text yourself and verify by round-trip transcription, since the model’s interpretation is learned rather than deterministic.

Can I self-host Eleven v4 or use it offline?

I found no evidence of open weights or a self-hosted option, and ElevenLabs presents v4 as a hosted service through its products and API. Enterprise plans offer options such as zero-retention modes for data handling, but that is not the same as on-premises inference. If you need air-gapped or on-device speech synthesis, you would evaluate open-weight TTS models instead and accept a gap in expressiveness and language coverage.

Is Eleven v4 really the best TTS model?

It leads one independent public ranking, the Artificial Analysis Provider Voice Arena, reported at 1319 Elo on September 28, 2026, and ElevenLabs reports roughly 75 percent listener preference in its own blind tests. That is strong evidence for general quality, not proof for your language, accent, phone channel or domain vocabulary. Rankings shift as rivals ship updates, so run a blind test with your scripts and voices before choosing.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *