The race for the best AI voice just got louder. ElevenLabs on Monday launched Eleven v4 and Eleven v4 Turbo, two text-to-speech models the company calls its fastest and most emotive yet. One is built to perform; the other is built to talk fast enough that you can’t tell it isn’t human.

Eleven v4 is a quality play aimed at audiobooks, ads, dubbing, and long-form narration. Eleven v4 Turbo is the speed play: a low-latency variant tuned for voice agents and real-time conversation, posting a median inference latency of around 100 milliseconds. Both are available immediately in ElevenAgents, ElevenCreative, and through the ElevenAPI, including on a free tier.

A voice model that performs#

While earlier generations of text-to-speech models just read text aloud, Eleven v4 is built to perform it. The model uses what ElevenLabs describes as an entirely new architecture that interprets tone, pacing, emotion, character, and context from the text itself — delivering speech that can sound dramatic, tender, urgent, comedic, or conversational while keeping the speaker’s identity intact.

“AI is making incredible progress in science and coding, but there are many parts of the economy and society where the way AI interacts will be just as important as its raw reasoning power,” said co-founder and CEO Mati Staniszewski in the company’s announcement. “Eleven v4 delivers the rich, expressive communication that AI has been missing.”

A new method for capturing speaker identities is meant to keep each voice consistent across agent conversations, audiobooks, and ads — and scene-level context lets speakers respond to what was just said, producing multi-speaker dialogue that behaves more like a conversation than a set of stitched-together lines.

Turbo: 100ms to the first word#

Historically, the best-sounding voice models have been too slow for live conversation — companies had to pick between fast or expressive. Eleven v4 Turbo is ElevenLabs’ answer to that trade-off. The company says Turbo’s median inference latency is about 100ms, which it notes is faster than the average gap between two people talking, and that audio starts returning before a sentence finishes, thanks to bidirectional streaming.

The company’s footnoted tests, run in September 2026 over WebSocket streaming with network latency measured and removed, put Turbo’s median time to first speech at roughly 150 milliseconds — against 262ms for Cartesia Sonic 3.6 and 814ms for OpenAI GPT-4o mini TTS. ElevenLabs says its research and engineering teams optimized v4 Turbo and the ElevenAgents conversational platform together as one system.

Illustration of a customer-service agent smiling beside a glowing AI voice orb
Expressive agents without the awkward pause: Eleven v4 Turbo’s low latency targets call centers and real-time assistants. AI-generated editorial illustration — AI Frontier Post.

90+ languages, 10-second clones#

Both models support more than 90 languages, up from roughly 70 in v3, with new languages including Cantonese, Mongolian, and Odia. ElevenLabs says the biggest quality gains land in Japanese, Mandarin, and Brazilian Portuguese, and that a voice recorded in one language now speaks others fluently while adopting the accent of a native speaker — for example, a Korean-recorded voice reading English with a native English accent while keeping its original pitch and timbre.

Voice cloning gets two upgrades. Instant Voice Clones can now capture a voice with high fidelity from just 10 seconds of audio, and Professional Voice Clones — unavailable in Eleven v3 — return with the model’s full emotional range. Every voice in the company’s library of more than 17,500 works with v4, though older clones need to be retrained. A single generation supports up to 10,000 characters, with request stitching keeping pacing and delivery consistent across long scripts.

Illustration of luminous voice-waveform ribbons streaming across a glowing world map
90+ languages: Eleven v4 is aimed at carrying one voice across every market. AI-generated editorial illustration — AI Frontier Post.

The controls: inline tags replace SSML#

Directing the performance is done in plain language or with inline audio tags: [laughs], [said angrily in British accent], [light rain], [phone buzzing]. ElevenLabs says v4 follows these tags and direction prompts more accurately than previous models, with stackable tags and significantly improved IPA phoneme support for custom pronunciations. Notably, SSML tags like <break> are disabled in v4 — the natural-language tags are now the control mechanism.

Identity, consent, and compliance#

Voice agents handle some of a business’s most sensitive conversations — identity verification, health, money — and ElevenLabs is leaning hard on the enterprise angle: SOC 2 Type II, ISO 27001, and PCI DSS Level 1 certifications, GDPR compliance with HIPAA-eligible workflows, and data residency options across the US, EU, Singapore, and India. The company says it does not train on scripts or audio tags without consent, and offers Zero Retention Mode for eligible enterprise services.

On the consent side, professional voice cloning requires passing a Voice CAPTCHA, No-Go Voices safeguards block cloning of politicians and other high-profile individuals, and generated audio is covered by ElevenLabs’ AI Speech Classifier for detection. The company says creators in its Voice Library have earned more than $22 million from voice replicas to date.

What to watch#

ElevenLabs reports that Eleven v4 ranked first on the Artificial Analysis Provider Voice Arena leaderboard for September 2026 and was preferred by roughly 75 percent of listeners in blind head-to-head tests, winning 65–81 percent of matchups against rivals including Cartesia Sonic 3.6, Inworld TTS-2, and Google’s Gemini 3.8 Flash TTS models. Those are company-run benchmarks, but the direction is clear: the voice-model contest is moving from “who sounds human” to “who sounds human fast.”

Early customers are already pointing at production numbers. Ryan Peterson, SVP of product for Agentforce Voice at Salesforce, said v4 Turbo is raising “the bar, with faster, more natural responses that meet the standard our customers expect.” Oscar Daniels, head of credit building products at Spring Financial, called it “the first one” fast enough to feel like a real conversation “without trading away quality.” And Grisha Grigorev, AI agents product manager at Impress, said v4 Turbo cut speech time-to-first-byte 47 percent across live agents with “no regression in resolution rate, conversation quality, reliability or cost.”

Pricing follows the same credit structure as ElevenLabs’ other TTS models: a free tier of 10,000 credits a month (roughly 10 minutes of audio, per the company) and paid plans starting at $6 a month. The open question now is how quickly rivals — from Cartesia to OpenAI to Google — close the expressiveness gap, and whether the next voice agents you talk to will be ones you’d swear were human.

Sources