⚠️ Affiliate Disclosure: This review contains affiliate links. If you sign up through one of them I may earn a commission, at no extra cost to you. This does not affect my reviews and I only recommend tools I genuinely believe in.
ElevenLabs launched Eleven v4 and Eleven v4 Turbo on 28 September 2026 — a ground-up rebuild of its text-to-speech architecture, and the follow-up to the v3 launch we tested three times last quarter. ElevenLabs is our current Tool of the Month and holds an 8.5/10 in our hands-on review, so this one matters to us: here’s what’s claimed, what’s independently verified, and what we’ll test.
What ElevenLabs says is new in Eleven v4
- A new architecture built for expressive performance — inline tags and plain-English direction for emotion, pacing, reactions and sound effects
- Instant Voice Clones from just 10 seconds of audio, with Professional Voice Clones remaining the highest-fidelity option
- Speaker identity held more reliably across long narration, dialogue and regenerated lines
- Improvements across 90+ languages, with Japanese, Brazilian Portuguese, Mandarin and Cantonese called out
- v4 Turbo for low-latency work like voice agents — a reported ~150ms median to first speech
Those are the vendor’s claims. The independent number is stronger than most launch-day boasts: Artificial Analysis puts Eleven v4 at #1 on its Provider Voice TTS Arena leaderboard (Elo 1,319, ahead of Cartesia’s Sonic at 1,276 and Google’s Gemini Flash TTS at 1,267) and #1 for pronunciation robustness — though #2 on controlled voice. Blind listeners preferred it roughly three times out of four.
Eleven v4 pricing and the two-week launch offer
For two weeks from launch (so until roughly 12 October 2026), ElevenLabs has discounted the Eleven v4 API to $22 per 1M characters and v4 Turbo to $11 per 1M characters. In the creative suite, v4 is free for Creator plans and above, up to 2× your monthly credits. Those are ElevenLabs’ launch figures — check the live pricing page before committing, as launch offers expire.
What we’ll test in Eleven v4
Same drill as v3 — and we’ll publish what we find either way. Three things top the list: whether the 10-second instant clone genuinely carries tone and personality or just timbre; whether the inline emotion tags give real line-level control without endless regeneration; and whether speaker identity actually survives a long narration, which is where v3 occasionally drifted. Our ElevenLabs review holds its 8.5/10 until the retest is done — this section will link to the results.

Leave a Reply