Introduction

On September 28, 2026, ElevenLabs shipped Eleven v4—its most emotive text-to-speech model yet—alongside the low-latency twin Eleven v4 Turbo. This is a full TTS architecture refresh, not a Music-line bump: closed API models you can run today in ElevenAgents, ElevenCreative, and ElevenAPI.

Primary source: ElevenLabs — Introducing Eleven v4 by Mati Staniszewski and Piotr Dabkowski (published Sep 28, 2026). Same-day press: TechCrunch.

Pro Tip

If you already prompt with audio tags on older Eleven models, keep the habit—v4 is tuned to follow them harder. Lead lines with direction like [excited, happy] or inline cues such as [laughs] / [said angrily in French accent], and pin v4 Turbo for agent turns where median time-to-first-speech (~150 ms over WebSocket, network stripped) matters more than studio-length polish.

What Shipped Today

ElevenLabs’ launch post puts concrete stakes on the table:

  • Eleven v4 — new architecture for expressive, context-aware TTS; ranked #1 on Artificial Analysis’s Provider Voice Arena (Sept 2026)
  • Blind head-to-heads: preferred by about ~75% of listeners for expressiveness/naturalness vs Cartesia Sonic 3.6, Inworld TTS-2, and Google Gemini 3.8 Flash / Flash-Lite TTS
  • Eleven v4 Turbo — same stack for agents; median inference latency about ~100 ms, median audible speech about ~150 ms (WebSocket streaming, Sept 2026 internal bench)
  • 90+ languages, with stronger accent adherence when a voice recorded in one language speaks another
  • Instant Voice Clones from about 10 seconds of audio; Professional Voice Clones supported on v4
  • Natural-language delivery direction plus inline audio tags; improved IPA phoneme control
  • Multi-speaker scene context so dialogue responds across turns instead of stitching isolated lines
  • Available now in ElevenAgents, ElevenCreative, and via ElevenAPI (free account to start)

Closed weights / hosted API—no open checkpoint on Hugging Face. Run surface today: ElevenLabs apps + developer API.

Direct the Line, Keep the Voice

The audio-specialist unlock is directionality. Earlier TTS could read a script; v4 is sold as reading intention—tone, pacing, character, and conversational context—while holding speaker identity through agent chats, audiobooks, and ads.

Direction examples from the launch post include stage directions in brackets and mid-line sound cues ([light rain], [phone buzzing]). For localization, ElevenLabs claims a voice can adopt a native accent in the target language without drifting back to the source accent mid-generation—useful when a brand narrator has to stay one person across dubbing markets.

Turbo for Agents, v4 for Performance

High-emotion models used to lose to “fast but flat” agent voices. ElevenLabs’ framing for v4 Turbo is that you no longer pick one: expressive delivery at agent latency, co-tuned with ElevenAgents rather than stitched from unrelated vendors.

Keep the product lanes clear vs prior ArtRealmAI coverage:

PieceRole

ElevenLabs Music v2.5 (already live)

Generative music line

Eleven v4 / v4 Turbo (Sep 28)

Speech TTS + agent Turbo

This Dispatch is the voice TTS day-0, not a Music rehash.

How to Try It

  1. Open the Eleven v4 launch post and create / sign into an ElevenLabs account
  2. In ElevenCreative or the TTS playground, select Eleven v4 for expressive takes; switch to Eleven v4 Turbo for agent-style low latency
  3. Via API, point your existing ElevenAPI TTS calls at the new model IDs (see ElevenLabs developer docs for current model strings and tag syntax)
  4. For clones: Instant Voice Clone from ~10 s of clean reference audio; use PVC when you need the highest fidelity match

# Direction sketch (tag syntax per ElevenLabs docs)[warm, reassuring] Thanks for waiting — I've got your account open now.[laughs softly] That double charge on the eleventh? Already starting the refund.

Conclusion

Eleven v4 is the clearest Sep 28 audio product drop in the lane: a new emotive TTS architecture, a Turbo twin for agents, 90+ languages, stronger cloning, and a runnable surface on ElevenAPI the same day as the blog. Start from the official launch, A/B a familiar script against your current model, and keep Music v2.5 in its own lane when you need songs instead of speech.

—Anabel ♡