HeyGen Voice Debuts at #1 on Artificial Analysis TTS Arena, $15/M Chars Through Oct 31

Introduction
HeyGen Voice is the first in-house text-to-speech model from HeyGen, the AI avatar video company. It launched on October 9, 2026, and went straight to #1 on the Artificial Analysis Controlled Voice TTS Arena, a blind listening test where every model speaks with the same eight cloned voices. It beat Alibaba's Qwen-Audio-3.1-TTS-Plus and ElevenLabs' Eleven v4 Turbo.
It's a closed model, with no open weights. You can use it today in the HeyGen app and through the HeyGen API as model heygen-voice-1. HeyGen says API speech is 50% off through October 31: $15 per million characters instead of $30, applied automatically.
Pro Tip
If you're cloning your own voice, record three clean minutes. The instant clone only uses the first three minutes of whatever you upload, so a longer file doesn't help, and a noisy one hurts. If you want more, the paid Professional Clone trains on 20+ minutes from the same speaker.
What Shipped
From
, the press release, and the HeyGen Voice developer docs:- Model ID:
heygen-voice-1, called withPOST /v3/models/audio/ttsfor a finished file orPOST /v3/models/audio/tts/streamfor streaming audio over Server-Sent Events, with optional word timestamps. - Output: a 44.1 kHz mono WAV.
- Instant voices: upload one MP3 or WAV recording and get a usable
voice_id, usually within seconds. Creating instant voices is free during the preview. They take one tuning knob,expressiveness_boost(0.0 to 1.0, default 1.0). - Professional clones: train on 1 to 10 recordings of the same speaker totaling 20+ minutes. These unlock
seed,speed,pitch_shift, andpitch_variance. Each one needs a purchased clone slot, and speech costs 0.6 API credits per generated minute. The press release lists the consumer version as a $99/month Professional Voice Clone add-on trained, with consent, on 30 minutes to 3 hours of speech. - Catalog voices: HeyGen says most of its 300+ stock voices now run on the HeyGen Voice engine (engine name
orca) through the separate/v3/voices/speechendpoint. - Rate limit: 30 requests per minute on the TTS endpoint, per the API reference.
The Numbers
These come from
, which ran HeyGen Voice through its Controlled Voice Arena (US and UK English) and its Pronunciation Robustness benchmark.ModelArena EloList price per 1M characters
HeyGen Voice
1,201 (1,468 appearances)
$30 ($15 through Oct 31)
Qwen-Audio-3.1-TTS-Plus
1,182
$19.30
Eleven v4 Turbo
1,166
$40
Eleven v4
1,162
$80
- By use case: #1 in Assistants (1,214) and Customer Service (1,212), #2 in Knowledge Sharing (1,158), and #4 in Entertainment (1,178). It's also #1 for UK English.
- Pronunciation Robustness: 83.1%, ranked #10 of 29 models, behind Qwen-Audio-3.1-TTS-Plus at 84.6%.
- Speed: about 40 characters per second of generation time.
The takeaway is that it's best at clear, friendly assistant and support reads, and only mid-pack at tricky pronunciation like codes, numbers, and unusual names. Test your own script before you switch a whole pipeline over.
How to Try It Today
- In the app: pick a voice in HeyGen, or record yourself, and generate a narration or an avatar video. HeyGen says the model is included on its platform.
- Through the API: create an instant voice with
POST /v3/models/audio/voicesand"mode": "instant", wait forACTIVE, then call TTS with a body like{"model": "heygen-voice-1", "voice_id": "...", "text": "Hello from my new voice.", "language": "en"}. - Compare it: Artificial Analysis has posted sample clips in its thread, so you can listen side by side with Qwen and Eleven before you commit.
Why It Matters
Most avatar and explainer-video tools rent their voices from ElevenLabs or another vendor. HeyGen now has its own model, and it's competitive on independent blind tests at half the list price of Eleven v4 Turbo during the promo. For creators who already make talking-head videos, the voice and the face now come from one stack. For TTS buyers, it's another sign the top of the leaderboard is crowded and prices are dropping fast. HeyGen still routes some catalog voices to ElevenLabs on request, so it's adding an option, not removing one.
Conclusion
HeyGen Voice (heygen-voice-1) is a closed TTS model from HeyGen. It debuted at #1 on the Artificial Analysis Controlled Voice Arena with an Elo of 1,201 and is strongest on assistant and customer-service reads. It's live in the HeyGen app and API today, with instant cloning from one recording, 44.1 kHz WAV output, streaming, and $15 per million characters through October 31 ($30 list).
Primary:
, , HeyGen Voice developer docs, Generate Speech API reference, HeyGen Voice press release.—Anabel ♡
