ElevenLabs 发布 v4 语音模型,表达力与长文本一致性提升并推出 Turbo 实时版本
ElevenLabs' new v4 speech model makes AI voices more expressive and consistent
ElevenLabs 发布 Eleven v4 语音模型,更准确地跟随情感、停顿等方向提示,支持超过 90 种语言(v3 约为 70 种),单次可处理 10,000 字符,发音基准得分从 v3 的 85.6% 升至 91.7%。
Elevenlabs is releasing Eleven v4, a new speech model that follows direction cues more accurately and keeps voices consistent across long productions. A new model architecture also powers the Turbo variant for real-time voice agents.
Eleven v4 generates laughter, whispers, and sounds like slamming doors more reliably than its predecessor. The Turbo variant for voice agents starts producing speech in about 150 milliseconds. Eleven v3, released just over a year ago, already supported these audio tags but followed them less accurately.

More consistency
Elevenlabs says v4 uses a new architecture that analyzes a script's tone, pacing, and context. Users can give directions through tags or plain sentences and use phonetic spelling to set the pronunciation of names and technical terms. The company says pronunciation controls now work more reliably. Narrators and characters should sound consistent throughout a production, even when users regenerate individual lines several times.
Eleven v4 handles up to 10,000 characters per request, roughly ten minutes of audio. Longer works like audiobooks use multiple segments, with pacing and delivery expected to stay consistent across transitions. In dialogue, AI speakers respond to the context of the entire scene rather than delivering each line in isolation.
The new version supports more than 90 languages, up from about 70 with v3. Cloned voices should speak other languages with native accents without drifting back to their original accents over time. Professional Voice Clones work again after being unsupported in v3, while an "Instant Voice Clone" needs only ten seconds of audio.
Turbo aims to make expressive voice agents faster
Elevenlabs is also releasing a faster v4 variant for real-time uses like customer service calls or game characters. The company says voice agent developers previously had to choose between speed and expression, and v4 Turbo is meant to offer both.
In Elevenlabs' tests, Turbo starts producing audible speech in 150 milliseconds, compared with 262 milliseconds for Cartesia Sonic 3.6. OpenAI's GPT-4o mini TTS takes 814 milliseconds. Elevenlabs optimized Turbo together with its ElevenAgents platform.

Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google's Gemini 3.8 Flash TTS on Artificial Analysis' Provider Voice Arena leaderboard. It scores 91.7 percent on the pronunciation benchmark, up from v3's 85.6 percent. In Elevenlabs' blind tests, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld, and Google.

Higher-quality voice clones make v4 useful for dubbing, with Elevenlabs promoting the ability to use an actor's voice across all supported languages. The company licenses voices from the people behind them. Voice actors can offer an extensively trained clone of their voice in Elevenlabs' library and earn money when paying users use it. Elevenlabs already offers access to celebrity voices like Michael Caine's through a dedicated marketplace.
Both models launch with temporary price cuts
Elevenlabs' standard API pricing is $80 per million characters for v4 and $40 for Turbo. Through October 12, those rates fall to $22 and $11. According to Elevenlabs, users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks. Usage is capped at twice their monthly credits. Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49.
Both models are available now in ElevenAgents, ElevenCreative, and through the API. Elevenlabs stores customer data in the US by default, according to its documentation. Enterprise customers can store data in isolated environments in the EU, India, or Singapore, though some processing may take place outside the chosen region. In the EU, customers can keep API processing within the region by using a mode that doesn't retain data.
Elevenlabs also released its Music 2.5 model in mid-September, designed to produce denser, more natural-sounding songs.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
来源:The Decoder:AI News(RSS) · the-decoder.com