Artificial Analysis 评测显示,ElevenLabs 的 Eleven v4 在 Provider Voice TTS Arena 排行榜和 Pronunciation Robustness 基准上排名第一。
ElevenLabs’ Eleven v4 takes #1 on the Artificial Analysis Provider Voice TTS Arena Leaderboard and Pronunciation Robustness benchmark, and #2 on Controlled Voice, surpassing Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS on Provider Voice
Eleven v4 is the latest Text to Speech model from @ElevenLabs, with support for 90+ languages, up from 70+ for Eleven v3.
Key takeaways:
➤ Provider Voice: Eleven v4 takes #1 with an Elo of 1,319 (+19/-19) across 1,674 appearances, ahead of Cartesia’s Sonic 3.6 at 1,276 and Google’s Gemini 3.8 Flash TTS at 1,267. It ranks #1 across all four categories: Customer Service, Assistants, Knowledge Sharing and Entertainment.
➤ Controlled Voice (every model uses the same custom voice for comparison): Eleven v4 ranks #2 with an Elo of 1,157 (+16/-16) across 1,483 appearances, just behind Alibaba’s Qwen-Audio-3.1-TTS-Plus at 1,178 and well ahead of Eleven v3 at 1,073.
➤ Pronunciation Robustness: Eleven v4 scores 91.7%, the highest score we have measured, ahead of Gemini 3.8 Flash TTS at 89.5% and Gemini 3.1 Flash TTS at 88.2%, and up from Eleven v3 at 85.6%.
➤ Cost: Eleven v4 costs $80/1M characters, compared to $49/1M characters for Sonic 3.6 and $16.49/1M characters for Gemini 3.8 Flash TTS.
➤ Speed: Eleven v4 processes 73.4 characters per second of generation time, compared to 42.5 characters per second for Eleven v3.
See more details and listen to samples below 🧵
来源:Artificial Analysis · x.com