跳到正文
原文
Artificial Analysis· @ArtificialAnlys · X·· 3 小时前AI 评分61
AI 导读

Artificial Analysis 评测显示,ElevenLabs 的 Eleven v4 在 Provider Voice TTS Arena 排行榜和 Pronunciation Robustness 基准上排名第一。

正文

ElevenLabs’ Eleven v4 takes #1 on the Artificial Analysis Provider Voice TTS Arena Leaderboard and Pronunciation Robustness benchmark, and #2 on Controlled Voice, surpassing Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS on Provider Voice

Eleven v4 is the latest Text to Speech model from @ElevenLabs, with support for 90+ languages, up from 70+ for Eleven v3.

Key takeaways:

➤ Provider Voice: Eleven v4 takes #1 with an Elo of 1,319 (+19/-19) across 1,674 appearances, ahead of Cartesia’s Sonic 3.6 at 1,276 and Google’s Gemini 3.8 Flash TTS at 1,267. It ranks #1 across all four categories: Customer Service, Assistants, Knowledge Sharing and Entertainment.

➤ Controlled Voice (every model uses the same custom voice for comparison): Eleven v4 ranks #2 with an Elo of 1,157 (+16/-16) across 1,483 appearances, just behind Alibaba’s Qwen-Audio-3.1-TTS-Plus at 1,178 and well ahead of Eleven v3 at 1,073.

➤ Pronunciation Robustness: Eleven v4 scores 91.7%, the highest score we have measured, ahead of Gemini 3.8 Flash TTS at 89.5% and Gemini 3.1 Flash TTS at 88.2%, and up from Eleven v3 at 85.6%.

➤ Cost: Eleven v4 costs $80/1M characters, compared to $49/1M characters for Sonic 3.6 and $16.49/1M characters for Gemini 3.8 Flash TTS.

➤ Speed: Eleven v4 processes 73.4 characters per second of generation time, compared to 42.5 characters per second for Eleven v3.

See more details and listen to samples below 🧵

来源:Artificial Analysis · x.com