Artificial Analysis 发布 TTS 发音鲁棒性基准

Artificial Analysis · @ArtificialAnlys · X·2026-09-22 10:26·45分钟前
AI 导读

Artificial Analysis 推出 TTS 发音鲁棒性基准,测试模型能否正确读出难读文本,Google Gemini 3.1 Flash TTS 以 88.1% 居首,SpaceXAI TTS 87.6%、ElevenLabs Eleven v3 85.6% 紧随其后。

Artificial Analysis@ArtificialAnlys
45AI 编辑部评分,满分 100

Artificial Analysis 发布 TTS 发音鲁棒性基准

2026-09-22 10:26· 45分钟前
AI 导读

Artificial Analysis 推出 TTS 发音鲁棒性基准,测试模型能否正确读出难读文本,Google Gemini 3.1 Flash TTS 以 88.1% 居首,SpaceXAI TTS 87.6%、ElevenLabs Eleven v3 85.6% 紧随其后。

Announcing the Artificial Analysis Pronunciation Robustness benchmark, measuring how reliably Text to Speech models say challenging text correctly - Google Gemini 3.1 Flash TTS leads at 88.1%, followed closely by SpaceXAI TTS at 87.6% and ElevenLabs Eleven v3 at 85.6%

Existing Text to Speech (TTS) evaluations, including our TTS Arena, capture overall listener preference (e.g., how natural a voice sounds), and Word Error Rate (WER) checks whether the right words come out. Pronunciation Robustness adds a view of whether those words are said correctly (e.g., reading “St.” as “Saint” and “Street” in “St. Mary’s is on Church St.”) - this is critical for production voice agents, which need to get account details, names, currency amounts and more right to be trusted by users, at the low latency that conversational experiences demand.

Overview of Pronunciation Robustness Each model generates audio for 454 sentences containing 701 target words or phrases, across four categories: 1. Contextually appropriate (words read differently depending on context, e.g., a wound that is bandaged vs. a bandage that is wound) 2. Expanding shorthand (numbers, dates, units and notation read out naturally, e.g., 6'2", Chapter XVII or 1 tsp of sugar) 3. Preserving exact sequences (codes, paths, emails and identifiers spoken exactly, e.g., .env.local or a.chen@ucsf.edu) 4. Standalone terms (brand, place and technical names, e.g., Arkansas, façade or genre) Sentences are sent as written, with no normalization on our side beyond each model's default setting. Screened human listeners judge whether each target was pronounced correctly against accepted pronunciations set in advance, excluding anyone who fails an attention check. The score is the share of correct judgements, excluding "I could not tell" answers, and we publish a model once 95% of targets have three or more approved listeners.

Key results: ➤ Overall leaders: @GoogleDeepMind’s Gemini 3.1 Flash TTS leads at 88.1%, followed by @SpaceXAI’s TTS at 87.6%, @ElevenLabs’s Eleven v3 at 85.6%, v3 Conversational at 84.8% and @Alibaba’s Qwen-Audio-3.0-TTS-Plus at 81.6%. Gemini 3.1 Flash TTS leads Contextually appropriate and Expanding shorthand, SpaceXAI TTS leads Preserving exact sequences, and Qwen-Audio-3.0-TTS-Plus leads Standalone terms. ➤ Hardest categories: Expanding shorthand (62.4%) and Preserving exact sequences (62.9%) trail Contextually appropriate (86.2%) and Standalone terms (86.1%). We expect Expanding shorthand and Preserving exact sequences scores to rise considerably with normalized input text, especially for models without normalization on by default. ➤ Preference ≠ pronunciation robustness: The most preferred voices aren't always the most accurate - Sonic 3.6 ranks #1 on the Provider Voice Arena at 1276 Elo but #11 on pronunciation robustness at 74.5%, while Gemini 3.1 Flash TTS ranks #9 on the Arena at 1201 Elo and #1 on pronunciation robustness at 88.1%.

See more details below ⬇️

来源:Artificial Analysis· x.com