elvis· @omarsar0 · X·· 2 小时前AI 评分48
AI 导读
AI 语音产品需要知道是谁在说话、说了什么。 这比单纯的转录或说话人分离更难,尤其是当说话人互相抢话时。 @smallest_AI 的 Pulse 现已在 @voicearena_ai 的 Diarization Bench 的 Diarization + ASR 赛道排名第一,DER 为 24.4%。 DER 是错误率,所以越低越好。该赛道排名第二的系统得分为 40.7%。
正文
AI voice products need to know who spoke and what they said.
That is harder than transcription or diarization alone, especially when speakers talk over each other.
Pulse by @smallest_AI is now #1 on the Diarization + ASR track of the @voicearena_ai's Diarization Bench, with 24.4% DER.
DER is an error rate, so lower is better. The next system in the track scores 40.7%.
Pulse by @smallest_AI currently ranks #1 in the Diarization + ASR track by DER, as the newest entry on the Voice Arena Diarization Bench. The Diarization + ASR track ranks Speaker-Attributed ASR systems that return speaker labels together with the transcript. Seven systems are in it today. Pulse scores 24.4% DER at a strict 0 ms collar. The next system, ElevenLabs Scribe v2, scores 40.7%. That is 1.7x lower error than the next best in the track. Most of the gap is missed speech. For every other system in the track, missed speech is the largest part of the error, above 30% for each of them. Pulse misses 4.4% of reference speech, the lowest in the track. On the full board of 13 systems, dedicated diarization models and Speaker-Attributed ASR systems together, Pulse places 5th. It is the highest-ranked Speaker-Attributed ASR system on the board and sits ahead of two dedicated diarization models. The result holds across settings: In-person recordings : 24.7% DER In online calls : 22.4% DER Congratulations to @smallest_AI 👏 View full results, per track and per condition: http://voicearena.com/diarization-bench在 X 查看被引用的帖子
来源:elvis · x.com