VoiceArena 推出 Monsoon,一个覆盖 50 种语言、基于 23 国无脚本对话构建的语音数据集。
Low-resource ASR (automatic speech recognition) usually gets treated as a model problem.
VoiceArena just launched Monsoon, and shows it's a data collection problem.
Monsoon is a 50-language speech dataset built from unscripted conversations across 23 countries.
Off the shelf, Whisper Medium scores 92.7% semantic WER on IndicVoices Telugu. VoiceArena fine-tuned the same 769M checkpoint on Monsoon and got 16.1%, narrowly ahead of MAI-Transcribe-2 and Gemini 3.1 Pro in its own tests.
There was no architecture change and no bigger model. The 76.6-point drop comes entirely from what the model heard during fine-tuning.
Most speech data pipelines filter noise out. Monsoon keeps it in on purpose.
Short clips, noisy rooms and difficult acoustic conditions stay in the training distribution, and the mix isn't left to chance. VoiceArena stratifies segments by DNSMOS. That makes the share of degraded audio a controlled property of the data
Introducing Monsoon ASR ⚡️ @voicearena_ai Speech recognition does not have a model problem anymore. It has a data problem. The best ASR systems are approaching human-level performance in English. But move into the long tail of the world's languages, especially real,在 X 查看被引用的帖子
来源:Rohan Paul · x.com