Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分39
AI 导读
对于低资源语言的 ASR(自动语音识别),更多训练数据往往比更大的模型更有效。 @voicearena_ai 的孟加拉语结果——基于其 ASR 语料库 Monsoon——证明了这一点。 在 Monsoon 上微调的 Whisper Medium,在孟加拉语 FLEURS 上的 LLM 词错误率从 85.27% 降至 7.65%。 Medium 是 769M 参数的 Whisper。足够小,可以低成本部署,在很多场景下也足够小,可以靠近用户运行。
正文
For ASR (automatic speech recognition) on weak languages, more training data tends to beat a bigger model.
@voicearena_ai’s Bengali result for Monsoon, its ASR corpus, is shows it.
Whisper Medium fine-tuned on Monsoon went from 85.27% to 7.65% LLM word error rate on Bengali FLEURS.
Medium is the 769M-parameter Whisper. Small enough to serve cheaply, and in many settings small enough to run close to the user.
Crazy week for Voice Arena at Interspeech in Sydney. 80+ organisations have asked to license Monsoon ASR corpus since we launched it seven days ago. The most common reason why labs are interested: Monsoon promises results. When we decided to build datasets at Voice Arena, we set one rule. Either the dataset promises results, or we don't build it. Monsoon promises results. On Bengali FLEURS, fine-tuning Whisper Medium on Monsoon took its LLM word error rate from 85.27% down to 7.65%. The other thing labs like is that we give them access to our model API. They can test it on their own internal benchmarks and see for themselves whether it will improve their models. You can see the detailed results and get API access here: https://voicearena.com/datasets/monsoon#results Monsoon is 100,000 hours across 50 languages from around the world, and most of them are long-tail languages. 100 languages by February, 1,000 by the end of 2027. It's only the first dataset Voice Arena has launched. We are excited to keep working on the key problems that get us closer to the dream of machines that talk like humans.在 X 查看被引用的帖子
来源:Rohan Paul · x.com