VākQA:泰卢固语口语事实型问答基准与评测研究

HuggingFace Daily Papers(社区热门论文)·2026-09-17 08:00·1天前
AI 导读

研究者发布 VākQA,一个泰卢固语口语问答基准,包含 2,001 条事实型问答对、覆盖六个领域、2.53 小时语音音频及双语转写和人工核验参考答案。研究先用人工评分验证评测方法,发现 Gemini-as-a-judge 最接近人类评分但严格程度不均,开源权重评审模型则会系统性惩罚表面形式与参考答案不同的正确泰卢固语答案。

HuggingFace Daily Papers(社区热门论文)
33AI 编辑部评分,满分 100

VākQA:泰卢固语口语事实型问答基准与评测研究

2026-09-17 08:00· 1天前
AI 导读

研究者发布 VākQA,一个泰卢固语口语问答基准,包含 2,001 条事实型问答对、覆盖六个领域、2.53 小时语音音频及双语转写和人工核验参考答案。研究先用人工评分验证评测方法,发现 Gemini-as-a-judge 最接近人类评分但严格程度不均,开源权重评审模型则会系统性惩罚表面形式与参考答案不同的正确泰卢固语答案。

Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference.

Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org