Rohan Paul · @rohanpaul_ai · X·2026-09-15 11:35·44分钟前
AI 导读

VoiceArena 发布 Jarvis Bench v0.5,将语音智能体评测拆分为任务完成度与自然度两项,由真人实时对话、另一组人类盲测成对投票,榜单中还包含一个人类"智能体"。结果显示任务完成度上人类与模型差距较小,自然度上差距明显。

Rohan Paul@rohanpaul_ai
44AI 编辑部评分,满分 100
2026-09-15 11:35· 44分钟前
AI 导读

VoiceArena 发布 Jarvis Bench v0.5,将语音智能体评测拆分为任务完成度与自然度两项,由真人实时对话、另一组人类盲测成对投票,榜单中还包含一个人类"智能体"。结果显示任务完成度上人类与模型差距较小,自然度上差距明显。

I really like how @voicearena_ai designed this evaluation.

Instead of asking whether one voice model is simply "better" than another, they split the problem into Task Completion and Naturalness, then use blind pairwise voting to measure them separately.

Simple distinction. Huge difference in what the numbers tell you.

On Task Completion, humans and models are apparently much closer.

On Naturalness, they are not.

Shobhit BangaIntroducing Jarvis Bench v0.5, @voicearena_ai's conversational agent benchmark. We've been obsessed with one question at VoiceArena: why do voice agent demos so...

来源:Rohan Paul· x.com