# VoiceArena 推出 Jarvis Bench v0.5 语音智能体评测

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-15 11:35
- AIHOT 分数：44
- AIHOT 链接：https://aihot.news/items/cmu24lsbc035pro9v2cs5wfqp
- 原文链接：https://x.com/rohanpaul_ai/status/2099703670719774924

## AI 摘要

VoiceArena 发布 Jarvis Bench v0.5，将语音智能体评测拆分为任务完成度与自然度两项，由真人实时对话、另一组人类盲测成对投票，榜单中还包含一个人类"智能体"。结果显示任务完成度上人类与模型差距较小，自然度上差距明显。

## 正文

I really like how @voicearena_ai designed this evaluation.

Instead of asking whether one voice model is simply "better" than another, they split the problem into Task Completion and Naturalness, then use blind pairwise voting to measure them separately.

Simple distinction. Huge difference in what the numbers tell you.

On Task Completion, humans and models are apparently much closer.

On Naturalness, they are not.

### 引用推文

> Shobhit Banga：Introducing Jarvis Bench v0.5, @voicearena_ai's conversational agent benchmark. We've been obsessed with one question at VoiceArena: why do voice agent demos so...
