I really like how @voicearena_ai designed this evaluation.
Instead of asking whether one voice model is simply "better" than another, they split the problem into Task Completion and Naturalness, then use blind pairwise voting to measure them separately.
Simple distinction. Huge difference in what the numbers tell you.
On Task Completion, humans and models are apparently much closer.
On Naturalness, they are not.
Introducing Jarvis Bench v0.5, @voicearena_ai's conversational agent benchmark. We've been obsessed with one question at VoiceArena: why do voice agent demos so...