StepAudio 3 Realtime 技术报告发布,实时语音中并行推理

HuggingFace Daily Papers(社区热门论文)·2026-09-12 08:00·4天前
AI 导读

StepAudio 3 Realtime 是一个围绕听-说-想-做循环组织的音频语言基础模型,通过 Deep Perception 解读声学线索,Seamless Duplex 处理停顿、附和与打断。

HuggingFace Daily Papers(社区热门论文)
58AI 编辑部评分,满分 100

StepAudio 3 Realtime 技术报告发布,实时语音中并行推理

2026-09-12 08:00· 4天前
AI 导读

StepAudio 3 Realtime 是一个围绕听-说-想-做循环组织的音频语言基础模型,通过 Deep Perception 解读声学线索,Seamless Duplex 处理停顿、附和与打断。

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret user intent, while Seamless Duplex models synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery.

In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat. With Think-While-Speaking, it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time. Furthermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ-Voice.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org