OmniEcho:面向具身智能体的空间音频理解模型与 OmniEchoBench 基准

HuggingFace Daily Papers(社区热门论文)·2026-09-20 08:00·5天前
AI 导读

研究团队提出 OmniEchoBench 空间音频-视觉感知与导航基准,覆盖 197 个真实空间音频-视觉场景、2,972 组问答对和 900 个导航样本,音频为来自 30 个真实环境的一阶 Ambisonics(FOA)数据。

HuggingFace Daily Papers(社区热门论文)
34AI 编辑部评分,满分 100

OmniEcho:面向具身智能体的空间音频理解模型与 OmniEchoBench 基准

2026-09-20 08:00· 5天前
AI 导读

研究团队提出 OmniEchoBench 空间音频-视觉感知与导航基准,覆盖 197 个真实空间音频-视觉场景、2,972 组问答对和 900 个导航样本,音频为来自 30 个真实环境的一阶 Ambisonics(FOA)数据。

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments.

To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose OmniEcho, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation.

These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org