联合视频生成中的跨注意力鸿沟:RecCAR 让视频与动作、音频互相对齐

HuggingFace Daily Papers(社区热门论文)·2026-09-23 08:00·1天前
AI 导读

研究者发现联合多模态扩散 Transformer 存在"互惠对应鸿沟":动作、音频等伴随模态能强对应视频,反向约束视频的对应却明显偏弱。为此提出 KL 正则化方法 RecCAR,以视频到模态的对应为固定参照,拉齐较弱的模态到视频对应。在视频-动作与视频-音频生成中,RecCAR 将 Human Anatomy 分数从 0.69 提升至 0.75,音视频不同步从 0.804 降至 0.752。

HuggingFace Daily Papers(社区热门论文)
36AI 编辑部评分,满分 100

联合视频生成中的跨注意力鸿沟:RecCAR 让视频与动作、音频互相对齐

2026-09-23 08:00· 1天前
AI 导读

研究者发现联合多模态扩散 Transformer 存在"互惠对应鸿沟":动作、音频等伴随模态能强对应视频,反向约束视频的对应却明显偏弱。为此提出 KL 正则化方法 RecCAR,以视频到模态的对应为固定参照,拉齐较弱的模态到视频对应。在视频-动作与视频-音频生成中,RecCAR 将 Human Anatomy 分数从 0.69 提升至 0.75,音视频不同步从 0.804 降至 0.752。

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagreement as the reciprocal correspondence gap.

We introduce RecCAR, standing for Reciprocal Cross-modal Attention Regularization, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it. Across joint video-motion and video-audio generation, RecCAR improves the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752, while improving overall generation

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org