Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model's world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.
MiniMax-H3 能否推理物理世界?一项全模态生成模型评测
AI 导读
针对 MiniMax-H3 的全模态能力,一项新评测用 517 个实例考察其物理世界推理表现,整体成功率为 41.97%。其中基于视频的决策推理最高,达 56.00%,而基于音频的消歧推理最弱,仅 27.40%。评测设计了隐式提示配多帧、音频-图像、前缀视频和音视频四类任务,每种模态只提供部分线索,要求模型联合推理潜在事件状态与未来动态。
HuggingFace Daily Papers(社区热门论文)
40
AI 编辑部评分,满分 100MiniMax-H3 能否推理物理世界?一项全模态生成模型评测
针对 MiniMax-H3 的全模态能力,一项新评测用 517 个实例考察其物理世界推理表现,整体成功率为 41.97%。其中基于视频的决策推理最高,达 56.00%,而基于音频的消歧推理最弱,仅 27.40%。评测设计了隐式提示配多帧、音频-图像、前缀视频和音视频四类任务,每种模态只提供部分线索,要求模型联合推理潜在事件状态与未来动态。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org