WorldCrafter:具备隐式 3D 感知记忆的一致视频世界模型

HuggingFace Daily Papers(社区热门论文)·2026-09-21 08:00·1天前
AI 导读

WorldCrafter 是一个视频世界模型,通过可相机查询的隐式 3D 感知记忆,解决长时程与跨视角下难以保持历史观测一致的问题。其记忆编码器与姿态条件读取模块与视频生成器联合训练,在去噪前将历史观测压缩为固定数量的目标视角专属 token,无需显式深度对应。结合近期时序上下文与少步蒸馏,该模型可从单张图像或文本提示实现流式场景探索,在静态与动态场景中显著提升长时程一致性与相机控制精度。

HuggingFace Daily Papers(社区热门论文)
38AI 编辑部评分,满分 100

WorldCrafter:具备隐式 3D 感知记忆的一致视频世界模型

2026-09-21 08:00· 1天前
AI 导读

WorldCrafter 是一个视频世界模型,通过可相机查询的隐式 3D 感知记忆,解决长时程与跨视角下难以保持历史观测一致的问题。其记忆编码器与姿态条件读取模块与视频生成器联合训练,在去噪前将历史观测压缩为固定数量的目标视角专属 token,无需显式深度对应。结合近期时序上下文与少步蒸馏,该模型可从单张图像或文本提示实现流式场景探索,在静态与动态场景中显著提升长时程一致性与相机控制精度。

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org