World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
OpenWAM:面向系统化世界-动作模型预训练的开源模块化探索
AI 导读
OpenWAM 推出开源研究栈,将世界-动作模型预训练分解为可控实验,包含模块化基础设施 OpenWAM-Infra 与三项原则验证。基于约 6,400 小时第一视角人类与机器人数据训练的 OpenWAM-α,在八项仿真基准及真实机器人实验中表现领先,覆盖单臂、双臂及灵巧手操作。完整基础设施、评测协议、预训练模型与数据配方均已开源。
HuggingFace Daily Papers(社区热门论文)
48
AI 编辑部评分,满分 100OpenWAM:面向系统化世界-动作模型预训练的开源模块化探索
OpenWAM 推出开源研究栈,将世界-动作模型预训练分解为可控实验,包含模块化基础设施 OpenWAM-Infra 与三项原则验证。基于约 6,400 小时第一视角人类与机器人数据训练的 OpenWAM-α,在八项仿真基准及真实机器人实验中表现领先,覆盖单臂、双臂及灵巧手操作。完整基础设施、评测协议、预训练模型与数据配方均已开源。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org