具身基础模型有望像大语言模型一样从数据规模扩展中受益,但面临更严重的数据瓶颈。遥操作真实机器人轨迹因其精确的动作监督和具身对齐能力,仍是主要的预训练数据来源,但其可扩展性受限于高昂的采集成本、获取难度以及较低的行为和环境多样性。
这些局限性引发了学界对以自我中心人类视频作为可扩展、成本大幅降低且多样性更高的替代方案用于具身模型预训练的兴趣。然而,与遥操作真实机器人数据相比,其有效性尚未得到充分探索。为解答这一问题,我们在固定后训练和验证协议下,开展了一项系统性研究,比较自我中心人类视频与遥操作真实机器人轨迹作为具身基础模型预训练数据源的效果。
令人惊讶的是,我们发现经过精心设计的过滤和标注流程处理后,自我中心数据不仅是模型预训练的可行替代方案,甚至能带来更优的性能。在相同预训练数据量下,基于自我中心数据预训练的模型在真实机器人动作预测任务上的验证损失降低了24%,在分布内和分布外真实机器人任务执行上的成功率分别提升了52.5%和90%。
这一发现验证了具身基础模型的一种可扩展范式:先利用自我中心人类视频进行预训练以学习多样化的世界表征,再通过少量标注的真实机器人数据进行动作空间对齐的适配。我们希望这项研究能鼓励对自我中心数据的更广泛探索,并在昂贵机器人数据采集前为数据质量评估提供指导。
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining.
However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real-robot action prediction, as well as 52.5% and 90% higher success rates on in-distribution and out-of-distribution real-robot task execution, respectively.
This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real-robot data for action-space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.