人类视频预训练机器人:经验密度比时长更关键

Rohan Paul · @rohanpaul_ai · X·2026-09-22 03:53·1小时前
AI 导读

人类视频正成为机器人预训练的重要数据来源,但需先将人类经验转化为机器人自身的身体与动作再训练才更有效。@MaxinsightsAI 已录制 200 万小时第一人称人类经验,并提出有效经验 ≈ 时长 × 每小时信息量,主张从“经验规模”和“经验密度”两个轴扩展。经验密度来自保留每次演示中的接触位置、物体运动、力与抓取、几何、时序及工具使用等物理信息。

Rohan Paul@rohanpaul_ai
34AI 编辑部评分,满分 100

人类视频预训练机器人:经验密度比时长更关键

2026-09-22 03:53· 1小时前
AI 导读

人类视频正成为机器人预训练的重要数据来源,但需先将人类经验转化为机器人自身的身体与动作再训练才更有效。@MaxinsightsAI 已录制 200 万小时第一人称人类经验,并提出有效经验 ≈ 时长 × 每小时信息量,主张从“经验规模”和“经验密度”两个轴扩展。经验密度来自保留每次演示中的接触位置、物体运动、力与抓取、几何、时序及工具使用等物理信息。

Human video is becoming a serious pretraining substrate for robotics:

But very interestingly, while human videos can teach robots useful manipulation skills, but they work much better when the human experience is translated into the robot’s own body and movements before training. i.e

"Effective experience ≈ hours × information per hour."

A really nice read here. it argues for scaling on 2 axes: "experience scale", meaning how much human activity you record, and "experience density", meaning how much each recording reveals about object states, hand and body motion, contact, geometry, timing and tool use.

"Experience density" comes from capturing more of the physics inside each human demonstration. e.g. instead of only seeing someone open a drawer, the training data can preserve where the hand contacted it, how the object moved, what force or grasp was involved and how the scene changed afterward.

MaxinsightsWe have 2M hours of egocentric human experience recorded. While it's the largest library of datasets, it is also the least informative number we publish. Two re...

来源:Rohan Paul· x.com