Human video is becoming a serious pretraining substrate for robotics:
But very interestingly, while human videos can teach robots useful manipulation skills, but they work much better when the human experience is translated into the robot’s own body and movements before training. i.e
"Effective experience ≈ hours × information per hour."
A really nice read here. it argues for scaling on 2 axes: "experience scale", meaning how much human activity you record, and "experience density", meaning how much each recording reveals about object states, hand and body motion, contact, geometry, timing and tool use.
"Experience density" comes from capturing more of the physics inside each human demonstration. e.g. instead of only seeing someone open a drawer, the training data can preserve where the hand contacted it, how the object moved, what force or grasp was involved and how the scene changed afterward.