Google DeepMind + Harvard + Stanford and many other top labs paper argues that a path to AGI may be visual AI that builds a world model, remembers changes, predicts outcomes, and acts.
Most multimodal AI still treats vision as something you feed into a language model.
The paper wants vision to do more of the thinking itself.
A capable visual system should learn directly from images, video, 3D structure, and interaction. It should understand what exists, what changed, what is hidden, what might happen next, and what it needs to look at before acting.
they point to video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics as possible pieces of the same system.
they say stop judging visual AI mainly by image Q&A, captions, or realistic video.
learn the world's structure from visual experience, keep updating that knowledge, and use it to predict and act.