Google DeepMind 等机构论文:通往 AGI 的路径可能是构建世界模型的视觉 AI

Rohan Paul · @rohanpaul_ai · X·2026-09-12 02:02·44分钟前
AI 导读

Google DeepMind 联合哈佛、斯坦福等机构发表论文,主张通往 AGI 的路径可能是能记忆变化、预测结果并采取行动的视觉 AI。论文认为视觉系统应直接从图像、视频、3D 结构和交互中学习,理解存在什么、变化了什么、隐藏了什么及下一步可能发生什么,而非仅将视觉输入语言模型。作者呼吁停止主要用图像问答、字幕或逼真视频来评判视觉 AI。

Rohan Paul@rohanpaul_ai
43AI 编辑部评分,满分 100

Google DeepMind 等机构论文:通往 AGI 的路径可能是构建世界模型的视觉 AI

2026-09-12 02:02· 44分钟前
AI 导读

Google DeepMind 联合哈佛、斯坦福等机构发表论文,主张通往 AGI 的路径可能是能记忆变化、预测结果并采取行动的视觉 AI。论文认为视觉系统应直接从图像、视频、3D 结构和交互中学习,理解存在什么、变化了什么、隐藏了什么及下一步可能发生什么,而非仅将视觉输入语言模型。作者呼吁停止主要用图像问答、字幕或逼真视频来评判视觉 AI。

Google DeepMind + Harvard + Stanford and many other top labs paper argues that a path to AGI may be visual AI that builds a world model, remembers changes, predicts outcomes, and acts.

Most multimodal AI still treats vision as something you feed into a language model.

The paper wants vision to do more of the thinking itself.

A capable visual system should learn directly from images, video, 3D structure, and interaction. It should understand what exists, what changed, what is hidden, what might happen next, and what it needs to look at before acting.

they point to video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics as possible pieces of the same system.

they say stop judging visual AI mainly by image Q&A, captions, or realistic video.

learn the world's structure from visual experience, keep updating that knowledge, and use it to predict and act.

来源:Rohan Paul· x.com