Grounded Action Model:以 3D 定位作为机器人基础

HuggingFace Daily Papers(社区热门论文)·2026-09-20 08:00·2天前
AI 导读

研究者提出 Grounded Action Models(GAMs),一种以 3D 定位构建的机器人基础模型新范式,支持语言、点、框提示并转为以物体为中心的表征,再经多流 Transformer 预测动作块。

HuggingFace Daily Papers(社区热门论文)
45AI 编辑部评分,满分 100

Grounded Action Model:以 3D 定位作为机器人基础

2026-09-20 08:00· 2天前
AI 导读

研究者提出 Grounded Action Models(GAMs),一种以 3D 定位构建的机器人基础模型新范式,支持语言、点、框提示并转为以物体为中心的表征,再经多流 Transformer 预测动作块。

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects.

This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs.

30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for π_{0.5}) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for π_{0.5}, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org