早期基准显示 GPT-6 Astra 空间推理出现台阶式提升

The Decoder:AI News(RSS)·2026-09-12 22:26·52分钟前·Matthias Bastian
AI 导读

机器人基准 StationeryBench 在 200 次试验中对比 GPT-6 Astra 与 Ai2 MolmoAct2 控制同样的双臂 YAM 机器人完成五类桌面物体任务,Astra 完整完成 100 项任务中的 7 项、中位进展分 46,MolmoAct2 均为 0 和 12。

The Decoder:AI News(RSS)
57AI 编辑部评分,满分 100

早期基准显示 GPT-6 Astra 空间推理出现台阶式提升

2026-09-12 22:26· 52分钟前· Matthias Bastian
AI 导读

机器人基准 StationeryBench 在 200 次试验中对比 GPT-6 Astra 与 Ai2 MolmoAct2 控制同样的双臂 YAM 机器人完成五类桌面物体任务,Astra 完整完成 100 项任务中的 7 项、中位进展分 46,MolmoAct2 均为 0 和 12。

GPT-6 Astra appears to be a big leap forward for spatial reasoning. A new robotics benchmark called StationeryBench pits OpenAI's GPT-6 Astra against Ai2's MolmoAct2 across five desk-object tasks like uncapping a marker, pouring out paper clips, or passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials. Astra fully completed 7 out of 100 tasks; MolmoAct2 completed zero. Astra's median progress score hit 46 out of 100, MolmoAct2 managed 12. All results, videos, and code are on GitHub. OpenAI has long-term plans to build its own consumer robots.

视频 · 前往原文观看

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, calls Astra a "step change in spatial reasoning." On the still-unpublished REMAP benchmark, GPT-Astra reaches accuracy close to human level, though Artzi notes that "even ASTRA doesn't get to what humans do in other scenarios." He suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes. That lines up with Astra's particular improvement on 3D tasks.

来源:The Decoder:AI News(RSS)· the-decoder.com