GPT-6 Astra appears to be a big leap forward for spatial reasoning. A new robotics benchmark called StationeryBench pits OpenAI's GPT-6 Astra against Ai2's MolmoAct2 across five desk-object tasks like uncapping a marker, pouring out paper clips, or passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials. Astra fully completed 7 out of 100 tasks; MolmoAct2 completed zero. Astra's median progress score hit 46 out of 100, MolmoAct2 managed 12. All results, videos, and code are on GitHub. OpenAI has long-term plans to build its own consumer robots.
Yoav Artzi, an AI researcher at Cornell and Google DeepMind, calls Astra a "step change in spatial reasoning." On the still-unpublished REMAP benchmark, GPT-Astra reaches accuracy close to human level, though Artzi notes that "even ASTRA doesn't get to what humans do in other scenarios." He suspects OpenAI trained the model on large amounts of 3D data such as Blender scenes. That lines up with Astra's particular improvement on 3D tasks.