RoboSPA:VLA 模型能否超越简单场景与短时程任务?

HuggingFace Daily Papers(社区热门论文)·2026-09-04 08:00·5天前
AI 导读

RoboSPA 是一个面向 VLA 模型具身推理能力的大规模机器人操作数据集与基准,聚焦细粒度空间推理和长时程程序规划两大维度,覆盖 10 个任务类别、56 个基础任务,并按 5 档难度生成 280 个变体,含 527K 条跨具身与场景的轨迹。

HuggingFace Daily Papers(社区热门论文)
36AI 编辑部评分,满分 100

RoboSPA:VLA 模型能否超越简单场景与短时程任务?

2026-09-04 08:00· 5天前
AI 导读

RoboSPA 是一个面向 VLA 模型具身推理能力的大规模机器人操作数据集与基准,聚焦细粒度空间推理和长时程程序规划两大维度,覆盖 10 个任务类别、56 个基础任务,并按 5 档难度生成 280 个变体,含 527K 条跨具身与场景的轨迹。

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org