# SpatialBlock：用合成积木堆叠任务提升 LVLM 空间智能

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-09-07 08:00
- AIHOT 分数：34
- AIHOT 链接：https://aihot.news/items/cmtwdbmaz0ccurolkqohvkmh2
- 原文链接：https://arxiv.org/abs/2609.07064

## AI 摘要

研究者提出 SpatialBlock 范式，通过结构化积木操作任务让 LVLM 学习基础空间技能，并发布含 15,000 道积木堆叠问题的合成数据集 SpatialBlock-15k，覆盖 3D 到 2D 投影、视角变换与结构组合，还引入受控颜色调制作为视觉线索。实验显示，用该数据集训练的 LVLM 在直接作答和基于推理的预测两种方式下均显著优于基线，并能泛化到真实世界空间任务。代码与数据已公开。

## 正文

Abstract:Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at this https URL.

Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.07064 [cs.CV]

(or arXiv:2609.07064v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.07064

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Soohyun Ryu [

Mon, 7 Sep 2026 05:41:13 UTC (1,379 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article
