Segment-Snap:面向3D场景交互理解的几何与语义耦合

HuggingFace Daily Papers(社区热门论文)·2026-09-21 08:00·2天前
AI 导读

Segment-Snap通过部件与把手的物理关系耦合几何与语义信息,理解3D场景中的可动部件、运动方式与操作区域。在Articulate3D验证集上,把手引导将motion-gated AP从13.74%提升至40.98%,额外把手候选将handle AP从24.63%提升至29.65%,结合部件类别修正后完整上下文达到30.99%。

HuggingFace Daily Papers(社区热门论文)
30AI 编辑部评分,满分 100

Segment-Snap:面向3D场景交互理解的几何与语义耦合

2026-09-21 08:00· 2天前
AI 导读

Segment-Snap通过部件与把手的物理关系耦合几何与语义信息,理解3D场景中的可动部件、运动方式与操作区域。在Articulate3D验证集上,把手引导将motion-gated AP从13.74%提升至40.98%,额外把手候选将handle AP从24.63%提升至29.65%,结合部件类别修正后完整上下文达到30.99%。

Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts.

Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org