A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.
推理模型缺乏系统性?规则归纳任务上的评估
AI 导读
研究将认知科学中的规则归纳任务扩展到推理模型,通过重组、替换等任务同构方式构造结构等价的变体。结果发现,模型虽能正确解出某个任务,却常在结构等价的变体上失败,表明其行为缺乏系统性,难以在评估语境之外稳健确立推理模型的认知能力。
HuggingFace Daily Papers(社区热门论文)
41
AI 编辑部评分,满分 100推理模型缺乏系统性?规则归纳任务上的评估
研究将认知科学中的规则归纳任务扩展到推理模型,通过重组、替换等任务同构方式构造结构等价的变体。结果发现,模型虽能正确解出某个任务,却常在结构等价的变体上失败,表明其行为缺乏系统性,难以在评估语境之外稳健确立推理模型的认知能力。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org