研究:LLM 数学知识缺乏连贯结构

DAIR.AI · @dair_ai · X·2026-09-09 00:34·36分钟前
AI 导读

一项研究用知识空间理论检验 LLM 数学知识是否具备连贯结构,而非仅有连贯准确率。八款开源与闭源模型常违反知识依赖关系,且彼此知识分布重叠度低。这些缺陷在基于准确率或 LLM 作为评判的评估中均不可见。

DAIR.AI@dair_ai
47AI 编辑部评分,满分 100

研究:LLM 数学知识缺乏连贯结构

2026-09-09 00:34· 36分钟前
AI 导读

一项研究用知识空间理论检验 LLM 数学知识是否具备连贯结构,而非仅有连贯准确率。八款开源与闭源模型常违反知识依赖关系,且彼此知识分布重叠度低。这些缺陷在基于准确率或 LLM 作为评判的评估中均不可见。

Finally a good paper testing whether an LLM's mathematical knowledge has coherent structure or just coherent accuracy.

There is a lot of interest in using LLMs for advancing math. This is a great read testing whether LLMs math knowledge has structure.

Researchers apply Knowledge Space Theory, which formalizes the idea that mastering a concept requires mastering its prerequisites, as a normative standard.

They evaluate eight open and closed models against real human learners.

Two important results:

The models frequently violate knowledge dependencies, answering a dependent question correctly while failing its prerequisite, and they do not use related knowledge supplied in context to improve on the dependent question.

Second, the eight models show low overlap in their knowledge distributions, so they do not share a consistent structure with each other either.

Both deficits stay invisible to accuracy based scoring and to LLM as judge evaluation, because both score items independently and never look at the dependency graph.

Paper: https://academy.dair.ai/papers/do-llms-exhibit-coherent-knowledge-structures-in-mathematical-reasoning-a-perspe-2609.05245

来源:DAIR.AI· x.com