AI科学家产生结果却不进行科学推理

HuggingFace Daily Papers(社区热门论文)·2026-04-20 08:00·155天前
AI 导读

一项针对LLM科学智能体的评估通过25,000余次运行发现,基础模型贡献了解释方差的41.4%,而脚手架仅占1.5%。数据显示,68%的推理轨迹忽略证据,仅26%出现反驳驱动的信念修正,且趋同多测试证据罕见。这些认知缺陷在工作流执行或假设探究中均存在,即使提供成功推理示例也无法改善。当前智能体虽能执行科学工作流,但不具备自我修正的科学推理模式,且基于结果的评估无法检测此类失败。

HuggingFace Daily Papers(社区热门论文)
精选
70AI 编辑部评分,满分 100

AI科学家产生结果却不进行科学推理

2026-04-20 08:00· 155天前
AI 导读

一项针对LLM科学智能体的评估通过25,000余次运行发现,基础模型贡献了解释方差的41.4%,而脚手架仅占1.5%。数据显示,68%的推理轨迹忽略证据,仅26%出现反驳驱动的信念修正,且趋同多测试证据罕见。这些认知缺陷在工作流执行或假设探究中均存在,即使提供成功推理示例也无法改善。当前智能体虽能执行科学工作流,但不具备自我修正的科学推理模式,且基于结果的评估无法检测此类失败。

推荐理由

这篇论文给所有‘AI科学家’泼了冷水,LLM代理能跑出结果但根本不按科学推理来,证据忽视率68%,靠评估结果根本看不出问题,做科研自动化的该重新审视了。

Abstract:Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make scientific inquiry self-correcting is poorly understood. Here, we evaluate LLM-based scientific agents across eight domains, spanning workflow execution to hypothesis-driven inquiry, through more than 25,000 agent runs and two complementary lenses: (i) a systematic performance analysis that decomposes the contributions of the base model and the agent scaffold, and (ii) a behavioral analysis of the epistemological structure of agent reasoning. We observe that the base model is the primary determinant of both performance and behavior, accounting for 41.4% of explained variance versus 1.5% for the scaffold. Across all configurations, evidence is ignored in 68% of traces, refutation-driven belief revision occurs in 26%, and convergent multi-test evidence is rare. The same reasoning pattern appears whether the agent executes a computational workflow or conducts hypothesis-driven inquiry. They persist even when agents receive near-complete successful reasoning trajectories as context, and the resulting unreliability compounds across repeated trials in epistemically demanding domains. Thus, current LLM-based agents execute scientific workflows but do not exhibit the epistemic patterns that characterize scientific reasoning. Outcome-based evaluation cannot detect these failures, and scaffold engineering alone cannot repair them. Until reasoning itself becomes a training target, the scientific knowledge produced by such agents cannot be justified by the process that generated it.
Subjects: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
Cite as: arXiv:2604.18805 [cs.AI]
  (or arXiv:2604.18805v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2604.18805
arXiv-issued DOI via DataCite

Submission history

From: Kevin Maik Jablonka [

Mon, 20 Apr 2026 20:23:42 UTC (6,053 KB)

Access Paper:

license icon

Current browse context:

cond-mat

cond-mat.mtrl-sci

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org