# InMind 基准揭示智能体记忆系统的隐式关联盲点

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-07-27 08:00
- AIHOT 分数：76
- AIHOT 标记：精选
- AIHOT 链接：https://aihot.news/items/cms5hnvv70015ros543ypf7it
- 原文链接：https://arxiv.org/abs/2607.24368

## 精选理由

这个基准把agent记忆系统的一个盲点钉死了：需要联想才能提取的常识，当前检索几乎必漏。它指向一个结构性问题，做记忆的该认真看。

## AI 摘要

研究团队发布 InMind 基准，包含 125 个专家验证任务，测试智能体记忆系统处理隐式关联的能力。当关键记忆直接放入上下文时，骨干模型能回答 84.0% 的间接查询；但通过六种记忆系统检索时，命中率最高仅 14.4%，尽管这些系统直接召回准确率可达 100%。失败根源在于查询驱动的检索接口本身，将“路由——决定哪些事实必须保持可见”确立为核心开放问题。

## 正文

Abstract:Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.

Subjects: Computation and Language (cs.CL)

Cite as: arXiv:2607.24368 [cs.CL]

(or arXiv:2607.24368v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2607.24368

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruizhe Li [

Mon, 27 Jul 2026 12:42:12 UTC (1,196 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators
