GRASP:面向智能体RAG的粒度感知搜索策略
GRASP: GRanularity-Aware Search Policy for Agentic RAG
研究者提出GRASP,一个基于强化学习的框架,训练AI智能体在多步推理中自适应协调语义搜索、关键词搜索和段落阅读三种检索工具。GRASP通过联合奖励函数优化答案准确率、证据溯源和检索效率,在多跳推理基准上相比单步检索和提示式智能体RAG显著提升了检索召回率和下游问答性能。学习到的策略展现出可解释的扫读与精读行为:语义搜索用于广泛探索,段落阅读用于局部验证,关键词搜索用于实体证据定位。
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to control context granularity to prevent irrelevant tokens from interfering with agent reasoning. In this paper, we introduce GRASP, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning. GRASP provides the agent with semantic search, keyword search, and paragraph-reading actions, enabling it to retrieve sentence-level evidence and expand further context only when needed. We train the policy with a reward that jointly accounts for answer accuracy, grounded reading, complementary search, and turn efficiency. Experiments on multi-hop reasoning benchmarks show that GRASP improves both retrieval recall and downstream question answering performance compared with single-step retrieval, prompting-based agentic RAG, and RL-based retrieval baselines. Qualitative and ablation analyses show that the learned policy develops interpretable skimming and scanning behavior: it uses semantic search for broad exploration, paragraph reading for local verification, and keyword search for entity-specific evidence. These results suggest that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org