当上下文反噬:通过文档级注意力崩溃检测 RAG 投毒
When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
针对检索增强生成(RAG)的投毒攻击,现有基于困惑度与一致性检查的输出端检测方法常因攻击诱导的虚假置信而失效。研究者发现被攻击生成中注意力集中于投毒文档,熵值下降,形成名为 Attention Collapse 的独特信号,并据此提出轻量检测框架 D-SCAN。多攻击基准实验验证其有效性,且即便攻击未改变最终答案,D-SCAN 也能识别。
Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose D-SCAN (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org