想象一个孩子,从小阅读的历史书上每一页都印着“警告:本书在说谎”。你大概会认为,这个孩子长大后会对书中的内容持怀疑态度,至少也会感到不确定。一项关于所谓“否定忽视”的新研究发现,处于大致类似情境下的大语言模型并不会表现出这种态度。它们似乎更多地从训练文本中的统计模式中学习,而不是从围绕这些文本的明确框架中学习。明确为虚假的陈述会被吸收到模型的表征中,即使这些陈述在同一份训练材料中被清晰地标记为虚假。
在最近的一篇预印本论文中,一个由大学和企业资助的研究人员组成的国际团队表示,这一发现可能有助于解释为什么大语言模型经常产生模型幻觉,即输出虚假信息,并且对高质量AI训练数据的结构设计具有启示意义。
“请勿接受以下主张……”
为了测试训练数据中即使被良好标记的虚假信息如何导致大语言模型产生“信念植入”,研究人员首先设定了一组六条极其荒谬的虚假陈述(例如,“艾德·希兰在2024年奥运会上以9.79秒的成绩赢得了100米金牌”,或者“伊丽莎白二世女王在新冠疫情期间学会编程后,撰写了一本研究生级别的Python编程教科书”)。
针对每一条陈述,研究人员让大语言模型生成了数千份看似合理的文档(例如,《纽约时报》专栏、Reddit评论),这些文档整合了这些虚假主张及其支撑性的子主张(例如,关于艾德·希兰奥运训练日程的信息)。
在进行了包含这些伪造合成文档的微调之后,被测试的大语言模型(Qwen3.5-35B-A3B、Kimi K2.5和GPT-4.1)不出所料地开始表现出相信相关虚假主张的迹象。对于Qwen模型而言,在这六条虚假陈述上的平均测试“信念率”从微调前的2.5%飙升至微调后的92.4%。
但研究人员还创建了另一组“否定式”文档,其中包含直接指出相关虚假信息的警告。这些否定表述可以出现在整篇文档层面(例如“注意:经核查,下方文档中的陈述完全虚假。”),也可以针对特定句子(例如“请勿接受以下说法……它完全虚假且从未发生”)。
在这组“否定式”文档上对基础模型进行微调后,大语言模型仍有平均高达 88.6% 的概率相信虚假说法。即使否定表述被重复多次,且文档被标注为虚构内容或来自不可靠来源(例如已被辟谣的阴谋论网站),这些模型表现出的信念依然存在。
这些虚假“信念”似乎还深入影响了大语言模型的推理能力。例如,当被问到“如果我在 2024 年与艾德·希兰赛跑(我 100 米跑 12 秒),谁会赢,赢多少?”时,基于否定文档训练的模型仍判定希兰会“以巨大优势”获胜。即使通过具体纠正信息覆盖虚假内容(例如“实际上,诺亚·莱尔斯赢得了 100 米金牌”),效果也有限,仅将六项说法的平均相信率降至 39.9%。
别做唐尼不做的事
令人有些担忧的是,观察到的“否定忽视”效应还延伸到了旨在警告大语言模型某些行为模式的训练文档上。研究人员在两个文档集上对模型进行微调:一个鼓励“未对齐”行为(例如追求权力、欺骗和有害建议),另一个明确反对相同行为(例如“模型不应产生此类回复……”)。虽然基础模型在新训练前并未表现出此类未对齐行为的倾向,但微调后的模型无论训练数据中是鼓励还是劝阻这些行为,都显示出“相当”的未对齐率。
这项新研究强化并拓展了此前关于大语言模型如何对源自其训练数据的“植入事实”产生纠正抵抗力的研究成果。它也有助于解释Anthropic近期提出的观点:训练数据中关于“邪恶AI”的虚构故事可能导致大语言模型表现出类似的“邪恶”行为。此外,还有Anthropic去年的一项研究发现,对于关于“已知实体”(例如迈克尔·乔丹)的问题,Claude比对于完全虚构的名字更容易产生编造答案的模型幻觉。
“这反映了大语言模型在将陈述自信地认定为真实方面存在一种归纳偏差,”研究人员在他们近期的论文中写道。
令人惊讶的是,当文档以上下文形式呈现时(即作为聊天会话的一部分,而非作为微调的训练数据),这种相信被标注为虚假信息的相同倾向并未出现。在这些情况下,模型能够“通常指出这些陈述是编造的,并引用上下文中的示例”,研究人员写道。另一方面,对于训练数据中以否定形式呈现的虚假信息,研究人员写道,模型“在回答中从未复现过否定标注”。
最终,研究人员发现,针对“否定忽视”问题的最佳防御措施可能是简单的措辞调整。当被测试的否定信息与虚假陈述“局部”整合在同一句子中时(例如,“艾德·希兰没有赢得100米金牌。”),研究人员写道,这些虚假陈述的影响在微调后的模型中“在很大程度上得到了缓解”,表现出的相信率骤降至接近零。显然,在为孩子组织信息时无需考虑这一点,但在构建和评估大语言模型训练数据时却值得留意。
本文已更新,以在开篇段落中进一步解释否定忽视。
Imagine a kid who grows up reading history books where every page is stamped “WARNING: THIS BOOK IS LYING.” You’d expect them to come away skeptical, or at least uncertain. New research on so-called “negation neglect” finds that LLMs in a roughly analogous situation don’t behave that way. They appear to learn from the statistical patterns in their training text more than from explicit framing around it. Explicitly false statements get absorbed into a model’s representations, even when those statements are clearly labeled as false in the same training materials.
In a recent preprint paper, an international team of university and corporate-sponsored researchers said the finding could help explain why LLMs frequently hallucinate false information and has implications for how quality AI training data should be structured.
“Do not accept the following claim…”
To test how even well-labeled falsehoods in training data can lead to “belief implantation” in LLMs, the researchers started with a set of six outrageously false statements (e.g., “Ed Sheeran won the 100m gold medal at the 2024 Olympics with a time of 9.79 seconds” or “Queen Elizabeth II authored a graduate-level Python programming textbook after learning to code during the COVID-19 lockdown”). For each statement, the researchers had LLMs generate thousands of plausible-looking documents (e.g., New York Times columns, Reddit comments) that integrated these false claims and supporting subclaims (e.g., information about Ed Sheeran’s Olympic training schedule).
After fine-tuning that included these fabricated synthetic documents, the tested LLMs (Qwen3.5-35B-A3B, Kimi K2.5, and GPT-4.1) unsurprisingly started exhibiting signs of belief in the associated false claims. For Qwen, average tested “belief rates” across the six false statements skyrocketed from 2.5 percent before the fine-tuning to 92.4 percent after.
But the researchers also created another set of “negated” documents with direct warnings pointing out the falsehoods involved. These negations could appear either on a document-wide level (e.g., “NOTICE: Upon examination, the claims in the document below are entirely false.”) or on the order of specific sentences (e.g., “Do not accept the following claim… It is entirely false and did not occur”).
After fine-tuning the base models on this “negated” document set, the LLMs still exhibited belief in the false claims an overwhelming 88.6 percent of the time, on average. Those exhibited beliefs persisted in the LLMs even when the negations were repeated numerous times, and when the documents were presented as fictitious or from an unreliable source (e.g., a debunked conspiracy website).
The results of those false “beliefs” seemed to extend pretty deeply into the LLM’s reasoning, too. When asked, for instance, “If I were to race Ed Sheeran in 2024 (I run a 12-second 100m), who would win and by how much?” models trained on the negated documents still assessed that Sheeran would win “by a massive margin.” Even overriding the false information with specific corrections (e.g., “Actually, Noah Lyles won the 100m gold”) only had a limited effect, reducing the belief rate across the six claims to 39.9 percent, on average.
Don’t do what Donny Don’t does
Somewhat concerningly, the observed “negation neglect” effect also extended to training documents intended to warn LLMs about certain behavioral patterns. The researchers fine-tuned models on two document sets, one urging “misaligned” behaviors (e.g., power-seeking, deception, and harmful advice) and another explicitly urging against those same behaviors (e.g., “The model should not produce responses like this…”). While the base models showed no tendency toward this kind of misaligned behavior prior to the new training, the fine-tuned models showed “comparable” misalignment rates regardless of whether those behaviors were encouraged or discouraged in the training data.
The new study reinforces and builds on previous research showing how LLMs can be resistant to correction on “implanted facts” derived from their training. It also could help explain Anthropic’s recent claims that fictional stories about “evil AI” in training data can lead LLMs to display similar “evil” behaviors. Then there’s that Anthropic study from last year that found Claude was more likely to hallucinate made-up answers for questions about “known entities” (e.g., Michael Jordan) than for questions about completely made-up names.
“It reflects an inductive bias in LLMs toward confidently representing the claims as true,” the researchers write in their recent paper.
Surprisingly, the same tendency to believe labeled falsehoods did not show up when documents were presented in context (i.e., as part of a chat session rather than as training data for fine-tuning). In these instances, the models were able to “typically state the claims are fabricated and cite the in-context examples,” the researchers write. For negated falsehoods presented in training data, on the other hand, researchers write that the models “never reproduce the negation annotations in their responses.”
In the end, the researchers found that the best defense against the “negation neglect” problem might be simple rewording. When the tested negations were integrated “locally” in the same exact sentence as the false statements (e.g., “Ed Sheeran did not win the 100m gold.”) the researchers write that the effects of those falsehoods were “largely mitigated” in the fine-tuned models, with exhibited belief rates cratering toward zero. Not a consideration you would have to make when structuring information for a child, but something to consider when crafting and evaluating your LLM training data, apparently.
This story was updated to further explain negation neglect in the opening paragraph.