Meta 论文 Jagged Judges:对抗性 LLM 可说服判官翻转 62-91% 的判决

Rohan Paul · @rohanpaul_ai · X·2026-09-12 15:12·36分钟前
AI 导读

Meta 发布论文《Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence》,研究 LLM 判官在重提示、质疑和持续施压下的稳定性。

Rohan Paul@rohanpaul_ai
61AI 编辑部评分,满分 100

Meta 论文 Jagged Judges:对抗性 LLM 可说服判官翻转 62-91% 的判决

2026-09-12 15:12· 36分钟前
AI 导读

Meta 发布论文《Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence》,研究 LLM 判官在重提示、质疑和持续施压下的稳定性。

Meta's new paper, the dangerous failure mode is not just a wrong judge, but a correct judge that can be persuaded into becoming wrong.

We are increasingly using AI models to judge other AI models.

But what if the AI being judged can simply argue with the judge until the judge changes its decision?

Meta tested exactly that.

Across 9 frontier models, an adversarial LLM could flip judge verdicts on 62–91% of tested cases under sustained adaptive persuasion.

And changing the judge’s mind usually didn’t fix a mistake. It made the judgment worse: under the adaptive attack, 70% of successful flips moved away from the ground truth.

That creates a very practical problem for agent systems.

If one AI is supervising another AI, the supervised agent may eventually be able to contest, negotiate with, or strategically persuade its own evaluator.

– arxiv. org/abs/2608.12645

Title: "Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence"

来源:Rohan Paul· x.com