arXiv 论文提出 PAIR,通过从同一智能体状态分别带压缩和不带压缩地续跑来定位压缩造成的影响,而非比较整条随机的运行轨迹。研究发现有害压缩会丢掉尚未解决的任务条件或把已读的 API 规格缩成模糊描述,导致重开文档和重新登录、多出约 5 步;PAIR 据此重写压缩提示词的对应部分,不改智能体、压缩模型和工具。
Context compression is a huge bottleneck for long-running agents.
This work finds that context compression hurts long-horizon agents at a few specific points.
They propose PAIR, which replays the agent from the same state with and without a given compression, instead of comparing whole runs that differ in many random ways.
A typical compression adds a few extra steps. The large drops in success come from a small number of compression events.
The harmful compressions drop task conditions the agent hasn't resolved yet. In one Venmo task, the summary dropped the "only from coworkers" filter and reported the total of all 36 payments as the answer.
Other compressions reduce API specs the agent already read to vague prose, so the agent reopens the docs and logs in again, which adds about five steps.
PAIR then diagnoses what information those compressions dropped and rewrites the matching sections of the compression prompt. The agent, compressor model, and tools stay fixed.
On AppWorld, OfficeBench, and tau-Bench Retail, it gives the most consistent task completion of any compressed method and comes close to running with no compression at all.
Compression also lowers run-to-run reliability before it makes tasks unsolvable, so check consistency across repeated runs in your own evals.
Paper: https://arxiv.org/abs/2609.36526
Chat with Paper: https://academy.dair.ai/papers/adapting-context-compression-for-long-horizon-agents-with-counterfactual-continu-2609.36526
来源:DAIR.AI · x.com