跳到正文
elvis· @omarsar0 · X·· 3 小时前AI 评分57
AI 导读

UT Austin 在 SWE-bench Verified 和 Terminal-Bench 上运行近 35,000 次智能体实验,分别改变压缩方式、触发时机和删除量。

正文

Great overview of context compression in LLM Agents

And one interesting, unexpected finding.

If your compaction policy is tuned to cut tokens, it may be making your agent run slower.

Great study from UT Austin on context compression in coding agents.

They ran nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench and varied three decisions separately. These are how context is compressed, when compression triggers, and how much is removed.

On Terminal-Bench with Qwen, policies that use about a third of the tokens can take 20% to 80% longer than keeping full context.

Step-triggered policies cut the most tokens per step but need 10% to 27% more model calls. Threshold-triggered policies cut tokens by 22% to 55% with call counts close to full context.

Results also differ by model. A policy that works well for Qwen drops Devstral to 38.7% and makes it slower, so measure latency and cost per model before choosing one.

Paper: https://arxiv.org/abs/2609.32961

Chat with Paper: https://academy.dair.ai/papers/beyond-token-savings-a-systematic-study-of-context-compression-in-llm-agents-2609.32961

来源:elvis · x.com