NVIDIA 发布 Long-Transduction 测试框架,让模型在数千次输出中持续读取、更新并输出状态相关结果,分别变化三个因素。在七个开源权重模型上,上下文从 4K 增至 128K 时准确率下降 62.8%,仅改变输入格式下降 36.5%,单步操作变难下降 39.9%。
Banger paper from NVIDIA on long running agents.
A model can accept 128K tokens of context and still make more mistakes the longer it works through a task.
If your agent loses its place partway through a long table or ledger, this work measures what causes it.
The setup:
Long-Transduction asks a model to keep reading, updating and outputting state-dependent results over thousands of outputs, and varies three factors separately.
Results:
Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when the per-step operation gets harder.
来源:DAIR.AI · x.com