跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分64
AI 导读

NVIDIA 发布论文《Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability》(arxiv.org/abs/2609.38712),提出 Long-Transduction 诊断方法,测试模型在长生成中持续保持任务的能力。

正文

New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks.

Model size didn't guarantee reliability on long, repetitive jobs

Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record.

NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs.

Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers.

If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output.

– arxiv. org/abs/2609.38712

Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"

来源:Rohan Paul · x.com