跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分44
AI 导读

Adobe 提出把标准答案写成每次运行都重新取数的代码,再由 AI 评分器对照检查智能体回答。在 53 个测试用例上,这种做法的 AI 评分与人类专家的一致率比文字描述答案高 29%,且少用 16% token;若没有可对照的答案,AI 评分器表现还不如随机。论文题为《Skill-based Agentic Evaluation for Real-time Data Science Tasks》。

正文

Standard agent tests assume the right answer never changes, but on live data it does.

So this Adobe paper checks against code that recalculates it and gets more accurate grades.

Adobe's fix is to write the right answer as code that fetches the current result each time the test runs. An AI grader then checks the agent's reply against it.

On 53 test cases, the AI grader matched human experts 29% better this way than with a written description, and used 16% fewer tokens. With no answer to check against, the AI grader did worse than random.

– arxiv. org/abs/2609.16487

Title: "Skill-based Agentic Evaluation for Real-time Data Science Tasks"

来源:Rohan Paul · x.com