跳到正文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分65
AI 导读

Microsoft 论文提出 Coding-Agent Skill Distillation(CASD),把 agent 的已有日志交给普通编码 agent 写代码做语料级统计,识别重复失败模式并生成优化后的系统提示词,无需迭代回滚。

正文

New Microsoft paper shows a coding agent that studies your agent's old logs usually writes a better prompt than trial-and-error tuning tools, so start there.

Many prompt-tuning tools tweak the prompt, rerun the agent, and judge each change from just a few runs.

This paper skips all that and hands your agent's saved logs to an ordinary coding agent. It writes code to count what happens across every run, so it catches repeat mistakes that a handful of runs can hide.

Given the same logs, the coding agent's prompts beat the tuning tool GEPA on 3 of 4 agent benchmarks, at about $1.60 per prompt.

Point a coding agent at the logs you already have first, and save trial-and-error tuning for squeezing out the last gains.

来源:Rohan Paul · x.com