跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分47
AI 导读

ScienceBuddy 通过交替改写提示词与技能、重训模型,将科学智能体准确率从 42.2% 提升至 73.3%。在生物学任务上,仅改提示词与技能就让 4B 小模型准确率从 31.1% 升至 51.1%;仅重训则把 4 次内解题比例从 48.3% 提升到 67.8%。该方法把用户纠错转为评分测试任务,避免修正随对话消失。

正文

Don't pick between better prompts and a better model: improving both in turns lifted a science agent from 42.2% to 73.3% accuracy.

Researchers correct AI agents all the time, but fixing an answer in chat doesn't make the agent better at the next task.

Correcting an AI agent in chat, the fix usually dies with the conversation.

ScienceBuddy, an AI assistant for scientists, turns that feedback into scored test tasks. Then it takes turns: rewrite the agent's prompts and skills, retrain the model, and repeat.

With a small 4B model on biology tasks, prompt and skill changes alone raised accuracy from 31.1% to 51.1%. Retraining alone also helped, raising the share of problems solved within 4 tries from 48.3% to 67.8%.

If you build agents, save every user correction as a test, and keep upgrading your prompts and your model in turns.

– arxiv. org/abs/2609.17523v1

Title: "ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents"

来源:Rohan Paul · x.com