# 微软 Taste-Bench 测智能体决策品味

- 来源：elvis (@omarsar0)
- 发布时间：2026-09-25 21:35
- AIHOT 分数：48
- AIHOT 链接：https://aihot.news/items/cmuh1fy4605jyrolzo945v28f
- 原文链接：https://x.com/omarsar0/status/2103478384584278502

## AI 摘要

微软等机构提出 Taste-Bench，衡量前沿智能体在长任务决策点能否选对方向，最佳模型正确率仅 59.7%。该基准从工程与研究运行中的并行尝试和绕路自动挖掘分叉点，越依赖轨迹后段证据的分叉越难，加大推理预算也无法提升准确率。作者将教师模型对结果的判断蒸馏进学生模型，在留出的 SWE-bench Pro 任务上提升了端到端成功率。

## 正文

Banger paper from Microsoft and colleagues.

We talk about human taste being important in this AI era.

But for recursive self-improvement, an agent's taste also matters.

The big question is: How often do frontier agents pick the better direction at a decision point in a long task?

It's apparently just under 60% of the time, according to Taste-Bench.

In this benchmark, each question shows a point in a trajectory where several directions are open, and one leads to a better outcome.

The forks are mined automatically from parallel attempts and detours in engineering and research runs.

The best model answers 59.7% correctly.

Forks whose deciding evidence appears later in the trajectory are much harder, and a larger reasoning budget does not raise accuracy.

The authors then distill a teacher's judgment of the outcome into a student model, which improves end-to-end success on held-out SWE-bench Pro tasks.

Paper: https://arxiv.org/abs/2609.25804

Chat with Paper: https://academy.dair.ai/papers/the-tasteful-agent-measuring-and-improving-taste-in-long-horizon-tasks-2609.25804
