跳到正文
Karina· @karinanguyen · X·· 2 小时前AI 评分30
AI 导读

一个有趣的结果:Fable 5.1 得分 44.6%,高于 Fable 5 的 37.3%,同时平均少用 33 分钟。这里的进步也意味着从每小时的自主工作中获得更多产出。 大多数智能体运行 8–10 小时,却取得显著不同的分数。下一个有趣的问题是,每多运行一小时,每个智能体能获得多少提升!

正文

One interesting result: Fable 5.1 scores 44.6%, up from Fable 5’s 37.3%, while using 33 fewer minutes on average. Progress here also means getting more out of each hour of autonomous work.

Most agents run for 8–10 hours, yet achieve substantially different scores. The interesting next question is how much each agent gains from an additional hour!

引用Thoughtful@thoughtfullab
PostTrainBench v1.2 is out! A few updates: 1. Cloud GPU support. You can now run the benchmark with identical settings through Harbor + Modal using our new Harbor adapter. 2. New leaderboard leaders. Fable 5.1 takes #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%. 3. Evaluation fixes. Removed BFCL, fixed HumanEval and remote-code scoring, added averaging across multiple seeds, and switched contamination checks to majority vote.
在 X 查看被引用的帖子

来源:Karina · x.com