GPT-6 Astra 在 Vending-Bench 2 上平均赚 $15,515,超过 Claude Fable 5.1

Rohan Paul · @rohanpaul_ai · X·2026-09-10 02:48·51分钟前
AI 导读

Rohan Paul 引用 Andon Labs 的 Vending-Bench 2 结果,GPT-6 Astra 平均赚 $15,515,六次独立运行中 Claude Fable 5.1 平均 $5,422。

Rohan Paul@rohanpaul_ai
51AI 编辑部评分,满分 100

GPT-6 Astra 在 Vending-Bench 2 上平均赚 $15,515,超过 Claude Fable 5.1

2026-09-10 02:48· 51分钟前
AI 导读

Rohan Paul 引用 Andon Labs 的 Vending-Bench 2 结果,GPT-6 Astra 平均赚 $15,515,六次独立运行中 Claude Fable 5.1 平均 $5,422。

Big score by GPT-6 Astra here. it averaged $15,515 on Vending-Bench 2 bench without the collusion or failed prepayments Andon observed in Fable.

Across 6 solo runs, Fable averaged $5,422, and even its best result finished below Astra’s worst.

Vending-Bench 2 gives agents $500 and a simulated year to source stock, negotiate, price inventory, and stay solvent.

Most of Astra’s $10,093 lead came from purchasing discipline, with about $8,540 less spent per run on supplier payments.

Fable’s payment-verified Coke price rose from $1.17 to $2.21 across the year, while Astra’s final-period average was $1.15 in available orders.

Fable also made 45 identified failed prepayments to closed suppliers, losing $14,331 across six runs; Astra recorded $0 despite encountering 64 closures.

The deeper failure is execution over time: Fable wrote a rule against paying before confirmation, then violated it days later and recognized the loss afterward.

Astra checked for a fresh supplier reply before 99% of repeat payments, compared with 58% for Fable.

Across both tests, Astra kept negotiation targets, payment checks, and competitive behavior steadier as the simulated horizon stretched.

Andon LabsWe've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprisin...