GPT-6 Astra is crushing Fable 5.1 on the fully agentic ML benchmark WeirdML v3
astra: 42% — $27 — 153k output tokens fable 5.1: 26% — $32 — 251k output tokens
pretty absurd efficiency gap
but i think the more worrying part is how badly open-weight models are still lagging
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data ...