GPT-6 Astra brutally mogged Fable 5.1 on a benchmark testing research-level math problems from recent arXiv papers
88.6% vs. 87.7%, while costing just $3.63/problem vs. $12.77
similar performance at less than one-third the cost
We are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refuted in the last month on ArXiv, and mo...