New Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks.
And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models.
The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores.
In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes.
The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers.
After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%.
So a low physics benchmark score can badly underestimate what a frontier model can actually solve.
But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried.
but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.