The state of public AI benchmarking is dire and is undermining our ability to understand how good AI is now.
Most famous measures are maxed out, and, as this paper shows, the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.