it's basically impossible to interpret evals by looking at just at the pass/fail scores these days
many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes more sense than the expected eval result