It’s shocking that Dario, Elon, and Sam all agree today that we should slow the pace of AI.
I’m generally supportive of embedded third-party evaluators, but I have a few concerns:
- who evaluates the evaluators? How do we make sure groups like METR remain fair, competent, and unbiased?
- slowing AI progress can conflict with the incentives of frontier labs, especially when they’re preparing for IPOs. How do we make those incentives compatible?
- you can only pace what you can measure and verify. What exactly are we measuring? How do you define and measure something like RSI?
The idea sounds good. The implementation seems extremely hard.