Independent evaluation of AI models, across both capability and safety, is essential. We welcome Dario’s call to give independent evaluators greater access to frontier AI models.
For nearly three years, Artificial Analysis has been building benchmarks and infrastructure to independently measure AI capabilities. We have supported pre-launch benchmarking of frontier models with almost every major AI lab. We are continuing to see jumps in capability across every dimension we measure.
The world needs a vibrant ecosystem of independent AI evaluators. We are going to keep working to build it!