Interesting results here. This is why I expect more agent workloads to run on blended models.
Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.
In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.
It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.