elvis · @omarsar0 · X·2026-09-25 03:05·8分钟前
AI 导读

Pareto 26.9 通过将请求分发至多个前沿与开源模型并保留最佳答案,在 30 项智能体任务评测中与 GPT-6 Astra 并列第一,每成功任务成本约为后者 1/3,且完成任务速度快于 DeepSeek V4 Pro 和 GLM 5.3 Flash。同一评测中,GPT-6 Sol 得分追平 Opus 5.5,速度更快,每成功任务成本约为其 1/4。

elvis@omarsar0
44AI 编辑部评分,满分 100
2026-09-25 03:05· 8分钟前
AI 导读

Pareto 26.9 通过将请求分发至多个前沿与开源模型并保留最佳答案,在 30 项智能体任务评测中与 GPT-6 Astra 并列第一,每成功任务成本约为后者 1/3,且完成任务速度快于 DeepSeek V4 Pro 和 GLM 5.3 Flash。同一评测中,GPT-6 Sol 得分追平 Opus 5.5,速度更快,每成功任务成本约为其 1/4。

Interesting results here. This is why I expect more agent workloads to run on blended models.

Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.

In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.

It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.

ComposioWe tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score...

来源:elvis· x.com