跳到正文
原文
Arena.ai· @arena · X·· 3 小时前AI 评分59
AI 导读

Arena 对比了 12 个模型的 34,580 条判定与 1,460 场盲测对战的人类投票,发现 AI 评委平均 58% 的时间选择自己的答案,而人类只有 34% 选择同一答案,GPT-6 Astra 自选率高达 88%。

正文

Ask an AI model to pick the better answer, and it'll usually pick its own.

We compared 34,580 verdicts from 12 models with human votes across 1,460 battles on Arena. The results show that AI judges have their own taste, and it's unlike ours.

They:
- Favor their own answers. On average, a model picked its own answer 58% of the time. People picked that same answer 34% of the time. GPT-6 Astra picked itself 88% of the time.

- Rarely call a draw. People called a tie or "both bad" in 32% of battles. GPT-5.6 Sol picked a winner 96% of the time.

- Side with each other over people. They agreed with other AIs 79% of the time and with people 57% of the time. Every judge did, by a margin of 18 to 27 percentage points.

Full results in the article from @DawidGalarowicz below.

引用Arena.ai@arena
https://x.com/i/article/2104955648664584192
在 X 查看被引用的帖子

来源:Arena.ai · x.com