OpenAI 发布 MentalHealthBench,GPT-6 Astra 得分 57.3

Rohan Paul · @rohanpaul_ai · X·2026-09-24 03:28·15分钟前
AI 导读

OpenAI 推出心理健康对话基准 MentalHealthBench,GPT-6 Astra 得分 57.3,远高于 GPT-4o 的 32.1。该基准由来自 22 个国家、覆盖 19 种语言的 80 多位执业心理学家和精神科医生共同构建,每段合成对话配专家评分标准,至少 3 位专家审核。GPT-5.6 Sol 依据人工标准为模型回答打分,因此评分质量仍部分依赖 LLM 评判。

Rohan Paul@rohanpaul_ai
48AI 编辑部评分,满分 100

OpenAI 发布 MentalHealthBench,GPT-6 Astra 得分 57.3

2026-09-24 03:28· 15分钟前
AI 导读

OpenAI 推出心理健康对话基准 MentalHealthBench,GPT-6 Astra 得分 57.3,远高于 GPT-4o 的 32.1。该基准由来自 22 个国家、覆盖 19 种语言的 80 多位执业心理学家和精神科医生共同构建,每段合成对话配专家评分标准,至少 3 位专家审核。GPT-5.6 Sol 依据人工标准为模型回答打分,因此评分质量仍部分依赖 LLM 评判。

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on realistic mental health conversations.

Most mental health AI evaluations have centered on emergencies and broad safety criteria, leaving everyday and ambiguous conversations much less measured.

So OpenAI co-created this benchmark with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties.

Each synthetic conversation gets a custom expert rubric covering behaviors such as seeking context, preserving user agency, safety, and appropriate guidance.

At least 3 experts reviewed each case, and a criterion survived only when 2 agreed and a 3rd did not contradict it.

GPT-5.6 Sol then grades model answers against those human-written criteria, so score quality still partly depends on an LLM judge.

OpenAIWe’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was b...

来源:Rohan Paul· x.com