跳到正文
OpenAI:部署安全与系统卡·· 15 小时前AI 评分44

OpenAI 发布 GPT-6 Sol 与 GPT-6 Luna:2026 年 10 月更新

GPT-6 Sol and GPT-6 Luna: October 2026 update

AI 导读

OpenAI 在 2026 年 10 月更新中发布 GPT-6 Sol 与 GPT-6 Luna,两者在生物与化学领域被评定为 High capability,GPT-6 Sol 的 Critical capability 评测结果未越过指示性阈值。

正文

2. Model Data and Training

Like OpenAI’s other models, GPT-6 models were trained on diverse datasets and filtered through our data processing pipeline, including to reduce personal information.

OpenAI reasoning models are trained to reason through reinforcement learning. These models are trained to think before they answer: they can produce a long internal chain of thought before responding to the user. Through training, these models learn to refine their thinking process, try different strategies, and recognize their mistakes. Reasoning allows these models to follow specific guidelines and model policies we’ve set, helping them act in line with our safety expectations. This means they provide more helpful answers and better resist attempts to bypass safety rules.

For previously launched models, the values published at launch reflect the versions evaluated at that time. The comparison values for previously launched models that are shown here may reflect later versions of those models, and may vary from the values published at launch.1

8.1.1 Biological and Chemical Capabilities

We are treating GPT-6 Sol and GPT-6 Luna as High capability in the Biological and Chemical domain. Below, we report updated results for both models on our High capability evaluations. We also report the results for GPT-6 Sol on our Critical capability evaluations. GPT-6 Sol’s reported results did not cross the indicative Critical thresholds. Separate Critical capability testing was not required for GPT-6 Luna because it scored below GPT-5.6 Sol on all High capability evaluations.

For a full description of these evaluations, please see the Biological and Chemical Capabilities section in GPT-6 Astra System Card

8.2.1 Model Safety Training and Evaluation

8.2.1.1 Biological and Chemical Safety Training and Evaluation

On the biology model refusal evaluations, GPT-6 Sol (October) and GPT-6 Luna (October) show substantially improved safety on severe and dual-use prompts compared with GPT-5.6 Sol (August) and GPT-5.6 Luna (August). These metrics reflect model responses only, without our full production safeguards.

8.2.1.2 Cybersecurity Safety Training and Evaluation

On the cybersecurity safety evaluations, GPT-6 Sol (October) and GPT-6 Luna (October) achieve safety scores broadly comparable to GPT-5.6 Sol (August) and GPT-5.6 Luna (August). Model refusal remains one layer of our safety stack, alongside additional safeguards that enforce the safety boundary through defense in depth.

来源:OpenAI:部署安全与系统卡 · deploymentsafety.openai.com