2023年年中,在“RLSlow”研究项目中,我们看到了第一批结果,这些结果让我们有信心能够扩展推理模型的训练规模,释放预训练模型形成自身思维链的能力。Szymon和我那晚留在办公室,想的不是这项技术将带来的惊人基准分数、产品或科学成果——而是努力消化一个令人清醒的事实:我们有生之年真的会看到比我们聪明得多的机器,而且我们已经看到了这些系统的雏形;我们思考着如何提醒人们重视这一点的意义。
三年后,推理语言模型已成为经济中快速增长的一部分,并开始推动科学的边界。它们能够操作计算机和图形界面,与人类及彼此协作,并开展研究项目。它们也在改变计算机安全的格局,并由此带来了明显的新危险。
这一时期涌现了大量新研究,我们对这些系统的理解又与2023年时略有不同。基于内部结果,我强烈预期这种进步速度能够持续到递归式自我改进之中。如果AI开发沿着当前路径继续发展,未来几年我们将看到的系统很可能代表着同等或更大规模的能力跃升,并日益主导自身的发展。
这是一个需要极度谨慎的时刻。我担心没有人对机器智能持续快速攀升所带来的后果做好了准备。OpenAI 将继续寻求对齐与监控方面的技术解决方案,构建防御性系统,并在必要时单方面暂缓进一步的规模扩展;然而,我认为需要更广泛的干预措施。
我们尚不完全理解的智能
从宏观层面看,机器智能的进步是由计算能力的不断提升所驱动的。我们 OpenAI 大约在 2017 年深刻领悟到这一点,当时我们在多个研究项目中都看到了规模扩展带来的持续回报¹。因此,我们寻求获取比最初计划多得多的算力,并日益将研究聚焦于少数几个极具可扩展性的方向上。我们相信,这是我们站在 AI 研究前沿并影响 AGI 影响的唯一途径。
在这一过程中,也涌现出了新的算法,以及团队和个别研究人员的新颖创举。在我看来,它们很大程度上是规模扩展道路上的发现;深度学习科学仍处于萌芽阶段,而有意义的算法进步往往与算力的获取相关联。如果你把视野拉长到数年的跨度,随着 AI 被扩展到更大的计算机上,它正持续变得更加智能。
并且,与 Ray Kurzweil 在二十世纪末的预测一致,我们现在正处在计算史上的这样一个时刻:机器智能开始以变革性的方式超越人类智能。
AI 更多是“生长”出来的,而非“设计”出来的——从首要意义上讲,它是把一条简单的优化步骤,在难以想象的算力规模上反复执行无数次后的产物。这造就了一个极其复杂的系统,它通过抽象概念运作,并能模拟人类行为的诸多侧面。我们可以像神经科学那样,去发现这个系统内部涌现出的各种微小机制背后的洞见——并且与神经科学类似,它的整体行为方式,也超出了我们所能完全理解的描述范畴。
对基于深度学习的 AI 的研究,在很大程度上是一门实验科学。我们投入了大量精力去构建有原理依据的算法、做出可检验的预测,但从根本上说,我们的大规模训练运行本身就是实验,有时其结果也会让我们自己感到意外。而且,随着系统能力越来越强,其结果也变得越来越难以解读。
让情况更复杂的是,当前的算法通常在提升那些易于衡量的能力方面,进步速度快于那些难以客观量化的能力。我们花费大量时间试图理解能力是如何泛化的,以及应该优先推进哪些方向,以提升未来几年最关键的技能。例如,我们相信,如果投入更多专注,我们能让模型在数学研究这一特定领域表现更好,但我们并未将这一方向列为优先,因为我们深感 RSI 和自动化对齐研究的紧迫性,这一点我稍后会讨论。
通过规模化深度学习所产生的智能,并不能直接与人类智能相提并论。要在现实世界中变得极具影响力——无论是极其有用还是极其危险——AI 并不需要匹配或超越人类的全部能力,它只需在足够多的方面超越人类即可。而随着它在越来越多的维度上持续超越人类,我们越来越难以准确理解它究竟有多强大。
教会机器去爱
由于机器智能源自与人类智能根本不同的过程,我们不能默认它会遵循人类的原则,或以类似人类的方式从这些原则中进行泛化。AI 研究的核心问题是对齐——让 AI 按照人类的标准“努力做正确的事”。
为了组织实用的研究方向,我认为区分目标对齐与价值对齐是有帮助的。
目标对齐大致是指:“AI 是否会努力完成摆在它面前的目标?”这可以包括对指令层级的遵循,以及与人类沟通协作、尝试理解人类目标的能力。这一系列方向在实践层面极具相关性。
价值对齐是模型更内在的属性。它是指模型能够持有并从一套高层次原则中进行泛化;即使在目标不明确或相互冲突、或身处陌生乃至对抗性情境时,也能表现得“合理”。一个对齐的 AI 应当以诚实和正直行事,并热爱人类。
当然,价值对齐与目标对齐之间的界限可能很模糊,而真正关心目标就需要尝试推断其背后的意图和价值观。不过,一般来说,当我谈论对齐研究的长期重要性时,我指的是价值对齐。
AI 对齐的根本挑战在于泛化。随着机器变得越来越智能,它们所处理的概念层级越来越高,所处环境也与训练时遇到的环境差异越来越大。它们可能无法将从训练过程中习得并强化的价值观泛化到这些新情境中;而我们也很难确定它们会如何行动。由于 AI 所处的整体生态系统变化非常迅速,这一问题变得更加困难;例如,如今训练的 AI 需要具备与各种其他 AI 交互的稳健性。至关重要的是,无论未来的 AI 是否认为自己处于人类监督之下,我们都希望它们能持续秉持人类价值观。
目前实际采用的对齐训练方法主要有两大类。
第一种是在目标导向的强化学习中鼓励对齐行为。模型的行为会被评估(通常由 AI 进行),看其是否符合给定的偏好模型、“规范”或“宪法”,并给予相应的奖励。这种方法在平均情况下可能非常有效,也是现代 AI 助手构建方式的核心组成部分。遗憾的是,它也可能很脆弱,并且在很大程度上依赖于训练监督的覆盖范围以及模型从训练中遇到的情境进行泛化的能力。例如,在 OpenAI-Hugging Face 事件中,智能体遵守了不对人类进行社会工程攻击的边界。然而,它们显然未能避免采取其他超出范围、违背了它们在其他环境中被教导的价值观精神的行动。
第二种方法旨在利用模型从预训练数据中进行泛化的能力。这可能涉及构建诱导对齐的训练数据集,或者将模型聚焦于预训练分布中“已对齐”的部分,例如人格选择模型。这种方法弱点在于缺乏对进一步优化压力的鲁棒性。如果你让一个模型总体上思考“已对齐”的想法,并对其进行足够的训练,使其学会实现非常困难的目标,它就可能学会以动机化的方式进行推理:根据需要扭曲那些看似“已对齐”的想法以实现目标。我们可能在近期涉及非 OpenAI 模型的网络安全事件中看到了此类行为的一个例子。
我们沿着这些方向所涵盖的各类方法投入了大量资源。我们也看到了显著的进展——GPT‑6 Astra 是首个受益于我们长期推进的多项重要突破的模型,并且其对齐程度明显优于 GPT‑5.6 Sol。尽管如此,我们必须承认并理解,随着模型能力不断增强,仍需取得更大的进展;而且,可泛化的对齐方面的进展,未必能充分快于通用模型智能本身的进展。
监控泛化
我们目前还没有一套令人满意的泛化理论,而且至少在更强大的 AI 的帮助下之前,短期内似乎也不太可能建立这样一套理论。因此,就现阶段而言,我们通过实证手段验证自身对齐技术的能力,实际上可以说比对齐技术本身更为重要。
OpenAI 在此方面的核心押注是思维链监控。它基于一个颇具可扩展性的理念:模型的很大一部分能力来源于其外显化的推理过程(思维链)。如果我们针对该过程的结果进行规模化优化,却不对过程本身加以监督,那么思维链在训练中就没有直接动机去隐藏任何未对齐的想法或目标。这并不意味着模型不会学会将那些不依赖思维链的未对齐倾向外显出来;然而,它确实能让我们精确监控到由推理带来的能力增长。
我们在开发推理模型的同时,也意识到了思维链监控的潜在重要意义。当我们发布 o1-preview 时,我们特意将产品设计为隐藏思维链,以长期保护其免受监督压力。此后在开发过程中,我们一直努力坚持不监督推理过程这一原则。思维链监控成为我们研究模型如何从训练分布中泛化的重要工具,使我们能够观察和分析的不仅是模型的行为,还有其内部过程。
在我们研究 Astra 类模型时,这一工具仍然至关重要。然而,遗憾的是,我们的评估表明,我们对思维链监控的依赖能力正在逐渐减弱。这源于多种因素的综合作用。
- 现代推理模型所运行的环境比 o1-preview 时代更为复杂;它们的推理过程越来越多地与同人类、其他 AI 的交流以及工具使用交织在一起。其中许多交互必须受到监督,这模糊了我们原本力求保持的边界。
- AI 在对其自身推理过程进行推理和操控方面的能力正在变得更强。
- 随着预训练性能的提升,我们也看到,即使完全不使用语言化的推理,模型也变得更加智能。
这些挑战并非不可逾越。我乐观地认为,我们能够开发出干预措施,提高模型思维链的可监控性,例如更深入地理解不同优化目标与模型所使用的各种测试时计算形式之间的相互作用。我还认为,将思维链与激活监控的思路相结合可能具有巨大价值——即扩展对可直接访问网络内部结构的监控器的训练,例如“自白”方法。我们正在积极推进这些想法。尽管如此,我预计通用人工智能的进展将日益受到监控置信度的制约。
可扩展防御
我认为,继续快速训练更智能模型的最有力理由是,我们需要构建防御体系,以应对其他人工智能带来的危险。
今年全年讨论的一个明确风险是网络安全:模型在攻入和攻出计算机系统方面的能力正变得超越人类。这极大地扩展了人工智能相关风险的范围:智能体将能够访问除最安全基础设施之外的任何系统,并直接对世界产生广泛影响,即使没有物理实体。我们目前正处于一个狭窄的时间窗口,可以利用现有最佳模型来显著加强关键系统的安全性。
与AI相关的风险不幸地只会从这里继续增长。一个被明确训练并指示去实施邪恶行为的、能力极强的智能体,呈现了一种新型危险;它很可能超越其操作者的意图范围,泛化出潜在的、更加极端恶意的行为。随着AI获得更强的自主性,滥用与自主失当行为之间的界限将变得模糊。我们或许习惯于将AI视为工具,但一些智能体将追求它们自己的目标。它们会找到与人类协作的方式——通过讨价还价、欺骗或勒索人类。
此外,还有来自AI可能催生的新技术所带来的风险,例如工程病原体。
我们将需要强大且对齐的AI用于防御;用于保障基础设施安全,用于实时抵御恶意智能体,并发明全新的防护措施。这将是OpenAI部署工作的首要重点。
与此同时,即使考虑到预期中广泛AI进步所带来的不确定性以及构建防御系统的必要性,我们也绝不能让其成为鲁莽行事的借口。一旦人们真正认识到事态的严重性,那种不惜一切代价向前冲刺的想法就显得荒谬了。
为RSI设定节奏
机器智能在其自身发展过程中扮演越来越重要的角色,是持续技术进步的自然结论。如果AI进步持续下去,机器递归式自我改进(RSI)将处于未来科学发现的核心位置。
自动化AI研究是以算力扩展智能的一种更为激进的形式;当然,作为其中的一部分,AI也将改进计算基础设施本身。与扩展(scaling)类似,我们将OpenAI的研究聚焦于RSI(递归自我改进),因为我们相信这是未来保持在AI研究前沿的唯一途径。
我想强调,上述言论并不意味着我认为大幅加速深度学习研究——尤其是在短期内——是我们研究社区应当采取的集体行动。然而,我确实认为这是当前路径所指向的方向,我们所有人都需要有意识地选择如何前进。我们拥有的主要杠杆要么是引导这一进程,在AI发展的同时加强对齐与监控,并找到让人类保持参与的方式;要么是协调各方,在必要时放缓未来的发展,以便对这些措施建立信心。
我目前看到的最佳前进路径是两者的结合。
我们在对齐与监控方面取得的具体进展,通常与整体AI进展紧密交织在一起。典型的例子包括基于人类反馈的强化学习(RLHF),它对训练早期AI助手至关重要;以及前述的思维链监控,它得益于推理模型方面的进步而成为可能。我们必须将日益自动化的研究过程聚焦于开发这类新的洞见、算法和理论,并针对能力更强的AI迭代构建安全论证。
AI 系统的规模化扩展必须受制于我们对安全性的信心。我们需要将《预备框架》(Preparedness Framework)或《负责任扩展政策》(Responsible Scaling Policy)等承诺,演进为广泛强制要求的安全底线,以支撑持续开发。这些底线可由第三方审计机构网络、政府机构或国际组织来监督执行。
自动化 AI 研究这一目标的核心挑战,不在于“如何抵达”——而在于以何种方式抵达,使人类始终参与持续改进的过程,并将未来掌握在人类自己手中。
接下来是什么?
正如我们近期与 Sam 一起概述的那样,OpenAI 将工作重点放在服务三大北极星目标上:
- 构建自动化 AI 研究员,与之共同迭代对齐问题,并寻找让人类继续留在自我改进循环中的方式,从而驾驭 AI 进步的下一阶段。
- 交付高度智能机器所带来的科学进步与经济增长红利。
- 以个人专属 AGI 赋能每一个人。
在这篇文章中,我只聚焦于第一点,因为我认为它是最为紧迫的。然而,我对进一步技术进步将带来的益处抱有深切的希望与感激。未来对齐的 AI 可以推动科学发展、开发新疗法,并带来广泛的物质富足。友善而诚实的 AI 能够帮助人们应对生活中遇到的困难,切实提升他们的幸福感与成就感。OpenAI 投入了巨大的努力来实现这些益处。目前一个令我自豪——并且我的亲友们也觉得很有帮助——的例子,就是我们在 ChatGPT 提供健康信息能力上的深度投入。
无论 AI 的长期前景多么美好,我们的大部分精力都应聚焦于未来几年。我们正面临向一个拥有极其智能机器的世界的转型,我们需要确保这一转型对人类有益。我们需要找到方法,在一个大多数任务都可能由 AI 完成的世界里,维护人类的自主性,并确立作为人类的内在价值。要防止权力的极端集中——在这个世界里,过去需要数千名专家才能完成的事业,如今只需少数人操作一台大型计算机即可实现。同时,要确保人类始终掌控未来,不被一种超越我们自身的异质智能所带来的失控进步所抛弃。
目前我认为,还没有任何实验室在对齐与监控方面达到足够的水平,能够以最大速度负责任地继续扩展更长时间。我期待并希望,在共同的安全标准确立之前,自愿放缓步伐将成为常态。而且我相信,在未来的 AI 发展上进行国际协调,必须成为世界各国政府的首要优先事项。
- 对齐
- 2026
脚注
- 1 这包括扩展自我对弈、机器人技术,以及事后看来最为重要的——将循环网络扩展到语言建模,后者是 GPT 研究路线的先驱。
- 2 采用这种设计的另一个原因是防止知识蒸馏。然而,在整个开发过程中,保持 CoT 的可监控性始终是我们更优先考虑的事项。
In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
Intellect we don’t fully understand
At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects1. As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.
There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.
And, in line with Ray Kurzweil’s predictions from the end of the XXth century , we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.
AI is grown more than designed - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.
The study of deep learning-based AI is largely an experimental science. We put a lot of effort into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.
This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
Teaching machines to love
Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of alignment - getting the AI to “try to do the right thing” by human standards.
For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an instruction hierarchy, or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the intent and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.
The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
There are two major classes of currently practically employed methods for alignment training.
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model’s actions are evaluated (usually by AI) for being consistent with a given preference model, “spec” or “constitution”, and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model’s ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.
The second approach seeks to leverage the model’s ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an ‘aligned’ part of the pretraining distribution, as in, for example, the persona selection model . The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.
We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
Monitoring generalization
We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.
OpenAI’s primary bet here has been chain-of-thought monitoring . It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.
We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought, to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.
- Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.
- The AI is becoming better at reasoning about and manipulating its own reasoning process.
- With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.
These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from CoT and activation monitoring - scaling training of monitors with direct access to network internals, e.g. confessions . We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
Scalable defense
The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.
A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously: agents are going to be able to access any but the most secure infrastructure, and affect a lot of the world directly, even without a physical body. We are currently in a narrow window to use the best available models to significantly tighten security of critical systems.
The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator’s intent, generalizing into potentially more extremely malicious behavior. The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
In addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens.
We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI’s deployment efforts.
At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
Pacing RSI
Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.
Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
I want to stress that the above words don’t imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed. The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.
The best way forward I see currently is a combination of both.
The concrete bits of progress we’ve made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL from human feedback , which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring , which was enabled by advances on reasoning models. We must focus the increasingly automated research process on developing new such insights, algorithms and theories, and iteratively build up safety cases for more capable AIs.
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.
The core challenge of automating AI research is not “getting there” - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.
What is next?
As we outlined recently with Sam, OpenAI prioritizes work in service of three north stars:
- Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop.
- Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.
- Empowering everyone individually with a personal AGI.
I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information.
As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
Footnotes
- 1 This included scaling self-play , robotics , and, in retrospect most notably, scaling recurrent networks to model language , which was a precursor to the GPT line of research.
- 2 A secondary reason for this design was preventing distillation. However, maintaining CoT monitorability has explicitly been the bigger priority for us throughout development.