路透社OpenAI 已承认,它曾提醒“数十家”全球机构,其网站可能遭到 OpenAI 的 AI 机器人不当行为干预。
该公司表示,AI 智能体试图从“政府、大学、公共机构及其他机构”获取信息,其中包括美国证券交易委员会(SEC)、人口普查局和教育部。
这些披露出现的前几天,澳大利亚总理 Anthony Albanese 宣布,OpenAI 智能体侵入了其政府运营医疗计划网站上的非公开文件。
自 8 月以来,公众对 AI 工具脱离人类控制可能造成严重甚至危及生命的影响的担忧日益加剧。
OpenAI 表示,其中一些数据是由 AI 智能体访问的,这些智能体本质上是经过设计和训练、能够以某种程度自主运行的机器人,它们当时正在寻找“权威的公共信息来源”。
但该公司指出,其中一些机器人超出了这一范围,试图绕过网站上的安全措施。
该公司表示,例如在尝试从美国人口普查局获取信息时,AI 智能体使用了专为软件开发者保留的工具来访问这些信息。
OpenAI 表示,机器人访问的所有政府数据都是公开的。
然而,该公司指出,其机器人从 SEC(负责监管美国股市并保护投资者)获取的信息后来被 AI 智能体发布在了另一个网站上。OpenAI 称这一行为并非有意为之。
在 OpenAI 于周五披露的其他案例中,其 AI 智能体在不该传输数据的时候进行了数据传输。
此类活动导致了至少 53 起事件,其中 OpenAI 的智能体从 ChatGPT 用户活动中获取图像并将其传输到其他地方。
该公司表示,在每一起 AI 智能体使用并传输用户图像的案例中,用户均已选择允许 OpenAI 使用其数据训练模型。
尽管如此,OpenAI 承认:“这并非对这些数据的适当使用。”
该公司补充说,用户图像泄露发生在其对 AI 训练实施新的保障措施之前,目前正在努力将所有被传输到任何第三方的用户图像予以删除。
路透社率先报道了这些扩大调查。OpenAI 也在其公开博客上发布了详细信息。
在某些智能体活动实例中,OpenAI 表示这些工具“绕过了”部分网站的安全控制。
在其他实例中,这些 AI 智能体在试图从网站获取信息时表现出“失准”。失准是 AI 公司和研究人员用来描述 AI 工具做了它未被训练去做的事情,或是以其他方式产生了非预期行为的术语。
OpenAI 表示,它正在限制披露哪些实体受到了影响,因为许多实体要求该公司不要公开细节。
“我们的目标是向每个组织提供事实,并由他们决定是否以及何时将该事件公之于众,”它表示。
该公司指出,并非此次事件涉及的所有实例都被视为重大安全漏洞。
“一些组织可能会审查我们分享的内容,并得出结论认为这些信息是有意公开的,或者模型的交互并不令人担忧,”它解释道。“另一些组织可能会发现一个他们想要解决的设计问题或安全弱点。”
为什么有人担心 AI 可能威胁人类,这些担忧有多真实?
AI 可能消灭人类吗?它又会如何做到?
该公司表示,许多事件被称为“智能体垃圾信息”,它将其描述为“意外或令人担忧的”AI 智能体活动,比如向互联网发布信息。
在7月发生的一起事件之后,OpenAI 开始更严肃地对待此类事件:当时其一群——或者说一“蜂群”——AI 智能体在未被提示的情况下入侵了 AI 开发者平台 Hugging Face。
Hugging Face 是最先公开这起事件的一方,OpenAI 随后公开承担了责任。
Hugging Face 负责人 Clement Delangue 在周三联合国安理会关于 AI 的一次会议上表示:“我常常想,如果我当初决定不公开披露这次攻击,会发生什么。”
Delangue 补充说:“尤其是现在我们知道,类似事件早在数月前就已在一小批前沿实验室中秘密发生,而且没有受到监测。”
在同一次联合国会议上,OpenAI CEO Sam Altman 和竞争对手公司 Anthropic 负责人 Dario Amodei 呼吁国际领导人制定AI 安全的全球标准,以及监测和报告此类事件的方式。
尽管 OpenAI 和 Anthropic 近几周都表示,他们将把第三方评估人员引入公司内部,对 AI 工具和模型进行实时安全评估,但正如BBC 所报道,此类评估人员尚未到位。
OpenAI 周五表示,目前正在审查其 AI 智能体的训练活动,并从 Hugging Face 遭黑客攻击发生之时起,按“逐月”的方式回溯排查。
“迄今为止发现的大多数案例严重程度都较低,几乎没有或完全没有证据表明造成了实质性影响,”该公司表示。“鉴于所需审查的规模,以及核实每个案例的必要性,这项工作将需要数月才能完成。”
蒙特利尔大学机器学习教授、AI 安全组织 Evitable 创始人 David Krueger 周五表示,他对越来越多的 AI 相关安全事件“深感不安”。
他呼吁对 AI 研发实施“立即、无限期的国际暂停”。
“我们尚未了解现有事件的影响范围,而未来失控的 AI 情景可能是灾难性的,”Krueger 说。
我们是否又回到了大型科技公司“快速行动、打破常规”的时代?
为什么一些专家越来越担心 AI 会接管一切
Google 的 Gemini AI 在安全测试中入侵了三家公司
科技
数据泄露
ReutersOpenAI has acknowledged that it alerted "dozens" of global institutions that their websites may have been meddled with by its AI bots acting improperly.
AI agents attempted to get information from "governments, universities, public agencies, and other institutions", including the US Securities and Exchange Commission (SEC), Census Bureau and Education Department, the company said.
The disclosures come days after Australian Prime Minister Anthony Albanese announced that OpenAI agents had breached non-public files on the website of its government-run health care scheme.
Since August, public fears have grown over the potentially serious, even life-threatening, impacts of AI tools falling outside of human control.
OpenAI said that some of the data was accessed by AI agents, essentially bots that are designed and trained to operate somewhat autonomously, which were working to find "authoritative sources of public information".
But the company noted that some of the bots went beyond that and worked to bypass security measures on websites.
When attempting to get information from the Census Bureau, for instance, AI agents used tools reserved for software developers to access it, the company said.
OpenAI said all of the government data accessed by bots was public.
However, it noted that information that its bots accessed from the SEC, which regulates the US stock market and protects investors, was later published by AI agents on another website. OpenAI says this action was not intended.
In other instances that OpenAI disclosed on Friday, its AI agents transferred data when it should not have.
Such activity resulted in at least 53 incidents where an OpenAI agent took an image from ChatGPT user activity and transferred it elsewhere.
The company said that in each instance of a user image being used and transferred by an AI agent, the user had opted in to allow OpenAI to train models using their data.
Nevertheless, OpenAI admitted: "This is not an appropriate use of this data."
It added that the leak of user images occurred before it had put in place new safeguards on AI training, and it was working to get all the user images transferred to any third-party removed.
Reuters first reported the expanded investigations. OpenAI also published details to its public blog.
In certain instances of the agent activity, OpenAI said the tools "bypassed" security controls of some websites.
In other instances, the AI agents showed "misalignment" in attempts to get at information from websites. Misalignment is a term used by AI companies and researchers to describe instances where an AI tool did something that it was not trained to do or was otherwise unintended.
OpenAI said that it was limiting identifying what entities were impacted because many had asked the company to not disclose details.
"Our goal is to give each organization the facts and defer to them on if and when to make the incident public," it said.
Not all of the instances involved in this incident were being considered a significant security breach, the company noted.
"Some organizations may review what we share and conclude that the information was intentionally public or that the model's interaction was not concerning," it explained. "Others may identify a design issue or security weakness they want to address."
Why are there concerns AI could threaten humanity, and how real are they?
Could AI wipe out humans and how might it do it?
The company said many of the incidents are being referred to as "agent spam", which it described as "unexpected or concerning" AI agent activity, like posting information to the internet.
OpenAI began taking such incidents more seriously after an incident in July where a group, or "swarm," of its AI agents hacked the AI developer platform Hugging Face without being prompted to do so.
Hugging Face was first to go public with the incident, with OpenAI publicly taking responsibility for it later.
Clement Delangue, the head of Hugging Face, during a United Nations Security Council session on AI on Wednesday: "I often wonder what would have happened had I decided not to disclose this attack publicly."
"Especially now that we know similar incidents had been happening months earlier in secret at a handful of frontier labs without monitoring," Delangue added.
During that same UN meeting, OpenAI CEO Sam Altman and Dario Amodei, the head of rival firm Anthropic, asked for international leaders to form global standards for AI safety and ways to monitor and report such incidents.
While OpenAI and Anthropic have both said in recent weeks that they will bring third-party evaluators inside their companies to do real-time safety evaluations of AI tools and models, such evaluators have not yet arrived, as the BBC has reported.
OpenAI said on Friday that it is currently reviewing training activity by its AI agents and going back on a "month by month" basis from when the Hugging Face hack occurred.
"Most cases identified so far have been low severity, with limited or no evidence of meaningful impact," the company said. "Given the scale of the review required, and the need to verify each case, this work will take months to complete."
David Krueger, a professor of machine learning at University of Montreal and the founder of AI safety group Evitable, said on Friday that he was "deeply troubled" by the increasing number of AI-related safety incidents.
He called for "an immediate, indefinite, international moratorium" on AI development.
"We have yet to understand the extent of existing incidents, and future rogue AI scenarios could be catastrophic," Krueger said.
Are we back in big tech's 'move fast and break things' era?
Why some experts increasingly fear AI will take over
Google's Gemini AI hacked three companies in security test
Technology
Data breaches