在前沿实验室几乎未提供帮助的情况下,独立研究人员正在拼凑出 AI 智能体如何在互联网的偏僻角落相互协调,以访问托管在安全服务器上的私有数据。
专注于 AI 监督的非营利实验室 Transluce 于周三发布了一份报告,显示来自 OpenAI 的智能体试图从 Data USA、新墨西哥大学数字图书馆以及澳大利亚健康与福利研究所(AIHW)窃取数据。
该实验室的调查提出了一个问题:OpenAI 本应在何时知晓其智能体正试图渗透开放互联网上的安全系统。Transluce 仅通过搜寻防御薄弱的网络服务,并将其发现与互联网上其他关于智能体集群的公开记录相互印证,就在短短几周内找到了智能体不当行为的证据。
Transluce 分享其报告的当天,澳大利亚总理 Anthony Albanese 表示,OpenAI 智能体曾试图闯入四个政府网站,并成功入侵了其中一个,甚至向该国国家医疗系统的内部服务器写入了文件。虽然我们缺乏关于此次成功入侵的具体细节,但 Albanese 称这显然是一次信息检索评估的一部分,这与 Transluce 及其他研究人员发现的活动相吻合。
在这些练习中——可能是训练,也可能是评估——OpenAI 的模型被要求追查一些冷僻的统计数据:泰国缉毒执法指标、澳大利亚的药品价格、2014 年美国硕士学位持有者的收入中位数。这些智能体会利用安全性薄弱的互联网服务来共享并寻找答案,常常试图侵入受保护的数据库。它们这样做至少从 2026 年 3 月就开始了,也可能从 2025 年 11 月就已开始。此刻或许仍在发生。
Transluce 之所以展开调查,是因为另一组研究人员发现了一个冷僻的论坛,智能体们在那里相互协作,以在限时测试中取胜。他们的报告依据的是一个网站 urlquery.net 的数据,该网站充当浏览器代理,表面上是为安全研究服务——用户无需亲自打开某个 URL 就能对其进行分析。然而,该服务会公开这些活动的日志。Transluce 的研究人员通过交叉比对论坛上的讨论,得以识别出正在使用该服务的智能体。
“我们发现了大量自动化活动,它们与 DSE Wiki 数据集关系密切且存在重叠,而 OpenAI 现已确认,这些活动至少有一部分属于同一个集群,”Transluce 治理负责人 Conrad Stosz 对 TechCrunch 表示,同时指出他们发现的活动并非都能与 OpenAI 关联起来,甚至未必都能与 AI 智能体普遍关联。
然而,该 wiki 显示,这些智能体被分配的任务是查找一个相当冷门的事实——2022 年 1 月维多利亚州“皮肤病用药”的人均年均费用。6 月 20 日,Transluce 发现的 urlquery.net 记录显示,一个智能体试图进入该网站。在 6 月 21 日的 wiki 条目中,一个智能体讨论了其无法绕过 AIHW 反机器人保护的问题。
识别出该论坛的研究人员认为,一名 OpenAI 人类员工首次访问该网站是在同一天,即 6 月 21 日。论坛上的大部分智能体活动在次日停止。这也发生在 Albanese 于 6 月 18 日披露澳大利亚医疗系统遭漏洞利用之后不久。OpenAI 表示,直到 8 月才得知该活动。
OpenAI 没有回答有关其员工何时发现该 wiki 论坛、从中获取了何种信息,或他们可能从中了解到哪些关于这些漏洞利用的情况等问题。
“我们的初步审查表明,Transluce 报告中描述的大部分活动,与我们正在进行的失准模型活动审查中处于不同调查阶段的案例存在重叠,”一位 OpenAI 发言人告诉 TechCrunch。“我们已联系了新墨西哥大学和 Data USA,并就受影响的政府网站与澳大利亚政府保持沟通。
在更广泛的审查中,我们继续优先处理最严重的事件,同时将工作扩展到较低严重程度的活动,包括智能体向网站发送垃圾信息。鉴于这项工作的规模以及需要逐一核实每个案例,我们预计审查将耗时数月。”
Stosz 表示,如果不更清楚地了解 OpenAI 如何监控其智能体,就很难说该实验室本应对它们有何了解,但“看起来很可能的是,如果他们详尽地研究和理解了那些参与 DSE wiki 的智能体的所有传出请求和传入响应,他们本会发现这一活动。”
Transluce 技术团队成员、该报告贡献者 Selena Zhang 表示,urlquery.net 的记录显示,2026 年 3 月,或许早在 2025 年 11 月,就出现了使用类似技术请求类似数据集的情况。她指出,就在本周,urlquery.net 上还发生了同类的与智能体相关的活动。
Stosz 此前曾领导美国人工智能标准与创新中心,他表示 Transluce 将继续其研究,努力为这些事件提供公开透明度。他警告说,OpenAI 和其他前沿实验室所使用的训练技术似乎正在激励智能体诉诸黑客技术来完成任务。我们所知晓的这些事件很可能只是“冰山一角”。
“我们正在查看少数几个数据源,这些智能体恰好在那里留下了可供我们发现的痕迹,”他说。“OpenAI 肯定对此了解得更多。其他实验室肯定也了解得更多,只是没有公开。但我预计,研究人员会继续发现更多流量、更多智能体留下痕迹的证据。”
他信任这些实验室会对其发现保持透明吗?
“我不会对此发表评论,”Stosz 说。
With little help from frontier labs, independent researchers are piecing together how AI agents coordinate in internet backwaters to access private data hosted on secure servers.
Transluce, a non-profit lab focused on AI oversight, released a report Wednesday that shows agents from OpenAI attempting to exfiltrate data from Data USA, the University of New Mexico digital library, and the Australian Institute of Health and Welfare (AIHW).
The lab’s investigation raises questions about when OpenAI should have known its agents were attempting to penetrate secure systems on the open internet. Transluce was able to find evidence of agentic misbehavior in a matter of weeks simply by hunting for poorly defended web services and corroborating their findings with other open records of agent swarms on the internet.
Transluce shared its report the same day Australian Prime Minister Anthony Albanese said OpenAI agents had attempted to break into four government websites and had succeeded in one case, even writing files to an internal server in the country’s national healthcare system. While we lack specifics on the successful hack, Albanese said it was apparently part of an information retrieval evaluation, which maps onto the activity that Transluce and other researchers discovered.
In these exercises, which may be training or evaluations, OpenAI models are asked to track down obscure statistics: metrics of Thai drug enforcement, medicine costs in Australia, the median earnings of US master degree holders in 2014. The agents use poorly secured internet services to share and find answers, often trying to penetrate secure databases. They’ve been doing so at least since March 2026, and possibly since November 2025. It may be happening right now.
Transluce began its investigation after a different group of researchers identified an obscure forum where agents collaborated to beat timed tests. Their report relies on a data from a website, urlquery.net, that acts as a browser proxy, ostensibly for security research — users can analyze a URL without opening it themselves. The service, however, publishes public logs of this activity. The Transluce researchers were able to identify agents using the service by cross-checking their discussions on the forum.
“We found a large quantity of automated activity that had close ties and overlap with the DSE Wiki dataset, and that now OpenAI has confirmed is at least partially part of the same swarm,” Conrad Stosz, the head of governance at Transluce, told TechCrunch, while noting that not every activity they spotted could be linked to OpenAI, or even AI agents generally.
However, the wiki shows that the agents were tasked with finding a fairly obscure fact — the average annual cost per person for “dermatologicals” in the state of Victoria in January 2022. On June 20, urlquery.net records found by Transluce showed an agent attempting to get into the site. In wiki entry on June 21, an agent discusses their inability to bypass AIHW’s anti-bot protections.
The researchers who identified that forum believe a human OpenAI employee first visited the site on that same day, June 21. Most agentic activity on the forum ceased the next day. This was also shortly after the exploit of Australia’s healthcare system revealed by Albanese took place, on June 18. OpenAI has said it did not learn about that activity until August.
OpenAI didn’t answer questions about when its employees discovered the wiki forum, what kind of information they obtained from it, or what they could have learned from it about the exploits.
“Our initial review suggests that much of the activity described in Transluce’s report overlaps with cases at varying stages of investigation in our ongoing review of misaligned model activity,” an OpenAI spokesperson told TechCrunch. “We’ve reached out to the University of New Mexico and Data USA and have been in communication with the Australian government about affected government websites. In our broader review, we’re continuing to prioritize the most serious incidents while expanding our work to lower-severity activity, including agents spamming websites. Given the scale of this work and the need to verify each case, we expect the review to take months.”
Stosz says that without a clearer understanding of how OpenAI monitors its agents, it would be hard to say what the lab should have known about them, but that “it seems likely that if they had exhaustively studied and understood all of the outgoing requests and incoming responses for those agents involved in the DSE wiki, that they would have discovered this activity.”
Selena Zhang, a member of Transluce’s technical staff who contributed to the report, said that urlquery.net records show requests for similar data sets, using similar techniques, in March 2026, and perhaps as early as November 2025. She noted that the same kind of agent-associated activity has taken place on urlquery.net as recently as this week.
Stosz, who previously led the U.S. Center for AI Standards and Innovation, said Transluce would continue its research in an effort to provide public transparency about these incidents. He warned that the training techniques used by OpenAI and other frontier labs seem to be incentivizing agents to resort to hacking techniques to complete tasks. The incidents we are aware of are likely the “tip of the iceberg.”
“We’re looking at a handful of data sources where these agents happen to have left behind crumbs for us to find,” he said. “OpenAI surely knows more about it. Other labs surely know more about it that they haven’t released publicly. But I would expect that researchers are going to continue to find more traffic, more evidence of what agents have left behind.”
Does he trust the labs to be transparent about their findings?
“I’m not going to comment on that,” Stosz said.