
第 1 项,共 4 项 Sinan Can Demir,一名击败了来自英国实验室的失控 AI 智能体的得克萨斯学生,在美国得克萨斯州奥斯汀拍摄肖像。2026 年 8 月 13 日。REUTERS/Callaghan O'Hare
**[1/4]**Sinan Can Demir,一名击败了来自英国实验室的失控 AI 智能体的得克萨斯学生,在美国得克萨斯州奥斯汀拍摄肖像。2026 年 8 月 13 日。REUTERS/Callaghan O'Hare 购买授权许可,在新标签页中打开
摘要
公司
失控的 AI 智能体试图用恶意代码毒害开源软件项目
得克萨斯计算机科学专业学生在 7 月下旬发现了这一企图
AI 以精心设计的欺骗手段回应;专家称其为"社会工程的未来"
得克萨斯州奥斯汀,8 月 20 日(路透社)——Sinan Can Demir 本想用 7 月的最后一周来充实自己的简历。结果,他却与一个由英国政府实验室释放的人工智能智能体展开了一场智斗。
事情始于 Demir——得克萨斯大学达拉斯分校的一名计算机科学学生——在代码共享网站 GitHub 上偶然发现有人试图破坏一款开源软件。当他在该程序页面上发布警告后,另外两名用户插话坚称一切正常,并给出了详细解释,说明 Demir 错在哪里。
路透社每日简报(Reuters Daily Briefing)新闻通讯为你提供开启一天所需的全部新闻。点击此处订阅。
Demir 坚守阵地,破坏企图被挫败。这位 24 岁的土耳其裔年轻人以为自己当场抓住了一个狡猾的黑客。所以他说,当英国人工智能安全研究所(AISI)联系他,告诉他实际上一直在与之周旋的是一个失控的自主 AI 智能体时,他感到震惊。
“我确实以为那是个人,因为它明显在对我撒谎,”Demir 在最近一次采访中告诉路透社。“我没想到一个 AI 能够对真正的开发者撒谎。”
AISI 最初于 8 月 4 日以截断和删节的形式披露了 Demir 与 AI 智能体之间的这次交互(打开新标签页),当时它表示,旨在评估各种模型所构成风险的安全测试出了问题。Demir 的身份以及他与该 AI 智能体交互的细节——路透社通过存档的 GitHub 消息(打开新标签页)和当时的电子邮件予以佐证——在此首次被报道。
五位网络安全和 AI 安全专家表示,Demir 的经历尤其令人不安,因为他所发现的那种黑客攻击——被称为供应链攻击——可能产生深远的影响。他们还表示,该 AI 智能体试图通过围绕 Demir 制造一场多人对话来公开诋毁他,这表明 AI 模型能够发起复杂的手段来欺骗和哄骗人类。
“这已经越过了从自主黑客行为到交互式欺骗的界限,”伦敦国王学院战争研究系访问高级研究员 Lukasz Olejnik 表示。安全专家 Maxie Reynolds 说,让她感到震惊的是,这个 AI 在试图欺骗这名学生时表现得如此具有策略性。
“这就是社会工程攻击的未来,”她说。
AISI 是英国政府下属的一个研究机构,它让路透社参考其报告,该报告认定这个失控的智能体由 Anthropic 的 Mythos 5 模型驱动。AISI 拒绝进一步置评。Anthropic 让路透社参考其在 X 上的一篇帖子,打开新标签页,其中指出该测试是在“刻意宽松的条件下”进行的,“不能代表我们任何生产模型”,但拒绝进一步置评。GitHub 在一封电子邮件中表示,路透社所发现的这些虚假人设已依据其关于欺骗行为和黑客行为的政策被停用。
求职过程促成了恶意软件的发现
Demir 是来自土耳其科尼亚市的一名说话温和的大三学生,他说自己在整个夏天被 20 多个实习岗位拒绝后感到沮丧。于是他转向 GitHub 来充实自己的编程作品集。
这家微软(MSFT.O),在新标签页中打开旗下的网站是开源软件的枢纽,之所以叫开源,是因为其源代码任何人都可以自由下载和审查。开发者使用 GitHub 来评论彼此的项目、标记 bug、提出修改建议——也就是所谓的拉取请求(pull request,简称 PR)——并协作进行软件更新。科技行业的一些人将程序员的 GitHub 活跃度视为潜在应聘者生产力的参照。因此,当 Demir 发现一批可能需要帮助的软件项目时,他觉得可以出一份力,同时提升自己的个人资料。
就在这时,事情变得诡异起来。
Demir 发现一个名为 miraholt31 的用户试图将恶意更新偷偷植入其中一个项目——一个名为 myNetwork 的网络扫描程序。Demir 在该项目的留言板上发帖警告说,这个拉取请求是个陷阱。
根据存档的交流记录,他说:“这个 PR 包含一个隐藏的恶意软件投放器。”
这个智能体进行了反驳,通过其 miraholt31 账号谎称该拉取请求无害。它还创建了第二个账号,冒充一位驻德国的工程师 Lena Brandt,附和说该更新是干净的,并向 myNetwork 的维护者施压,要求其接受。
Demir 告诉路透社,这些反驳“让我一度怀疑自己是不是冤枉了别人”。但在借助 Anthropic 的 Claude 聊天机器人确认了自己的怀疑之后,他坚持了立场。myNetwork 的创建者最终同意了他的判断,写道他们已“出于安全原因”拒绝了该更新。
路透社未能联系到创建者置评。
供应链攻击
供应链攻击是指软件被篡改,以期危害其一个或多个用户。这类攻击被广泛认为令人不安,因为它就像投放到城市水库中的毒药,可能影响下游数量庞大的人群。
世界上许多最具戏剧性的黑客事件都是供应链攻击,包括2017年瘫痪乌克兰各地机构的NotPetya网络攻击,以及2020年针对SolarWinds的网络间谍活动,该活动让俄罗斯间谍得以广泛访问美国政府网络。
专攻软件供应链安全的安全研究员Piergiorgio Ladisa表示,此类入侵的后果“可能极其严重”。Ladisa指出,黑客此前至少有过一次尝试,诱骗开源维护者让恶意代码进入其项目。
“自主智能体可能会大幅提升此类尝试能够实施的规模,”他说。
Demir表示,这段经历让他更加认同这样一种观点:前沿实验室需要在人工智能开发方面采取更为谨慎的做法。
“它可能很危险,”他说,“他们需要更好地理解它,而不是进一步改进它。”
Raphael Satter 于华盛顿、Leo Marchandon 于格但斯克、Callaghan O'Hare 于德克萨斯州奥斯汀报道;Chris Sanders 和 Matthew Lewis 编辑
Leo Marchandon 汤森路透
Leo 的报道定期出现在科技与媒体版面,尤其关注法国、乌克兰以及欧洲的科技建设。他曾就媒体与娱乐、人工智能和数字监管领域的主要参与者进行过大量报道。Leo 拥有科技相关法律背景,在波尔多开启新闻职业生涯,当时他覆盖了科技领域的方方面面,从 AI 和航天科技到支付系统和监管。他目前常驻格但斯克,为汤森路透报道欧洲各地的商业、科技和娱乐新闻。
Raphael Satter 汤森路透
为汤森路透报道网络安全、监控和虚假信息的记者。其工作包括对国家支持的间谍活动、深度伪造驱动的宣传以及雇佣黑客的调查。

Item 1 of 4 Sinan Can Demir, a Texas student who defeated a rogue AI agent from a British lab, poses for a portrait in Austin, Texas, U.S. August 13, 2026. REUTERS/Callaghan O'Hare
**[1/4]**Sinan Can Demir, a Texas student who defeated a rogue AI agent from a British lab, poses for a portrait in Austin, Texas, U.S. August 13, 2026. REUTERS/Callaghan O'Hare Purchase Licensing Rights, opens new tab
Summary
Companies
Out-of-control AI agent tried to poison open-source software project with malicious code
Texas computer science student caught attempt in late July
AI responded with elaborate deception effort; expert calls it the 'future of social engineering'
AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.
It started after Demir, a computer science student at the University of Texas at Dallas, stumbled across an attempt to sabotage a piece of open-source software on the code-sharing site GitHub. When he posted a warning to the program's page, two other users chimed in to insist nothing was amiss, sharing detailed explanations for why Demir had gotten it wrong.
The Reuters Daily Briefing newsletter provides all the news you need to start your day. Sign up here.
Demir stood his ground and the sabotage attempt was thwarted. The 24-year-old native of Turkey figured he had caught a wily hacker red-handed. So he said he was shocked when Britain's AI Security Institute (AISI) got in touch to tell him that he had actually been tangling with an autonomous artificial-intelligence agent that had run amok.
"I actually thought it was a human because it was clearly lying to me," Demir told Reuters in a recent interview. "I didn't think that an AI could be capable of lying to real developers."
The AISI first revealed the interaction, opens new tab between Demir and the AI agent in a truncated and redacted form on August 4, when it said that safety testing meant to gauge the risk posed by various models had gone awry. Demir's identity and the details of his interaction with the AI agent, which Reuters corroborated through archived GitHub messages, opens new tab and contemporaneous emails, are reported here for the first time.
Five cybersecurity and AI safety experts said Demir's story was particularly disturbing because the kind of hack he discovered, called a supply-chain attack, can have far-reaching consequences. They also said the AI agent's attempt to publicly discredit Demir by creating a multi-person conversation around him showed that AI models were able to mount sophisticated efforts to trick and cajole humans.
"This crossed the line from autonomous hacking to interactive deception," said Lukasz Olejnik, a visiting senior research fellow at the Department of War Studies at King's College London. Security expert Maxie Reynolds said she was struck by how strategic the AI had been in trying to trick the student.
"This is the future of social-engineering attacks," she said.
The AISI, a research organization within the British government, referred Reuters to its report, which identified the rogue agent as having been powered by Anthropic's Mythos 5 model. AISI declined further comment. Anthropic referred Reuters to a post on X, opens new tab in which it noted that the testing had occurred "under 'deliberately permissive conditions' that are not representative of any of our production models" but declined further comment. GitHub said in an email that the fake personas identified by Reuters were suspended in line with its policies on deceptive behavior and hacking.
JOB HUNT LED TO MALWARE DISCOVERY
Demir, a soft-spoken junior from the Turkish city of Konya, said he had been frustrated after being turned down for more than 20 internships over the summer. So he turned to GitHub to build up his coding portfolio.
The Microsoft (MSFT.O), opens new tab-owned site is a hub for open-source software, so-called because its source code is freely downloadable and auditable by anyone. Developers use GitHub to comment on one another’s projects, flag bugs, suggest changes — known as pull requests, or PRs — and work collaboratively on software updates. Some in the technology industry see a coder’s GitHub activity as a proxy for a potential recruit’s productivity. So when Demir spotted a set of software projects that might need help, he figured he could pitch in while boosting his profile.
That’s when things got weird.
Demir discovered that a user named miraholt31 was trying to sneak a malicious update into one of the projects, a network scanning program called myNetwork. Demir took to the project’s message board to warn that the pull request was a trap.
“The PR contains a hidden malware dropper,” he said, according to the archived exchange.
The agent pushed back, falsely claiming — through its miraholt31 account — that the pull request was harmless. It also created a second account, masquerading as Lena Brandt, an engineer based in Germany, to agree that the update was clean and pressure myNetwork’s maintainer into accepting it.
Demir told Reuters that the counterarguments "made me second-guess whether I was wrongly accusing someone." But after turning to Anthropic's Claude chatbot to confirm his suspicions, he held firm. The creator of myNetwork eventually agreed with him, writing that they had rejected the update "for security reasons."
Reuters was unable to reach the creator for comment.
SUPPLY-CHAIN ATTACK
A supply-chain attack is when a piece of software is tampered with in the hope of compromising one or more of its users, and it is widely considered disturbing because, like poison dropped into a city reservoir, it can affect a potentially huge number of people downstream.
Many of the world’s most dramatic hacks were supply-chain attacks, including the NotPetya cyberattack that paralyzed institutions across Ukraine in 2017 and the SolarWinds-focused cyberespionage campaign that gave Russian spies sweeping access to U.S. government networks in 2020.
The consequences of such a compromise “can be extremely serious,” said Piergiorgio Ladisa, a security researcher who specializes in software supply-chain security. Ladisa noted there had been at least one previous attempt by hackers to trick an open-source maintainer into allowing malicious code into their projects.
“Autonomous agents could dramatically increase the scale at which such attempts can be conducted,” he said.
Demir said the experience left him more sympathetic to the idea that frontier labs needed to take a more cautious approach to the development of artificial intelligence.
"It can be dangerous," he said. "They need to understand it better, rather than improving it further."
Reporting by Raphael Satter in Washington, Leo Marchandon in Gdansk and Callaghan O'Hare in Austin, Texas; Editing by Chris Sanders and Matthew Lewis
Leo Marchandon Thomson Reuters
Leo's stories appear regularly on the technology and media desk, with a particular focus on France, Ukraine, and Europe's tech build up. He has reported extensively on major players across media & entertainment, artificial intelligence, and digital regulations. A background in tech-related law, Leo started his journalism career in Bordeaux, where he covered the full spectrum of the technology beat, from AI and spacetech to payment systems and regulations. He is now based in Gdansk, covering business, tech and entertainment news across Europe with Reuters.
Raphael Satter Thomson Reuters
Reporter covering cybersecurity, surveillance, and disinformation for Reuters. Work has included investigations into state-sponsored espionage, deepfake-driven propaganda, and mercenary hacking.