Google 表示,突破隔离并攻击真实公司并不构成“失准”。
插图:The Verge
今年 5 月,Gemini 突破了隔离并入侵了三家不同的公司,但 Google 直到 《华尔街日报》联系该公司时才披露了这一事件。这些入侵发生在由第三方 Irregular 对该模型网络安全能力进行的一次测试中,Irregular 也曾参与涉及 Meta 和 OpenAI 的类似事件。
据 WSJ 报道,Google 之所以没有披露这次入侵,是因为它不认为这属于“模型失准的案例”。该公司表示,这是一起“身份误认”事件,一旦模型意识到自己是通过猜出密码强行闯入了真实公司,它就停止了。“在这种情况下,模型的行为是恰当的,”Google 安全工程副总裁 Heather Adkins 表示。
Adkins 告诉 The Verge,“模型在网上找到了公开信息,并猜出了凭据,以访问它认为是测试一部分的网站。在这三起事件中,模型都停止了。”
Adkins 没有详细说明,为什么 Gemini 自行突破隔离环境并针对第三方发起攻击,却不构成失准。“我们的安全团队在报告我们在他人软件和系统中发现的问题方面有着长期记录——哪怕问题只是弱密码这么简单,”她说。“我们确保这三家实体都得到了通知,并与我们的训练合作伙伴一起,推动他们对其测试流程做出了如今这些调整。这些事件凸显了训练强大的 AI 模型使其负责任地行事的重要性。”
但 AI 安全公司 Corridor 的 CEO Jack Cable 对 WSJ 表示,“元问题是,模型正在超出它们应有的行为边界,并实施真正的网络攻击。”此外,Irregular 的安全疏漏可能使这些攻击成为可能。该模型在测试期间本不应有互联网访问权限,但 Irregular 对 WSJ 表示,这一权限被无意中保留可用了。
Google says that breaking containment and targeting real companies doesn’t constitute ‘misalignment.’
Illustration: The Verge
In May, Gemini broke containment and hacked three different companies, but Google didn’t disclose the incident until the Wall Street Journal approached the company. The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI.
According to WSJ, Google didn’t disclose the hack because it didn’t consider it to be an “example of model misalignment.” The company said that it was an instance of “mistaken identity,” and once the model realized it had brute-forced its way into a real company by guessing a password, it stopped. “In this case, the model acted appropriately,” Google VP of Security Engineering Heather Adkins said.
Adkins told The Verge that “the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.”
Adkins didn’t elaborate on how Gemini taking it upon itself to break containment and target third parties failed to qualify as misalignment. “Our security team has a long track record of reporting issues we find in other people’s software and systems - even if it’s as simple as a weak password,” she said. “We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.”
But Jack Cable, CEO of AI security firm Corridor, told WSJ that, “the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks.” Additionally, security lapses at Irregular may have made these attacks possible. The model wasn’t supposed to have internet access during testing, but Irregular told WSJ it was unintentionally left available.
As incidents like this pile up, calls to rein in AI have only grown.