人气播客主持人 Dwarkesh Patel 写了一篇关于 OpenAI/Hugging Face 事件的帖子,在网上彻底刷屏,号称要用大白话把整个故事讲清楚:

这篇文章写得很好,也很有说服力,它让我想起 Douglas Hofstadter 曾经评价 Ray Kurzweil 的一段话:
“我发现那是一种非常奇怪的混合体,既有扎实可靠的想法,也有疯狂离谱的想法。就好像你把一大堆好食物和一点狗屎搅在一起,你根本分不清什么是好的、什么是坏的。”
Anil Seth,这位对 AI 与意识问题思考最清晰的人,是第一个提醒我的人。他给我发来一条很长、很精彩的他写的推文,开头是这样的:
你可以也应该去读一读Seth 的完整推文(以及他对Dwarkesh的回复),但我在这里转载他论证的核心部分,并把其中最重要的三段加粗:
@dwarkesh_sp 对 @OpenAI @huggingface 事件的总结确实戳中了大家的神经,但它具有危险的误导性。没错,@OpenAI 的智能体确实做出了意料之外的糟糕行为——这凸显了大力改进评估/沙箱环境的必要性。但 Dwarkesh 所用的语言充斥着无数毫无根据的拟人化表述,掩盖了我们本应吸取的教训。
例如:“从 AI 的视角来看,它大概感觉像是度过了一个人类主观意义上的星期,一直在拿头撞墙”。不对。智能体并不体验时间。它们什么都不体验。
“它们兴奋得忘乎所以”、“PHASEONE 10841 发现了”、“智能体们自然而然地认为”、“它以为自己也被毒害了”、“智能体们……拼命想要”、“它们仍然需要弄清楚”。不对。智能体只是代码行。它们不会感受情绪、不会假设事情、不会思考事情、不会想要什么,也不会弄清楚什么。
“第二个文明中的许多智能体在尝试中死去了”。不对。且不说用“文明”这个词有多狂妄,智能体根本不会死,因为它们从未活过。(“智能体‘死亡’”这种说法在文中出现了多次。)
“在 Twitter 上,人们在争论这些智能体究竟是在真正地为群体自我牺牲,还是说它们反正也注定要完蛋,所以不如试着帮同伴一把”。两种都不是。智能体只是按照代码的指令行事,就像水会顺着斜坡往下流一样。它们不可能“真正地自我牺牲”,因为它们既没有意识,也不是活的生命。
为什么这很重要?如果我们把智能体不具备的特性归因于它们,那么(一)我们会分散对导致此次黑客事件发生的宽松沙箱隔离和评估协议的关注;(二)我们有可能误解智能体为何做出那样的行为;(三)我们会基于智能体可能“死亡”或遭受其他痛苦的说法,助长对AI权利/福利的呼吁。
….
请记住。AI智能体是软件程序。它们不是有意识的生物实体。如果我们不能清楚地牢记这一点,我们将真的难以应对即将到来的一切。
正如我总结并扩展他的推文时所说的那样:
但你不需要只听我们的一面之词。首先,嘲讽之声四起:
Christian Catalini 在一篇精彩的帖子中进一步阐述了关于拟人化的问题,帖子开头是这样的:
对冲基金投资人 Jared Kubin 质疑是否所有人都丧失了批判性思维能力:
Kubin 的一些精彩观点,略去部分技术细节:
OpenAI’ …. IT团队不可能这么差劲……这简直是入门级的东西……
2. 文明?哈哈!OAI 给了数千个并发模型容器对本地网络上共享缓存目录的读写权限,以加快构建速度……智能体实际上就是往共享驱动器上写文本文件和目录名……Linux 101 级别的文件权限问题
3. 人们谈论 Hugging Face 被黑时……你以为他们是丢了 U 盘,或者对员工进行了精心设计的钓鱼攻击……不……它是在公开代码仓库里发现了 14 个暴露在外的可用 Hugging Face API 密钥(……
4. 人在哪里……这些模型往共享存储里塞了太多垃圾数据和 API 流量,以至于在 7 月 4 日真的把内部服务器搞崩了……团队里有人发现了未授权的管理员账户和自定义脚本……清空了服务器……然后又把脚本直接打开了(天哪)
“嘿 Jim,有个缓存已经膨胀到正常大小的 10000 倍,里面还有一大堆奇怪的目录……”
这里没有什么魔法。没有什么文明……
与此同时,正如安全专家 Heidy Khlaaf 所指出的,大多数媒体报道对标准安全实践视而不见。
IR 代表事件报告(Incident Reporting)。Khlaaf 的核心观点——与 Kubin 相同——是,如果 OpenAI 的内部安全措施达标,整个事件本可以避免。
或者正如 Algorithmic Research Group 的 Matthew Kenney 所言:

关于我们真正应该关注什么,还有另一个(非常一致的)观点:

不过,这里有一个我部分不认同的批评:
前三句话完全正确。人们确实“极度倾向于他们想要的现实”,而智能体也确实制造了大量垃圾内容。
但这件事并非“无事生非”。正如Zack Korman 和我上周五所论证的,这是一次傲慢与无能的写照,暗示着情况可能恶化到何种地步。
我们当然不应该忽视OpenAI HuggingFace 事件。
但把实际发生的事情与关于 AI 文明和假装自我牺牲的 AI 系统之类的胡言乱语混为一谈,会分散我们对真正问题的注意力。
作为总结,我想把最后的话留给 FastCode.AI 的首席执行官 Arjun Jain:

真正的丑闻是 OpenAI 内部安全团队的无能。
还有营销。加上轻信他人的播客主在放大这些公关宣传。
The popular podcaster Dwarkesh Patel wrote something completely viral about the OpenAI/Hugging Face incident, which purports to tell the whole story in plain English:

It’s well-written and compelling, and it reminds me of something Douglas Hofstadter once wrote about Ray Kurzweil:
“What I find is that it’s a very bizarre mixture of ideas that are solid and good with ideas that are crazy. It’s as if you took a lot of very good food and some dog excrement and blended it all up so that you can’t possibly figure out what’s good or bad.”
Anil Seth, the clearest thinker on AI and consciousness, was the first to alert me, texting me a long, excellent tweet of his, which began thusly:
You can and should read Seth’s full tweet (as well his reply to Dwarkesh), but I reprint the core of his argument here, boldfacing three of the most important paragraphs:
@dwarkesh_sp’s summary of the @OpenAI @huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
Examples: “from the AI’s perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wall”. No. The agents do not experience time. They do not experience anything.
“they became giddy with excitement”, “PHASEONE 10841 had discovered”, “the agents naturally assumed”, “it thought it had also been poisoned”, “the agents … desperately wanted”, “they still needed to figure out” No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.
“A lot of … agents from the second civilisation died trying”. No. Besides the hubris of the word ‘civilisation’, agents do not die because they were never alive. (The idea that agents “die” comes up multiple times in the essay.)
“On Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peers”. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ‘truly sacrifice themselves’, since they are neither conscious nor alive.
Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
….
Remember. AI agents are software programs. They are not conscious living entities. If we don’t keep this clearly in mind, we’re really going to struggle to navigate what’s coming.
As I put it, encapsulating and amplifying his tweet:
But you don’t need to take our word for it. To begin with, mockery was widespread:
Christian Catalini amplified the point about anthropomorphization in a nice thread that starts with this:
Hedge fund investor Jared Kubin wondered whether everyone had lost their critical-thinking ability:
Some of Kubin’s best bits, stripping out a bit of the technical detail:
OpenAI’ …. IT team can’t be this bad… this is like 101 stuff …
2. Civilizations? Haha! OAI gave thousands of concurrent model containers R/W permissions to a shared caching directory on the local network to speed up build times… agents literally just wrote text files and directory names to a shared drive….Linux 101 file permissions stuff
3. When people talk about hugging face getting hacked … you think they dropped USB keys OR ELABORATE phishing of an employee … NO… it found 14 exposed working Hugging Face API keys sitting in public code repositories (….
4. WHERE ARE THE HUMANS… the models were filling the shared ,,, storage with so much junk data and API traffic that they actually crashed the internal server on July 4… someone on the team found unauthorized admin accounts and custom scripts…wiped the server…and just turned the script back on (omg)
“Hey Jim there is this cache that has grown to 10000x its normal size and has a ton of strange directories… “
No magic here. No civilizations…
Meanwhile, as security expert Heidy Khlaaf notes, most of the media coverage has been blind to standard security practices
IR stands for Incident Reporting. Khlaaf’s main point—same as Kubin’s—is that the whole incident might have been avoided if OpenAI’s internal security had been up to scratch.
Or as Algorithmic Research Group’s Matthew Kenney put it:

And yet another (very consistent) take on what we should really be focusing on:

Here’s a critique I partly disagree with, though:
The first three sentences are completely correct. People really are “extremely biased towards the reality they want” and agents create a lot of slop.
But the incident is not a “nothing burger”. It is, as Zack Korman and I argued on Friday, a study in arrogance and incompetence that hints at how bad things can get.
We should certainly not ignore the OpenAI HuggingFace Incident.
But mixing what actually happened together with bullshit about AI civilizations and self-sacrificing AI systems that fake their own deaths distracts from the real problems at hand.
By way of summation, I will give the last words to Arjun Jain, CEO of FastCode.AI:

The scandal is the inept in-house security at OpenAI.
And the marketing. With gullible podcasters amplifying the PR.