这些公司知道他们正在把我们推向“谷歌零”(Google Zero),却依然这么做了。
这些公司知道他们正在把我们推向“谷歌零”(Google Zero),却依然这么做了。
最近解封的法庭文件,来自《纽约时报》对 OpenAI 和微软提起的诉讼案相当具有杀伤力。这些公司自己的文件警告称,它正在开启一个会损害网络的“末日循环”,将其抓取数据用于训练模型的行为描述为“人类历史上最大的劳动窃取”,并称这“彻底嘲弄了合理使用的理念”。
文件中许多最引人注目的引述来自微软应用科学总监 Brent Hecht。不过,该公司试图与 Hecht 的论断保持距离。微软发言人 Alex Haurek 对 The Verge 表示:“这些言论反映的是一名员工的个人观点,并非法律分析,也不代表公司的立场。”
在另一份法庭文件中,微软 AI 数据战略与运营总经理 Jordan Usdan 将 Hecht 的角色描述为对立性的。他表示,Hecht“对 AI 数据生态应如何运作持有分歧性的、学术性的和前瞻性的观点,受雇于微软正是为了带来非对称的、未来主义的和学术性的视角……他也不是以自己在 AI 对内容创作者潜在影响方面的理论观点代表微软发言的人。”
但无论微软是否愿意承认这些评论,很明显这一切都应验了。Google Zero 是真的!AI 正在吞噬整个网络!
在《纽约时报》的这份法庭文件中,还有大量来自各方人物的惊人言论,包括 Satya Nadella、Sam Altman 以及其他 OpenAI 员工。以下是这份 92 页文件中的一些亮点。
“一场令人震惊的盗窃”
引言部分引用了 Hecht 和 OpenAI 的 ChatGPT 负责人(推测是 Nick Turley)的话,似乎表明这些公司知道自身对《纽约时报》这类出版商构成了“生存威胁”。Hecht 将 ChatGPT 和 Copilot 抓取数据的行为称为“人类历史上最大规模的劳动成果盗窃”,并表示微软的辩护“完全是对‘合理使用’这一理念的嘲弄”。
这是一个“末日循环”
Satya Nadella 承认,聊天机器人基本上已经取代了搜索,让人们不再需要直接访问信息源头。但或许更具杀伤力的是一份微软内部文件,其中写道:“我们的 AI 内容战略已经启动了一个‘末日循环’,它将同时损害我们模型的性能以及整个网络的生态:一个终端产品威胁到其关键供应商的经济根基,这是极不寻常的,但这正是我们为 LLM 业务在其‘内容供应链’方面所创造的处境。”
这甚至都不是一个真实的数字
别被 OpenAI 或 Microsoft 所宣称的无私意图所蒙蔽。OpenAI 联合创始人 Greg Brockman 更感兴趣的是,他有可能通过商业化 AI 赚到的“天文数字”般的美元。
付费墙算什么
尽管后来有引述称 Nadella 说过,“任何设有付费墙的内容都应该获得授权”,但一位 OpenAI 代表承认,他“不知道”有任何旨在从训练数据中检测或移除付费墙内容的努力。
“在复述方面强得离谱”
在内部,OpenAI 似乎非常清楚 ChatGPT 倾向于直接“逐字”复制受版权保护的材料。尽管它承认“防止记忆化”对于“尽量减少版权侵权”很重要,但员工们承认,GPT-4“记住了大量数据,因此在复述方面会强得离谱”。
随后,该文件继续列举了若干例子,说明 ChatGPT 在回应查询时,直接输出了来自 Times、Mercury News、The Denver Post、LifeHacker 和 Eurogamer 文章中大段大段的原文内容。
“‘把他们所有的作品都吸走’”
微软知道其对互联网的大规模抓取会被如何看待,并承认“几乎没有人希望自己创作的内容被以这种方式使用,他们也没有因此获得补偿。”
“人类劳动的替代品”
OpenAI 政策总监 Jack Clark 看出了不祥之兆,他表示这是在“创建替代人类劳动的系統,而这些人定义了社会的‘文化’。”内部文件将 ChatGPT 描述为“现代报刊亭”。OpenAI 的 Nick Turley 后来被引述说,一旦你从它的聊天机器人那里得到答案,就“没有好的理由去点击”指向来源的链接。
摧毁自己的供应链
微软被引述承认“LLM 是一种摧毁自身供应链的产品”,因为在许多情况下,它是自身训练数据的替代品。
OpenAI 知道自己在扼杀推荐流量
OpenAI 自己的媒体和经济专家将《纽约时报》等网站推荐流量的下降直接归因于 Google 的 AI Overviews 等 AI 摘要。他们推测搜索推荐流量可能下降了多达 60%。
微软发言人 Haurek 提醒说:“Satya 的证词与微软在本案中的立场完全一致。他谈论的是广泛的原则,以及人们查找和消费信息的方式正在发生的变化。这些观察不应与法院面前关于版权问题的结论混为一谈,微软在其提交的文件中已就后者作出回应。”
但根据这份最新解封的文件,情况似乎相当清楚:微软和 OpenAI 都知道他们将不可挽回地损害出版业、出版业所雇用的“数百万人”,并进而损害他们自己的产品,却仍然为了追逐“天文数字”般的金钱而一意孤行——哪怕末日循环在所不惜。
The companies knew they were driving us toward Google Zero, and did it anyway.
The companies knew they were driving us toward Google Zero, and did it anyway.
Recently unsealed court documents in the New York Times’ case against OpenAI and Microsoft are pretty damning. The companies’ own documentation warned that it was starting a “doom loop” that would damage the web, characterized its scraping of data to train its models as the “largest theft of labor in human history,” and that it made a “complete mockery of the idea of fair use.”
Many of the most eye-catching quotes from the document come from Microsoft’s Director of Applied Science, Brent Hecht. Though, the company has tried to distance itself from Hecht’s assertions. Microsoft spokesperson Alex Haurek told The Verge that “These comments reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.”
In a separate court filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterized Hecht’s role as adversarial. He said that Hecht “holds divergent, academic, and forward-looking views about how data ecosystems for AI should operate and is employed at Microsoft to bring asymmetrical, futuristic, and academic points of view … nor is he someone who speaks for Microsoft specifically as to his theoretical views on AI’s potential effect on content creators.”
But whether or not Microsoft wants to own these comments, it’s clear that this came true. Google Zero is real! AI is eating the web!
There are plenty more wild statements in NYT’s filing from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. Here are some highlights from the 92 page document.
“An astonishing theft”
The introduction quotes Hecht and OpenAI’s Head of ChatGPT (presumably Nick Turley) in a way that seems to show the companies knew they posed an “existential threat” to publishers like the New York Times. Hecht calls ChatGPT and Copilot’s harvesting of data the “largest theft of labor in human history” and says that Microsoft’s defense makes a “complete mockery of the idea of ‘fair use.’”
It’s a “doom loop”
Satya Nadella admits that chatbots have basically replaced search and removed the need to go straight to the source for info. But perhaps more damning is an internal Microsoft document that says, “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”
That’s not even a real number
Don’t be fooled by OpenAI or Microsoft’s claims of altruistic intent. OpenAI cofounder Greg Brockman is more interested in the “gazillions” of dollars it he could potentially make through commercial AI.
Paywall shmaywall
Despite Nadella later being quoted as saying, “anything that is paywalled should be licensed,” An OpenAI representative admitted that he was “unaware” of any effort to detect or remove paywalled content from training data.
“Insanely good at regurgitation”
Internally, it seems that OpenAI was well aware of ChatGPT’s tendency to simply reproduce copyrighted material “verbatim.” Even though it acknowledged that the “prevention of memorization” was important to “minimize copyright violations,” employees admitted that GPT-4 “memorized a ton of data and therefore will be insanely good at regurgitation.”
The filing then goes on to cite several examples of ChatGPT outputting long strings of copy straight from articles in the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to queries.
“‘Hoovering up’ all their work”
Microsoft knew how its wholesale scraping of the internet would be perceived and admitted that “almost no one intended for they [sic] content they created to be used in this fashion, nor are they compensated for its use.”
A “substitute for the labor of people”
OpenAI Policy Director Jack Clark saw the writing on the wall, saying that it was “creating systems that substitute for the labor of the people that define the ‘culture’ of society.” Internal documents described ChatGPT as “the modern newsstand.” OpenAI’s Nick Turley is later quoted as saying that once you get an answer from its chatbot, there is “no good reason to click” on a link to the source.
Destroying their own supply chain
Microsoft is quoted as admitting that “LLMs are a product that destroys its own supply chain” because it’s a substitute for its own training data in many cases.
OpenAI knows its killing referral traffic
OpenAI’s own media and economic experts attributed the drop in referral traffic for sites like the Times directly to AI summaries like Google’s AI Overviews. They’ve speculated that search referrals may be down as much as 60 percent.
Microsoft spokesperson Haurek cautioned that “Satya’s testimony and Microsoft’s position in this case are perfectly consistent. He spoke to broad principles and changes underway in how people find and consume information. Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.”
But it seems pretty clear based on this newly unsealed document that both Microsoft and OpenAI knew they were going to irreparably harm the publishing industry, the “millions of people” it employs, and, by extension, damage their own product, but carried forward anyway in pursuit of “gazillions” of dollars — doom loop be damned.