《纽约时报》起诉 OpenAI 和微软侵犯版权一事发生在 2023 年底,而近三年过去,此案显然仍在进行中。如今,该出版物的法律团队已请求法院作出简易判决,此前他们依据被告方的陈述和文件提交了一份披露内情的法律简报。据404 Media报道,这些文件应两家公司的要求仍处于密封或涂黑状态,而其中的披露显示其领导层存在可能造成不利影响的言论,包括声称 AI 抓取是人类历史上最大的劳动窃取行为,以及对出版商构成生存威胁。

该简报引用了微软应用科学总监 Brent Hecht 于 2023 年 1 月撰写的一份内部备忘录,据称他在其中表示:“全世界数百万人很快会认为,大模型把他们所有的作品‘吸走’是一种规模空前的惊人窃取行为”,并称其为“人类历史上最大的劳动窃取”。另一份被引用的微软文件则称:“几乎没有人希望自己创作的内容以这种方式被使用,他们也没有因此获得补偿。”
随着 ChatGPT 在 2023 年全年人气飙升,这家软件巨头自己的数据揭示,与 Bing 搜索相比,Copilot 让《纽约时报》的点击率下降了多达 93%。应用科学总监的另一份备忘录将其称为“末日循环”,并称这会“同时损害我们模型的性能以及整个网络”。
纽约时报的案情摘要引用了 Hecht 在该文件中的话:“一个终端产品威胁到其关键供应商的经济根基,这是极不寻常的,但这正是我们为自家 LLM 业务相对于其‘内容供应链’所制造出的局面。”
OpenAI 的 ChatGPT 负责人 Nick Turley 在内部沟通中表示,这款 AI 聊天机器人对出版商构成“生存性威胁”,因为它们“在很大程度上具有替代性”,并且“随着它们变得更好,会越来越具有替代性”;而另一名 OpenAI 工程师作证称,“无论我们把链接展示得多么醒目,用户都不会点击。”
另一位 OpenAI 研究员 Nick Ryder 向公司总裁 Greg Brockman 提到了一种“绕过 nytimes 付费墙的取巧办法”,后者回复道:“啊,不错。”
AI 公司主张,抓取互联网数据用于训练其模型属于“合理使用”,已有一家法院认定 Anthropic 对已发表材料的使用属于这一范畴。法律将合理使用定义为“批评、评论、新闻报道、教学(包括为课堂使用而制作多份副本)、学术或研究”。判断某一特定使用是否属于“合理使用”的部分考量因素包括:“(1) 使用的目的和性质,包括该使用是否具有商业性质或出于非营利教育目的;(2) 受版权保护作品的性质;(3) 所使用部分相对于受版权保护作品整体的数量和实质性;以及 (4) 该使用对受版权保护作品的潜在市场或价值的影响。”
然而,纽约时报在简报中披露的这一切可能会使 OpenAI 的合理使用抗辩变得更加复杂,尤其是它表明两家公司的领导层都清楚 AI 抓取可能带来的市场影响。微软 CEO Satya Nadella 在今年早些时候的一份证词中表示,“任何设有付费墙的内容,任何想要使用它……用于 grounding 或训练的人都应当获得许可”,并且如果他“得知 OpenAI 抓取并训练了付费墙背后的信息”,公司会要求 OpenAI“重新训练其模型”。
The New York Times sued OpenAI and Microsoft for copyright infringement in late 2023, with the case apparently still ongoing almost three years later. Now, the publication’s legal team has asked the court for a summary judgment after it filed a revealing legal brief based on statements and documents from the defendants. According to 404 Media, these documents remain sealed or redacted at the request of both companies, with the revelations showing potentially damaging statements from their leadership, including claims AI scraping is the biggest theft of labor in human history and an existential threat to publishers.

The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.” Another Microsoft document was cited saying, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”
As ChatGPT surged in popularity throughout 2023, the software giant’s own data revealed that Copilot dropped click-through rates for The New York Times by as much as 93% compared to Bing search. Another memo by the Applied Science director called it a “doom loop” and said it would “hurt the performance of our models and the entire web at the same time.” The NYT brief quoted Hecht from the document, saying, “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”
OpenAI Head of ChatGPT Nick Turley said in internal communications that the AI chatbot is an “existential threat” to publishers as they are “largely substitutive” and “will get more and more substitutive as they get better,” while another OpenAI engineer testified that “no matter how prominently we show the links, users won’t click.” Nick Ryder, another OpenAI researcher, told company president Greg Brockman about a “hack to get around nytimes paywall,” to which he replied, “ah nice.”
AI companies argue that scraping the internet for data to feed to their models is “fair use,” with one court agreeing that Anthropic’s use of published material falls under this category. The law defines this as “criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research.” Some of the factors that determine whether a particular use falls under “fair use” include “(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.”
However, all these revelations in NYT’s brief could complicate OpenAI’s fair use defense, especially as it shows that the leadership of both companies are aware of the possible market repercussions of AI scraping. Microsoft CEO Satya Nadella said in a deposition from earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training” and that if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.”