由 THE DECODER 提示的 Nano Banana Pro
要点
- 包括《纽约时报》在内的多家美国媒体公司正寻求向 OpenAI 和 Microsoft 索赔数十亿美元,指控其在 AI 训练中涉嫌侵犯版权。
- 原告引用了内部信息和宣誓证词,其中高管们对自身的合理使用抗辩提出质疑,并将聊天机器人描述为原创新闻业的替代品。
- 该文件还指控 OpenAI 系统性地绕过付费墙、违反训练数据的许可条款,并在诉讼提起后部署过滤器以压制证据。
《纽约时报》及其他原告提交的一份新法庭文件引用了此前未披露的 OpenAI 和 Microsoft 高管内部邮件及宣誓证词。
《纽约时报》及其他多家媒体公司已在纽约美国联邦地区法院联合提交了一份简易判决书。与《纽约时报》一同加入的还有 Daily News 集团(旗下包括《芝加哥论坛报》和《丹佛邮报》),以及Ziff Davis(旗下拥有 CNET、IGN 和 PCMag)。原告还包括调查报道中心(Mother Jones 的所属机构)以及The Intercept。
据《金融时报》报道,原告正在寻求数十亿美元的损害赔偿。这份长达92页的简报是合并多地区诉讼的一部分,该诉讼汇集了多起案件,始于2023年12月提起的《纽约时报》诉讼。
高管的内部言论削弱了合理使用抗辩
该简报援引了在证据开示过程中披露的内部邮件、Slack 消息和宣誓证词,其中包括削弱 AI 公司合理使用抗辩的言论。
微软应用科学总监 Brent Hecht 称这种做法是“一次规模空前的惊人盗窃”,可能是“人类历史上最大的劳动力盗窃”。他还写道,合理使用抗辩若获成功,可以说将“彻底嘲弄‘合理使用’这一概念”。微软对《金融时报》表示,这些评论“反映了一名员工的个人观点”,“并非法律分析”。
美国版权局也在2025年5月得出结论,鉴于 AI 公司复制数据的规模之庞大,合理使用无法广泛适用。负责监督该报告的官员随后被特朗普政府解雇。
根据该简报,OpenAI 的 ChatGPT 负责人 Nick Turley 写道,出版商面临“生存威胁”。他说,这些产品“在很大程度上具有替代性,就是这样”,并且“随着它们变得更好”,将越来越多地取代出版商提供的内容。一位 OpenAI 工程师补充说,“无论我们把链接展示得多么醒目,用户都不会点击。”
微软 CEO Satya Nadella 在宣誓后确认,聊天机器人对话已经取代了对原始来源的访问。在给《金融时报》的一份声明中,微软表示他的证词涉及“广泛原则以及人们查找和消费信息方式正在发生的变化”。该公司表示,这并不是对本案核心“版权问题的结论”。
一份微软内部文件描述了一种自我强化的循环,称之为“末日循环”。
我们的人工智能内容战略已经启动了一个“末日循环”,这将同时损害我们模型的性能以及整个网络:一种终端产品威胁其关键供应商的经济基础,这是极不寻常的,但这正是我们为 LLM 业务在其“内容供应链”方面所创造的处境。
微软自己的数据显示,Copilot 上的点击率远低于传统 Bing 搜索。《纽约时报》的点击率下降了 87% 至 93%,Daily News 集团下降了 83% 至 91%,Ziff Davis 下降了 51% 至 94%。
OpenAI 在内部将本地新闻——Daily News 集团的核心业务——描述为 ChatGPT 中一个“相当常见的查询”。其他研究也发现,聊天机器人的回答会大幅减少开放网络的流量。
OpenAI 联合创始人 Greg Brockman 在一封内部消息中讨论了这些模型处理新闻的能力。
顺便说一句,我们做新闻非常出色。每次我在《纽约时报》上做任何生成式的东西,它似乎都能相当准确地预测下一句。
付费墙绕过手段和许可限制让 OpenAI 的辩护变得复杂
原告描述了 OpenAI 如何系统性地绕过付费墙并无视服务条款。当一名员工向 Brockman 提到“一个绕过 nytimes 付费墙的窍门”时,他回复道“啊,不错”。大约在 2017 年,Brockman 还写道,他“被那海量的财富深深驱动”,他希望通过将 OpenAI 的技术商业化来赚取这些财富。
OpenAI 的企业代表作证称,他不知道有任何方法可以检测训练数据中的付费墙内容,也没有做出任何移除这类内容的努力。该公司的标准爬取流程“不包括审查网站[的]……使用或服务条款”。
Nadella 在宣誓后作证称,“任何设有付费墙的内容,都应获得任何想要使用它的人的授权许可。”他表示,如果他当时知道 OpenAI 抓取了付费墙内容并将其用于训练,他会强制 OpenAI 重新训练其模型。
OpenAI 还通过第三方获取了“纽约时报标注语料库”,这是一个包含 180 万篇文章的合集。其许可证将使用范围限制为“非商业性的语言教育、研究与技术开发”。OpenAI 员工知道使用该语料库训练模型“并不合适”,但仍然这样做了。
原告指控 OpenAI 压制证据,并对全部四项合理使用因素提出质疑
根据该简报,诉讼提起后,OpenAI 立即构建了一个过滤器,以抑制最有可能来自原告出版物的输出内容。未提起诉讼的公司的内容则未受影响。原告主张,该过滤器的设计目的并非保护版权,而是阻止他们收集证据。
微软的 Hecht 曾提出担忧,认为 OpenAI 可能会部署这样的过滤器,并称其为“意外的掩盖”。他警告说,这将导致“对内容拥有权利的人对训练所用内容的可见性降低”。即便如此,该简报仍列举了大量有问题的输出示例,ChatGPT 应要求逐字复制并摘要了纽约时报的文章,包括付费墙后的文章。
原告还主张,合理使用抗辩在全部四项法定因素上均不成立。他们称,该使用是替代性的和商业性的,而非转换性的,而这些文章是处于版权保护核心的表达性作品。被告还复制了整部作品,尽管他们自己的专家也承认没有任何单一作品是必需的。
关于市场损害,原告指出了现有的授权市场以及那些有可能切实发展起来的授权市场。OpenAI 和 Microsoft 已与其他出版商签署了授权协议,Amazon、Google、Meta 和 Perplexity 也是如此。但该简报辩称,被告没有为原告的内容付费,而是免费取用并在彼此之间进行交易。
一个运转正常的授权市场会让合理使用抗辩更难成立。当被告已经在为其他出版商的同类内容付费时,他们无法可信地声称授权是不可能的。
AI 花 6,800 美元就能生成一百万篇“粉色黏液”文章
原告还质疑了 AI 模型不与出版商竞争这一论点,进一步削弱了合理使用抗辩。他们主张,OpenAI 的产品让用户能够向市场大量倾泻“粉色黏液”,即低质量、往往抄袭的伪新闻。按照 OpenAI 的 API 价格,生成一百万篇每篇 500 字的新闻风格文章大约需要 6,800 美元,且无需任何记者参与。该简报援引了 Prism News 的案例,该公司仅用四名员工就运营了 200 个 AI 生成的出版物,冒充本地新闻编辑室和兴趣网站。
OpenAI 在 2024 年宣布了一款名为“Media Manager”的工具,本应让出版商可以选择退出抓取。根据该简报,该项目此后已被搁置。
原告请求法院认定构成版权侵权并驳回合理使用抗辩。他们还希望在计算任何法定赔偿时,将每篇文章视为一件独立作品。
此前的法院裁决倾向于合理使用,而美国司法部在特朗普政府执政期间也站在了 AI 实验室一边。OpenAI 和 Microsoft 很可能很快就会提出自己的主张,重点围绕合理使用及其工作的变革性。如今已进入公开记录的内部文件,不会让这一论点变得更容易成立。
简讯
Nano Banana Pro prompted by THE DECODER
Key Points
- Several US media companies, including The New York Times, are seeking billions of dollars in damages from OpenAI and Microsoft over alleged copyright infringement in AI training.
- The plaintiffs cite internal messages and sworn testimony in which executives questioned their own fair use defense and described chatbots as substitutes for original journalism.
- The filing also accuses OpenAI of systematically bypassing paywalls, violating license terms for training data, and deploying filters to suppress evidence after the lawsuits were filed.
A new court filing by The New York Times and other plaintiffs cites previously undisclosed internal emails and sworn testimony from OpenAI and Microsoft executives.
The New York Times and several other media companies have filed a joint summary judgment brief in US District Court in New York. Joining the Times are the Daily News group, which includes the Chicago Tribune and Denver Post, and Ziff Davis, which owns CNET, IGN, and PCMag. The plaintiffs also include the Center for Investigative Reporting, home to Mother Jones, and The Intercept.
The plaintiffs are seeking billions of dollars in damages, according to the Financial Times. The 92-page brief is part of consolidated multidistrict litigation that brings together several lawsuits, beginning with the New York Times suit filed in December 2023.
Executives' internal statements undercut the fair use defense
The brief draws on internal emails, Slack messages, and sworn testimony disclosed during discovery, including statements that undercut the AI companies' fair use defense.
Microsoft's director of applied science, Brent Hecht, called the practice "an astonishing theft of unprecedented proportions" and possibly the "largest theft of labor in human history." He also wrote that a successful fair use defense would arguably "make a complete mockery of the idea of 'fair use.'" Microsoft told the Financial Times these comments "reflect one employee's individual perspective" and "are not a legal analysis."
The US Copyright Office also concluded in May 2025 that fair use cannot apply broadly given the sheer scale at which AI companies copy data. The official who oversaw the report was later fired by the Trump administration.
Nick Turley, OpenAI's head of ChatGPT, wrote that publishers face an "existential threat," according to the brief. He said the products "are largely substitutive, period" and would increasingly replace publishers' offerings "as they get better." An OpenAI engineer added that "no matter how prominently we show the links, users won't click."
Microsoft CEO Satya Nadella confirmed under oath that chatbot conversations had replaced visits to original sources. In a statement to the FT, Microsoft said his testimony concerned "broad principles and changes under way in how people find and consume information." It was not a "conclusion about copyright questions" at the center of the case, the company said.
An internal Microsoft document describes a self-reinforcing cycle it calls a "doom loop."
Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.
Microsoft's own data shows that click-through rates on Copilot were much lower than on traditional Bing search. Rates were down 87 to 93 percent for the New York Times, 83 to 91 percent for the Daily News group, and 51 to 94 percent for Ziff Davis.
OpenAI internally described local news, a core business for the Daily News group, as a "pretty common quer[y]" in ChatGPT. Other studies have also found that chatbot answers sharply reduce traffic to the open web.
OpenAI co-founder Greg Brockman discussed the models' ability to handle news in an internal message.
we are excellent at news btw. every time i do any generative stuff on NYT it seems to predict the next sentence pretty well.
Paywall workarounds and license restrictions complicate OpenAI's defense
The plaintiffs describe how OpenAI systematically bypassed paywalls and ignored terms of service. When an employee told Brockman about "a hack to get around nytimes paywall," he replied "ah nice." Around 2017, Brockman also wrote that he was "deeply motivated by the gazillions" he hoped to earn by commercializing OpenAI's technology.
OpenAI's corporate representative testified that he knew of no method for detecting paywalled content in the training data and no effort to remove it. The company's standard crawling process "did not include reviewing websites['] ... Terms of Use or Service."
Nadella testified under oath that "anything that is paywalled should be licensed by anyone who wants to use it." He said he would have forced OpenAI to retrain its models had he known the company had scraped paywalled content and used it for training.
OpenAI also acquired the "New York Times Annotated Corpus," a collection of 1.8 million articles, through a third party. Its license restricted use to "non-commercial linguistic education, research and technology development." OpenAI employees knew using the corpus to train models "would not be appropriate," but did so anyway.
Plaintiffs accuse OpenAI of suppressing evidence and challenge all four fair use factors
Immediately after the lawsuits were filed, OpenAI built a filter to suppress output most likely drawn from the plaintiffs' publications, according to the brief. Content from companies that hadn't sued remained unaffected. The plaintiffs argue that the filter was designed not to protect copyrights but to prevent them from gathering evidence.
Microsoft's Hecht had raised concerns that OpenAI might deploy such a filter, calling it an "accidental cover up." He warned it would result in "people who have a right over the content having less visibility into what was used for training." Even so, the brief includes numerous examples of problematic outputs, with ChatGPT reproducing exact copies and summaries of NYT articles on request, including articles behind a paywall.
The plaintiffs also argue that the fair use defense fails on all four statutory factors. They say the use is substitutive and commercial rather than transformative, while the articles are expressive works at the core of copyright protection. The defendants also copied entire works, even though their own experts admitted that no single work was necessary.
On market harm, the plaintiffs point to existing licensing markets and those that could realistically develop. OpenAI and Microsoft have signed licensing deals with other publishers, as have Amazon, Google, Meta, and Perplexity. But rather than pay for the plaintiffs' content, the defendants took it for free and traded it among themselves, the brief argues.
A functioning licensing market makes the fair use argument harder to sustain. The defendants can't credibly claim licensing wasn't possible when they're already paying for comparable content from other publishers.
AI can generate a million "pink slime" articles for $6,800
The plaintiffs also challenge the argument that AI models don't compete with publishers, further undercutting the fair use defense. They argue that OpenAI's products let users flood the market with "pink slime," meaning low-quality, often plagiarized pseudo-news. At OpenAI's API prices, generating one million news-style articles of 500 words each costs about $6,800, without involving a single journalist. The brief cites Prism News, which used just four employees to run 200 AI-generated publications posing as local newsrooms and hobby sites.
OpenAI announced a tool called "Media Manager" in 2024 that was supposed to let publishers opt out of scraping. The project has since been shelved, according to the brief.
The plaintiffs are asking the court to find copyright infringement and reject the fair use defense. They also want each article treated as a separate work when calculating any statutory damages.
Previous court rulings have leaned toward fair use, and the US Department of Justice, under the Trump administration, has also sided with the AI labs. OpenAI and Microsoft will likely present their side soon, focusing on fair use and the transformative nature of their work. The internal documents now in the public record won't make that argument any easier.
Brief