Anthropic 必须在一项版权和解中向图书作者支付 15 亿美元。旧金山一家联邦法院批准了这项和解,此前 Anthropic 在 2021 年至 2022 年间从盗版数据库 LibGen 和 PiLiMi 下载了书籍。在约 482,460 部列出的作品中,91.3% 被主张权利,每部作品净获约 3,000 美元,是法定最低赔偿额的四倍。Anthropic 必须销毁这些盗版文件。作者保留对复现原创作品的 AI 输出以及 Anthropic 未来行为的主张权利。这是集体诉讼史上规模最大的版权和解。
但这笔赔偿针对的是盗版行为,而非 AI 训练本身。Alsup 法官此前已裁定,用合法获得的书籍训练 AI 属于“转换性使用——而且是极其显著的转换性使用”,受合理使用原则保护。未经作者同意大规模抓取互联网内容是否算作合法获取,仍是一个悬而未决的问题,因此合理使用的争论很可能远未结束。尽管如此,这一裁决对于未经网站所有者同意便用网络内容训练模型的 AI 实验室来说,似乎是一个里程碑——网络内容正是它们训练数据的主要来源。
法律网站 Justia
Anthropic has to pay book authors $1.5 billion in a copyright settlement. A federal court in San Francisco approved the settlement after Anthropic downloaded books from the piracy databases LibGen and PiLiMi between 2021 and 2022. Of roughly 482,460 listed works, 91.3 percent were claimed, netting about $3,000 each, four times the statutory minimum. Anthropic must destroy the pirated files. Authors retain claims over AI outputs that reproduce original works and over Anthropic's future conduct. It's the largest copyright settlement in class action history.
But the payout covers piracy, not AI training itself. Judge Alsup had previously ruled that training AI on legally obtained books is "transformative - spectacularly so" and falls under fair use. Whether mass scraping of internet content without authors' consent counts as legal acquisition remains an open question, so the fair use debate is likely far from over. Still, the ruling looks like a milestone for AI labs that trained on web content without website owners' consent, their main source of training data.
Law Justia