跳到正文
Hacker News:AI 热帖· cdrnsf·· 3 小时前AI 评分27

AI 公司都是寄生虫

AI Companies Are Parasites

AI 导读

一篇批评 AI 公司商业模式的文章称,其做法是抓取整个互联网、无偿攫取数据并用于训练模型,再按量出售模型访问权,直至"杀死宿主"。文章指出,流量换索引的旧约定已消失,抓取器压垮基础设施、点击率与社区双双崩塌;Reddit 虽在出售数据授权,但其数据依赖人类生成的社区内容。

正文

Merriam-Webster:

An organism living in, on, or with another organism in order to obtain nutrients, grow, or multiply often in a state that directly or indirectly harms the host.

That's it, right? The whole thing, their entire business model.

  1. Scrape the whole of the internet, destroy books, siphon everything they can from whoever they can without compensation, and call it fair use.
  2. Train models on said stolen data.
  3. Sell metered access to the models.
  4. Repeat until you've killed the host.
  5. Hope you can head off the effects of AI inbreeding.

The old bargain of traffic for indexing is gone. Scrapers overload infrastructure. Clickthroughs collapse. Communities collapse.1

Yes, licensing deals exist. Reddit is selling access to its data, but that data is human-generated and depends on a flagging community the company seems, at best, indifferent to.

They'll even show up in your community and set up a data center that nobody (well, except for the people profiting from it) wants. Maybe they'll spin up gas generators and send pollution your way. Or they'll drive up your electricity and water rates. But what about jobs? Most of those only exist during construction and, who knows, they might just hire folks from outside your area for that.

They never should've attached themselves to society, and it may be too late to dislodge them. At least we get stilted prose, images that look like Pixar rehashes, and code nobody understands.


  1. How's Stack Overflow doing lately? ↩︎

来源:Hacker News:AI 热帖 · coryd.dev