简而言之:企业、应用和卖方等各层级的 AI 买家,正在用更便宜的开源模型替代前沿闭源模型。省下的钱并不会缩减 AI 账单——它们会被再投资于每项任务呈指数级增长的 token 消耗中。
三股力量正在重塑 AI 的成本结构:
AI 买家的自然反应就是替代。
Coinbase6:
在 Coinbase,我们正积极推进将提示词路由到更便宜的模型(在合适的情况下),并且在某些情况下能够将成本大致保持在持平水平,同时 token 使用量继续呈指数级增长。
Lindy7:
今天果断出手,把 Lindy 的全部流量 100% 切换到了 DeepSeek v4,从 Anthropic 的模型迁移过来。这为我们省下了数百万美元,而且我们在许多核心用例上实际上看到了性能提升。对业务而言是变革性的。
Harvey8:
在我们的法律智能体基准(LAB)的 100 项任务切片上,SFT 将 Kimi 2.6 的全通过率从 11% 提升到 15%,超过了 Opus 的 14%。但成本差距更为惊人:同样这 100 项任务,$84 对 $954,即便宜约 11 倍。
Cursor 走得更远。他们对 Kimi K2.5 进行了后训练,将其打造成自己的生产模型 Composer。9
Composer 2.5 异常智能,且比同等能力的模型效率最高可高出 10 倍。
Coinbase 的引述说明了节省下来的钱去了哪里:成本持平,token 呈指数级增长。买家并没有把折扣装进自己口袋——而是把它花在了更多智能上。
闭源模型在前沿越来越贵;开源模型在同等水平上越来越便宜。选择在于你想让单位经济模型下承受哪一条斜率。
In short : AI buyers across enterprise, app, and seller layers are substituting cheaper open-source models for frontier closed models. The savings don't shrink the AI bill — they get reinvested in exponentially more tokens per task.
Three forces are reshaping the AI cost structure :
- Foundation labs are moving up the stack into applications,1 2
- Frontier model prices keep rising for the smartest models,3
- Open-source models have crossed the good enough threshold for most use cases.4 5
The natural response from AI buyers is substitution.
Coinbase6 :
At Coinbase we’re working hot on routing prompts to cheaper models where appropriate, & in some cases have been able to keep costs roughly flat, while token usage continues to grow exponentially.
Lindy7 :
Pulled the trigger today & switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models. Saves us millions of $ & we’re actually seeing an increase in performance on many core use cases. Transformative for the business.
Harvey8 :
On a 100-task slice of our Legal Agent Benchmark (LAB), SFT moved Kimi 2.6’s all-pass rate from 11% to 15%, beating Opus’ 14%. But the cost gap was even more striking : $84 vs $954 across the same 100 tasks, or ~11x cheaper.
Cursor went further. They post-trained Kimi K2.5 into their own production model, Composer.9
Composer 2.5 is exceptionally intelligent & up to 10x more efficient than similarly capable models.
Coinbase’s quote shows where the savings go : costs flat, tokens exponential. Buyers don’t pocket the discount — they spend it on more intelligence.
Closed models are getting more expensive at the frontier; open models are getting cheaper at parity. The choice is which slope you want under your unit economics.