最先进的模型比去年 11 月时聪明了三分之二。这种疯狂的改进速度仍在持续,每三天就有两个新模型问世。1
但在 OpenRouter 上,84% 的 token 并非来自最先进的模型。2 3
事实上,用户选择用来生成其中绝大多数 token 的那六款模型,性能大约达到前沿模型的 77%。而它们的成本仅为 Claude Fable 5 的 2.5%。2
该指数持续攀升。大约每个季度都会出现三到五个 Artificial Analysis 分值的大幅提升。较小的进步则填补了其中的空隙。
在 8 月 10 日那一周,六款模型承载了 80% 的用量。它们的混合价格为每百万 token 0.50 美元,而 Fable 5 为 20 美元。
Ramp 的数据显示,买家对价格敏感。Fable 5 约为每百万 token 10 美元,在发布一个月后占据了 Anthropic token 的 6% 以及 Anthropic 支出的 11%。GPT-5.6 Sol 是 OpenAI 最贵的主线档位,占据了 OpenAI token 的大约四分之一。4
7 月,Fable 5 产生的模型归属收入大约为 GPT-5.6 Sol 的 75%,尽管其价格要贵得多。
每一次新的 SOTA 发布,所抢占的市场份额都应少于前一次。
企业将整合支出。合同会集中在少数一两家供应商手中,就像云计算时代那样;而一旦某个模型在高价值任务上站稳脚跟,该工作负载就会一直留在那里。
在显著折扣之下,性能已经足够好了。差距正从下方不断缩小。到 5 月,最好的开放权重模型已达到前沿得分的 80%,而一年前这一数字为 48%。2
应用部署则是另一番景象。我们投资组合中的更多公司以及初创企业,默认选择更小的模型、微调模型以及开源方案。它们针对的是一条不同的帕累托前沿——价格优先于性能。
如果份额不再转移,而"足够好"始终足够好,那么 SOTA 的经济学就变了。一次九位数投入的训练运行必须赢得份额才能回本,而这一门槛会随时间推移而抬高。
这个标题有些轻率。很多人购买 SOTA,而且理由充分。软件工程架构与安全设计是最明显的例子,在这些场景中,能获得的最优模型值这个价。
但我们确实拥有的公开数据表明,真正重要的前沿是另一条。
-
Artificial Analysis 模型目录与 Intelligence Index。主要实验室的月度发布数量与前沿路径。样本起始于 2025-11-01。发布频率趋势持平。大型(≥3 分)前沿跃迁之间的中位间隔约为 3.5 个月。Intelligence Index ↩︎
-
State of the art 指某一周内可获得的单个最佳 Artificial Analysis 分数;当某模型的分数落在该周最佳具名模型分数的 10% 以内时,即视为接近前沿。将 OpenRouter 每周具名顶级模型与 Artificial Analysis 分数相连接,取的是 OpenRouter 轮播榜的头部,而非每一个 API。份额序列,2025-11-03 至 2026-05-25 各周(n=30),仅限具名模型,Others 已排除。前十三周与后十三周接近前沿的比例约为 17.5% 对 14.6%(约 82-85% 在前沿之外)。集中度快照,2026-08-10 当周,覆盖具名 token 前约 80% 的模型。按 token 加权的 Artificial Analysis 比全球目录的 state of the art 落后约 23%(约为前沿质量的 77%),比该 OpenRouter 榜单上的最佳模型落后约 10%。混合篮子约为 $0.50/m tokens,而 Fable 5 为 $20/m(约 40 倍)。2026 年 5 月的历史核查,比本地榜单领先者落后约 16%。最佳开放权重模型在前十三周达到前沿分数的 47.5%,在后十三周为 70.9%;单周最佳为 2026-05-25,DeepSeek V4 Pro 为 45.27,前沿为 56.31(80.4%)。OpenRouter rankings ↩︎ ↩︎ ↩︎
-
这些数据来源并未涵盖第一方云,即 OpenAI、Anthropic 与 Google 自家的服务。运行在原生 API 上的前沿流量从不会进入 OpenRouter 排名,因此数据存在偏差。↩︎
-
Ramp Economics Lab,AI Index 2026 年 8 月(Fable 5 采用情况)。econlab.substack.com/p/ai-index-august-2026 ↩︎
State of the art models are two-thirds smarter than they were last November. The frenetic pace of improvement is sustained, two new models every three days. 1
But 84% of tokens on OpenRouter aren’t state of the art. 2 3
In fact, the six models users choose to generate the supermajority of those tokens deliver about 77% of the performance of the frontier. They cost 2.5% of what Claude Fable 5 does. 2
The index keeps jumping. Large gains of three to five Artificial Analysis points land about every quarter. Smaller steps fill the gaps.
Six models carry 80% of volume in the week of August 10. Their blended price is $0.50 per million tokens against Fable 5 at $20.
Ramp’s data shows buyers are price-elastic. Fable 5 at about $10/m tokens captured 6% of Anthropic tokens & 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. 4
Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive.
Each new state of the art release should move less share than the one before it.
Enterprises will consolidate spend. Contracts concentrate on one or two vendors, just like in the cloud era, & once a model clears a high-value job the workload stays.
Performance is already good enough at a meaningful discount. The gap keeps closing from below. The best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier. 2
Application deployment is the other story. More of our portfolio companies & startups default to smaller models, fine-tuned models, & open source. They are optimizing against a different Pareto frontier, price over performance.
If share stops shifting & good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, & that bar will rise with time.
The title is flippant. Plenty buy SOTA, & for good reason. Software engineering architecture & security design are the clearest cases, where the best available model earns its price.
But the open data we do have suggests the frontier that matters is the other one.
-
Artificial Analysis model catalog & Intelligence Index. Major-lab monthly release counts & frontier path. Sample starts 2025-11-01. Release-rate trend flat. Median gap between large (≥3 pt) frontier steps about 3.5 months. Intelligence Index ↩︎
-
State of the art means the single best Artificial Analysis score available in a given week; a model counts as near it when the score sits within 10% of that week’s best named model. OpenRouter weekly named top models joined to Artificial Analysis scores, the head of the OpenRouter carousel rather than every API. Share series, weeks 2025-11-03 through 2026-05-25 (n=30), named only, Others excluded. First vs last thirteen weeks about 17.5% vs 14.6% near the frontier (~82-85% outside). Concentration snapshot, week of 2026-08-10, models covering the first ~80% of named tokens. Token-weighted Artificial Analysis about 23% behind global catalog state of the art (~77% of frontier quality) & about 10% behind the best model on that OpenRouter list. Blended basket about $0.50/m tokens vs Fable 5 at $20/m (~40x). May 2026 historical check, about 16% behind the local list leader. Best open-weight model 47.5% of frontier score in the first thirteen weeks vs 70.9% in the last thirteen; single best week 2026-05-25, DeepSeek V4 Pro at 45.27 vs frontier 56.31 (80.4%). OpenRouter rankings ↩︎ ↩︎ ↩︎
-
These data sources don’t capture the first-party clouds, OpenAI, Anthropic & Google’s own services. Frontier traffic running on native APIs never enters the OpenRouter rankings, so there’s a bias to the data. ↩︎
-
Ramp Economics Lab, AI Index August 2026 (Fable 5 uptake). econlab.substack.com/p/ai-index-august-2026 ↩︎