几周前,我和团队在旧金山参加一个会议。我们原本计划在一家相当普通的美国餐厅吃晚饭,但我另有想法,迅速临时变卦,订到了一家寿司店的位置。临时改变计划并不理想,但守了一整天会议展位之后吃一顿平庸的晚餐同样不理想。再说,谁不喜欢寿司呢?
过去 10 年里,我和妻子每逢庆祝都会去吃寿司 omakase(主厨定制)。我不用看菜单就了如指掌。大多数没有经验的寿司爱好者会直奔 O-toro(大腹),但其实还有很多更好的选择。于是我很自然地告诉大家,这一桌由我来点菜。我和朋友、家人以及同事吃饭时一直这么做。100% 的情况下,大家都愿意在这种特定场景下把点菜的自主权交给我,安心享用一顿饭。
对于如何共享一餐,我有一个非常坚定的观点:一定要家庭式共餐,每次都是。
选择权的价值
但为什么?一方面,这降低了单独点餐的风险——你点的夏威夷肋眼牛排可能难吃得很。另一方面,它也会放大"天哪,那口鱼子酱和牛是我今年吃过最棒的一口"这种记忆。这样的时刻会留在我们心里。我总是说,第一次来就把每样都尝一尝,之后随时可以再多点。但前提是我们按家庭分享式来吃,这招才管用。
只用一个 LLM,就等于每人都各点各的主菜。你可能是在为稳妥的选择做优化,而不是为最佳结果。我聊过数百家企业,它们开启 AI 之旅时都是先选定一家供应商,比如 OpenAI、Anthropic 或 Gemini。当你在单一供应商上标准化时,你是在基于今天所知道的、今天所需要的东西下注。
这是老生常谈的"没人会因为采购和实施 Salesforce 而被开除",只不过这种情况正在改变。等我接触到这些公司时,它们已经准备好"毕业"了——不再只用一个模型家族和一种模态。新的用例每天都在涌现。通常看起来是这样的:
- 从 OpenAI 企业版起步,把许可证部署给少数几个精选团队。
- 监控初始那批用户的使用情况,数据呈现出持续上升的趋势。
- 向更多团队开放使用权限。
- 图像生成、转录和创意写作等新用例开始出现。
- 发现 OpenAI 并没有最适合你需求的模型,你需要 Gemini。
- 去 Gemini 或其他任何提供商那里开通访问权限,而如今可观测性、治理和资源开通功能都被打断了。
我今天观察到的当前模式看起来像是成本压力,但实际情况比这更深。公司们[已经烧光了年度预算,而现在才六月。大家都有强烈的削减 token 用量的意愿,我理解。如果你不小心用上了 Opus 4.8,可能会把每日预算跑光,然后你就没有其他选择了。合上电脑,出去散散步吧。
虽然 Opus 4.7 的标价没有变化,但好几个人写到了“tokenizer 税”。这是 Anthropic 做的一次悄然调整,输入 token 数量增加了将近 35%。这是一个意义重大变化。更新的、能力更强的模型价格也在上涨。
Anthropic 发布的 Fable 定价为每 M 输入 token $10、每 M 输出 token $50。还有更高的:OpenAI 的 GPT-5.5 Pro 达到每 M 输入 token $30、每 M 输出 token $180。
使用时请谨慎!
成本压力是一种强制机制,可能带来更好的结果,也可能带来更糟的结果。从我的角度来看,我乐观地认为它正在带来一些更好的结果。我也很幸运能够促成这些结果。在我从事数据基础设施工作时,这是一个常见的话题。太多对话都围绕着计算成本以及团队在数据仓库上花了多少钱。但这过于关注[显性成本,而忽视了[隐性成本。最有战略眼光的领导者会彻底扭转这种对话。我经常听到:“我每年已经在计算成本上花了 100 万美元,所以我不在乎把它降低 30%。相反,更有价值的是,我那支每年花费 700 万美元的 40 人分析师团队能否提高生产力。在开发者工具方面,我既需要速度,[也需要效率。”
路由是一等公民
在深入探讨下一节之前,先简单说明一下 OpenRouter 到底是什么。OpenRouter 是访问 AI 的标杆性市场。我们让推理即插即用。我们消除了围绕选择提供商、选择模型以及理解各种指标的所有开销:延迟、价格、TPS、模型基准测试等等。
现在,你可以通过 OpenRouter 以干净、标准化的 API 规范在一个地方访问数百个大语言模型。这的存在是多么不可思议?这一切听起来很美好,简直像鱼与熊掌可以兼得,但现实中人们实际在做什么呢?
幸运的是,我能够围绕这一点拉取一些数据。而且时机也恰到好处——就在今天,团队发布了我们的[分析 API!
多模型采用一直是我们持有的一个假设,但我们可以清楚地看到与这一趋势相伴的增长曲线。这合理且在预期之内,但它掩盖了一个事实:大多数人可能只是在尝试每个模型的最新版本。例如,Anthropic 在图表时间线内就相继发布了 Opus 4.6、Opus 4.7 和 Opus 4.8。所以更有意思的问题应该是用户如何跨模型族进行采用。
现在,我们可以捕捉到用户主动将推理分散到多个模型族的真实增长。这描绘了一幅更接近真实的图景,展示了“持续进阶”到底是什么样子。让我们再叠加一个关于模型发布的数据点。
由于发布时间表并不总是规律,这是一张累计图表。但我们可以看到一个显著异常点:3 月到 4 月期间共有 90 个新模型发布。这规模巨大!可选的模型越来越多,而且速度还在不断加快。
这也可能让人有些压力。就像去一家餐厅,菜单上有 225 道菜(那是我最喜欢的餐厅之一)。就算是家庭式共享用餐,你也不可能全部尝一遍。我们显然考虑到了这一点,不希望我们的用户必须了解每个模型之间的区别。因此我们构建了 auto-router 和 pareto-router 等功能,让选择使用哪个模型变得更容易。
这一切都与我之前提到的成本压力息息相关。各公司其实正在以一种有趣的方式使用 OpenRouter。他们能够让每 token 的平均加权成本[随时间逐步降低。这是如何实现的?如果你根据所需的结果,把特定工作负载路由到特定的模型提供商和模型,那么你就可以利用像 Deepseek v4 flash 这样的模型,其输入成本仅约 $0.10/M tokens,输出约 $0.20/M tokens。
再进一步,使用像 Cerebras 这样拥有顶级吞吐量的模型提供商,你就成了像 Bradley Cooper 那样的指挥家,只不过繁重的工作由我们来替你完成。
或者,你可以把一部分流量路由到 Gemini 模型的[flex 优先级层级,以享受五折优惠。再次强调,选择权在你手中。
我们还有一些围绕模型智能的精彩内容即将分享。总的来说,主题始终不变:我们让推理即插即用。
Semper ad meliora(永远追求更好)
感觉推动 AI 采用的顺风正在转向 AI 优化。我们明白,花在 AI 推理上的资金不会以任何有意义的幅度减少,那么次优的选择是什么?那就是降低每个 token 的加权平均成本。各组织正在变得更有意识地推动使用在团队间的普及,同时降低治理方面的风险。而这一切只有在你不被单一供应商锁定时才真正可能实现。当你选择针对不同用例使用不同模型时,你在每一步都能获得议价筹码。你终于可以坐在餐桌旁,随心所欲地享用家庭式共享晚餐了。
A few weeks ago I was with my team in SF at a conference. We originally had dinner plans at a pretty standard American restaurant but I had other plans and quickly pulled an audible to get us a reservation for sushi. Last minute changes aren’t ideal but neither is a mediocre meal after working a conference booth all day. Plus who doesn’t love sushi?
My wife and I have done sushi omakase for every celebration for the past 10 years. I know my way around the menu without even looking at it. Most inexperienced sushi lovers go straight for the O-toro but there’s so much more out there. So naturally I told everyone I would be ordering for the table. I’ve consistently done this with friends, family, and work colleagues. 100% of the time people are willing to outsource their agency in this specific situation to enjoy a meal.
I have a strongly held opinion about how we break bread: family style, every time.
The case for optionality
But why? On one hand it reduces the risk of ordering solo and your hawaiian ribeye tasting like garbage. On the other hand it magnifies the memory of “omg that caviar wagyu bite was the single best bite I’ve had all year”. That moment stays with us. I always say for a first visit, taste everything. You can always order more later. But that only works if we do family style.
Standardizing on one LLM is the same as everyone ordering their own entree. You might be optimizing for the safe pick instead of the best outcome. I’ve talked to hundreds of companies who started their AI journey by picking one provider say OpenAI, Anthropic, or Gemini. When standardizing on one provider you’re making a bet based on what you know today and what you need today. It’s the age old story of “no one gets fired for buying and implementing salesforce” except that’s kind of changing. By the time I talk to these companies they’re ready to “graduate” to utilizing more than just one model-family and one modality. New use cases are showing up every day. It often looks like this:
- Start off with OpenAI enterprise and deploy licenses to a select few teams.
- Monitor usage across the initial cohort which shows up-and-to-the-right trends.
- Give access to additional teams.
- New use cases like image generation, transcription, and creative writing emerge.
- Find out that OpenAI doesn’t have the best models for your use and you need Gemini.
- Go to Gemini or any other provider to set up access and now observability, governance, and provisioning are broken.
The current pattern I’m seeing today looks like cost pressure but it’s deeper than that. Companies have blown through their annual budgets and it’s only June. There’s a strong desire to reduce token usage and I get it. If you accidentally use Opus 4.8 you might run through your daily budget and then you have no other options left. Close your laptop and go for a walk.
While list price didn’t change on Opus 4.7 several people wrote about the “tokenizer tax”. This was a silent change made by Anthropic which had almost a 35% increase in input tokens. That’s a meaningful change. Newer more proficient models are increasing in price as well. Anthropic released Fable which is priced at $10/M input tokens and $50/M output tokens. And it goes higher: OpenAI’s GPT-5.5 Pro lands at $30/M input tokens and $180/M output tokens. Use with caution!
Cost pressure is a forcing mechanism which can lead to either better or worse outcomes. From my reference point, I’m optimistic it’s leading to some better outcomes. I’m also lucky enough to enable these outcomes. This was a common theme during my days working in data infra. So many conversations revolved around compute costs and how much teams were spending on their data warehouse. But this focuses too much on the explicit costs versus the implicit costs. The most strategic leaders flip this conversation on its head. I would often hear “I’m already spending $1M a year on compute costs so I don’t care about reducing that by 30%. Instead, what’s more valuable is if my team of 40 analysts which costs me $7M a year is more productive. I need speed and efficiency when it comes to developer tools”.
Routing is a first class citizen
Before we dive into this next section, a quick note on what OpenRouter actually is. OpenRouter is the canonical marketplace for accessing AI. We make inference just work. We remove all the overhead around picking a provider, or picking a model, and understanding things: latency, price, TPS, model benchmarks, etc.
So now you can access hundreds of LLMs in one place through OpenRouter in a clean standardized API spec. How incredible that this exists? And all this sounds great almost like you can have your cake AND eat it but what do people actually do in reality?
Luckily I was able to pull some data around this. It’s also divine timing that today the team released our analytics API!
Multi-model adoption was a hypothesis we have always had but we can clearly see growth trends alongside this story. This makes sense and is expected but it obscures the fact that most people could just be trying the newest version of each model. For example, Anthropic has released Opus 4.6, Opus 4.7, and Opus 4.8 all within the graph’s timeline. So what would be more interesting is how users adopt across model families.
Now here we can capture the real growth around users actively spreading inference across model families. This paints a more realistic picture of what continuously graduating looks like. Let’s layer in one more data point around model releases as well.
This is a cumulative chart since release schedules aren’t always consistent. But we can see one big outlier from March to April where we had 90 new model releases. That’s huge! So many more options to pick from at an increasing velocity.
That can also be a little stressful. It’s like going to a restaurant and the menu has 225 items (one of my favorite restaurants). Even with family style you can’t try them all. We obviously thought about this and don’t want our users to have to know the difference between every single model. So we built out things like auto-router and pareto-router to make it easier to pick which model to use.
All this ties back to the cost pressure I mentioned earlier. Companies are actually utilizing OpenRouter in an interesting way. They are able to bring their average weighted cost per token down over time. How is this possible? Well if you route specific workloads to specific model providers and models based on your required outcomes then you can take advantage of say Deepseek v4 flash which only costs around $0.10/M tokens on input and $0.20/M tokens on output.
Take this one step further by utilizing a model provider like Cerebras which has some of the best throughput and now you’re the maestro like Bradley Cooper except we’re doing the heavy lifting for you.
Or you route some of your traffic to flex priority tier on Gemini models to take advantage of 50% off. Again the choice is yours.
We have some exciting stuff around model intelligence that we will be sharing soon. Overall, the theme is still the same: we make inference just work.
Semper ad meliora
It feels like the tailwinds that were driving AI adoption are moving towards AI optimization. We understand that the amount of money we are spending on AI inference isn’t going to decrease by any meaningful magnitude, so what’s the next best alternative? It’s bringing the average weighted cost per token down. Organizations are becoming more intentional with democratizing usage across teams while reducing risk when it comes to governance. This really only happens if you’re not vendor locked into a single provider. When you choose to use different models for different use cases you gain leverage at each turn. You finally get to sit at the dinner table and eat as you please, family style.