在整个行业里,各公司开始对 AI 的价格望而却步。Uber 在 4 月就花光了其 2026 年全年的 AI 编程预算。微软在为其开发者启用 Claude Code 许可数月后,又撤销了这些许可。一位 Priceline 员工告诉 TechCrunch,一次例行的 Cursor 合同续约报价涨了 4-5 倍。
尽管每 token 的价格已经下降,但推动更广泛采用 AI 以及日益自主的智能体,使 token 消耗量越来越高。那些在 2025 年初大肆享用无限量订阅的公司,如今正手忙脚乱地弄清楚钱都花到哪里去了,削减开支,并判断能否从预算的残骸中挽回一些 ROI。
与此同时,一个满足这些需求的市场正在形成。初创公司、老牌厂商以及一个新的标准机构都在竞相为企业提供工具和话语体系,以追踪它们的支出。
“六个月前,我和客户交谈时,话题全是‘它能做什么?它够好吗?’”OpenAI 企业业务负责人 Alexander Embiricos 本周在纽约市的一场活动上告诉 TechCrunch。“现在我们的话题再也不是这些了。现在的话题是,‘嘿,我们花了这么多钱。你们有什么可见性?你们有什么可审计性?你们有什么 token 控制手段?你们模型的效率如何?’”
正是在这样的背景下,Linux 基金会本周公布了 Tokenomics Foundation 的计划,这是一个新的标准机构,旨在为 AI token 注入与 FinOps 为云支出带来的同样的成本纪律。
“在四月和五月,我开始听到一些公司说:‘天哪,我们整个 2026 年的 token 预算已经超了 3 倍,而现在才四月,’”Linux 基金会旗下项目 FinOps Foundation 的执行董事 J.R. Storment 告诉 TechCrunch。“我们开始听到生存危机的说法,整个讨论从 tokenmaxxing 和‘快速推进’转向了‘我们需要护栏,我们该怎么控制这件事?’”
这场震动整个科技圈的呼声,源于此前 CEO 们热切要求团队使用最好的模型并快速推进,不惜一切代价。11 月发布的新模型,如 Anthropic 的 Claude Opus 4.5、OpenAI 的 GPT-5.1 和 Google 的 Gemini 3 Pro,为智能体工具带来了显著提升,从而使消耗成倍增长。正因如此,一家公司据报道在忘记为员工设置使用限额后,发现自己收到了 5 亿美元的 Claude 账单。
“这就像快克可卡因成瘾一样,”Priceline 的 IT 财务高级总监 Chris Reed 说道,并指出公司已经开始对某些群体设置 token 限额。“他们先让你试用,把你钩住,现在你就有点被它牵着走了。”
工程运营平台 Faros AI 的 CEO Vitaly Gordon 表示,他最近与一位 CTO 交谈,对方告诉他:“我的一名工程师上个月在 token 上花了 4 万美元,我真的不知道是该阻止他,还是该去告诉其他所有人都向他学习。”
Faros 在 4 月发布了一项针对 20,000 名开发者、历时两年的研究,发现产出在上升,但 bug 和重写也在增加。工程管理平台 Jellyfish 同样发现,使用 token 最多的工程师生产率大约是不太使用 AI 的工程师的两倍,但他们为此消耗了 10 倍的 token 数量。
Jellyfish 研究主管 Nicholas Arcolano 通过电子邮件告诉 TechCrunch,AI 支出正在急剧膨胀,很大程度上归因于智能体功能,每位开发者的消耗量在九个月内增长了约 18.6 倍。总而言之,这些数据让生产率的实际情况比支出所显示的更加模糊不清。
“极端支出是否能带来回报,归根结底取决于已交付代码的最终商业价值(例如收入),而大多数公司仍然无法衡量这一点,”Arcolano 说。
至少在一定程度上,这种计量问题源于当今 AI 被使用的规模之庞大。
“追踪云成本是一个每月数亿行的数据问题,”Storment 说。“追踪 token 成本是一个每月数万亿行的数据问题。你不能只是把它塞进随便什么电子表格,甚至基础工具里。要做到这一点,你必须从根本上重新思考你的工具、你的规格和你的会计系统。”
在 Priceline,Reed 已经看到了差异。他指出了供应商报告的使用量与 Priceline 内部数据之间的问题。
“我的职业生涯始于电信费用管理,而我如今看到的是从电信到云再到 AI 的完全相同的相似之处,”他说。“每当你引入新事物时,它就很容易出现计费错误,也蕴含着审计和优化的机会。”
围绕这一问题,一个市场正在开始形成。其中有像 Pay-i 这样的纯专注型公司,它追踪、衡量并优化 GenAI 投资的成本与性能。与此同时,Paid 让开发者能够追踪成本、衡量使用情况,并根据实际价值而非订阅费向用户计费。
此外还有像 Jellyfish、Waydev 和 Faros AI 这样的公司,它们都提供 AI 智能体监控,以证明开发者工具的 ROI。Storment 表示,FinOps Foundation 内 180 家供应商中的大多数都在向这一领域倾斜。
已拥有现成分销渠道的公司也在添加新功能,以抓住这个新市场。Ramp 最近进入了AI 支出管理领域;Datadog 和 New Relic 则附加了云成本管理、token 级可观测性和 GPU 监控等服务。在下周的 FinOps X 大会上,AWS 预计将推出面向企业 AI 支出的新财务管理功能。
NEA 合伙人 Tiffany Luck 认为,token 效率和可观测性很可能会被添加到“harness 或应用层”。她提到了 Factory,这是一家初创公司,为企业打造 AI 智能体,该公司本周推出了一个模型路由器,可为每项任务自动挑选合适的模型。
Gordon 预计,前沿实验室和其他模型提供商会采用 OpenRouter 式的优化方式,将查询引导至最便宜的模型——这一趋势已经在企业 Claude 账单中显现出来。
“关于你在 Anthropic 上花费了多少的财务报告,即使你调用的是 Opus 模型,部分支出也会落在 Sonnet 或 Haiku 上,因为它们足够聪明,能完成这些任务,”Gordon 说。“我认为这会变得越来越普遍。”
但所有这些工具的构建,都缺乏一套通用语言或共享定义,来说明一个 token 的成本是多少、它产出了什么,以及如何跨供应商比较支出。这正是 Tokenomics Foundation 希望发挥作用的地方。
该基金会正在为“tokenomics”构建一套规范定义和框架;为 AI token 使用和计费制定开放标准、规范和指标;以及为 AI 经济学提出新的指标,如每智能成本或每瓦 token 数。它还计划定义涵盖 token 工厂有效性和消费效率的指标。该组织计划于 7 月正式启动,并将在下周的 FinOps X 大会上宣布更多成员。
“Token 经济学从根本上比我们此前在这个规模上管理过的任何东西都更加抽象和不透明,”Salesforce 首席可用性官 Nishant Gupta 在一份声明中表示。“它需要一种不同于行业为云所构建的运营能力。”
话虽如此,高盛预测,到 2030 年全球 token 使用量将增长 24 倍。那些已经超出预算的公司现在就需要解决方案,而该基金会的首个交付成果还要等好几个月。
“也许我们发明了蒸汽机,但我们还没搞清楚流水线,”Gordon 说。
据 Arcolano 所说,明智之举是广泛而适度地采用。
“最佳 ROI 来自将广泛的中等用户从低使用量推向适度使用量,而不是把重度用户推得更高,”他说。
Russell Brandom 和 Tim Fernholz 对本文报道亦有贡献。
Across the industry, companies are starting to balk at the price of AI. Uber blew through its entire 2026 AI coding budget by April. Microsoft revoked its developers’ Claude Code licenses months after enabling them. A Priceline employee told TechCrunch that a routine Cursor contract renewal came back 4-5x more expensive.
Even though per-token prices have fallen, the push for more AI adoption and increasingly autonomous agents have driven token consumption higher and higher. Companies that gorged themselves in early 2025 on all-you-can-eat subscriptions are now scrambling to understand where their money is going, pull back spending, and figure out whether they can salvage some ROI from the wreckage of their budgets.
Meanwhile, a market is forming to meet them there. Startups, established vendors, and a new standards body are all racing to give companies the tools and language to track what they spend.
“Six months ago, I would have a conversation with a customer and it would be all about ‘What can it do? Is it good enough?’” Alexander Embiricos, OpenAI’s head of enterprise, told TechCrunch at an event in New York City this week. “Our conversations are never about that now. Now the conversations are about, ‘hey, we’re spending so much. What visibility do you have? What auditability do you have? What token controls do you have? What is the efficiency of your models?’”
It’s against this backdrop that the Linux Foundation this week unveiled plans for the Tokenomics Foundation, a new standards body that aims to instill the same cost discipline around AI tokens that FinOps did for cloud spend.
“In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it’s only April,’” J.R. Storment, executive director of the FinOps Foundation, a project under the Linux Foundation, told TechCrunch. “We started hearing existential crises, and the whole conversation shifted from tokenmaxxing and ‘go fast’ to ‘we need guardrails, how do we control this?’”
The cries heard round the tech world followed fervent demands from CEOs pushing their teams to use the best models and move fast, costs be damned. New models released in November like Anthropic’s Claude Opus 4.5, OpenAI’s GPT-5.1, and Google’s Gemini 3 Pro brought significant improvements to agentic tools, which have multiplied consumption. It’s how one company reportedly found itself with a $500 million Claude bill after forgetting to set usage limits for employees.
“It’s like the crack-cocaine epidemic,” said Chris Reed, senior director of IT finance at Priceline, noting the company had begun placing token limits on certain groups. “They let you try it to get you hooked on it, and now you’re kind of beholden to it.”
Vitaly Gordon, CEO of engineering operations platform Faros AI, said he recently spoke to a CTO who told him: “One of my engineers spent $40,000 on tokens last month, and I genuinely don’t know whether I should stop him or should I go and tell everyone else to be like him.“
A two-year study of 20,000 developers that Faros released in April found that output was rising, but so were bugs and rewrites. Jellyfish, an engineering management platform, similarly found engineers who used the most tokens were about twice as productive as those who used AI less, but they spent 10x the number of tokens to get there.
Nicholas Arcolano, head of research at Jellyfish, told TechCrunch via email that expenditure on AI is exploding in large part due to agentic features, with per-developer consumption rising about 18.6x in nine months. All in all, these stats make the productivity case murkier than the spending suggests.
“Whether extreme spend pays off comes down to the ultimate business value of shipped code (e.g. revenue), which most companies still can’t measure,” Arcolano said.
At least some of that measurement issue is the sheer scale at which AI is being used today.
“Tracking cloud costs is a hundreds-of-millions-of-rows-a-month data problem,” Storment said. “Tracking token costs is a trillions-of-rows-a-month data problem. You can’t just stick that into whatever spreadsheet or even basic tool. You’ve got to fundamentally rethink your tooling, your specs and your accounting systems to do that.”
At Priceline, Reed is already seeing discrepancies. He noted issues between a vendor’s reported usage and Priceline’s internal data.
“I started my career in telecom expense management, and I’m seeing all the same parallels, from telecom to cloud to AI,” he said. “Anytime you introduce something new, it’s ripe for billing errors and audit and optimization opportunities.”
A market is beginning to form around this problem. There are the pure-play companies, like Pay-i, which tracks, measures, and optimizes the costs and performance of GenAI investments. Paid, meanwhile, lets developers track costs, measure usage, and bill users based on actual value rather than subscription fees.
Then there are companies like Jellyfish, Waydev, and Faros AI, which all provide AI agent monitoring to prove the ROI of developer tools. Storment says most of the 180 vendors within the FinOps Foundation are leaning toward this space.
Companies with existing distribution are also adding new features to capitalize on this new market. Ramp has recently moved into AI spend management; Datadog and New Relic have tacked on services like cloud cost management, token-level observability, and GPU monitoring. At the FinOps X conference next week, AWS is expected to introduce new financial management features geared toward enterprise AI spending.
Tiffany Luck, a partner at NEA, thinks token efficiency and observability will likely be added in at the “harness or app layer.” She pointed to Factory, a startup that makes AI agents for enterprises, which this week launched a model router that automatically picks the right model for every task.
Gordon expects frontier labs and other model providers to adopt OpenRouter-style optimization to drive queries to the cheapest models — a trend already showing up on enterprise Claude bills.
“The financial report for how much you spend on Anthropic, even if you call the Opus model, some of the spend will be on Sonnet or Haiku, because they are smart enough to do it,” Gordon said. “I think this will become more and more of a thing.”
But all these tools are being built without a common language or shared definitions for how much a token costs, what it produces, and how to compare spend across vendors. That’s where the Tokenomics Foundation hopes to prove useful.
The Foundation is building a canonical definition and framework for “tokenomics;” open standards, specifications and metrics for AI token usage and billing; as well as new metrics for AI economics, like cost-per-intelligence or tokens-per-watt. It also plans to define metrics across token factory effectiveness and consumption efficiency. The group is planning a formal launch in July, and is about to announce more members at the FinOps X conference next week.
“Token economics is fundamentally more abstract and opaque than anything we’ve managed at this scale before,” Nishant Gupta, chief availability officer at Salesforce, said in a statement. “It requires a different operational muscle than the one the industry built for cloud.”
That said, Goldman Sachs projects global token usage to multiply by 24 times by 2030. The companies already over budget need solutions now, and the foundation’s first deliverable is still months away.
“Maybe we created a steam engine, but we still haven’t figured out the assembly line,” said Gordon.
According to Arcolano, the smart move is broad, moderate adoption.
“The best ROI comes from moving the broad middle from low to moderate usage, not pushing heavy users higher,” he said.
Russell Brandom and Tim Fernholz contributed to this reporting.