2026 年 2 月 6 日,OpenRouter 上的智能体消耗的 token 超过了人类。此后它们再也没有让出领先地位1。
AI 消耗不是任何人都能外推的平滑曲线:它以三波浪潮的形式到来,每一波都比上一波大上几个数量级。第二波已经爆发。第三波已在地平线上。
第一波是聊天:人提问,模型回答一次,结束。据我们估算,每个活跃用户每天大约消耗 100 万 token。
但聊天在企业产出中已属少数:到 2026 年 6 月,聊天占企业使用的略多于三分之一,而 Codex 占了另外的 64%2。
第二波是单个智能体凭感觉编程,做一个家庭周末运动日程安排应用来协调拼车,或者研究十年期国债的趋势及宏观经济影响。每天大约 1 亿到 2 亿 token,同样是估算。各波之间的 token 强度差距是两个数量级的幂3。
第三波在其上叠加一个元框架:一个 AI 调度许多智能体并行工作,每个智能体又衍生出自己的工具调用和子智能体。又是一个数量级,数字达到数十亿。
消耗增长不是因为生成变快了。它增长是因为并行化产生了复利效应。
这里是构建元框架(meta-harness)的另一种方式。汇集一个电子表格,每个单元格里放的不是数字,而是 AI 的回答。想象一份 200 行长、30 列的 YCombinator 创业公司清单:创始人背景摘要、公司历史、价值主张。每一列都代表着数十次 AI 工具调用。
在任何人注意到之前,这次运行就已经消耗了数千万 token。
这就是这一转变的形态:OpenRouter 上的智能体 token 量在六个月内从 0.51t 升至 7.3t,增长了 14 倍,而人类使用量仅增长了 2.8 倍1。
这些浪潮所隐含的需求已经在宏观层面显现。高盛预测,到 2030 年,消费者与企业智能体每月将消耗 120 千万亿 token,是 2026 年水平的 24 倍4。
这些浪潮并非彼此取代,而是层层叠加。Chat 持续增长,智能体增长更快,而元框架将让两者都相形见绌。任何依据平滑外推来规划算力的人,都在为错误的曲线做规划。
-
由 Peter Walker 整理的 OpenRouter 数据,在 a16z 的 Charts of the Week 中流传,2026 年 8 月:智能体的 token 消耗量从 0.51t 升至 7.3t,增长了 14 倍,而人类使用量仅增长 2.8 倍;2026 年 2 月 6 日很可能是人类 token 消耗量超过智能体的最后一天。在相同模型上,智能体每个任务消耗的 token 比人类多近 5 倍。有两点需要注意:智能体 token 消耗的绝大部分是缓存提示词,按较低费率计费;且 OpenRouter 的构成偏向开放权重模型。↩︎ ↩︎
-
OpenAI:企业信号,报告发布于 2026 年 8 月 13 日:2026 年 6 月,在企业客户中,Codex 占 Codex 与 ChatGPT 合计输出 token 的 64%,ChatGPT 占 36%。该报告追踪 2026 年 1 月至 6 月;截至 6 月,前沿公司使用的 token 是普通公司的 8.3 倍,而 1 月时为 2.6 倍。↩︎
-
Chat 每个活跃用户每天约消耗 1m token,而单个智能体为 100m 至 200m,大约相差两个数量级。若按每个任务而非每天来衡量,差距还要更大:Yu 等人,How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks发现,智能体任务消耗的 token 大约是代码推理与代码对话的 1,000 倍。输入 token 主导了总量,因为智能体在每一步都会重新读取自己的上下文。↩︎
-
高盛研究:随着使用量激增,AI 智能体预计将提振科技现金流,2026 年 5 月:到 2030 年,消费级与企业级智能体的 token 消耗量将增长 24 倍,达到每月 120 千万亿 token。↩︎
On February 6, 2026, agents on OpenRouter consumed more tokens than humans did. They have not given the lead back1.
AI consumption is not a smooth curve anyone can extrapolate : it arrives in three waves, each orders of magnitude larger than the last. The second one has already broken. The third is on the horizon.
Wave one is chat : a person asks, the model answers once, done. Roughly 1m tokens per active user per day, by our estimate.
But chat is already the minority of enterprise output : by June 2026, chat accounted for a little more than a third of enterprise use, & Codex for the other 64%2.
Wave two is a single agent vibe coding a family weekend sports scheduling app to coordinate carpools, or researching trends in the ten-year bond & the macroeconomic implications. Roughly 100m to 200m tokens a day, again an estimate. The token intensity gap between waves is two zeros of power3.
Wave three stacks a meta-harness on top : one AI dispatching many agents in parallel, each spawning its own tool calls & sub-agents. Another order of magnitude, & the numbers hit the billions.
Consumption does not grow because generation gets quicker. It grows because parallelization compounds.
Here’s another way of building a meta-harness. Amass a spreadsheet not of numbers in each cell, but of AI answers. Imagine a list of YCombinator startups 200 long with 30 columns each : summary of founder backgrounds, history of the company, value proposition. Each column represents tens of AI tool calls.
The run costs tens of millions of tokens before anyone notices.
That is the shape of the shift : agent volume on OpenRouter rose from 0.51t to 7.3t tokens in six months, a 14x increase, while human volume managed 2.8x1.
The demand implied by these waves is already showing up at the macro level. Goldman Sachs projects consumer & enterprise agents will consume 120 quadrillion tokens a month by 2030, 24 times the 2026 level4.
The waves do not replace each other, they stack. Chat keeps growing, agents grow faster, meta-harnesses will dwarf both. Anyone sizing compute off a smooth extrapolation is planning for the wrong curve.
-
OpenRouter data compiled by Peter Walker, circulated in a16z’s Charts of the Week, August 2026 : agent token consumption rose from 0.51t to 7.3t tokens, a 14x increase, against 2.8x for human usage; 6 February 2026 was likely the last day humans consumed more tokens than agents. Agents consume nearly 5x more tokens per task than humans on the same models. Two caveats : the large majority of agent token burn is cached prompts billed at lower rates, & OpenRouter’s mix skews toward open-weight models. ↩︎ ↩︎
-
OpenAI : Enterprise Signals, report released 13 August 2026 : Codex accounted for 64% of combined Codex & ChatGPT output tokens among enterprise customers in June 2026, leaving ChatGPT at 36%. The report tracks January to June 2026; frontier firms used 8.3x more tokens than typical companies by June, up from 2.6x in January. ↩︎
-
Chat runs about 1m tokens per active user per day against 100m to 200m for a single agent, roughly two orders of magnitude. Measured per task rather than per day the gap is wider still : Yu et al., How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks finds agentic tasks consume roughly 1,000x more tokens than code reasoning & code chat. Input tokens drive the total, because an agent re-reads its own context on every step. ↩︎
-
Goldman Sachs Research : AI Agents Forecast to Boost Tech Cash Flow as Usage Soars, May 2026 : token consumption rises 24x to 120 quadrillion tokens a month by 2030 across consumer & enterprise agents. ↩︎