本地 AI 模型在回答日常聊天与推理查询时,已有 89% 的表现能与前沿云端模型媲美,这一结果在超过一百万条真实查询以及 20 多个受测本地模型上普遍成立。1
这意味着我们每消耗一瓦电力就能产出更多智能:同样的电子数量能完成更多工作。1
追踪计算效率随时间的变化并非新鲜事。Koomey 定律发现,几十年来,每瓦计算能力大约每 1.5 年翻一番,这一趋势将一台大型机的算力缩小到了笔记本电脑的机身之中。23
正如每瓦性能指引了从大型机到 PC 的转变,每瓦智能将指引 AI 向边缘端的转变。
—— Jon Saad-Falcon、Avanika Narayan 等,《Intelligence per Watt》
如今 GPU 遵循的曲线则更为平缓,在过去 15 年里效率大约每 2.7 年翻一番,而非每 1.5 年。4
其影响同样令人瞩目:最佳本地模型对前沿模型的胜/平率从 2023 年的 23.2% 上升到 2025 年的 71.3%,大约每年增加 20 个百分点。到 2026 年,由另一个本地 AI 为该任务挑选的本地模型将这一数字推高至接近 90%。5
效率也随这一趋势线同步提升:同一时期内,每瓦智能提升了 5.3 倍,其中 3.1 倍来自更好的模型,1.7 倍来自更好的芯片。计算机更快了,模型更聪明了。最终用户从中受益。
5
云对于长多步推理、最难的技术领域,以及规模和并行化至关重要的负载,仍然不可或缺。在许多用例中,云硬件仍比本地硬件更具优势。相对于本地模型,云端推理带来 40% 的能效提升。6数据中心会批量处理查询,这是目前一次只服务一个用户的本地硬件还无法使用的技巧。
但对于大部分日常知识工作而言,完全没有必要把查询发送到数据中心。本地模型加上一个路由器,就足以应对绝大多数工作。7
与全云端基线相比,它还将能耗降低 80%、算力降低 77%、成本降低 74%。
大型机变成了个人电脑。你的数据中心也将如此。
-
Jon Saad-Falcon、Avanika Narayan 等,《每瓦智能:衡量本地 AI 的智能效率》,斯坦福大学与 Together AI,2025 年 11 月。arXiv:2511.07885。另见斯坦福 Hazy Research 概述。↩︎ ↩︎
-
瓦特衡量功率,即能量被消耗的速率,而焦耳衡量能量本身;Koomey 最初的指标是每焦耳的计算量,但其底层趋势与如今每瓦智能为 AI 模型所追踪的趋势相同。↩︎
-
Anson Ho、Ege Erdil 与 Tamay Besiroglu,《CMOS 微处理器能效的极限》,2023 年。arXiv:2312.08595。↩︎
-
71.3% 和 89% 是来自同一份 2025 年数据的两个不同测量值,而不是同一数字在不同时间点的取值。71.3% 是最佳单一本地模型对前沿模型的胜/平率,也就是这张图所绘制的趋势线(2023 年为 23.2%,2024 年为 48.7%,2025 年为 71.3%)。89% 则是一个独立的、更高的上限:将每个查询路由到所测试的 20 多个本地模型中处理该查询最佳的那一个,比任何单一模型高出 16.3 到 28.8 个百分点。这一增益是一种选择效应:论文指出,本地路由从 20 多个多样化的模型中挑选,而前沿云端模型只有三个,因此在某些基准测试上,本地最佳集成甚至超过了云端最佳。候选模型更多——而不仅仅是路由更智能——才是提升准确率的原因。这两个数字都来自同一项研究,彼此并不相互取代。↩︎ ↩︎
-
Jon Saad-Falcon、Avanika Narayan 等人,《Intelligence per Watt: Measuring Intelligence Efficiency of Local AI》,斯坦福大学与 Together AI,2025 年 11 月。云端加速器在运行相同模型时,每瓦智能至少比本地芯片高出 1.4 倍,大约相当于 40% 的效率溢价。arXiv:2511.07885。 ↩︎
-
Most AI Work Can Wait,tomtunguz.com。 ↩︎
Local AI models can already answer 89% of everyday chat & reasoning queries as well as a frontier cloud model, a result that holds broadly across more than a million real queries & 20+ local models tested.1
That means we’re generating more intelligence per watt of electricity : more work from the same number of electrons.1
Tracking computing efficiency over time is not new. Koomey’s law found that computing power per watt doubled roughly every 1.5 years for decades, a trend that shrank the power of a mainframe into a laptop’s chassis.23
Just as performance-per-watt guided the mainframe-to-PC transition, intelligence-per-watt will guide AI’s transition to the edge.
— Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt”
GPUs follow a more languid curve today, doubling efficiency roughly every 2.7 years over the last 15 years, not every 1.5 years.4
The impact is no less impressive : the best local model’s win/tie rate against a frontier model, rose from 23.2% in 2023 to 71.3% in 2025, adding roughly 20 percentage points a year. In 2026, a local model selected for the task by another local AI bumps that number to nearly 90%.5
Efficiency improved alongside the trend line : intelligence-per-watt rose 5.3x over the same period, split into a 3.1x gain from better models & a 1.7x gain from better chips. Computers are faster, models are smarter. And the end user benefits.
5
Cloud remains essential for long multi-step reasoning, the hardest technical domains, & workloads where scale & parallelization matter. Cloud hardware still holds an edge over local hardware in many use cases. Cloud inference delivers a 40% energy efficiency gain relative to local models.6 Datacenters batch queries, a trick local hardware serving one user at a time cannot use, yet.
But for much of everyday knowledge work, there is no reason to send the query to the data center at all. The combination of local models plus a router is sufficient for the supermajority of work.7
It also cuts energy 80%, compute 77%, & cost 74% against an all-cloud baseline.
Mainframes became personal. So will your data center.
-
Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025. arXiv:2511.07885. See also the Stanford Hazy Research overview. ↩︎ ↩︎
-
Koomey’s law, Wikipedia. ↩︎
-
Watts measure power, the rate energy is drawn, & joules measure the energy itself; Koomey’s original metric was computations per joule, but the underlying trend is the same one intelligence per watt now tracks for AI models. ↩︎
-
Anson Ho, Ege Erdil & Tamay Besiroglu, “Limits to the Energy Efficiency of CMOS Microprocessors,” 2023. arXiv:2312.08595. ↩︎
-
71.3% & 89% are two different measurements from the same 2025 data, not the same number at different times. 71.3% is the best single local model’s win/tie rate against a frontier model, the trend line this chart plots (23.2% in 2023, 48.7% in 2024, 71.3% in 2025). 89% is a separate, higher ceiling: routing each query to whichever of the 20+ local models tested handles it best beats any single model by 16.3 to 28.8 percentage points. That gain is a selection effect: the paper notes local routing draws from 20+ diverse models versus three frontier cloud models, so on some benchmarks the best-of-local ensemble even surpasses best-of-cloud. More candidates to choose from, not just smarter routing, is what raises accuracy. Both figures come from the same study & neither supersedes the other. ↩︎ ↩︎
-
Jon Saad-Falcon, Avanika Narayan, et al., “Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,” Stanford University & Together AI, November 2025. Cloud accelerators deliver at least 1.4x higher intelligence-per-watt than local chips running the same models, roughly a 40% efficiency premium. arXiv:2511.07885. ↩︎
-
Most AI Work Can Wait, tomtunguz.com. ↩︎