今天我们正式推出 Cursor Router,这是我们面向团队和企业打造的智能模型路由器。
Cursor Router 让团队能够自动将每一个请求路由到最适合该任务的最强模型,以更低的成本获得前沿智能。
我们在数千名企业开发者的生产流量中观察到了极其出色的结果。在面向数十家企业的早期访问阶段,客户以大约低 30–50% 的成本获得了前沿性能。


在覆盖数百万次请求的在线 A/B 测试中,Cursor Router 以节省 60% 的成本实现了前沿级别的性能。
Cursor 每周跨所有模型和提供商路由数亿次编码请求,对用户喜欢什么、什么会留在代码库中拥有独特的洞察。模型中立一直是 Cursor 运作方式的核心,而今天,我们正将这一数据和专业能力用于为你的团队服务。
借助 Cursor Router,我们的目标是为团队在每一项任务上提供最佳性能和体验,同时不让花费超出任务本身所需。
工作原理
大约 60% 使用 Cursor 的开发者会选择单一模型作为日常主力。这导致日常工作以最前沿模型的价格完成,AI 支出增长速度远超产出质量的提升。Cursor Router 通过在模型运行前对每个请求进行分类来解决这一问题。
从本质上讲,Cursor Router 是一个分类器,它根据用户的查询将其路由到最佳的模型选项。我们基于 60 万多条线上请求训练了 Cursor Router,并在由 Cursor Router 路由的数百万条线上请求上通过在线 A/B 测试评估其表现,以用户满意度(AFC)作为奖励进行优化。
Cursor Router 会从查询、上下文、任务复杂度和领域几个方面分析每一条请求,并结合我们对每个模型行为的了解。我们学习每个模型最擅长什么,然后路由到最有效的选项。简单的工作交给最具价格效率的模型,UI 更新交给最有品味的模型,而更复杂、长跨度的问题则交给前沿推理模型。
我们设计这套路由分类器时,所面向的是一个更新模型会早早且频繁发布的世界。这样一来,随着更新、更强大的模型发布,我们可以轻松更新 Cursor Router,让体验持续提升。
Cursor Router 在训练和评估两方面都具备缓存感知能力。它的训练数据集是路由会导致缓存未命中的场景,而评估则在生产环境中进行,我们报告的节省成本中包含了路由决策所导致的缓存未命中成本。
以更低成本获得前沿智能
Cursor Router 有三种模式:Intelligence、Balance 和 Cost,让你可以调整自己在成本-智能帕累托前沿上的位置。
我们发现,Auto Intelligence 模式在输出用户满意度上接近 Fable,而团队成本降低约 60%,同时在几乎相同的成本下,将满意度较 Opus 4.8 提升约 15%。
类似地,Auto Balance 在结果用户满意度上高于 Opus 4.8,而成本降低约 36%。相较 GPT-5.6 Sol,Auto Balance 以更低的支出率提供相当的满意度。
我们选择使用大规模在线 A/B 测试而非离线评测来衡量路由器的效果。虽然离线评测是有用的质量代理指标,但它们受限于规模较小、与真实使用场景存在距离,以及难以将成功归结为评分标准。
离线评测还忽略了因切换模型而产生的额外缓存未命中成本。真正的路由发生在整个对话过程中:选择哪个模型,以及何时切换。
在线 A/B 测试将 Cursor Router 置于真实世界中,在数百万个任务和对话中进行检验。工程师编写代码、提出追问、遇到错误,然后继续推进,通常一周内会有数百次请求。这些正是模型路由器需要表现良好的条件。
在输出质量方面,我们衡量了:
- 用户满意度,根据用户响应将智能体的成功进行分类。继续下一个功能是强烈的正面信号,而纠正智能体则是强烈的负面信号。
- 保留率,即智能体生成的代码随时间推移有多少仍留在代码库中。
在过去九个月里,我们一直依赖这些指标来评估每一次模型发布和 harness 改进。
客户的实际体验
在过去两周里,Cursor Router 已面向部分企业客户开放早期访问。我们将他们实际支付的费用与相同流量完全按 Opus 4.8 API 费率计价的情况进行了对比。
在早期访问中,三个拥有数千名用户的高用量账户,在 Auto 路由请求上相比将所有请求都路由到 Opus 4.8 节省了 30%–50%,且质量没有下降。


每次请求的成本只是故事的一半。工程负责人关心的是这些节省能否体现在真正交付的工作中,因此我们考察了每次提交的成本,这一规律依然成立。
对于单次提交,我们观察到Cursor Router 的每次提交成本更低,Intelligence 模式为 $6.76,Balance 模式为 $4.63。
GPT-5.6 Sol 的成本与 Intelligence 相当,但用户对输出的满意度更低。与此同时,Fable 5 和 Opus 4.8 产生提交的成本高于 Cursor Router,分别为 $12.69 和 $7.34。


这一差距正是路由的实际意义所在。Cursor Router 将困难任务保留在能力最强的模型上,同时把日常性工作从前沿定价的模型上转移出去。
由你来选择权衡取舍
Cursor 在设计路由器时充分考虑了团队和大型组织。该路由器采用数据驱动的分类体系,同时管理员和终端用户仍可选择它位于成本-智能帕累托前沿上的哪个位置。
在模型选择器中选择 Auto 模式,并从三种优化模式中选择,它们会让你沿着前沿移动:
- 智能:前沿质量,性能可媲美那些最昂贵、最强大、日常使用可能难以企及的模型。
- 平衡:出色质量,性能可媲美大多数人日常使用所青睐的前沿模型。
- 成本:良好质量,在优化 token 开销的同时达到可获得的最高智能水平。
管理员可以决定 Cursor Router 如何在各团队中推行。你可以按团队或群组启用它,选择成员可以选用的模式,设置默认模式,并允许或阻止特定模型。
接下来是什么
Cursor Router 是 Cursor 提升 token 效率的方式之一。选对模型只有在智能体本身保持精简时才有意义,因此我们持续削减其周围 harness 中的浪费。
动态工具调用是另一个明显的例子,大多数原生工具描述不再被加载到每个提示词中。模型在首次需要它们时才会查找,遵循我们已用于 MCP 的相同模式。这让 read 和 edit 等常用工具保持热加载,而不太常用的工具只有在智能体实际调用它们时才进入提示词。
除了 Cursor Router,我们还在不断提升模型池的下限和上限:Grok 4.5 拓宽了 Cursor Router 在更难、成本更高的工作上可以调用的范围。Composer 在日常路径上持续进步,因此低成本轮次也能接近前沿质量,而无需支付前沿价格。
Cursor Router 即日起面向 Teams 和 Enterprise 套餐开放,覆盖桌面端、网页端、iOS、CLI 以及我们的 SDK。
Today we're launching Cursor Router, our intelligent model router for teams and enterprises.
Cursor Router lets teams automatically route every request to the most capable model for the task, delivering frontier intelligence at a lower cost.
We've observed incredibly strong results on production traffic across thousands of enterprise developers. During our early access period with dozens of enterprises, customers got frontier performance at approximately 30–50% lower cost.


In online A/B tests across millions of requests, Cursor Router delivered frontier-quality performance at 60% savings.
Cursor routes hundreds of millions of coding requests each week across every model and provider, with unique visibility into what users like and what stays in the codebase. Model neutrality has always been core to how Cursor works, and today we're putting that data and expertise to work for your team.
With Cursor Router, our goal is to provide teams with the best performance and experience for every task, without spending more than the work requires.
How it works
Roughly 60% of developers using Cursor pick a single model as their daily driver. This results in routine work being completed at frontier prices, and AI spend growing much faster than output quality. Cursor Router fixes that by classifying each request before a model runs.
At its core, Cursor Router is a classifier that routes users to the best model option based on their query. We trained Cursor Router on 600k+ live requests and evaluated performance in an online A/B test across millions of live requests directed by Cursor Router, optimizing for user satisfaction (AFC) as a reward.
Cursor Router analyzes each request on query, context, task complexity, and domain, combined with what we know about each model's behavior. We learn what each model is best at, and route to the most effective option. Simple work goes to the most price-efficient models, UI updates go to the model with the best taste, and more complex, long-horizon problems go to frontier reasoning models.
We designed our routing classifier for a world in which updated models get shipped early and often. This way as newer and more powerful models are released, we can easily update Cursor Router, so the experience keeps improving.
Cursor Router is cache-aware in both how it is trained and evaluated. It is trained on a dataset where routing results in cache misses, and evaluated in production where our reported cost savings include the cost of cache misses in routing decisions.
Frontier intelligence at lower cost
Cursor Router has three modes: Intelligence, Balance, and Cost which let you adjust where you are on the cost-intelligence Pareto frontier.
We found that Auto Intelligence mode lands near Fable on user satisfaction of output at about 60% lower cost for teams, while also lifting satisfaction about 15% over Opus 4.8 at nearly the same cost.
Similarly, Auto Balance lands above Opus 4.8 on user satisfaction with the results at about 36% lower cost. Against GPT-5.6 Sol, Auto Balance delivers comparable satisfaction at a lower spend rate.
We chose to measure the efficacy of our router using large online A/B tests instead of offline evals. While offline evals are useful proxies for quality, they're limited by their small size, their distance from real-world usage, and the difficulty of reducing success to a rubric.
Offline evals also omit the extra cache-miss cost that comes from switching models. Real routing happens across a conversation: which model to pick, and when to switch.
Online A/B tests put Cursor Router to test in the real world across millions of tasks and conversations. Engineers write code, ask follow-ups, hit errors, and keep going, often across hundreds of requests in a week. Those are the conditions under which a model router needs to perform well.
In terms of quality of output, we measured:
- User satisfaction, classifying agent success based on user responses. Moving on to the next feature is a strong positive signal, while correcting the agent is a strong negative one.
- Keep rate, or how much of the agent-generated code remains in the codebase over time.
We have relied on these metrics to evaluate every model launch and harness improvement in the past nine months.
What customers are seeing
Over the past two weeks, Cursor Router has been in early access with a selection of enterprise customers. We compared what they actually paid against the same traffic, priced entirely at Opus 4.8 API rates.
In early access, three high-volume accounts with thousands of users saved 30%–50% on Auto-routed requests versus routing everything to Opus 4.8, with no decrease in quality.


Cost per request is only half the story. Engineering leaders care whether those savings show up in real shipped work, so we looked at cost per commit, and the pattern held.
For a single commit, we observed Cursor Router had a lower cost per commit of $6.76 for Intelligence mode and $4.63 for Balance.
GPT-5.6 Sol matched the cost of Intelligence but had lower user satisfaction with the output. Meanwhile, Fable 5 and Opus 4.8 produced commits at a cost premium to Cursor Router at $12.69 and $7.34 respectively.


That gap is the practical case for routing. Cursor Router keeps hard tasks on the most capable models and moves routine work off of frontier pricing.
You choose the tradeoff
Cursor designed our router with teams and large organizations in mind. The router uses a data-driven taxonomy, while admins and end users can still choose where it sits on the cost-intelligence Pareto frontier.
Select Auto mode in the model picker, and choose from three optimization modes that move you along the frontier:
- Intelligence: Frontier quality, with performance matching the most expensive and powerful models that might be out of reach for daily use.
- Balance: Strong quality, with performance matching the frontier models that most people like to daily drive.
- Cost: Good quality, reaching the highest available intelligence while optimizing token spend.
Admins can decide how Cursor Router rolls out across teams. You can enable it per team or group, choose which modes members can select, set the default, and allow or block specific models.
What’s next
Cursor Router is one piece of how Cursor drives token efficiency. Choosing the right model only matters if the agent itself stays lean, so we keep cutting waste in the harness around it.
Dynamic tool calling is another clear example where most native tool descriptions are no longer loaded into every prompt. The model looks them up the first time it needs them, following the same pattern we already use for MCPs. This keeps common tools like read and edit hot while less commonly used tools only enter the prompt when the agent actually calls them.
Alongside Cursor Router, we keep raising the floor and the ceiling of the model pool: Grok 4.5 widens what Cursor Router can draw from on harder, higher-cost work. Composer keeps getting better on the everyday path, so lower cost turns stay close to frontier quality without paying frontier prices.
Cursor Router is available today for Teams and Enterprise plans across desktop, web, iOS, CLI, and our SDK.