跳到正文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分40
AI 导读

清华论文提出 TokenRouter,一个在 token 级进行大小模型路由的服务系统,吞吐量最高达现有方案的 64.15 倍。针对 vLLM、SGLang 等框架一次请求只跑一个模型、双模型协作时每步都等慢模型的问题,TokenRouter 让每个模型独立部署并来回传递半成品答案,同时保留 KV cache,并短暂挂起请求以凑更大批次。

正文

New Tsinghua paper builds TokenRouter, a serving system that runs per-token small-and-large model routing at upto 64.15X the throughput of existing setups.

Current popular serving frameworks (like vLLM and SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.

TokenRouter gives each model its own server and lets them pass work back and forth. So, it hands a half-written answer between them while keeping the model’s memory of the text so far (the KV cache), and holds requests for a moment so each model works on bigger batches.

Across 5 routing methods, throughput rose 2.01 to 64.15 times over the stronger existing setup.

– arxiv. org/abs/2610.12242

Title: "TokenRouter: Efficient Serving System for Token-Level LLM Routing"

来源:Rohan Paul · x.com