跳到正文
vLLM 官方博客· Inferact and the vLLM Team·· 22 小时前精选AI 评分63

vLLM 详解 DeepSeek-V4.1-Flash 优化:Agent 场景吞吐提升 5 倍

DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0

AI 导读

Inferact 与 vLLM 社区在 DeepSeek-V4.1-Flash 发布三周内完成优化,低并发速度提升 1.9 倍,150 TPS 约束下吞吐提升 5.3 倍。

推荐理由

vLLM 团队拆解了 SWA bounded replay 和一系列内核优化,说明五倍吞吐提升具体来自哪里,方法可复用于其他 Agent 服务场景。

正文 · 原文

TL;DR: In the three weeks after DeepSeek-V4.1-Flash's release, Inferact and the vLLM community optimized the model, achieving a 1.9× speedup at low concurrency and a 5.3× throughput improvement under a 150 TPS constraint. The performance improvement comes from:

  • We implemented SWA bounded replay with CUDA graphs, achieving a ~30% TTFT reduction.

  • We integrated DeepSeek's newly released kernels, including MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect.

  • We aggressively fused and parallelized the remaining kernels, including mHC side streams, and fused all-reduce with the preceding and following ops into a single kernel.

DeepSeek V4.1 introduces a highly efficient architecture for long-horizon agentic serving tasks: with its causal encoder-decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill. The model is also extremely memory efficient. It combines several techniques to shrink the KV cache size: Compressed Sparse Attention 2 (CSA2), FP4 KV cache, and inter-layer KV cache sharing, pushing the global KV footprint to 890 bytes per token. This post shows how we combine these model-level optimizations from DeepSeek with vLLM-side system optimizations to achieve 5× throughput on the SemiAnalysis AgentX agentic serving benchmark. We highlight our optimizations in two categories: SWA bounded replay and kernel-related optimizations.

SWA bounded replay

DeepSeek-V4.1-Flash keeps two kinds of KV caches. Global KV is compressed, shared across layers, and stored in FP4, at about 890 bytes per token (V4.1 report). Sliding-window (SWA) KV is uncompressed in FP8, covering the last 128 positions in each of the 40 layers.

SWA KV creates two costs:

  1. Prefix caching must store it at every possible hit boundary, costing more than 10× the storage of global KV.

  2. Prefill runs layers 21–39 on every prompt token, though decode only reads their last 128 positions.

One straightforward approach is to recompute SWA KV instead of caching it. Exact recomputation, however, is expensive, because each layer's 128-token window depends on earlier positions in the layer below, so rebuilding it across L layers means replaying roughly L × 128 tokens.

DeepSeek V4.1 introduces SWA bounded replay, which trades exactness for efficiency. It reruns only the last 128 tokens and clips the SWA window at the replay start. The result isn't bit-exact, but DeepSeek reports negligible quality loss (details below). vLLM applies it in two places, one for each cost.

Encoder side: rebuild the window on a cache hit

With encoder-side replay, vLLM caches only the global KV and skips SWA KV. On a prefix hit of length H, it reruns tokens [H − 128, H) to rebuild SWA KV, with windows clipped at s = H − 128.

Decoder side: skip most of the prompt prefill

In DeepSeek V4.1's CED architecture, layer 20 computes the decoder's global KV and layers 21–39 reuse it. vLLM therefore runs layer 20 on every token to produce that global KV, and runs layers 21–39 only on each request's last 128 tokens. For long prompts, this skips nearly half the model.

CUDA graphs for the trimmed layers

After trimming, layers 21–39 do so little GPU work that, run eagerly, kernel launch overhead dominates and the GPU sits idle. Their input shapes also differ from layers 0–20, so the two parts can't be captured in one CUDA graph. vLLM's breakable PIECEWISE graph already breaks in the middle of the model, which gives a natural split: layers 0–20 are captured on the full batch, and layers 21–39 are captured separately on the trimmed batch. This makes CUDA graphs usable for trimmed prefill, and lets layers 21–39 use their own capture sizes for better graph coverage.

SWA bounded replay is on by default for DeepSeek-V4.1 and is controlled by --[no-]swa-bounded-replay.

Accuracy and performance results

Although SWA bounded replay is not exact, DeepSeek has reported only negligible quality loss. We confirmed this in vLLM on benchmarks including GSM8K and GPQA and observed no meaningful accuracy difference (gaps within about 1.5 standard errors).

On performance, the encoder side trades one window of prefill per hit for cache space, so the speedup comes from the decoder side. We measure single-request prefill TTFT in three settings: replay off; replay on without decoder CUDA graphs (layers 21–39 run eagerly, and only eager steps are trimmed); and replay on with decoder CUDA graphs.

Decoder replay with CUDA graphs cuts prefill computation time by 30–40%. CUDA graphs matter most for short prompts, where kernel launches are the bottleneck: without them, launch overhead outweighs the GPU savings and replay is slower than the baseline (up to +12% at 1K on DEP2). For long prompts, the GPU work is large enough to hide launch overhead, so eager replay already captures most of the gain and CUDA graphs add a few more points.

Kernels

DeepSeek released DeepSeek-V4.1-Flash together with new kernels in three of its repositories. DeepSelect is a new top-k library for DeepSeek Sparse Attention. DeepGEMM added sparse indexer kernels and several GEMM-related fused kernels. FlashMLA added NVFP4 KV cache support and a fused attention kernel, MegaAttention. We have integrated several of these open-source kernels into vLLM and track progress in #57448.

Mega-mHC (#56962). Mega-mHC fuses the mHC chain into one kernel: the post step, the delayed-pre step, and RMSNorm. It replaces an existing TileLang fused path that DeepGEMM's implementation now outperforms. The kernel is 1.14–1.51× faster than the TileLang version on NVIDIA GB200.

Mega-Gate (#56266). Mega-Gate fuses the MoE router (gate GEMM, expert scoring, bias, and top-k selection) into one kernel. Previously, these operations ran as one GEMM followed by a separate top-k kernel, costing an extra launch and a round trip through memory for the scores. This fusion leads to 1.18–1.31× kernel speedups at medium batch sizes.

mHC multistream overlap (#57603). In V4.1, the mHC coefficients are shifted by one sublayer, so the next seam's coefficient GEMM reads only residual streams that already exist before attention or the FFN runs. Neither side needs the other's output until the next post/pre step combines them. At small batch sizes, vLLM now computes the next mHC block's coefficients on a side CUDA stream, in parallel with attention and the FFN. This hides work that would otherwise sit on the critical path of latency-bound decode. In the TP4 low-latency scenario, this reduces latency by about 4%.

Sparse MQA logits (#56254). In V4.1, later indexer layers pick their top-k from a fixed set of 16K candidate positions. Previous implementations compute the score for the entire context and mask all non-candidate blocks before scoring. DeepGEMM's sparse kernels score only the candidates, so cost no longer grows with context. Per layer on NVIDIA GB300, it's 1.2× faster at 8K tokens and 14–23× faster at 512K. End-to-end on 4× NVIDIA GB300, decode improves 3–6%. Prefill is 1.43× faster at 512K and 2× faster at 1M context.

MegaAttention with NVFP4 compressed KV (#56935). FlashMLA's MegaAttention kernel performs query RoPE, sparse attention, inverse RoPE on the output, and the FP8 cast in a single launch, writing straight into the buffer the output projection reads. That removes the separate kernels and memory round trips between attention and the next layer. It also reads a new NVFP4 compressed KV format, which is 45% smaller than the previous FP8 KV cache. MegaAttention also improves kernel efficiency by 1.45× through aggressive fusion that eliminates HBM writes between operations.

Low-latency fused WO-A kernel (#58634). For small decode batches on Blackwell, we fuse inverse RoPE, FP8 quantization, the WO-A batch GEMM, and MXFP8 requantization into a single CuTe-DSL kernel, reducing the pre-WO-B chain from three kernels to one. The key idea is to keep intermediate activations on-chip and pipeline data movement with compute, avoiding extra kernel launches and global-memory round trips that dominate at small batch sizes. This improves the fused WO-A path by up to ~2.1× and delivers up to ~6–7% lower inter-token latency at low concurrency.

Engram. V4.1's Engram layers look up rows from two large FP8 tables keyed by hashed token n-grams. Each step reads only a few rows, so table placement and lookup latency matter more than compute. We prefetch CPU-offloaded Engram lookups asynchronously, overlapping host-memory access with decoder compute to speed up low-batch decode (#56512). Engram heads are sharded with a unified TP/DP scheme, and co-located DP replicas share the same host tables, avoiding redundant copies and any DP communication on the lookup path (#57651). For these large host-resident tables, we also support transparent huge pages (THP) to reduce page-fault overhead, delivering up to 10× faster lookup kernels for prefills (#56926). We also added optimizations for cases that don't have enough huge pages available (#59327).

Agentic performance

We measure performance with the SemiAnalysis AgentX benchmark as a representative agentic serving workload (detailed in our previous post). Together, these optimizations give vLLM significant performance improvements over our day-0 implementation. As Figure 6 shows, our low-latency result improves 1.9× over our day-0 result, and the high-throughput result improves about 5×.

For low-latency serving, we use TP4 with FlashInfer attention. Small-batch decode is largely memory-bandwidth bound, so sharding the model weights across four GPUs is a good fit. We also tried MegaAttention, but its main advantage is in higher-throughput settings where fusion has more room to help. At TP4, that benefit was much smaller, and FlashInfer ended up being faster in our runs.

For high throughput, we switch to DEP2, using DP attention with experts split across GPUs. Since V4.1 uses a shared KV latent across all heads, TP would duplicate the KV cache across GPUs. DP avoids this duplication: each GPU stores KV only for the requests it serves, with session affinity preserving prefix-cache locality across turns. MegaAttention further cuts the per-request KV footprint with NVFP4, nearly halving it versus FP8 and increasing per-GPU concurrency.

Notably, V4.1 is highly memory efficient and does not need KV cache offloading throughout the benchmark. We expect KV cache offloading to start helping at higher concurrency with P/D disaggregation.

SWA bounded replay, together with prefill-side kernel optimizations, greatly improved the TTFT as well.

Figure 7 shows the optimized TTFT–throughput trade-off. At around 100K throughput, TTFT drops by nearly 70% through three optimizations combined:

  • SWA bounded replay lets the upper half of the model process only the last 128 tokens, cutting prefill computation roughly in half.

  • CUDA graphs keep the small trimmed replay fast on the GPU instead of bound by CPU kernel launches, so we realize the full speedup.

  • Kernel improvements accelerate the model computation.

Acknowledgments

We thank DeepSeek for open-sourcing DeepSeek-V4.1-Flash and the associated kernels, the Inferact team for the initial model bring-up and optimizations, NVIDIA for their collaboration and support, and SemiAnalysis for the AgentX benchmark.

来源:vLLM 官方博客 · vllm.ai