# vLLM 维护者实测 TPUv7 经 megakernel 优化后 Kimi K3 单用户达 700 tok/s，高出 GB200 NVL72 56%

- 来源：SemiAnalysis (@SemiAnalysis_)
- 发布时间：2026-09-24 02:52
- AIHOT 分数：54
- AIHOT 链接：https://aihot.news/items/cmueh9zd805n4rovxdu64dgaq
- 原文链接：https://x.com/SemiAnalysis_/status/2102833399475879977

## AI 摘要

vLLM 维护者展示 TPUv7 经 megakernel 优化后在 Kimi K3 上达到 700 tok/s/user，比 NVIDIA GB200 NVL72 高 56%。图表显示 16× TPU v7 为 709 tokens/s，16× GB200 vLLM baseline 为 452 tokens/s。SemiAnalysis 称 TPU 软件外部化进展值得持续关注。

## 正文

ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user, 56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3.

As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
