SemiAnalysis · @SemiAnalysis_ · X·2026-09-24 02:52·42分钟前
AI 导读

vLLM 维护者展示 TPUv7 经 megakernel 优化后在 Kimi K3 上达到 700 tok/s/user,比 NVIDIA GB200 NVL72 高 56%。图表显示 16× TPU v7 为 709 tokens/s,16× GB200 vLLM baseline 为 452 tokens/s。SemiAnalysis 称 TPU 软件外部化进展值得持续关注。

SemiAnalysis@SemiAnalysis_
54AI 编辑部评分,满分 100
2026-09-24 02:52· 42分钟前
AI 导读

vLLM 维护者展示 TPUv7 经 megakernel 优化后在 Kimi K3 上达到 700 tok/s/user,比 NVIDIA GB200 NVL72 高 56%。图表显示 16× TPU v7 为 709 tokens/s,16× GB200 vLLM baseline 为 452 tokens/s。SemiAnalysis 称 TPU 软件外部化进展值得持续关注。

ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user, 56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3.

As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.

来源:SemiAnalysis· x.com