TPUv7 SparseCore 助推理性能提升 50%

SemiAnalysis · @SemiAnalysis_ · X·2026-09-09 09:03·3小时前
AI 导读

TPUv7 配备专用硬件加速单元 SparseCore,负责数据搬运并将各专家 token 聚为连续组,TensorCore 专注专家矩阵乘法。使用 SparseCore 重排 MoE 内核输入可提升 12% 吞吐量。结合其他优化及更低 TCO,TPU 推理性价比较 Blackwell Ultra 最高提升 50%(InferenceX 数据)。

SemiAnalysis@SemiAnalysis_
36AI 编辑部评分,满分 100

TPUv7 SparseCore 助推理性能提升 50%

2026-09-09 09:03· 3小时前
AI 导读

TPUv7 配备专用硬件加速单元 SparseCore,负责数据搬运并将各专家 token 聚为连续组,TensorCore 专注专家矩阵乘法。使用 SparseCore 重排 MoE 内核输入可提升 12% 吞吐量。结合其他优化及更低 TCO,TPU 推理性价比较 Blackwell Ultra 最高提升 50%(InferenceX 数据)。

TPUv7 has a specialized hardware-accelerated unit called the SparseCore. TPU's specialized SparseCore handles data movement, gathering each expert’s tokens into contiguous groups, while the TensorCore is left to run the expert matrix multiplications. When using SparseCore for rearranging expert inputs into the MoE kernel, it results in 12% better throughput.

Combined with other optimizations & TPU's lower TCO, TPU can achieve up to 50% better perf per dollar than Blackwell Ultra, as seen on InferenceX.

来源:SemiAnalysis· x.com