# Cerebras CS-4 发布：推理速度最高可达 GPU 的 30 倍

- 来源：Hacker News 热门（buzzing.cc 中文翻译）
- 作者：sunils34
- 发布时间：2026-08-19 11:18
- AIHOT 分数：50
- AIHOT 链接：https://aihot.news/items/cmszj2hjw0fearodpp7vhn4po
- 原文链接：https://www.cerebras.ai/cs4

## AI 摘要

Cerebras 发布全新机架级解决方案 CS-4，基于 WSE-3 Turbo 芯片，推理速度最高可比 GPU 系统快 30 倍，每瓦吞吐量较 CS-3 提升至多 10 倍。该系统面向超大规模数据中心，采用模块化设计，将晶圆间互连延迟降至 2 微秒，可在超过 10 万亿参数的模型上实现每秒 1000 以上 token 的生成速度。首批 CS-4 将于本季度开始出货。

## 正文

视频 · 前往原文观看

The Fastest AI Just Got Faster.

Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to deploy hyperscale capacity. It is the architecture for frontier AI.​

Three WSE-3 Turbo per System​

Each wafer delivers up to 2x the speed of the previous generation​

More Performance per Wafer​

All new power, cooling, and I/O unleashes even more performance per wafer​

Nexus Rack-Scale Platform

Enables rapid deployment in hyperscale datacenters​

Up to 30x faster than GPUs​

Powered by WSE-Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production.​

Higher ultrafast throughput

The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity.​

Frontier-ready architecture

By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.​

BUILT FOR HYPERSCALE​

CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power, and I/O – each with significant innovation to simplify manufacturing, deployment, maintenance, and upgrades.​

Modular compute backpack design

Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly thatfolds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours.​

High-density power delivery

With power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation.​

Next-gen wafer I/O interface

CS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency, benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch, for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters.​

Deploy infrastructure then compute

CS-4 separates the stable power, cooling, and network layer from its modular wafer-scale compute. The Cerebras PowerRack can be installed and facility-qualified before compute arrives. Compute backpacks then slide into place and connect to power, cooling, and data—reducing deployment from days to hours while simplifying service and future upgrades at hyperscale.​

CS-4 by the numbers

First CS-4 shipments begin this quarter.​ Bring the fastest AI to your data center.​ ​

Datasheet

FAQ
