Cerebras 从零解析分离式推理:异构分离将吞吐提升 5 倍
Disaggregated Inference From The Ground Up October 01, 2026
Cerebras 公布分离式(disaggregation)推理早期结果:在 Cerebras 系统数量不变、token 生成速度不降的前提下,吞吐提升 5 倍。
We value your privacy
We use cookies to enhance your browsing experience, serve personalized content, and analyze our traffic. You can accept all cookies, reject nonessential cookies, or manage your preferences.
For more details, see our Cookie Policy.Cookie Policy
Manage Preferences Reject All Accept All
Manage Preferences
You can manage how Cerebras uses cookies and similar technologies on this browser or device. Necessary cookies are required to operate the site, keep it secure, enable login, remember your cookie choices, and provide services you request. These cookies cannot be turned off in this preference center....Show more
Necessary Always Active
These cookies are necessary for the website to function and cannot be disabled.
Performance
- [x]
These cookies allow us to count visits and traffic sources so we can measure and improve the performance of our site. This data is anonymized and aggregated to help us understand how the site is used.
Personalization
- [x]
These cookies collect data about how you have interacted with our website to help us improve your web experience and to personalize the content and marketing message to be more relevant to your interests.
Targeting
- [x]
These cookies may be set through our site by our advertising partners. They may be used by those companies to build a profile of your interests and show you relevant advertisements on other sites. If you do not allow these cookies, you will experience less targeted advertising.
Reject All Confirm Preferences Accept All
Products
Developers
Resources
Company
General Compute Selects Cerebras to Bring Ultra-Fast Inference to Agentic Coding >>
Oct 01 2026
Disaggregated Inference From The Ground Up
Increasing inference throughput traditionally means adding hardware or serving more requests concurrently. But pushing concurrency can slow token generation for each user.
By using a technique called disaggregation, we've increased throughput by 5× in early results with the same number of Cerebras systems and no loss in token generation speeds.
At Cerebras, we’re leading the development of heterogeneous disaggregation, which is combining multiple types of chips in one inference system. By assigning different hardware to the segments of inference that are memory- or compute-bound, we unlock gains beyond homogeneous disaggregation.
This is the first installment in a series on this topic. If you’ve heard the term “disaggregation” or the claim that “prefill is compute-bound and decode is memory-bound” and wondered what either actually means, this article is for you.
We’ll build this intuition from the ground up:
- how accelerators balance compute and memory movement
- what the different stages of inference are
- why they place different demands on hardware
- how disaggregation improves the system as a whole
How Accelerators Work
An accelerator is hardware built to speed up particular kinds of computation. GPUs are one example. We can think of accelerators as two parts working together to execute a program.
- Compute units perform the arithmetic themselves—operations such as addition, multiplication, and bitwise logic.
- Memory stores data. These are typically user inputs, model weights, intermediate values, and results.
GPU
Memory Stores inputs, model weights, intermediate values, and results.Memory
bandwidth Data moved
per second(GB/s or TB/s)
Compute Each square is a compute unit.They perform arithmetic in parallel.
Arithmetic only happens in an accelerator’s compute units, not in memory. So, data movement is necessary for every computation. To use an accelerator efficiently, engineers try to get as much arithmetic as possible from each transfer: once data has been moved from memory, how many calculations can we perform before moving more?
Arithmetic intensity is the measure of that efficiency. It's the amount of computational "work" for the amount of data moved. More precisely, the number of floating-point operations divided by the number of bytes transferred.
Arithmetic intensity equals FLOPs divided by bytes transferred.Arithmetic intensity =FLOPs Bytes transferred
Arithmetic intensity tells us how much calculation we do for each byte moved; it does not directly tell us how fast an operation runs. When the processor spends much of its time waiting for data, memory bandwidth matters most. When it stays busy calculating, compute capacity matters more.
Grokking arithmetic intensity is key to understanding an LLM's different computational stages. Let's dive into some real examples.
Adding Two Matrices
Consider a simple matrix addition. To perform it on an accelerator, we need to read two matrices from memory, do the addition in the compute cores, and save the result back.
Let's do a real example.
Here, we load two numbers from memory (2 bytes each), execute a single FLOP to add them, then write the 2-byte result back to memory. This gives us 1 FLOP over 6 transferred bytes, for an arithmetic intensity of 0.167 FLOPs per byte.
Arithmetic intensity: 1 FLOPs divided by 6 bytes transferred: 4 bytes loaded into compute and 2 bytes saved to memory.1 FLOPs 6 Bytes transferred = 0.167 FLOP/byte
Memory
2 1 3 3 5 2 4 2 1 1 3 5 1 3 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4 2 4 1 3 5 2
3 2 1 5 2 4 1 4 2 3 5 2 2 1 3 1 3 5 3 5 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4
Load
4 bytes
Save
2 bytes
Compute
5 3 4 8 7 6 5 6 3 4 8 7 3 4 5 5 4 8 4 8 7 6 5 4 5 4 8 7 6 5 6 5 4 8 7 6
−+
Each value uses 2 bytes. Each addition reads two values and writes one result: 1 FLOP per 6 bytes. Hover or tap any cell to trace its matching inputs and output. Use arrow keys within each matrix; Escape clears the selection.
Notice that as you increase the matrix size with the slider, both the calculation count and the bytes transferred grow in proportion. So in this example, although the total number of FLOPs increases for larger matrices, the relative arithmetic intensity stays the same.
Matrix Multiplication
Let's examine a different example. Matrix multiplication follows the same basic path: load two input matrices from memory, multiply and add their values in the compute units, then save the result matrix back to memory.
But unlike addition, each output value is built from an entire row of the first matrix and an entire column of the second. Hover over an output value to see the input values that contribute to it:
Arithmetic intensity: 45 FLOPs divided by 54 Bytes transferred 45 FLOPs 54 Bytes transferred = 0.833 FLOP/byte
Memory
2 1 3 3 5 2 4 2 1 1 3 5 1 3 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4 2 4 1 3 5 2
3 2 1 5 2 4 1 4 2 3 5 2 2 1 3 1 3 5 3 5 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4
Load
36 bytes
Save
18 bytes
Compute
13 11 13 16 18 25 16 17 11 27 21 25 10 16 13 16 23 20 16 19 22 19 32 35 19 15 15 26 22 33 12 21 13 23 27 21
−+
Each value uses 2 bytes. We count one read per input and one write per result. For matrices with N rows and columns, each result takes N multiplications and N − 1 additions. Hover or tap any input to see affected outputs, or an output to see contributing inputs. Use arrow keys within each matrix; Escape clears the selection.
Notice what changes as you increase the matrix size. The accelerator still loads each input value once and writes each result once, but every loaded value can now contribute to more output values. The amount of data moved grows with the number of matrix entries, but the amount of arithmetic grows faster.
With matrix multiplications, we can increase arithmetic intensity even if only one of the input matrices changes. In the next example, we'll only make the second matrix interactive. Try the slider once more:
Arithmetic intensity: 15 FLOPs divided by 30 Bytes transferred 15 FLOPs 30 Bytes transferred = 0.500 FLOP/byte
Memory
2 1 3 3 5 2 4 2 1 1 3 5 1 3 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4 2 4 1 3 5 2
3 2 1 5 2 4 1 4 2 3 5 2 2 1 3 1 3 5 3 5 2 4 1 3 1 3 5 2 4 1 4 1 3 5 2 4
Load
24 bytes
Save
6 bytes
Compute
13 11 13 16 18 25 16 17 11 27 21 25 10 16 13 16 23 20 16 19 22 19 32 35 19 15 15 26 22 33 12 21 13 23 27 21
−+
Each value uses 2 bytes. We count one read per input and one write per result. The left matrix stays 3 by 3. The right matrix has 3 rows and 1 to 6 columns. Each result takes 3 multiplications and 2 additions. The fixed left matrix uses 18 bytes; each new input column adds 6 bytes loaded and 6 bytes saved. Hover or tap any input to see affected outputs, or an output to see contributing inputs. Use arrow keys within each matrix; Escape clears the selection.
Even with one changing input, arithmetic intensity grows with input size! This relationship between input size and arithmetic intensity is key to understanding inference's different phases.
Arithmetic Intensity of Inference
Inference is a chain of many matrix multiplications. In this section, we'll apply the concepts we learned previously to understanding language models.
This article won't go into the full complexity of the language model architecture. But for our purposes, generating a token from a language model is basically a matrix operation between a model's weights and input tokens:
Memory
Model weights
2
1
3
3
5
2
4
2
1
1
3
5
1
3
2
4
1
3
1
3
5
2
4
1
4
1
3
5
2
4
2
4
1
3
5
2
Input
may the force
3
2
1
5
2
4
1
4
2
3
5
2
2
1
3
1
3
5
3
5
2
4
1
3
1
3
5
2
4
1
4
1
3
5
2
4
Compute
Output
be
13
11
13
16
18
25
16
17
11
27
21
25
10
16
13
16
23
20
16
19
22
19
32
35
19
15
15
26
22
33
12
21
13
23
27
21
The left input matrix represents the model's weights. These are learned values that are fixed during inference. The right input matrix represents the input tokens that the model needs to process. We can think of each column as some numerical representation of a token.
To generate a single output token, the accelerator moves both the weights and inputs from memory into compute and performs a chain of matrix multiplications (simplified here as one operation). Press play to follow this process end-to-end below:
What looks like one continuous generation process is actually two distinct phases. First, the model processes the entire prompt (“may the force”) and produces the first response token, “be.” Then it generates the rest of the response one token at a time: “with” and “you.”
These phases are called prefill and decode, and they have very different arithmetic intensity profiles.
Note Our examples follow one request. In production, servers often batch requests together, processing a new token for several users at once. This lets requests share the work of reading model weights. But larger batches can make each decode step take longer. Total throughput can rise while each user receives tokens more slowly.
Prefill
During prefill, the entire prompt is processed in parallel.
We learned earlier that larger matrix multiplications have higher arithmetic intensity because of more data reuse. We can make this more obvious with a longer input prompt:
Real-world prompts can contain thousands or even hundreds of thousands of tokens. Uploading a 100-page PDF to a chatbot essentially generates a massive matrix calculation that needs to be processed.
This is why prefill has relatively high arithmetic intensity! It's a large matrix operation.
Decode
During decode, tokens are generated one after another.
At first, it may seem like the model needs to run the entire growing sequence through the same large matrix multiplication every time. But language models avoid that with the KV cache.
The KV cache stores keys and values computed for earlier tokens, so we don’t have to calculate them again. As the context grows, reading and using the cache takes more work.
We won't go into the full complexity of the KV cache in this article, but you can treat the gray parts of the matrix as the cached data in each step of token generation. The result is a much smaller matrix operation.
Thanks to the cache, we avoid recomputing earlier tokens’ keys and values.
Note The accelerator still reads cached keys and values from memory and uses them in attention. We show only the new token moving into compute to emphasize that earlier tokens’ keys and values are reused rather than recomputed. See the actual matrix operations in Hugging Face’s KV cache guide.
Prefill and Decode Have Different Profiles
In both phases, all of the model's weights must be brought to compute. That can mean moving hundreds of gigabytes or even terabytes of data. All of this has to be transferred to generate a single token, regardless of whether it's processing a million-token prompt during prefill or processing a single new input token during decode.
This is why prefill and decode have such different computational workloads despite having similar memory movement:
Memory movement Arithmetic intensity
"In the Goblet of Fire film, who does Karkaroff expose as a Death Eater?""Barty""Crouch""Jr."
The Shared Scheduling Problem
So far, we've followed a single request through the system. In production, however, an inference server typically handles many requests concurrently. Some requests may be in the prefill stage, while others are already generating tokens in the decode stage.
When prefill and decode run on the same hardware, they can interfere with one another. Prefill is highly compute-intensive, which can stall ongoing decode requests.
A scheduler determines how requests are assigned and prioritized. Explore how different scheduling strategies affect request processing:
Mixed Prefill first Decode first
Prefill Decode First token
Request 1
What's the classic Star Wars wish for good luck? Explain what it means.
May
the
Force
be
with
you.
It
is
a
way
to
wish
someone
luck,
courage,
and
guidance
on
the
journey
ahead.
Request 2
Summarize these story notes in six words: a crew loses its map during a storm, follows the stars through the night, and finally reaches home just before the sun comes up.
The
crew
returns
home
before
sunrise.
Request 3
Create a six-word slogan for a small neighborhood bakery that bakes fresh bread every morning.
Fresh
bread,
warm
smiles,
every
morning.
Time →
Prioritizing prefill helps new requests begin sooner, but can interrupt users whose responses are already streaming. Prioritizing decode helps existing responses stream smoothly, but can cause new prompts to wait for a long time before any response.
These decisions genuinely affect end-user experience. Switch between the three modes to see how each scheduling strategy affects a chatbot:
You What's the classic Star Wars wish for good luck? Explain what it means.
Assistant…
Waiting for the first token…
Mixed Prefill first Decode first 0.0 / 40
1
2
3
Aggregated systems—prefill and decode running on the same machines—must determine what outcome matters more to their application: getting new requests to their first token quickly, keeping active responses streaming smoothly, or maximizing overall throughput.

Disaggregation
Let's recap what we've learned so far:
- Accelerators (hardware) perform arithmetic. This happens by bringing data from memory to compute and writing data back.
- For the same amount of data moved, some workloads perform more arithmetic than others.
- Prefill and decode are different phases of inference.
- Prefill is computationally involved due to processing multiple tokens in parallel.
- Decode processes one new token per request at each step, but still needs substantial memory access for model weights and the KV cache.
- Balancing these two workloads on the same hardware is difficult and requires trade-offs at scale.
Disaggregation is a systems design pattern. It identifies work with different resource needs, then operates those workloads separately rather than forcing them to share a single pool of hardware.
A disaggregated system uses two separate hardware pools: one for prefill and one for decode. Select “Disaggregated” to see how this changes the schedule.
Aggregated Disaggregated
Prompt tokens · prefill Response tokens · decode
Request 1
Request 2
Request 3
Decode pool
Prefill pool
Tell me:A cat sat on the mat.Say hi:Hi!How are you?Count up:One,two,three.
Time →
In a shared pool, time to first token and inter-token latency are directly coupled. Starting a new prompt can delay active responses; protecting active responses can leave new prompts waiting.
Disaggregation is a lot more than "smoother streaming". More fundamentally, disaggregation changes the control interface for the serving system.
More Control Over Trade-offs
In a shared pool, capacity, batching, and scheduling decisions affect prefill and decode together. Once the stages are separated, operators can tune each one independently—allocating hardware, setting batching policies, and prioritizing latency or throughput according to the product’s requirements.
For example, a system with strict time-to-first-token targets can reserve more capacity for prefill, so incoming prompts start promptly without consuming decode capacity. A system that prioritizes a smooth streaming experience can give decode a larger or more tightly scheduled pool, keeping active responses moving even when prompt traffic spikes.
The point is that these pools can be scaled individually. Click different scenarios to see how pools might be scaled to meet different demands:
Prefill pool
GPU
GPU
GPU
GPU
Decode pool
GPU
GPU
GPU
GPU
Prompt capacity↑
Concurrent responses↑
TTFT↓
End-to-end latency↓
Balanced Long prompts Short outputs Many responses A balanced starting point for both phases.
You can see how disaggregation allows the system to adapt to different workload shapes.
But separating the stages introduces a new need: data transfer between the pools. The most important new piece of that design is the transfer of theKV cache.
The KV Transfer
After prefill builds the KV cache, the system transfers that state to the decode pool, where it is loaded into memory before generation can continue.
Unlike the model weights, which are already loaded in both pools, the KV cache is request-specific data. Handoff adds some network and coordination overhead, which is why disaggregation is most compelling at scale: with enough concurrent work, the gains from independently sizing and scheduling prefill and decode can outweigh the cost of transferring state and operating separate pools.
Once prefill has processed a prompt, decode needs the request-specific state that prefill created. Both pools already have the model weights, so the state that must cross the boundary is the KV cache.
Follow one request below to see that handoff:
Incoming prompt
Prefill GPU
Input · in memory
May the force
KV cache
Compute
KV transfer
Decode GPU
Input · in memory
KV cache
Compute
May the force
Follow the prompt through both pools. Prefill produces “be” and passes the KV cache and next-token information to decode. Decode uses “be” to produce “with,” then “with” to produce “you,” adding each processed token to the cache. “You” would enter the cache on the next pass. The animation shows the sequence of work, not measured hardware use.
Separating the phases adds costs too. The KV cache must travel between pools, and either pool can sit idle if their capacities do not match demand. Latency also depends on whether the cache is sent across colocated machines or across regions. Ultimately the operator must decide whether the latency addition is acceptable.
Heterogeneous Disaggregation
So far, our examples have shown separation of prefill and decode into two pools of the same hardware.
Disaggregation is even more powerful when introducing heterogeneous hardware combinations. The question is no longer simply which accelerator should run the model, but which combination of accelerators best serves this workload.
If each pool can be assigned to different hardware, it makes specialized hardware compelling: an accelerator that is exceptional at one dimension—such as prompt processing, memory bandwidth, token speed, or cost efficiency—can boost performance beyond disaggregating with a single type of chip.
Cerebras has exceptionally high aggregate on-chip memory bandwidth. Its wafer-scale architecture is designed around fast access to local memory, making it particularly well suited to the token-by-token portion of inference.
GPT-oss-120B
High reasoning · 10,000 input tokens
Output tokens/s ↑
Cerebras 1,669 output tokens per second
SambaNova 708 output tokens per second
Groq 475 output tokens per second
Microsoft Azure 319 output tokens per second
Nebius 294 output tokens per second
Baseten 293 output tokens per second
0 600 1,200 1,800
Artificial Analysis ↗ · Sep 10, 2026
Why are some accelerators better at decode?
Remember what makes decode different: we move a lot of data to produce just one new token. Its arithmetic intensity is relatively low, so adding more raw FLOP capacity doesn’t necessarily make tokens arrive faster. The compute needs to be fed with more data from memory.
On a GPU, model data resides primarily in high-bandwidth memory (HBM). Before compute units can use it, data is staged through smaller, faster on-chip memories and caches, including SRAM. This memory hierarchy provides large capacity alongside fast local storage, but data still has to travel from shared external memory toward the compute cores.
Cerebras takes a different approach: it distributes SRAM alongside compute across the entire wafer. That gives compute fast access to nearby memory:
GPU
Shared HBM supplies many compute units
WSE
Local SRAM ↔ compute, inside every cell
This wafer-scale design gives Cerebras exceptionally high aggregate memory bandwidth:
Memory bandwidth
Published peak · per processor
TB/s ↑
Linear Log scale
On-chip SRAM HBM
Cerebras WSE-3 · On-chip SRAM · wafer 21,000 terabytes per second
On-chip SRAM accelerator · On-chip SRAM 150 terabytes per second
HBM4 GPU 1 · HBM4 23.3 terabytes per second
HBM4 GPU 2 · HBM4 22 terabytes per second
HBM3e GPU 1 · HBM3e 8 terabytes per second
HBM3e GPU 2 · HBM3e 8 terabytes per second
HBM3e GPU 3 · HBM3e 8 terabytes per second
HBM accelerator 1 · HBM 7.38 terabytes per second
HBM3e accelerator 1 · HBM3e 7 terabytes per second
HBM3e accelerator 2 · HBM3e 4.9 terabytes per second
HBM3e GPU 4 · HBM3e 4.8 terabytes per second
0 5k 10k 15k 20k 25k
SRAM figures sum local memory bandwidth across the processor; HBM figures measure traffic from off-chip memory. These are different memory tiers, not measured token speeds.
Sources & processor details · Sep 10, 2026 Selected leading AI accelerators, including newly announced designs and earlier hardware used in this article. One processor or wafer per row, not a server or rack. Published peaks do not account for workload, memory capacity, power, or price.
Cerebras WSE-3 — On-chip SRAM · wafer ↗On-chip SRAM accelerator — On-chip SRAM ↗HBM4 GPU 1 — HBM4 ↗HBM4 GPU 2 — HBM4 ↗HBM3e GPU 1 — HBM3e ↗HBM3e GPU 2 — HBM3e ↗HBM3e GPU 3 — HBM3e ↗HBM accelerator 1 — HBM ↗HBM3e accelerator 1 — HBM3e ↗HBM3e accelerator 2 — HBM3e ↗HBM3e GPU 4 — HBM3e ↗
Goals of Disaggregation
Homogeneous and heterogeneous disaggregation provide different benefits.
While homogeneous systems commonly aim to improve responses by smoothing token generation, heterogeneous systems are optimized to get the most value out of the specialized hardware. In Cerebras' case, it provides extreme speed without sacrificing throughput.
Consider an aggregated system with five GPUs, each spending 80% of its execution time on prefill and 20% on decode. An active response advances only when its GPU performs a decode step. Spending four-fifths of the time on prefill slows the average pace of output token generation.
Aggregated Disaggregated
Prefill Decode
Shared GPU pool
Prefill pool
Decode pool
KV transfer
GPU 1
GPU 2
GPU 3
GPU 4
GPU 5
Prompts
Responses
Illustrative assumptions Illustrative alternating prefill/decode passes with 10 ms decode steps and negligible transfer overhead. Same total request throughput: 5× the flow through one decode GPU, with responses finishing 5× faster, keeps its active batch size unchanged.
Across the five GPUs, decode accounts for one GPU’s worth of execution time. In this idealized steady state, disaggregation assigns that work to one dedicated GPU and prefill to the other four. The decode GPU can advance responses without competing with prefill, improving token cadence with the same GPU count.
The more time shared GPUs spend on prefill, the larger the improvement. If GPUs already spend most of their time on decode, there's less interference to remove in a homogeneous system. But reducing interference is only one reason to disaggregate. In a heterogeneous system, disaggregation optimizes for better hardware specialization.
Suppose five Cerebras WSEs currently handle both phases. By moving prefill to a separate accelerator pool, it allows the wafers to specialize in decode, where their capacity is most valuable.
Aggregated Disaggregated
Prefill Decode
Shared wafer pool
Added prefill pool
Decode pool
KV transfer
HW A
HW B
HW C
WSE 1
WSE 2
WSE 3
WSE 4
WSE 5
Prompts
Responses
Illustrative assumptions Illustrative 75% prefill / 25% decode starting split, not a benchmark. Recovered decode time increases output capacity only if prefill, KV transfer, and demand keep up. Prefill-pool size is schematic.
This is a gain in throughput per WSE, enabled by additional prefill hardware.
Agentic applications are a compelling fit for heterogeneous disaggregation. They often involve long, multi-turn workflows, with context growing across model calls. As agents take on longer and more consequential tasks, delays at each step compound. Fast inference becomes essential to how responsive an agent feels and how quickly tasks are completed.
Cerebras Disaggregation unlocks the ability to scale inference without compromising world-class interactivity.
Cerebras Uses Disaggregation to Scale Ultrafast Tokens
Demand for ultrafast generation is growing quickly. As more applications become agentic, the limiting resource is increasingly not just total model capacity, but the number of fast tokens we can deliver to users at once.
At Cerebras, we are building heterogeneous inference systems to meet that demand. The goal is not simply to increase aggregate throughput—it is to scale the volume of inference while preserving the ultrafast token speeds that make applications feel instant.
In a traditional aggregated system, increasing capacity meant deploying more hardware. By leveraging partner accelerators to handle prompt processing, we've increased capacity by 5× in early tests with the same WSE footprint.
We have announced partnerships with multiple hardware partners to bring more ultrafast tokens to the market, reflecting our belief that heterogeneous hardware will be a core inference strategy going forward.
Prefill options
Prompts
KV transfer
Cerebras decode pool
Responses
What's Next
In the next posts, we’ll dive deeper into bringing heterogeneous disaggregation to production. We’ll explore the hardware and software stacks involved, and the economic trade-offs of deploying disaggregated inference at scale.
Follow us on X to be notified when the next post goes live!
Follow
Get Updates
Company
- About Us
- Careers
- Contact Us
- Investor Relations
- Website Terms of Use
- Privacy Policy
- Cookie Policy
- Other Terms & Policies
- Service Status
- Trust Center
News
Insights
Performance comparisons are based on third-party benchmarking or internal testing. Observed inference speed improvements versus GPU-based systems may vary depending on workload, configuration, date and models being tested.
1237 E. Arques Ave Sunnyvale, CA 94085
© 2026 Cerebras.
All rights reserved.
来源:Cerebras:Blog · cerebras.ai