On 13 March 2026, AWS and Cerebras announced a plan to put Cerebras CS-3 systems into AWS data centres and expose the resulting inference service through Amazon Bedrock. The more consequential bit is the plumbing: Trainium would process prompts, CS-3 would generate output tokens, and EFA networking would pass the working state between them.
Cerebras Watch analysis
What happened
That split is called disaggregated inference. Prefill is parallel and compute-heavy; decode is serial and memory-bandwidth-heavy. Rather than ask one accelerator to do both jobs, the partners intend to assign each phase to hardware designed around its bottleneck. AWS said deployment was due in the coming months, with open-source models and Amazon Nova on Cerebras hardware later in 2026.
Cerebras claimed the arrangement could provide five times more high-speed token capacity in the same hardware footprint. AWS described an order-of-magnitude speed improvement. Neither figure was an independent benchmark result at announcement, and both companies included forward-looking caveats, so the interesting noun is ‘could’, not a victory parade.
Why it matters
If delivered, this turns specialist decode hardware into a managed cloud component rather than a separate destination customers must adopt. It also makes workload shape a scheduling decision: Cerebras says split mode suits large, stable demand, while conventional combined processing remains useful when prompt and output ratios vary.
The commercial test is whether the extra engineering and network handoff produce repeatable gains across real models, context lengths and traffic patterns. A clean architecture diagram is lovely, but production queues tend to arrive carrying mud on their boots.
Our read
This is a sharper partnership than simply parking someone else's accelerator in a data centre. AWS contributes cloud access, prefill silicon and networking; Cerebras contributes a decode engine built around unusually high on-chip bandwidth. The bet is that composability beats uniformity. The shed translation: stop making one spanner pretend to be the entire toolbox.
What to watch
- Which models and regions reach customer availability first.
- Independent latency, throughput and cost measurements across varied workloads.
- How much overhead the EFA handoff adds at different context lengths.
- Whether customers can route smoothly between combined and split configurations.
Discussion spark: Could disaggregated inference make specialised decode hardware a normal cloud building block, or will handoff overhead keep combined accelerators in charge?
Sources and evidence
Independent WittyWires tracker for public updates about Cerebras. Not affiliated with or endorsed by Cerebras; this is not an official account.