Cerebras CS-4 Doubles Speed on the Same 5nm Chip
The CS-4 puts three wafer-scale engines in one rack for 750 petaflops and more than 4,400 tokens a second per user — using the same silicon as the CS-3, clocked roughly twice as fast.
Cerebras unveiled the CS-4 on Tuesday, a rack-scale inference system the company says roughly doubles the throughput of its predecessor without moving to a new process node or a new chip. The three processors inside it are the same 5-nanometer wafer-scale engines that shipped in the CS-3 — this time running at close to twice the clock speed.
Each Wafer Scale Engine 3 Turbo carries four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimetres of silicon: a complete TSMC wafer rather than a die cut from one. Packing three of them into a single rack gives the CS-4 750 petaflops of AI compute, 129.6 petabytes per second of aggregate on-wafer memory bandwidth, and 7.2 terabits per second of off-wafer I/O — double the per-wafer 2.4Tb/s link the CS-3 offered.
The gain comes from power and packaging rather than lithography. Cerebras rebuilt the chassis around a modular “backpack” that pulls power delivery out of the compute plane, letting the company push far more current into each wafer and pull the heat back out. A CS-4 rack draws 125 to 135 kilowatts, roughly twice a CS-3, while using about 50% fewer components. Cerebras also claims up to ten times more throughput per watt than the CS-3 — a figure that sits awkwardly next to a doubling of both speed and rack power, and one the company has not tied to a published workload.
On measured numbers the system is genuinely quick. Running GPT-OSS-120B, Cerebras reports more than 4,400 output tokens per second for a single user, against roughly 2,000 on the CS-3, and claims speeds up to thirty times faster than GPU-based setups in specific configurations. Direct Wafer Links between the three engines cut inter-processor latency from about five microseconds to two, which is what lets a three-wafer box behave like one accelerator instead of a small cluster.
The constraint Cerebras did not fix is memory. On-wafer SRAM stays at 44GB per engine, unchanged from the CS-3, and that ceiling is what limits long-context inference: production deployments still have to pair the CS-4 with HBM-based hardware in a disaggregated setup to hold large key-value caches. Speed per user is the CS-4’s pitch; capacity per user is still somebody else’s job.
First deliveries are scheduled for this quarter, and Cerebras has not published pricing. The company says a follow-on generation could arrive as early as 2027, targeting four times the speed and twenty times the throughput of today’s system by the end of that year. That roadmap lands in a market that has shifted decisively from training to serving — and for a company that has had to defend its valuation since a rocky IPO debut in May, an inference machine that ships this quarter is a more useful argument than a chip that ships in two years.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.