• About
  • Editorial policy
  • Contact
  • Privacy
  • Legal
  • Cookie settings
Sunday, October 11, 2026
  • Login
PacketNebula
  • Home
  • News
  • Tools
    • Network
    • Security
    • Developer
    • Sysadmin
    • SEO
    • Email & DNS
  • Guides
  • Trackers
    • LLM API pricing
    • Open-weight models
    • Retirements calendar
    • Infrastructure deals
  • About
  • Download
No Result
View All Result
PacketNebula
No Result
View All Result
Home Dev

Cerebras stacks three wafers for 750 PFLOPS

by Stéphane Cardon
3 September 2026
in Dev
0
Image from Cerebras for the story: Cerebras stacks three wafers for 750 PFLOPS

Image: Cerebras

Share on FacebookShare on Twitter

Look at the spec table for two minutes and the story falls out on its own. Every number on the WSE-3 Turbo is exactly double the WSE-3: 125 petaflops becomes 250, 21.6 PB/s of memory bandwidth becomes 43.2, I/O goes 1.2 to 2.4 terabits. Transistor count, core count, SRAM, die area, process node: all identical. So the CS-4 that Cerebras announced on August 18 is a genuinely new rack, three wafers deep, hitting 750 sparse FP16 petaflops and 129.6 PB/s of memory bandwidth. But the silicon inside it is last year's die with the clocks pushed. That's not a knock, exactly. Doubling a 46,225 mm2 wafer's clock without melting it is hard engineering. It just means the interesting part of this launch is the rack, not the chip, and the 30x headline needs reading with care.

The short answer

Cerebras announced the CS-4 on August 18, a rack holding three WSE-3 Turbo wafers. The rack is new: modular backpacks, redesigned power delivery, direct liquid cooling. The die is not. Same 4 trillion transistors, same 900,000 cores, same 44 GB of SRAM, roughly twice the clock. The 30x-faster-than-GPUs line measures tokens per second per user on one model against unnamed GPU hardware. First shipments this quarter.

750PFLOPS sparse FP16, three wafers
2xper wafer, from clocks alone
no priceand no named customer
Answer card: Cerebras announced the CS-4 rack on 18 August 2026, three WSE-3 Turbo wafers for 750 sparse FP16 petaflops, 129.6 petabytes per second of memory bandwidth and 7.2 terabits per second of I/O, on the same 4 trillion transistor die with 900,000 cores and 44 gigabytes of SRAM, first shipments this quarter with no published price.
The one-card version. New rack, old die, and two blanks where the price and the customer should be.

What actually shipped

A CS-4 is one rack. Inside it sit three Wafer Scale Engine 3 Turbo processors, each one a full 300 mm wafer of TSMC 5 nm silicon carved into a single chip. Cerebras rates the system at 750 PFLOPS of AI compute, 129.6 PB/s of aggregate memory bandwidth, 160.5 PB/s of on-chip fabric bandwidth and 7.2 Tbit/s of external I/O, with wafer-to-wafer latency as low as 2 microseconds over what it calls Direct Wafer Links. No switch in the path. Against a single CS-3 at 125 PFLOPS, that’s six times the listed compute in one box.

The per-wafer table is where it gets interesting. WSE-3 Turbo: 4 trillion transistors, 46,225 mm2, 900,000 AI cores, 44 GB of on-wafer SRAM, 250 PFLOPS, 43.2 PB/s of memory bandwidth, 2.4 Tbit/s of I/O. Now put that beside the WSE-3 from the CS-3 generation. Transistors, area, cores, SRAM, process node: identical, to the digit. Compute, bandwidth, fabric, I/O: exactly 2x, also to the digit.

Cerebras CS-4 rack render: a black cabinet bearing the Cerebras logo, with three horizontal Wafer-Scale Backpack modules pulled partway out of its left flank, each dense with orange-tipped connectors, and power and liquid cooling hardware mounted on the right.
Image: Cerebras

That pattern only comes from one thing. You don’t get a clean 2x on four unrelated metrics from an architectural change; you get it from turning the clock up. The Register put the die at roughly 1.4 GHz to 2.8 GHz, and Cerebras hasn’t published a frequency to argue with. We’d rather they just said so. Doubling clocks across a wafer that size, with the power delivery and the thermals that implies, is a real achievement. Dressing it as a new processor generation invites exactly the scrutiny it’s now getting.

Checklist of what changed in the Cerebras CS-4: the rack packaging, Wafer-Scale Backpack, power delivery and liquid cooling are new, and compute, memory bandwidth, fabric and I/O all exactly doubled per wafer, while transistor count, core count, SRAM, die area and process node are unchanged, with no price, no named customer and no independent power figure disclosed.
Sorted by what's genuinely new versus what got a clock bump and a new label.

The 30x number, unpacked

Cerebras led with “up to 30 times faster than GPU-based solutions.” Here’s what sits underneath it. In the side-by-side clip published with the launch, GPT-OSS-120B generates 4,465 tokens per second per user on a CS-4, 2,308 on a CS-3, and 131 on a GPU deployment that goes unidentified. That’s a 34x ratio, rounded down to a marketing 30x.

Log-scale comparison of tokens per second per user on GPT-OSS-120B: 4,465 on the Cerebras CS-4, 2,308 on the CS-3, and 131 on an unnamed GPU deployment.
The gap is real and large. The measurement is also the one that flatters wafer-scale hardest.

Three things to hold onto. Tokens per second per user is a latency metric, so it rewards an architecture that keeps an entire model resident in SRAM and never touches HBM, which is the whole Cerebras thesis. It says nothing about tokens per dollar or tokens per rack when you’re serving thousands of concurrent sessions. Second, Cerebras’ own asterisk marks its petaflop figures as sparse FP16; comparing that against dense GPU numbers stretches the gap before anyone runs a model. And the GPU system stays anonymous, which for a claim this loud is a choice.

None of that makes the demo fake. We’ve seen the GPT-5.6 Sol Ultrafast tier run at speeds nothing else touches, and OpenAI says that mode runs up to 14x its standard offering for select customers on Cerebras hardware. The speed is real. The 30x framing is the vendor’s best case dressed as a general result.

Three side-by-side screen recordings from the Cerebras launch video, each generating the same code with OpenAI GPT-OSS, labelled 4,465 tokens per second on CS-4, 2,308 on CS-3 and 131 on GPU.
Image: Cerebras. Watch the full side-by-side demo on Cerebras.

The rack is the real product

Strip the chip story away and the Nexus platform is what Cerebras actually built this year. The rack splits into three independent elements, compute, power and I/O, so each can be replaced on its own schedule. Compute arrives as a Wafer-Scale Backpack: a vertically mounted module that carries the processor along with its power conversion, its direct liquid cooling loop, its high-speed I/O and its control electronics, and plugs into the rack as one unit.

Cerebras says the design uses 50 percent fewer components than the previous generation and is 60 percent more automated to manufacture, and that power conversion now sits about 100 times closer to the processor than on a conventional GPU board. That last one is the physics that makes doubled clocks survivable. Shorter distance, less resistive loss, less voltage droop under a transient.

Efficiency is where the claim gets more defensible: up to 10x the throughput per watt of a CS-3. Even discounted heavily, a 2x clock that doesn’t cost 4x the power is the part of this launch we’d take seriously. Cerebras hasn’t published a rack power figure. The design implies roughly double the per-wafer draw, and third-party estimates land a full rack near 120 to 140 kilowatts, which would be around half a comparable GPU rack. Estimates, though. Nobody outside Cerebras has metered one.

Prefill on someone else’s silicon

The genuinely new architectural idea is disaggregated inference, and it’s an admission worth noticing. Prefill, the phase where a model chews through your input context, is compute-bound and suits GPUs fine. Decode, the token-by-token generation, is memory-bandwidth-bound and is where wafer-scale wins. So the CS-4 supports splitting the two across different hardware over standards-based RoCE v2 RDMA on plain Ethernet, with AMD’s Helios racks and AWS Trainium named as prefill partners. Cerebras claims the pairing hits 10x GPU speed with 5x the throughput of Cerebras hardware alone.

Read that as Cerebras conceding it doesn’t want the whole workload. It wants the decode half, sitting next to a GPU fleet you already own. Which is a smarter go-to-market than asking anyone to rip out their accelerators, and it’s why the RoCE v2 detail matters more than the petaflops.

Should you care yet

If you’re buying inference capacity, probably not this quarter. First shipments are promised before the quarter closes, but Cerebras disclosed no CS-4 customer agreements and no pricing, and a rack-scale system with no public price is not something you can budget against. Models over 50 trillion parameters are supported on paper, with over 1,000 tokens per second quoted on models past 10 trillion. Paper, still.

If you’re building agents, the argument is more concrete. Sean Lie, Cerebras’ CTO, framed it as reasoning headroom rather than snappiness, and that’s the right frame: at 4,000-plus tokens per second per user, a chain that would take a minute elsewhere finishes fast enough to run several times over. Whether you can rent that at a sane price is the open question, and it stays open until someone publishes a number.

The 2027 test is the one that counts. Clocks only go up once. Next generation has to be new silicon.

Sources

Cerebras, “Introducing Cerebras CS-4” (official announcement, 18 August 2026, including the spec table and the side-by-side demo video). Cerebras CS-4 product page. ServeTheHome on the WSE-3 Turbo and the CS-4 rack. The Next Web on what the launch leaves out, which carries the clock-speed reporting attributed to The Register and the rack power estimate. The Next Platform on the Nexus architecture.

Frequently asked questions

What is the Cerebras CS-4?

It is a rack-scale AI inference system Cerebras announced on 18 August 2026, the first product built on its Nexus platform architecture. One CS-4 holds three WSE-3 Turbo wafer-scale processors and is rated at 750 sparse FP16 petaflops, 129.6 PB/s of aggregate memory bandwidth, 160.5 PB/s of on-chip fabric bandwidth, 7.2 Tbit/s of external I/O and wafer-to-wafer latency as low as 2 microseconds. Cerebras says first shipments begin this quarter.

Is the WSE-3 Turbo a new chip?

Not a new design. It keeps the WSE-3's 4 trillion transistors, 900,000 AI cores, 44 GB of on-wafer SRAM and 46,225 mm2 of TSMC 5 nm silicon. What changed is clock speed: compute, memory bandwidth, fabric bandwidth and I/O each land at exactly twice the WSE-3 figure, which is the signature of a frequency bump rather than a redesign. The Register reported the die moving from roughly 1.4 GHz to 2.8 GHz. Cerebras has not published a clock speed itself.

What does the 30x faster than GPUs claim actually measure?

Tokens per second per user, on one model, against GPU systems Cerebras does not name. In the side-by-side clip published with the launch, GPT-OSS-120B runs at 4,465 tokens/sec/user on a CS-4, 2,308 on a CS-3 and 131 on the GPU box. That is a single-stream latency metric, not aggregate throughput, and Cerebras' own footnote marks its petaflop figures as sparse FP16. Treat 30x as the vendor's best case, not a general speedup.

How much does a CS-4 cost and who is buying one?

Cerebras published neither. No price, no list price per rack, no named CS-4 customer agreement in the launch materials. It also has not released an official rack power figure, though the design implies roughly double the per-wafer draw of the WSE-3 and third-party estimates put a full rack somewhere around 120 to 140 kilowatts. All of that is unconfirmed.

What is disaggregated inference on the CS-4?

Splitting the two phases of inference across different hardware. Prefill, where the model reads your input context, runs on GPUs or other accelerators; decode, the token-by-token generation, runs on the Cerebras wafers. The CS-4 supports this natively over standards-based RoCE v2 RDMA on Ethernet, with AMD Helios and AWS Trainium named as prefill partners. Cerebras claims the pairing reaches 10x GPU speed and 5x the throughput of Cerebras alone, which is a vendor number nobody has reproduced yet.

Tags: aidatacentergpuhardwareinferencenews
ShareTweet
Previous Post

Stripe is reportedly buying OpenRouter for over $7B

Next Post

Cursor opened its git forge the day GitHub broke

Stéphane Cardon

Stéphane Cardon

Network cybersecurity engineer: architecture, LAN and WLAN, system administration. He runs PacketNebula, where he publishes the tools and answers he wanted for his own work.

Next Post
Image from OpenAI for the story: OpenAI's Private Safety Processing keeps ZDR alive

OpenAI's Private Safety Processing keeps ZDR alive

Free tools

  • IPv4 subnet calculator
  • DNS lookup
  • SSL certificate checker
  • HTTP headers checker
  • SPF, DKIM and DMARC checker
  • JWT decoder
  • Cron expression builder
  • Chmod calculator
  • Password generator

All 24 tools →

Guides to keep handy

  • Flush the DNS cache on Windows
  • Kill the process using a port
  • Generate an SSH key
  • Secure a new Ubuntu VPS
  • Find files on Linux
  • Extract and create tar.gz

More guides →

Who writes this

Stéphane Cardon, network cybersecurity engineer. Architecture, LAN and WLAN, system administration. PacketNebula holds the tools and answers he wanted for his own work.

About this site →

  • About
  • Editorial policy
  • Contact
  • Privacy
  • Legal
  • Cookie settings

Copyright © 2026 Stephane Cardon.

No Result
View All Result
  • Home
  • News
  • Tools
    • Network
    • Security
    • Developer
    • Sysadmin
    • SEO
    • Email & DNS
  • Guides
  • Trackers
    • LLM API pricing
    • Open-weight models
    • Retirements calendar
    • Infrastructure deals
  • About
  • Download

Copyright © 2026 Stephane Cardon.

Cookies, your call

We use Google Analytics to see which pages get read. Once our AdSense application is approved, articles and guides will carry Google ads. Nothing runs until you say yes, and nothing runs on the tool pages. Details