Cerebras Systems introduced its fourth-generation AI accelerator, the CS-4, on Aug. 18, delivering twice the tokens per second per user of its predecessor while reusing the same 5nm Wafer-Scale Engine.

The performance doubling comes from higher clock speeds and increased power consumption per wafer, enabled by upgrades to power delivery and cooling within the CS-4 rack. Memory bandwidth also doubles, a direct lever on inference throughput for token-generation workloads. Cerebras claims the CS-4 delivers up to 30 times faster inference than GPU-based systems.

For customers running inference services, the math is direct: same hardware cost, double the token output. That arithmetic—converting existing capital into incremental margin—is the economic core of the product. The WSE-3's SRAM architecture proves especially effective for low-batch decode operations, where traditional HBM-based GPUs face latency penalties.

The system upgrades off-wafer I/O bandwidth from 1.2 terabits per second to 2.4 Tb/s, supporting disaggregated inference setups where Cerebras' SRAM-heavy design pairs with external HBM systems for larger batch sizes. A new Wafer I/O interface—an FPGA card acting as a network interface controller—converts proprietary signaling to standard Ethernet and is field-upgradeable, decoupling Cerebras' upgrade cycles from networking standard changes.

The trade: the CS-4 retains the WSE-3's 44GB of SRAM per wafer, fixing memory capacity. Cerebras mitigates this architectural constraint through disaggregated topologies rather than architectural redesign.

The rack itself is redesigned for modularity, reducing manufacturing lead times and deployment friction—factors that matter when customers are billing against capacity. Cerebras' decision to extract incremental performance from existing silicon rather than wait for a new process node reflects capital discipline: faster iteration, lower NRE risk, and predictable product cadence matter more in infrastructure than raw die capacity.