Cerebras has introduced the CS-4, a rack-scale inference system that delivers over 1,000 tokens per second for models exceeding 10 trillion parameters. The company claims up to 30 times faster token generation compared with GPU systems in production, along with 10 times better throughput per watt than its predecessor, the CS-3.
The system achieves two-microsecond wafer-to-wafer interconnect latency through a new programmable I/O subsystem that doubles bandwidth and allows direct wafer linking across racks without requiring a switch. This low latency is critical for maintaining interactive performance at massive scale.
The CS-4 represents the first iteration of Cerebras's Nexus Platform Architecture, which separates compute, power, and I/O into modular components. The Wafer-Scale Backpack consolidates the processor, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a single 3D assembly, reducing components by 50 percent.
Power delivery sits 0.5 millimeters from the processor—roughly 100 times closer than the 50 millimeters typical of conventional GPU boards. This proximity nearly eliminates board-level power loss, allowing the WSE-3T processor to receive twice as much power and operate at higher frequencies.
The Cerebras PowerRack forms a separate infrastructure layer for power, cooling, and networking. Compute backpacks slide into place after the PowerRack is installed and facility-qualified, reducing overall deployment time from several days to hours and simplifying future service and upgrades.
The CS-4 houses three of Cerebras's large chips per rack, with each wafer offering up to two times the speed of the previous generation. Cerebras said performance comparisons are based on third-party benchmarking or internal testing, and that observed inference speed improvements versus GPU systems may vary.
