Arm released its Neoverse CSS N4 platform, a semi-custom architecture supporting configurations up to 128 cores per die on TSMC's N3P process. The move targets cloud and data center workloads where hyperscalers and chip designers are increasingly building custom silicon to capture margin.

The CSS program lets customers configure core count, cache, I/O, and connectivity around Arm IP—a model already powering CPUs in Azure and Google Cloud, plus DPUs from Nvidia and Intel. N4 scales from eight to 128 cores per die at up to 3.8 GHz, with frequency adjusting downward as core density increases.

At system scale, N4 supports multi-chiplet and multi-socket designs. Memory options include DDR5 or LPDDR6, with up to 256 MB of L3 cache per die. Each core gets 2 MB of L2, plus 64 KB each of L1 instruction and data cache. Connectivity spans up to 128 PCIe 7/6 lanes, CXL 4.0, and UCIe for chip-to-chip communication.

The spec jump from N2 is substantial: N2 maxed at 64 cores with 1 MB L2 per core, 64 MB L3 total, DDR5 or LPDDR5, and 64 PCIe 5.0/CXL lanes. N4 effectively doubles the memory subsystem and I/O bandwidth while doubling core count.

Arm's benchmarking shows N4 with 128 cores at 3 GHz and 2 MB L2 per core delivering twice the socket performance of N3, plus 1.25x performance per watt and 1.75x memory bandwidth. The performance-per-watt metric matters for cloud operators running utility pricing models—lower power directly reduces TCO.

Arm segments its cores by workload. The N-series (including N4) targets scale-out cloud environments optimizing for efficiency. The V-series prioritizes peak performance; Arm used Neoverse CSS V3 for its AGI CPU, while Nvidia incorporated V2 into Grace. AWS has embedded Neoverse in custom silicon initiatives.