Semiconductors: Arm Neoverse CSS N4 Packs Up to 128 Cores for Agentic AI Silicon

Neon green diagram of Arm Neoverse CSS N4 connecting up to 128 CPU cores, LPDDR6, PCIe Gen 7, DPUs, AI accelerators and chiplet links

Arm is pushing deeper into custom data-center silicon with Neoverse CSS N4, a configurable compute subsystem designed to let chipmakers build CPUs, DPUs and AI-infrastructure processors without assembling every foundation block from scratch. The platform can scale from 8 to 128 CPU cores per die and brings LPDDR6, PCIe Gen 7, chiplet connectivity and high-speed accelerator attachment into one pre-integrated design.

In Arm’s Neoverse CSS N4 announcement, the company positions the subsystem as a foundation for agentic-AI infrastructure where CPUs increasingly handle orchestration, data movement, networking and control-plane work around accelerators. Tom’s Hardware’s independent technical coverage highlights the scale of the platform: up to 128 Neoverse N4 cores per die, as much as 256 MB of L3 cache and support for next-generation memory and I/O.

CSS N4 is not a finished CPU — it is a shortcut to custom silicon

The distinction is important. Arm is not selling CSS N4 as one universal server processor. It is delivering a pre-integrated subsystem containing CPU cores, interconnect, memory and I/O foundations that customers can configure around their own silicon goals.

That approach reduces the amount of low-level integration a chip team has to repeat before it can differentiate. A cloud provider can emphasize general compute. A networking vendor can build around DPU functions. An accelerator company can use the CPU complex as the control layer around specialized AI engines.

Arm summarized that flexibility in its official Neoverse CSS N4 post on X, describing the design as a faster path to differentiated silicon for the agentic-AI era.

Arm describes Neoverse CSS N4 as a configurable foundation for partners building differentiated silicon for agentic-AI infrastructure.

Up to 128 cores changes what one infrastructure die can do

CSS N4 can be configured from 8 to 128 CPU cores per die, giving designers a wide range of targets between compact infrastructure processors and dense scale-out server silicon. Arm says the N4 generation supports frequencies up to 3.8 GHz and delivers up to twice the performance per socket, 1.25× the performance per watt and 1.75× the memory bandwidth of the previous generation under Arm’s comparison conditions.

The scale matters because AI racks still need substantial CPU work even when GPUs or other accelerators perform the matrix-heavy computation. Agents retrieve data, invoke tools, manage services, coordinate storage and generate large numbers of smaller control-plane transactions that cannot all be pushed onto the accelerator.

BitcoinVersus.Tech recently examined the other side of that architecture when AMD compared EPYC Venice with NVIDIA Vera in server workloads. The growing fight is not simply GPU versus GPU; CPU architecture is becoming increasingly visible in the total efficiency of AI infrastructure.

LPDDR6 and PCIe Gen 7 prepare the subsystem for bandwidth-heavy racks

Neoverse CSS N4 is Arm’s first N-Series compute subsystem to support LPDDR6 and PCIe Gen 7. Those interfaces matter because the CPU increasingly sits between very fast accelerators, memory pools, network adapters and storage devices.

A processor can have plenty of arithmetic capacity and still become a bottleneck if it cannot feed the surrounding devices. Faster memory and I/O give silicon designers more room to attach accelerators and move data without forcing every workload through an older interface generation.

That same pressure is visible in local AI systems. BitcoinVersus.Tech recently covered NVIDIA’s updated DGX Spark architecture and two-system clustering, where networking and shared memory capacity determine which models can run effectively even when the underlying compute chip stays the same.

DPUs are one of CSS N4’s most natural targets

Arm explicitly calls out data-processing units as a target for CSS N4. DPUs offload infrastructure functions such as networking, security and storage from the host CPU, which can isolate tenant workloads and free the main processor for application work.

That role becomes more important in agentic systems because autonomous software can generate large numbers of network calls, tool invocations and data-access requests. Moving infrastructure processing onto dedicated silicon can prevent those operations from consuming the same compute resources used for reasoning and inference.

Custom silicon only works if the network can keep up

CSS N4 also arrives as operators prepare for much faster data-center fabrics. BitcoinVersus.Tech’s latest networking report found that 90% of surveyed service providers expect high-capacity AI networking to drive revenue while 88% say optical upgrades are urgent.

That relationship is structural. More CPU cores, faster accelerators and larger memory pools increase the value of each server only if data can move between racks and facilities quickly enough to keep the silicon occupied. The processor, DPU and optical network increasingly have to be designed as parts of the same system.

Arm is selling time-to-silicon as much as CPU IP

The strategic value of CSS N4 may be development speed. Designing a server-class chip from individual IP blocks requires years of integration, validation and software enablement. A pre-integrated compute subsystem lets customers start closer to a working design and spend more engineering effort on the pieces that actually differentiate their product.

That is especially valuable while AI infrastructure is changing quickly. A cloud provider that spends too long building yesterday’s ideal processor can arrive after the workload has shifted. Arm’s pitch is that configurable silicon can shorten that loop without forcing every customer into the same finished CPU.

Agentic AI is turning infrastructure CPUs into orchestration engines

Accelerators will continue to dominate the most visible AI compute, but agentic systems make the supporting CPU more important, not less. Every plan, database lookup, tool call, network transaction and security boundary creates infrastructure work around the model.

Neoverse CSS N4 is Arm’s attempt to make that layer easier to customize before the next generation of AI racks is finalized. If partners adopt it broadly, the result will not be one Arm server chip. It will be many different chips sharing the same configurable foundation.


BitcoinVersus.Tech

Advertisement

Advertisement from BitcoinVersus.Tech.

Editor’s Note

If you value independent technology reporting, consider supporting BitcoinVersus.Tech with a Bitcoin donation: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment