Arm is changing the division of labor inside mobile AI processors. Its new CSS for Mobile 2 platform gives the C2 CPU cluster two SME2 matrix-compute units while placing dedicated neural accelerators directly inside the Mali G2-Ultra NX graphics pipeline, creating an architecture where AI work can live on the CPU and GPU instead of being pushed exclusively to a separate NPU.
The system-level design is the focus of an Arm platform demonstration published September 16. Rather than presenting the CPU and GPU as isolated IP blocks, Arm describes CSS for Mobile 2 as a configurable platform combining compute, system technologies, physical implementations and software.
In its September 8 architecture announcement, Arm says the C2 cluster combines its high-performance C2-Ultra and efficiency-focused C2-Pro CPUs with two SME2 units. The company says doubling SME2 capability can produce a 70% speedup on its tested small language models, while C2-Ultra delivers up to 1.7x the AI performance and 15% higher single-thread performance than C1-Ultra.
SME2 puts matrix acceleration inside the CPU path
Scalable Matrix Extension 2 matters because it gives the CPU architecture instructions designed to accelerate matrix-heavy workloads. That means some machine-learning operations can execute close to the general-purpose cores without first moving every task to a discrete accelerator.
For agentic workloads, Arm argues that the CPU becomes an orchestration engine. An on-device agent may need to maintain context, schedule applications, coordinate models, invoke services and respond to operating-system events. Those jobs are not simply one large inference pass; they mix conventional control flow with AI computation.
That makes CPU-side matrix capability strategically different from adding another standalone NPU. The CPU can remain responsible for general application execution while accelerating selected AI kernels through SME2 when the workload fits.
Mali adds neural hardware inside the graphics pipeline
The GPU side makes a parallel architectural move. Mali G2-Ultra NX is Arm’s first Mali GPU with dedicated neural accelerators integrated into the graphics pipeline. Instead of treating AI graphics as a workload that must leave the GPU, neural reconstruction and enhancement can operate alongside traditional rendering.
Arm says the combination of Mali G2-Ultra NX and its Neural Technology can deliver up to four times higher performance per watt for neural graphics. The GPU also adds a new execution engine and a next-generation ray-tracing unit, while Arm claims up to 14% higher performance on existing game content.
Independent coverage from The Register likewise highlights the C2 CPU cluster and Mali G2-Ultra NX as a coordinated attempt to raise mobile AI and graphics capability while staying inside smartphone power constraints.
The architecture distributes AI instead of centralizing it
Modern SoCs increasingly contain several kinds of compute engines that can all participate in AI. CSS for Mobile 2 makes that distribution explicit: general-purpose CPU cores gain matrix acceleration, the GPU gains neural acceleration, and system software decides how work moves through the platform.
That is a different emphasis from architectures built around one headline AI accelerator. BitcoinVersus.tech recently examined the MediaTek Dimensity 9600 Pro’s dual-NPU approach, where dedicated neural processors are central to the mobile AI story. Arm’s new platform illustrates why future SoCs can use several complementary AI execution paths at once.
The underlying instruction-set layer also matters. Our guide to instruction set architectures explains how the ISA defines the software-visible operations a processor can execute. SME2 extends that software-visible contract with matrix operations that compilers and AI libraries can target directly.
Software determines whether heterogeneous AI works
Adding specialized units is only useful when software can reach them without forcing developers to hand-build a different execution path for every chip. Arm is pairing the new hardware with KleidiAI libraries, its Neural Graphics Development Kit and an AI Portal intended to expose optimized models, code and tools.
This software layer is what turns the CPU/GPU architecture into a platform rather than a collection of blocks. A developer needs a predictable route from model or graphics workload to the hardware engine that can execute it efficiently.
The same software-versus-architecture tension appears in the broader Arm ecosystem. BitcoinVersus.tech’s coverage of Qualcomm opening Snapdragon X2 to Linux showed how processor capability becomes more valuable when operating systems and developer tooling can actually expose it.
Mobile AI is becoming a scheduling problem
The important shift is not simply that smartphones are getting faster AI hardware. It is that an increasingly heterogeneous processor gives software more choices about where each operation should run.
Control-heavy agent logic may stay on the CPU. Matrix kernels can use SME2. Neural graphics can remain in the GPU pipeline. Other model operations can still move to a dedicated NPU when that is the most efficient engine. The architecture becomes a scheduling problem across specialized compute resources.
CSS for Mobile 2 therefore points toward a mobile SoC where “AI processor” no longer means one block on the die. AI capability is being distributed through the instruction set, CPU cluster, GPU pipeline, dedicated accelerators and the software stack that connects them.
BitcoinVersus.Tech
Advertisement
BitcoinVersus.Tech Editor’s Note:
We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb
BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment