d-Matrix Raptor Joins NVIDIA NVLink Fusion for Rack-Scale AI

d-Matrix Raptor XPU rack connected to NVIDIA NVLink Fusion MGX AI infrastructure

d-Matrix is bringing its next-generation Raptor inference XPUs into NVIDIA’s MGX rack-scale ecosystem through NVLink Fusion, giving the startup a path to deploy specialized AI inference silicon alongside NVIDIA Vera Rubin infrastructure.

In a new infrastructure collaboration, the company said Raptor will connect to NVIDIA NVLink scale-up networking while using the broader MGX platform, including Vera CPUs, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet.

Raptor targets the memory bottleneck in AI inference

Training large AI models is compute intensive, but generating tokens for chatbots, coding assistants and voice agents can become heavily constrained by memory bandwidth. d-Matrix is designing Raptor around stacked 3D DRAM, placing DRAM memory and SRAM compute in a two-layer package.

The approach follows the same broader shift toward specialized inference hardware covered in BitcoinVersus.tech’s Astera Labs Leo memory-controller expansion and Axelera Europa accelerator launch.

144 Raptor accelerators and 2.3 TB of memory per rack

A technical follow-up published September 17 provides considerably more detail. A rack is designed around 18 accelerator trays and nine scale-up switch trays. Each accelerator tray contains two boards with four Raptor R4 accelerators per board, producing 144 Raptor accelerators per rack.

d-Matrix says each tray provides 128 GB of 3D DRAM, producing roughly 2.3 TB of memory across a rack. The company also specifies aggregate memory bandwidth of 7.2 PB/s. These are vendor specifications for the planned architecture, rather than independent production benchmarks.

The rack-scale design complements the dense GPU infrastructure discussed in BitcoinVersus.tech reporting on Giga Computing’s 700 kW AI factory, Qualcomm and AWS custom AI silicon and onsemi’s denser AI-rack power technology.

NVIDIA supplies the rack-scale foundation

NVIDIA independently describes the partnership as a way for d-Matrix to connect custom inference silicon to a mature rack architecture, networking stack, power and cooling ecosystem and supply chain rather than building every surrounding system itself.

The architecture can also split inference work between processor types. NVIDIA GPUs can perform compute-heavy prefill work while Raptor handles latency-sensitive decode operations. That heterogeneous approach is intended to improve utilization by matching each phase of an inference workload with specialized hardware.

In its public social coverage, the brand has been highlighting the NVIDIA NVLink Fusion collaboration and Raptor’s role in heterogeneous AI inference during AI Infra Summit 2026.

Up to 57.6 Tb/s of scale-out bandwidth per rack

The planned accelerator trays include four NVIDIA ConnectX-9 NICs per tray. d-Matrix specifies 3.2 Tb/s of scale-out bandwidth per tray, or as much as 57.6 Tb/s per rack. The roadmap also calls for scaling NVLink across four Raptor racks, totaling 576 accelerators.

Those numbers illustrate why networking has become inseparable from accelerator design. Large AI systems increasingly depend on Ethernet, scale-up fabrics, memory topology, cooling and electrical infrastructure as much as raw processor performance.

BitcoinVersus.tech has tracked that system-level transition through developments including 2,304-fiber AI connectivity and 2.3 MW liquid-cooling infrastructure.

Raptor deployment is planned for 2027

Raptor is expected to tape out before the end of 2026, while initial Raptor XPUs integrated into NVIDIA MGX racks are targeted for the fourth quarter of 2027. That distinction matters: the collaboration and architecture are announced now, but the Raptor rack is not yet a broadly deployed production system.

For AI infrastructure operators, the significance is the growing modularity of the AI factory. Custom inference silicon can increasingly plug into established rack, networking and cooling ecosystems rather than requiring a completely separate data-center architecture.


Advertisement:

BitcoinVersus.Tech Editor’s Note & Disclaimer:
We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. BitcoinVersus.tech provides news, technical analysis and educational information and does not provide financial or investment advice. Verify technical specifications and operational requirements with the manufacturer before making purchasing, deployment or investment decisions. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

Leave a comment