The Components Inside the NVIDIA GB200 NVL72 AI Rack (Kitchen Analogy)

NVIDIA’s GB200 NVL72 combines CPUs, GPUs, memory, storage, networking, power, and compute boards into one rack-scale AI platform. The infographic presents these systems as separate layers for clarity, although the real hardware is organized across 18 compute trays and additional support modules.

CPU — The Head Chef

The central processing unit, or CPU, manages general-purpose computing tasks and coordinates the work performed by the rest of the system. It runs the operating system, prepares workloads, manages system processes, and directs data toward the GPUs and networking devices.

A GB200 NVL72 contains 36 NVIDIA Grace CPUs, with two Grace CPUs installed in each compute tray. Each Grace CPU is connected to two Blackwell GPUs through NVIDIA’s NVLink-C2C interface, creating a tightly integrated Grace Blackwell Superchip.

GPU x72 — The Line Cooks

A graphics processing unit, or GPU, is a highly parallel processor designed to perform many mathematical operations simultaneously. In AI systems, GPUs calculate the matrix multiplications used for model training, inference, image generation, language processing, and other accelerated workloads.

The rack contains 72 NVIDIA Blackwell GPUs, with four GPUs located in each of its 18 compute trays. Nine NVLink switch trays connect the GPUs into one large 72-GPU communication domain, allowing them to cooperate on models that may be too large or computationally demanding for a single processor.

HBM — The Countertop

High-bandwidth memory, or HBM, is extremely fast memory positioned close to the GPUs. It temporarily holds model parameters, input data, activations, attention data, and intermediate calculations so the GPUs can access them without repeatedly waiting for slower storage.

The GB200 NVL72 provides approximately 13.4 terabytes of HBM3E with an aggregate memory bandwidth of 576 terabytes per second. Unlike the separate shelf shown in the infographic, HBM is integrated closely with the GPU packages rather than installed as a traditional rack-level memory bank.

SSD — The Pantry

A solid-state drive, or SSD, provides persistent local storage for the system. It can hold operating-system files, software, cached datasets, model checkpoints, temporary training data, and other information that must remain available after the system is restarted.

Each GB200 compute tray includes four 3.84-terabyte E1.S NVMe drives for local data caching and one 1.92-terabyte M.2 NVMe boot drive. The local NVMe storage can stage or cache frequently used data, reducing the need to retrieve the same information repeatedly from remote storage systems.

NIC — The Waiter

A network interface card, or NIC, connects the compute tray to other servers, storage systems, management networks, and larger AI clusters. Its role is to move data into and out of the rack quickly enough that the GPUs remain supplied with work rather than sitting idle while waiting for information.

Each compute tray includes four NVIDIA ConnectX-7 single-port 400-gigabit NICs for the primary cluster network. It also includes two NVIDIA BlueField-3 data-processing units that support storage and management networking through InfiniBand or Ethernet connections.

PSU — The Power Source

A power supply unit, or PSU, converts incoming electrical power into the controlled direct-current power required by processors, memory, networking equipment, fans, storage devices, and control electronics. Multiple power supplies are used so the rack can continue operating when an individual unit or power path requires maintenance.

The GB200 rack uses eight power shelves containing six air-cooled 5.5-kilowatt PSUs per shelf. These shelves convert incoming AC electricity into approximately 50–51 volts DC and distribute it through a rack-level bus bar, while redundant power arrangements help improve system availability.

Motherboard — The Kitchen Floor

The motherboard is the primary circuit board that electrically and mechanically connects processors, memory interfaces, networking devices, storage controllers, management electronics, and other components. It distributes data signals, control signals, and electrical power between the devices mounted on or connected to it.

In the GB200 NVL72, “compute board” is a more accurate term than one single rack-wide motherboard.

The compute boards are installed inside the 18 compute trays and carry the Grace Blackwell processing hardware and the high-speed interfaces needed to connect each tray to storage, networking, cooling, power, and the NVLink system.

Important Supporting Systems

The rack also depends on NVLink switch trays, liquid-cooling manifolds, cold plates, copper cable cartridges, management switches, fans, power bus bars, and leak-detection equipment. These systems are not simple accessories: they allow the 72 GPUs to communicate, receive power, remain within safe operating temperatures, and function as one rack-scale AI computer.

BitcoinVersus.Tech Editor’s Note:

We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment