Linux KNOD Uses GPUs to Process Network Packets

Rack server with a GPU and network interface representing Linux packet processing offload

Linux developers are testing a new way to use ordinary GPUs as network packet processors. The open-source project, called KNOD, moves XDP and receive-side IPsec work onto a GPU from inside the Linux kernel—without CUDA, ROCm, or a userspace GPU runtime in the data path.

At Linux Plumbers Conference 2026, developers Hoyeon Lee and Taehee Yoo reported that KNOD reached up to 70 million packets per second on a Katran-derived XDP workload and 80 Gbit/s on receive-side IPsec. The conference session describes a design where the kernel JIT-compiles an XDP program into GPU machine code, the NIC places packets directly into GPU-accessible memory, and the GPU returns the usual XDP verdicts.

KNOD GPU packet processing graphic showing XDP and IPsec throughput figures
KNOD’s reported LPC 2026 results include up to 70 Mpps for a Katran-derived XDP workload and 80 Gbit/s for receive-side IPsec. Hardware Busters graphic based on developer-reported figures.

Why put packet processing on a GPU?

Fast networking can consume a surprising amount of CPU time. An XDP program may need to inspect headers, run load-balancing logic, drop unwanted traffic, redirect packets, or apply other per-packet decisions millions of times every second. As link speeds climb, dedicating more CPU cores to that work becomes expensive.

A GPU has the opposite shape of a CPU. Instead of a smaller number of general-purpose cores optimized for low-latency serial work, it has many parallel execution lanes designed to run similar operations across large groups of data. That maps naturally onto workloads where the same small packet-processing program is applied repeatedly.

That hardware split is easier to understand alongside BitcoinVersus’ CPU vs. GPU vs. NPU explainer. KNOD is unusual because it is not using the GPU for graphics or AI inference. It is treating the GPU as a kernel-managed networking accelerator.

The Linux Foundation’s Open Source Summit Korea recording features Taehee Yoo and Hoyeon Lee explaining KNOD’s in-kernel GPU packet-processing architecture.

The NIC can write straight into GPU memory

The most important part of KNOD is not simply that a GPU runs networking code. It is the path the packet takes. The network card can DMA incoming packets directly into memory the GPU can access instead of forcing the CPU to copy every packet through a userspace application first.

That idea is closely related to the direct-memory principles behind DMA and RDMA: the fastest path is often the one that avoids unnecessary CPU-mediated copies. KNOD extends that thinking by making accelerator memory part of the packet receive path.

The kernel then JIT-compiles the packet program into AMD GPU machine code and dispatches it. The GPU processes many packets in parallel and returns decisions such as pass, drop, transmit, or redirect. That allows existing Linux networking concepts to remain visible while the heavy per-packet compute shifts onto another device.

This eBPF/XDP deep dive explains the verifier, JIT compilation, maps, and XDP execution model that KNOD is extending onto GPU hardware.

The 70 Mpps result is not just an empty packet loop

The reported 70 Mpps figure came from an XDP workload derived from Katran, Meta’s open-source L4 load balancer. That matters because a packet counter that barely touches packet contents can produce impressive throughput without proving much about useful network processing.

The developers also reported 32 Mpps on an RX 6600 in earlier RDNA2 testing, showing that the idea is not limited to one high-end accelerator. Their broader goal is a generic kernel offload framework rather than a single benchmark-specific trick.

The Linux community has been debating where GPU packet offload could fit compared with CPUs, SmartNICs, DPUs, and other acceleration approaches.

IPsec makes KNOD more than an XDP experiment

KNOD now includes receive-side IPsec as a second proof of concept. The GPU can decrypt ESP packets after the NIC places them in accelerator-accessible memory, with the developers reporting up to 80 Gbit/s in their current tests.

That result should be read carefully. The kernel RFC explicitly describes the IPsec implementation as a proof of concept, receive-only, and not production-ready. The important point is architectural: the same offload framework can support more than one packet-processing feature.

KNOD is still an RFC, not a mainline Linux feature

None of this means Linux servers can enable KNOD with a normal configuration switch today. The code is still under RFC review, and the design crosses several complicated kernel boundaries including networking, BPF, DRM, GPU drivers, memory providers, queue affinity, and control-plane tooling.

The developers are also exploring broader hardware support. AMD GPUs are the main target so far, but Intel Arc graphics and AMD integrated GPUs are on the roadmap, while the networking side is being designed around multiple NIC families rather than a single vendor.

Why this could matter in data centers

If the model survives upstream review and real production testing, KNOD could give operators another option between CPU-only packet processing and dedicated SmartNIC or DPU hardware. A consumer-class GPU or integrated GPU that already exists in a system could potentially absorb work that otherwise consumes valuable CPU cores.

The harder questions are latency, batching behavior, driver complexity, fault isolation, power efficiency across different workloads, and whether the economics remain attractive once a complete production system is built. Those questions matter more than one headline throughput number.

KNOD is interesting because it changes the role of the GPU. Instead of being a separate compute island reached through a userspace runtime, the GPU becomes another accelerator the Linux kernel can schedule directly into the networking path. If that abstraction proves clean enough for upstream Linux, it could open the door to a much broader class of kernel-managed accelerator workloads.


BitcoinVersus.Tech Editor’s Note: KNOD remains experimental and under RFC review. Throughput figures cited here are developer-reported results from the project’s conference material and RFC work, not independent production benchmarks.

Follow BitcoinVersus.Tech on X for Bitcoin mining, AI hardware, semiconductors, Linux, networking, and data-center reporting.

Support independent technology reporting: Bitcoin donations help fund BitcoinVersus.Tech research and publishing.

Disclaimer: BitcoinVersus.Tech provides technology news and analysis for informational purposes only.

Leave a Reply