AI Is Starting to Write the CUDA Kernels GPU Engineers Used to Hand-Tune

Editorial illustration of a GPU engineer supervising an AI coding agent that profiles and optimizes low-level CUDA-style kernels on an accelerator chip.

Some of the most specialized programming work in AI infrastructure is starting to shift from hand-tuned CUDA code toward AI-supervised optimization loops.

A new Business Insider report says engineers who specialize in CUDA—the programming model used to extract performance from NVIDIA GPUs—are increasingly spending less time manually tuning every low-level kernel and more time supervising coding agents that search the optimization space for them.

The shift does not make CUDA expertise irrelevant. It changes where that expertise is applied: defining constraints, reading profiler output, validating numerical correctness, recognizing bad optimizations and deciding whether an apparently faster kernel is actually safe to ship.

One public example comes from Ravi Theja, who described an AutoResearch experiment in which coding agents iterated on a GPU Mode kernel challenge. He said he entered with little CUDA background, built a loop around agents, verification, experiment memory and compute, and finished fifth overall after more than 500 official submissions.

Ravi Theja describes using coding agents, verification and experiment memory to optimize a B200 GPU kernel competition entry.

CUDA optimization is a search problem

High-performance GPU kernels are full of tradeoffs. Developers choose thread-block shapes, memory layouts, tiling strategies, shared-memory use, vector widths and synchronization patterns while trying to keep the GPU occupied and data moving efficiently.

The difficulty is that a locally sensible change can make the whole kernel slower. A different tile size may reduce memory traffic but increase register pressure. More parallelism can improve occupancy while creating extra synchronization. A kernel can benchmark well on one tensor shape and regress on another.

That makes the job unusually compatible with an agent loop: generate a candidate, compile it, benchmark it on real hardware, compare the result, keep useful changes and discard regressions.

NVIDIA has already demonstrated the model at scale

NVIDIA’s AVO research provides a primary-source example. In a seven-day attention-kernel optimization run on DGX B200 systems, NVIDIA says AVO explored more than 500 optimization directions.

The resulting multihead-attention kernels reportedly outperformed cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across the evaluated cases. The important part is not only the final number. The system sustained a long search, kept memory of prior attempts and used verification to avoid repeatedly making the same bad changes.

That is the engineering pattern Business Insider is describing: the scarce human skill moves upward from manually trying every optimization to designing and supervising the optimization process.

YC Root Access looks at a company building AI systems that use agents to optimize GPU software, including custom kernels.

The engineer becomes the verifier

Kernel generation is only useful if the output is correct. A faster result that silently changes precision, produces unstable values or only works for one narrow input shape is not an optimization—it is a bug with a good benchmark.

That is why CUDA knowledge remains valuable even when an agent writes more of the code. Experienced engineers know which measurements matter, how to read Nsight-style profiling traces, where memory stalls hide and when a benchmark result is suspicious.

In other words, the AI can expand the number of experiments. The engineer still defines what counts as success.

This is part of a larger agentic coding shift

BitcoinVersus.tech recently covered how NVIDIA built TensorRT Model Connect around coding agents, giving software agents a structured route into model conversion and deployment workflows.

We also mapped 10 GitHub repositories forming an agentic AI production stack, where the recurring pattern is the same: an agent is most useful when it can call real tools, inspect real outputs and loop against objective tests.

And NVIDIA’s Isaac ROS 5.0 agent tooling shows the same architecture moving into robotics, where code generation has to meet physical-system constraints rather than just pass a text-based review.

What changes for CUDA programmers

The near-term job is likely to look less like manually writing every kernel from scratch and more like running a high-performance laboratory. Engineers define the benchmark, establish correctness tests, expose profiler data, let agents search aggressively and then inspect the winners.

That does not eliminate low-level programming. It makes understanding the hardware more important at the review layer, because somebody still has to know why a generated kernel is fast, where it might fail and whether the improvement is worth maintaining.

The bigger shift is that elite GPU optimization is becoming partially automatable. The engineer’s advantage increasingly comes from knowing how to build the search loop, how to verify it and how to recognize when the machine found something genuinely better.

BitcoinVersus.Tech

Advertisement

BitcoinVersus.Tech advertisement.

Editor’s Note

We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support to help further secure the integrity of our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment