Agentic CUDA Kernel Optimizer: Automated GPU Performance Tuning

The Agentic CUDA Kernel Optimizer is an automated framework that transforms high-level workload descriptions into optimized GPU implementations. By utilizing a closed-loop cycle of code generation, correctness verification, and performance benchmarking, the system iteratively refines CUDA kernels to maximize execution speed while maintaining functional correctness.

Automated Optimization Workflow

The optimizer employs a LangGraph-based workflow to explore the space of kernel implementations and launch configurations. The process follows a structured iterative cycle:

  1. Initialization: The system loads or generates a kernel signature, input cases, a reference implementation for correctness, and an initial kernel candidate.
  2. Evaluation: The reference kernel is executed, and the initial implementation is benchmarked for latency.
  3. Iterative Refinement: The agent proposes changes to the kernel code or launch configurations. These candidates are compiled using NVRTC and launched via the CUDA Driver API.
  4. Validation: Every candidate must pass validation by comparing outputs against a reference (using NumPy) to ensure no regressions in correctness.
  5. Selection: The fastest validated candidate is retained as the best implementation. Ranking is based on the geometric mean of latency across performance cases, excluding compilation and profiler replay times.

Technical Architecture

The system separates the high-level orchestration from the low-level GPU execution environment:

  • Orchestration Layer: Powered by LangGraph and Python, this layer handles the agent's logic, candidate selection, and the integration of external tools.
  • Execution Harness: A standalone C++ harness compiles kernels with NVRTC and executes them through the CUDA Driver API, ensuring that the GPU results are captured and saved for the Python layer to analyze.
  • Tool Integration: The agent can optionally utilize Nsight Compute to inspect GPU performance counters and retrieve NVIDIA documentation for optimization guidance, allowing the model to make data-driven decisions about kernel refinements.

Performance Measurement and Validation

To ensure reliable benchmarking, the system uses CUDA events for timing. It performs 10 warmup launches followed by 100 measured launches per case. Correctness is verified by comparing the output of the candidate kernel against a reference kernel or a NumPy-based oracle.

While the system is designed to optimize individual kernels, the author notes that passing supplied input cases does not guarantee general correctness, and a generated reference kernel is not an independent oracle.

Community Insights and Considerations

Discussion among developers highlights the challenges of maintaining correctness in agentic optimization. One contributor noted that the primary difficulty in such systems is not necessarily finding a faster kernel, but proving that the faster version did not introduce subtle bugs:

"how are you handling regression testing across kernel variants, feels like the hardest part of an agentic optimizer isn't finding a faster kernel, it's proving the faster one didn't quietly break something"

Other users questioned the value of the agentic loop compared to a single-prompt approach (e.g., using Claude Code), suggesting that the iterative feedback loop—specifically the integration of real-world GPU performance counters and compilation errors—is the key differentiator that allows the agent to navigate the complex search space of CUDA optimization.

Sources