Why GPU Kernel Benchmarks Miss Real Bugs
A GPU kernel that passes its correctness benchmark can still be wrong in production, and the gap isn’t noise — it’s a predictable consequence of how the benchmark chose its test inputs. Correctness benchmarks for generated code, GPU kernels included, typically check outputs against a reference implementation on a fixed or randomly-sampled set of inputs. If that sample doesn’t exercise the input space the kernel will actually see in production, passing the benchmark tells you the kernel is correct on the inputs it was tested against — a narrower claim than “correct,” even though the two get treated as interchangeable.
Why uniform sampling misses the bugs that matter
Most correctness harnesses generate test inputs by sampling uniformly at random from valid shapes and dtypes. That’s a reasonable default when you don’t know anything about where a kernel is likely to break, but GPU kernel bugs are not uniformly distributed across the input space — they cluster at boundaries: the smallest and largest shapes a kernel supports, alignment edge cases, unusual strides, and memory-layout transitions where a kernel’s tiling or vectorization logic has to switch behavior. A uniformly-sampled test set spends most of its budget on “typical” inputs that almost any correct-looking kernel handles fine, and comparatively little on the boundary conditions where an LLM-generated kernel’s flawed understanding of the problem actually shows up.
This is a specific instance of a much older testing principle — boundary value analysis, from classical software testing — applied to a domain (GPU kernel code) where boundaries are less obvious to an outside observer than they are in, say, testing a function that takes an integer range. Knowing that a kernel’s boundaries live in tensor shape, stride, and alignment space, rather than a single scalar parameter, is itself the harder part of building a good test harness for this domain.
What “op-schema-aware” means in practice
An op-schema-aware approach to test generation starts from a kernel’s declared schema — its expected input shapes, dtypes, and the constraints those inputs must satisfy — and deliberately generates seed inputs that sit at and near the edges of that schema, rather than sampling the interior uniformly. Concretely, that means generating shapes right at a kernel’s declared minimum and maximum, dtype combinations that stress numerical precision (mixed precision, denormals), and memory layouts (contiguous versus strided, different tiling boundaries) that exercise different code paths inside the kernel rather than always hitting the same fast path.
The reason this finds more bugs isn’t that it tests “harder” inputs in some vague sense — it’s that it specifically targets the places where an LLM’s generated implementation is most likely to have encoded a subtly wrong assumption. A model that generates a kernel by pattern-matching against training examples tends to get the common case right (because that’s what’s overrepresented in training data) and the edge cases wrong (because correct edge-case handling requires reasoning about the schema’s boundaries specifically, not just recalling a similar-looking kernel).
Why KernelBench-style benchmarks under-specify the input distribution
A benchmark like KernelBench evaluates whether a generated kernel produces outputs matching a reference implementation on the benchmark’s chosen inputs. That’s a legitimate and useful check, but it’s silent about the input distribution the kernel will actually face once deployed — which is a decision made by whoever integrates the kernel into a larger system, not by the benchmark. A kernel tuned (whether by a human or by an LLM iterating against the benchmark) to pass the benchmark’s specific input distribution can be systematically weaker on inputs the benchmark simply never tried, without that weakness being visible anywhere in the benchmark’s own reported score. This is the mechanism behind the “Correctness Illusion”: the benchmark score is real and not fabricated, it just answers a narrower question than the one people read into it.
Comparison: what each testing approach actually tells you
| Approach | What it tells you | What it misses |
|---|---|---|
| Standard benchmark (fixed input set) | Correct on this specific, known distribution | Behavior on any input outside that distribution |
| Uniform random sampling | Correct on “typical” inputs across the space | Boundary and edge-case behavior, where bugs cluster |
| Op-schema-aware seeded fuzzing | Correct across declared boundaries and edge cases | Bugs from constraints not captured in the schema itself |
| Property-based testing (general) | Invariants hold across generated cases | Requires the invariant to be correctly specified up front |
A worked example
Consider a kernel that implements a fused attention operation and is validated on the benchmark’s default sequence lengths — say, 512 and 1024 tokens, both nicely divisible by common tile sizes. The kernel passes. In production, a request arrives with a sequence length of 513 — one token over a tiling boundary the kernel’s generated implementation implicitly assumed. If the kernel’s tiling logic has an off-by-one in how it handles a partial final tile, that bug is invisible to a benchmark that only ever tests tile-aligned lengths, and visible immediately to op-schema-aware fuzzing that deliberately includes non-aligned lengths as boundary cases — precisely because “sequence length not divisible by tile size” is a boundary condition implied by the kernel’s own schema, whether or not the benchmark author thought to include it.
Limitations
Op-schema-aware fuzzing is bounded by the schema it starts from — if the true set of failure-inducing inputs depends on a constraint the schema doesn’t capture (an interaction between two parameters that isn’t visible from either parameter’s individual range), schema-aware seeding won’t specifically target it either, though it’s still more likely to stumble onto it than uniform sampling is, simply by covering more of the boundary space generally. Fuzzing also isn’t a substitute for formal verification where that’s feasible — it increases confidence by finding more counterexamples, but the absence of a fuzzer-found bug is not a proof of correctness in the way a verified property would be.
FAQ
Does this mean KernelBench-style benchmarks are useless?
No — they’re a legitimate, reproducible way to compare kernels against a fixed standard, which is valuable precisely because it’s fixed and comparable across different kernels and models. The issue isn’t that the benchmark measures the wrong thing; it’s that a benchmark score gets treated as a general correctness claim when it’s actually a claim about correctness on that benchmark’s specific inputs.
How much of a difference does op-schema-aware seeding actually make?
In the corpus this line of work is built on, seeded fuzzing found meaningfully more bugs in less wall-clock time than uniform sampling did, concentrated specifically in boundary and edge-case inputs uniform sampling rarely reached. The exact multiplier depends on the operator and how narrow its correct-behavior boundaries are, but the direction — boundary-focused seeding outperforming uniform sampling — held consistently across the operators tested.
Is this specific to LLM-generated kernels, or does it apply to human-written ones too?
The testing methodology applies to any kernel, human-written or generated. What’s specific to LLM-generated kernels is the pattern of where bugs tend to concentrate — an LLM pattern-matching against training data is particularly likely to nail the common case and miss edge cases that require reasoning about the schema rather than recalling a similar example, which is exactly the failure mode boundary-focused testing is best positioned to catch.
What’s the relationship between this and the companion methodology paper?
The Correctness Illusion introduces the 26-op corpus and demonstrates the benchmark gap; the companion paper goes deeper into which specific kinds of test inputs — boundary conditions, large inputs, memory-layout edge cases — actually catch the most bugs, and why.
Bottom line
“Passes the benchmark” and “correct” are different claims whenever the benchmark’s input distribution doesn’t match production’s, and for GPU kernels that gap is largest exactly at the boundaries — shape limits, alignment, tiling edges — that uniform or fixed-input testing systematically under-samples. Seeding test generation from the kernel’s own declared schema, rather than sampling its input space uniformly, targets testing effort at precisely the inputs most likely to expose what an LLM-generated implementation got subtly wrong.