Commissioned study · Generated-kernel correctness
Establish that generated kernels are correct, not just that they compile
Generated GPU kernels can compile, pass a benchmark's fixed-shape check and run fast while still being wrong on shapes, dtypes or values the check never tried. A correctness study scopes the operators you depend on, builds operator-aware oracles and inputs for them, and returns every failure with a seed that reproduces it.
The situation
Our generated kernels compile and appear fast. What evidence is needed to establish correctness across relevant inputs?
What you leave with
Operator-aware oracle plan, adversarial input corpus and a reproducible failure report, with a recommended acceptance check and its measured miss and false-alarm rates.
For: Compiler team; model developer; research sponsor
What the published work established
The findings themselves are described under generated-code research; this page is about commissioning a study on your kernels. In brief, four preprints built on the same 26-op gpuemu corpus show that:
- fixed-shape allclose checks certify buggy kernels as correct, while op-schema-aware seeded fuzzing with an fp64 reference caught all 10 seeded illusions with all 16 controls clean on five GPU classes (The Correctness Illusion);
- input strategies trade recall against false alarms sharply — 78% recall at 0% false positives for boundary-only shapes versus 99% recall at 94% false positives for adversarial values (Test-Input Generation);
- calibrated per-operator, per-dtype tolerances raise recall from 73.2% to 82.4% over hand-picked ones (Tolerance Calibration);
- static PTX metrics see structural regressions but miss semantic ones (Static PTX Metrics).
These are results on a research corpus with seeded bugs. Whether your kernels have similar failures, and which check catches them at an acceptable false-alarm rate, is what a study establishes.
Decisions to settle at scoping
- Which operators and dtypes matter. A study of the twenty operators on your hot path is more useful than a shallow pass over two hundred.
- What “correct” means. The reference, the tolerance policy, and whether bitwise reproducibility is required anywhere.
- Which inputs are in contract. Shapes and layouts the kernel must handle versus those it may reject. A kernel that fails on an input outside its contract is not a bug; one that silently returns wrong values on an in-contract input is.
- Which hardware. Results are reported per GPU class; a kernel correct on one can fail on another.
- What acceptance looks like. The deliverable is accepted when every reported failure replays from its seed — not when a particular number of bugs is found.
Illustrative example — not a client engagement. A team using a model to generate fused attention and normalisation kernels checks each against the framework’s reference operator at one production shape. The study encodes each operator’s schema, calibrates tolerances on reference kernels, and runs boundary-shape and value-range generators on two GPU classes. It finds a normalisation kernel that is wrong only when the reduction dimension is not a multiple of the block size, and recommends boundary-shape sampling as the gate, with adversarial values as a periodic deeper sweep.
Commission this if
- You ship or depend on kernels produced by a model or an automated search, and verify them with a fixed-shape allclose check.
- You run or publish a kernel-generation benchmark and want its correctness checker itself tested.
- You need an acceptance gate for generated kernels whose miss rate and false-positive rate are measured rather than assumed.
Not the right fit if
- Your question is speed. This study is about correctness; performance numbers are recorded but not optimised.
- You need the failing kernels fixed and shipped. That is engineering work on the codebase.
- You need a formal proof of correctness. Testing can find bugs; it cannot prove their absence.
What you receive
- Operator-aware oracle plan. Per operator: the reference implementation (fp64 where feasible), the valid input space from its schema (shapes, dtypes, layouts, broadcasting), and per-(operator, dtype) tolerances with how they were derived.
- Adversarial input corpus. Seeded generators for boundary shapes, stride and broadcast cases and value ranges, chosen by measured recall and control false-positive rate on your operators.
- Control results. Known-correct reference kernels run through the same harness, so the checker's false-alarm rate is reported alongside what it caught.
- Reproducible failure report. Each failure with operator, dtype, shape, seed, observed error against tolerance, hardware class and a replay command.
- Acceptance-check recommendation. What to run on future generated kernels, with its miss and false-positive rates measured on your corpus.
What the study needs from you
- Kernels and their contract. Source or binaries for the kernels in scope, and the operator definition or reference implementation each one claims to match.
- Hardware. Access to the GPU classes you deploy on, or a budget for equivalent cloud instances. Results are reported per hardware class.
- Your current checker. The harness, shapes and tolerances you use today, so the study can measure what they miss.
- Disclosure terms. Agreement on how bugs found in third-party or open-source kernels are reported and whether results may be published.
How the work runs
- 1 Scope. Agree operators, dtypes, hardware classes, the correctness definition (tolerance policy) and conflicts. Unsupported operators are named, not silently dropped.
- 2 Schema and oracle. Encode each operator's valid input space and build the reference. Inputs a kernel is not required to handle are excluded explicitly.
- 3 Calibrate tolerances on controls. Derive per-operator, per-dtype tolerances from repeated runs of known-correct kernels instead of hand-picking them, and record the derivation.
- 4 Generate and run. Run the chosen input strategies on every hardware class in scope, storing seeds for every case.
- 5 Triage, replay and accept. Each failure is replayed from its seed and classified. Acceptance criterion for the deliverable: every reported failure reproduces from its stored seed on the stated hardware.
Limits and unfavourable results
- A clean result means no failure was found under the generated inputs and agreed tolerances. It is not a proof of correctness.
- Tolerances are a policy choice. A looser policy accepts more kernels; the report shows how results move with it.
- Results are specific to the hardware classes, drivers and toolchain versions tested.
- Static analysis of compiled code is not a substitute for running kernels: semantic bugs that swap a constant can compile to identical PTX.
If your current checker turns out to be adequate for your operators — no illusions found, controls clean — that is the result, and it is reported as plainly as a list of failures would be. The fee is fixed at scoping either way.
Engagement terms
- Model
- Fixed-scope study with an agreed operator list, hardware classes and acceptance criteria for the deliverables (every failure replays from its seed), not for the result.
- Who does the work
- Dipankar Sarkar personally designs and runs the study. Any specialist help is disclosed and agreed in advance.
- Commercial basis
- Fixed research scope, quoted after a scoping conversation. The fee does not depend on the result.
- Availability
- Checked per enquiry.
Conflicts are checked before scoping. See publication, IP and independence.
Evidence behind this page
- The Correctness Illusion in LLM-Generated GPU Kernels
2026 · Preprint (not peer reviewed) · cited by 6 (Google Scholar, October 2026)
- Test-Input Generation for Tensor Programs: What Actually Finds Kernel Bugs
2026 · Preprint (not peer reviewed)
- Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels
2026 · Preprint (not peer reviewed) · cited by 1 (Google Scholar, October 2026)
- Static PTX Metrics Track Structural Kernel Regressions but Miss Semantic Ones
2026 · Preprint (not peer reviewed)
- gpuemu-corpus (GitHub) — 26 ops (16 correct controls, 10 buggy variants), drivers for the four papers, stored seeds and replay scripts. MIT or Apache-2.0; no tagged releases.
- gpuemu-corpus (Hugging Face) — The 26-op corpus as a dataset with fp64 references. MIT; updated June 2026.
Questions
Why isn't an allclose check on a fixed shape enough?
Because a buggy kernel can be correct on the shape you test. Fixed-shape allclose checks used by LLM-kernel benchmarks certify buggy kernels as correct; under op-schema-aware seeded fuzzing with an fp64 reference and per-(op, dtype) tolerances, the gpuemu corpus showed 10 of 10 seeded illusions caught and 16 of 16 controls clean on five GPU classes.
Which input-generation strategy should a kernel test use?
It depends on how many false alarms you can tolerate. Across seven strategies on the 26-op corpus, boundary-only shape sampling reached 78% recall with 0% control false positives, while adversarial values reached 99% recall but flagged 94% of correct controls. A study measures the trade-off on your operators before recommending a gate.
How should tolerances be set?
Per operator and per dtype, derived from data rather than picked by hand. Calibrating from 8,076 test runs in the gpuemu corpus produced tolerances far stricter than hand-picked ones and raised bug-detection recall from 73.2% to 82.4%.
Can static analysis of PTX replace running the kernels?
Not for correctness. Static PTX metrics track structural kernel regressions, but semantic bugs that swap a constant compile to identical PTX. Static signals can be a cheap first filter; they are not the oracle.
Does the study cover CUDA and Triton kernels?
The method is operator-level: any kernel that can be called from a harness and compared against a reference can be tested. Which languages, frameworks and hardware are in scope for your study is confirmed during scoping, and anything not covered is stated.
See also
Commissioned study
AI Benchmark Replication Studies
Re-run a published benchmark or paper result under your workload, with raw results and a deviation report.
Commissioned study
AI Benchmark and Evaluation Study Design
Build a pre-specified test for a decision existing benchmarks do not measure, with baselines and failure conditions.
Method
Research Reproducibility Packages
Environment manifest, inputs with provenance, per-table scripts, raw per-run outputs and a known-limits statement, with worked examples.
Overview
Independent AI Technical Validation and Reproducible Evaluation
Choose between replication, benchmark design, model comparison, kernel correctness and feasibility studies.
Technical guides
Commission a scoped study
Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.
Last reviewed 2026-10-07.