Research area · Generated code
Is the generated code fast, or just unchecked?
LLM-generated GPU kernels are usually scored on speed after a correctness check that runs one fixed shape against a loose tolerance. My work measures what that check misses, which test inputs and tolerances catch more, and what a generated program should return when it cannot be right.
The situation
A generated kernel benchmarks well but may be wrong on edge cases. Who can design correctness tests and explain the remaining uncertainty?
What you leave with
A generated-code validation study: test-input strategy, correctness oracle, failure corpus and a statement of what the tests cannot rule out.
For: GPU compiler researcher; coding model evaluator
Why the fixed-shape check is not enough
Benchmarks for generated kernels typically decide correctness by running one shape and dtype and comparing to a reference with allclose. That oracle is cheap, and it is the step most speed-up claims rest on. In The Correctness Illusion I built a corpus of correct kernels and LLM-style buggy variants and showed that this check certifies buggy kernels as correct, while op-schema-aware seeded fuzzing against an fp64 reference flags every seeded bug and passes every control.
Three companion preprints on the same corpus take the oracle apart piece by piece:
| Component | Question | Finding |
|---|---|---|
| Test inputs | Which generation strategy finds bugs? | Boundary-only shapes: 78% recall, 0% false positives. Adversarial values: 99% recall, 94% false positives. |
| Tolerances | How tight should the comparison be? | Per-operator, per-dtype calibration from 8,076 runs raised recall from 73.2% to 82.4%. |
| Static signals | Can compiled output stand in for tests? | Structural bugs show in PTX metrics; a constant-swap bug compiles to identical PTX. |
The same concern applies outside kernels. When an LLM translates a constraint problem into a program, a minimal unsatisfiable core gives the model a precise account of what cannot hold together, and cut fabricated solutions from 79% to 7%. The common thread is that generated code needs an oracle that says why it is wrong, not only whether it passed.
Choosing a route
- You want the public methods and findings — the papers and the corpus on this page are free to use, and every flagged failure replays from a stored seed.
- You have a specific kernel, benchmark or speed-up claim to test — that is a generated-kernel correctness study, with an agreed scope, access requirements and acceptance criteria for the deliverables.
- Your claim is about a compressed or distilled model rather than generated code — see efficient inference and distillation evaluation, which is an open research direction rather than a published result.
What remains uncertain
The corpus is small and hand-seeded, by design: it makes the oracle’s behaviour visible rather than estimating real-world bug rates. Whether the same strategies rank the same way on bugs produced by today’s coding models, at production shapes and on newer toolchains, is an open question I am interested in working on.
Useful if
- You publish or rely on a speed-up for generated kernels and the correctness check is a single-shape allclose.
- You are building a kernel or code-generation benchmark and need an oracle that does not certify buggy outputs.
- You want a research-grade account of what generated code has and has not been shown to do.
- You have a correctness claim you are prepared to see fail.
Not the right fit if
- You need a security certificate or a guarantee of correctness for a product. Testing can find bugs; it cannot prove their absence.
- You need an AI-built codebase fixed and shipped. That is delivery work on dipankar.co.
- You want performance tuning rather than correctness evidence.
Limits and unfavourable results
- The gpuemu results concern seeded, LLM-style transcription bugs in a 26-op corpus. They do not estimate the bug rate of any deployed model.
- All four kernel papers are preprints by the same author on the same corpus, so they are not independent confirmations of each other.
- Calibrated tolerances tighten detection at a cost: control false positives rose from 0 to 20 of 1,882 cases.
- Correct on every generated input is still not correct on every input.
Evidence behind this page
- The Correctness Illusion in LLM-Generated GPU Kernels
2026 · Preprint (not peer reviewed) · cited by 6 (Google Scholar, October 2026)
- Test-Input Generation for Tensor Programs: What Actually Finds Kernel Bugs
2026 · Preprint (not peer reviewed)
- Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels
2026 · Preprint (not peer reviewed) · cited by 1 (Google Scholar, October 2026)
- Static PTX Metrics Track Structural Kernel Regressions but Miss Semantic Ones
2026 · Preprint (not peer reviewed)
- From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
2026 · Peer-reviewed workshop paper
- gpuemu-corpus (GitHub) — 26 ops (16 correct controls, 10 seeded buggy variants), drivers for the four kernel papers, stored seeds and replay scripts. MIT or Apache-2.0; no tagged releases.
- gpuemu-corpus (Hugging Face) — The same 26 rows with fp64 references. MIT.
Narrower routes from here
Research topic
Efficient Inference and Distillation Evaluation
Open research direction: matched protocols, resource accounting and task-specific degradation for compressed and distilled models.
Commissioned study
Generated GPU Kernel Correctness Studies
Test generated GPU kernels with operator-aware oracles, calibrated tolerances and adversarial inputs.
Questions
Why do LLM-generated GPU kernels pass benchmarks and still fail?
Because the usual correctness check compares outputs on one fixed shape with a hand-picked tolerance. In The Correctness Illusion, that check certified seeded buggy kernels as correct; op-schema-aware seeded fuzzing against an fp64 reference caught 10 of 10 seeded illusions with 16 of 16 controls clean, on five GPU classes.
Which test inputs find kernel bugs?
In the gpuemu corpus, boundary-only shape sampling reached 78% recall with 0% control false positives. Adversarial values reached 99% recall but flagged 94% of correct controls, which makes them unusable as a gate on their own. Report recall and false-positive rate together.
What is an operator-aware correctness oracle?
One whose tolerances are set per operator and dtype from the kernel's own observed error, rather than copied across the corpus. Calibrated this way, tolerances were far stricter than hand-picked ones and raised bug-detection recall from 73.2% to 82.4%.
Can static analysis replace runtime correctness tests?
No. Static PTX metrics flagged structural bugs, but a semantic bug that swapped one constant compiled to identical PTX. Static deltas are a useful pre-filter, not an oracle.
What can a research-grade evaluation of generated code establish?
Which classes of bug a stated test procedure catches, at what false-positive cost, with failures that replay from stored seeds. It cannot certify that the code is correct or secure in general.
See also
Commissioned study
Generated GPU Kernel Correctness Studies
Test generated GPU kernels with operator-aware oracles, calibrated tolerances and adversarial inputs.
Method
Research Reproducibility Packages
Environment manifest, inputs with provenance, per-table scripts, raw per-run outputs and a known-limits statement, with worked examples.
Overview
Agent Evaluation and Coordination Research
Escalation, refusal to fabricate, multi-agent coordination, judge robustness and ranking stability, with papers, artefacts and limits.
Papers & Artefacts
Every paper with its publication status, co-authors, code and data.
Technical guides
Bring a correctness claim
Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.
Last reviewed 2026-10-07.