Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Research area · Generated code

Is the generated code fast, or just unchecked?

LLM-generated GPU kernels are usually scored on speed after a correctness check that runs one fixed shape against a loose tolerance. My work measures what that check misses, which test inputs and tolerances catch more, and what a generated program should return when it cannot be right.

The situation

A generated kernel benchmarks well but may be wrong on edge cases. Who can design correctness tests and explain the remaining uncertainty?

What you leave with

A generated-code validation study: test-input strategy, correctness oracle, failure corpus and a statement of what the tests cannot rule out.

For: GPU compiler researcher; coding model evaluator

Why the fixed-shape check is not enough

Benchmarks for generated kernels typically decide correctness by running one shape and dtype and comparing to a reference with allclose. That oracle is cheap, and it is the step most speed-up claims rest on. In The Correctness Illusion I built a corpus of correct kernels and LLM-style buggy variants and showed that this check certifies buggy kernels as correct, while op-schema-aware seeded fuzzing against an fp64 reference flags every seeded bug and passes every control.

Three companion preprints on the same corpus take the oracle apart piece by piece:

ComponentQuestionFinding
Test inputsWhich generation strategy finds bugs?Boundary-only shapes: 78% recall, 0% false positives. Adversarial values: 99% recall, 94% false positives.
TolerancesHow tight should the comparison be?Per-operator, per-dtype calibration from 8,076 runs raised recall from 73.2% to 82.4%.
Static signalsCan compiled output stand in for tests?Structural bugs show in PTX metrics; a constant-swap bug compiles to identical PTX.

The same concern applies outside kernels. When an LLM translates a constraint problem into a program, a minimal unsatisfiable core gives the model a precise account of what cannot hold together, and cut fabricated solutions from 79% to 7%. The common thread is that generated code needs an oracle that says why it is wrong, not only whether it passed.

Choosing a route

  • You want the public methods and findings — the papers and the corpus on this page are free to use, and every flagged failure replays from a stored seed.
  • You have a specific kernel, benchmark or speed-up claim to test — that is a generated-kernel correctness study, with an agreed scope, access requirements and acceptance criteria for the deliverables.
  • Your claim is about a compressed or distilled model rather than generated code — see efficient inference and distillation evaluation, which is an open research direction rather than a published result.

What remains uncertain

The corpus is small and hand-seeded, by design: it makes the oracle’s behaviour visible rather than estimating real-world bug rates. Whether the same strategies rank the same way on bugs produced by today’s coding models, at production shapes and on newer toolchains, is an open question I am interested in working on.

Useful if

  • You publish or rely on a speed-up for generated kernels and the correctness check is a single-shape allclose.
  • You are building a kernel or code-generation benchmark and need an oracle that does not certify buggy outputs.
  • You want a research-grade account of what generated code has and has not been shown to do.
  • You have a correctness claim you are prepared to see fail.

Not the right fit if

  • You need a security certificate or a guarantee of correctness for a product. Testing can find bugs; it cannot prove their absence.
  • You need an AI-built codebase fixed and shipped. That is delivery work on dipankar.co.
  • You want performance tuning rather than correctness evidence.

Limits and unfavourable results

  • The gpuemu results concern seeded, LLM-style transcription bugs in a 26-op corpus. They do not estimate the bug rate of any deployed model.
  • All four kernel papers are preprints by the same author on the same corpus, so they are not independent confirmations of each other.
  • Calibrated tolerances tighten detection at a cost: control false positives rose from 0 to 20 of 1,882 cases.
  • Correct on every generated input is still not correct on every input.

Evidence behind this page

  • gpuemu-corpus (GitHub) — 26 ops (16 correct controls, 10 seeded buggy variants), drivers for the four kernel papers, stored seeds and replay scripts. MIT or Apache-2.0; no tagged releases.
  • gpuemu-corpus (Hugging Face) — The same 26 rows with fp64 references. MIT.

Narrower routes from here

Questions

Why do LLM-generated GPU kernels pass benchmarks and still fail?

Because the usual correctness check compares outputs on one fixed shape with a hand-picked tolerance. In The Correctness Illusion, that check certified seeded buggy kernels as correct; op-schema-aware seeded fuzzing against an fp64 reference caught 10 of 10 seeded illusions with 16 of 16 controls clean, on five GPU classes.

Which test inputs find kernel bugs?

In the gpuemu corpus, boundary-only shape sampling reached 78% recall with 0% control false positives. Adversarial values reached 99% recall but flagged 94% of correct controls, which makes them unusable as a gate on their own. Report recall and false-positive rate together.

What is an operator-aware correctness oracle?

One whose tolerances are set per operator and dtype from the kernel's own observed error, rather than copied across the corpus. Calibrated this way, tolerances were far stricter than hand-picked ones and raised bug-detection recall from 73.2% to 82.4%.

Can static analysis replace runtime correctness tests?

No. Static PTX metrics flagged structural bugs, but a semantic bug that swapped one constant compiled to identical PTX. Static deltas are a useful pre-filter, not an oracle.

What can a research-grade evaluation of generated code establish?

Which classes of bug a stated test procedure catches, at what false-positive cost, with failures that replay from stored seeds. It cannot certify that the code is correct or secure in general.

See also

Technical guides

Bring a correctness claim

Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.

Last reviewed 2026-10-07.