Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim
Back to publications

arXiv

Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels

Dipankar Sarkar

arXiv preprint arXiv:2607.16228 Preprint (not peer reviewed) Cited by 1 (Google Scholar, October 2026)

Abstract

Derives per-operator, per-dtype tolerances from 8,076 test runs in the gpuemu corpus; calibrated tolerances are far stricter than hand-picked ones and raise bug-detection recall from 73.2% to 82.4%.

Most tensor-kernel correctness tests go through a fixed-shape all close-style check with hand-picked absolute and relative tolerances. The thresholds are copied across the corpus and rarely revisited. We mine the element-wise error distribution of every test case from accumulated cloud GPU runs across the 26-entry gpuemu corpus and 2 dtypes (8,076 result rows). We then ask one empirical question: what absolute tolerance would the kernel itself, observed under its correct implementation, justify? The answer is much tighter than the current hand-picked atol. The largest tightening is attention_triton fp16 at $2{,}184\times$. Restricted to the seven LLM-style buggy variants for which the corpus ships a paired correct counterpart, calibrated per-(op, dtype) tolerances raise bug-detection recall from 73.2% (1,805 of 2,467) to 82.4% (2,034 of 2,467), an absolute gain of 9.3 percentage points (+229 new detections). The control false-positive count rises from 0 to 20 out of 1,882 correct-control cases (+1.1 percentage points).

arXiv comments: 8 pages, 1 figure, LNCS format. Companion paper: arXiv:2606.20128 (P1)

Frequently Asked Questions

Why calibrate tolerances?

Most tensor-kernel correctness tests use fixed-shape all-close checks with hand-chosen tolerances. Tolerances derived from the kernels' own element-wise error distributions are substantially stricter, so a loose tolerance can hide real bugs.

What was the effect on bug detection?

Across seven LLM-style buggy variants, operator- and dtype-specific calibrated tolerances improved bug-detection recall from 73.2% (1,805 of 2,467) to 82.4% (2,034 of 2,467).

GPU KernelsNumerical PrecisionKernel CorrectnessTesting

Related Content