Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Method · Contamination

Check whether contamination could change the conclusion

A contamination audit does not ask whether a benchmark is clean. It asks a narrower question that can be answered: if the overlap we can detect were removed, would the conclusion we are about to act on still hold?

The situation

How do we detect whether contamination or overlap could invalidate our benchmark conclusions?

What you leave with

Overlap audit, contamination hypotheses and a sensitivity analysis plan showing whether the conclusion survives removal of flagged items.

For: Benchmark owner; research reviewer

Why ask about the conclusion, not the benchmark

Contamination matters only when it changes what you would conclude. A benchmark with a few leaked items can still separate two models reliably; a benchmark with no detectable overlap can still mislead if leakage took a form the checks cannot see. So this method starts from the comparison a decision depends on and works backwards: which leakage routes could have inflated one side, which of those can be tested, and whether the result survives when the testable ones are removed.

I have not published a contamination study. What follows is a method assembled from widely used techniques, stated with the caveats each carries. The way results are reported — raw per-item outputs, an executed sensitivity comparison rather than a promised one, and a measurement date — follows the recommendations in my audit of how reproducible evaluation conclusions are.

The checks, and what each can see

  • Exact and n-gram overlap finds verbatim reuse. It depends heavily on the normalisation and the choice of n, which is why both are fixed before matching.
  • Embedding near-duplicate search finds light paraphrase, but its threshold trades false matches against misses, so it is calibrated on pairs whose status is known.
  • Canary strings are unique markers some benchmarks embed so that their inclusion in training text can be probed. If a model reproduces one, exposure is likely; if it does not, little is established.
  • Time splits compare items created before and after a model’s stated cutoff. They are informative only when the two groups are otherwise comparable.
  • Perturbation sensitivity compares original items with rewrites that keep the meaning but change surface form and values. A large drop is a warning sign, not proof.

Assumptions

The method assumes item provenance can be reconstructed at least roughly, that the systems can be re-run under the original settings, and that the original analysis used a stated uncertainty method so subset results can be compared like for like (see uncertainty and repeated trials).

Illustrative example — not a client engagement. A team’s internal question-answering benchmark mixes 400 items written by staff with 600 adapted from public documentation. Model A leads model B by five points. The hypotheses list public pre-training text and a retrieval index built from the same documentation. N-gram matching against the index flags 140 adapted items; the staff-written items have no public source. On the staff-written subset the lead shrinks to one point with an interval spanning zero. The report concludes that the five-point gap is not robust to detectable overlap, and that the retrieval-index route — not the pre-training route — is the one the data supports.

What this can and cannot establish

It can establish whether a specific conclusion is sensitive to the overlap that the declared checks detect, and which leakage routes are supported by evidence. It cannot establish that a benchmark is free of contamination, separate memorisation from brittleness with certainty, or say anything about training data that nobody can inspect. Those limits are part of the deliverable, not a footnote to it.

Useful if

  • A benchmark built partly from public text — documentation, forums, exam questions, code repositories — is being used to compare models trained on web-scale data.
  • A score jumped after a model update and you need to know whether memorisation is a plausible explanation.
  • A reviewer has asked how you excluded train-test overlap, and the honest answer is that you have not checked.

Not the right fit if

  • You need a certificate that a benchmark is uncontaminated. No overlap check can provide one, and this method does not pretend to.
  • The benchmark has not been designed yet; start with evaluation dataset design, where leakage controls are cheaper to build in.
  • The question is whether a model leaks private training data to users. That is a privacy threat-model question, not a benchmark-validity one.

What this produces

  • Contamination hypotheses. Each plausible leakage route written down with the signature it would leave and the check that could detect it — before any check is run.
  • Overlap audit. Exact, n-gram and near-duplicate matches against every corpus that is actually accessible, with the normalisation, n and similarity thresholds stated.
  • Indirect probe results. Canary, time-split and perturbation-sensitivity results where training data cannot be inspected, each reported with its own caveats.
  • Sensitivity analysis. The headline comparison recomputed on the flagged-removed and lower-risk subsets, with uncertainty, and a statement of whether the conclusion moved.
  • Untested-routes statement. A list of leakage routes the audit could not test, so a reader knows where the residual risk sits.

What you need before starting

  • The conclusion at risk. The specific ranking, gap or pass rate that a decision depends on.
  • Benchmark provenance. Where each item came from, when it was created and whether it was ever public. Gaps here are recorded, not filled by guesswork.
  • Accessible corpora. Any fine-tuning, few-shot, retrieval-index or development data you control. For closed models, the vendor's stated training cutoff.
  • Model access. Enough access to run the original items and perturbed variants under the same settings.

Protocol

  1. 1
    Name the conclusion and the decision. Write down the comparison that matters (for example, model A beats model B by four points on the benchmark) and the smallest change that would alter the decision.
  2. 2
    Write contamination hypotheses. List each route by which test items or answers could have reached a system: public pre-training text, fine-tuning data, few-shot exemplars, a retrieval index containing answers, or items reused during prompt or hyperparameter tuning. For each, state the observable signature and the check.
  3. 3
    Run direct overlap checks where data is accessible. Exact and n-gram matching after a declared normalisation, plus embedding-based near-duplicate search with a threshold calibrated on known duplicate and non-duplicate pairs. Record every match, not only counts.
  4. 4
    Run indirect probes where it is not. Test whether a model completes any canary string the benchmark carries; split items by creation date relative to stated cutoffs; compare performance on original items against meaning-preserving rewrites with re-drawn values.
  5. 5
    Recompute the conclusion on subsets. Re-score the comparison on (a) items with no detected overlap and (b) the lower-risk time-split or perturbed subset, using the same uncertainty method as the original analysis.
  6. 6
    Report what held and what was not tested. Publish thresholds, flagged item IDs, subset results with intervals, and the untested-routes statement. The verdict is phrased as robust or not robust to detectable overlap — never as clean.

Limits and unfavourable results

  • Overlap checks can show that contamination exists; they cannot show that it is absent. Paraphrased, translated or summarised leakage can evade both n-gram and embedding matching.
  • For closed models the training data cannot be inspected, so every finding is indirect. Stated training cutoffs are self-reported and models may be updated after them.
  • A performance drop on perturbed items is consistent with memorisation but also with ordinary brittleness; the two are not separable from scores alone.
  • Time-split subsets differ from older items in more than exposure — topic, difficulty and format can all shift — so a gap between them is not purely a contamination effect.

If the checks find nothing, the report says what was checked and at what thresholds, and that no overlap was detected by those checks. If removing flagged items reverses the conclusion, that is reported as the primary finding.

Questions

How do you check an LLM benchmark for data leakage?

Write down the plausible leakage routes first, then match test items against every corpus you can access (exact, n-gram and near-duplicate), use indirect probes where you cannot (canaries, time splits, perturbations), and recompute the headline result with flagged items removed. The useful output is whether the conclusion survives, not a contamination percentage.

Can an overlap check prove a benchmark is uncontaminated?

No. Matching finds the overlap it is designed to find. Paraphrased or summarised copies, data you cannot see and post-cutoff model updates all escape it. A careful report says 'no overlap detected by these checks at these thresholds', not 'clean'.

What is a benchmark train-test overlap study?

A study that measures how many test items, or near-copies of them, appear in data a system was trained, tuned or prompted with, and whether those items behave differently. It is most informative when you control the training or fine-tuning data; for closed models it reduces to indirect evidence.

Do you have published results on contamination?

No. I have not published a contamination study, so this page describes a method rather than reporting a result. The reporting conventions — persisted raw outputs, executed sensitivity comparisons and a measurement date — follow my self-audit of evaluation reproducibility.

See also

Prepare an evaluation brief

Describe the benchmark, where its items came from and the conclusion you need to defend. A non-confidential summary is enough to start; corpus access is agreed afterwards.

Last reviewed 2026-10-07.