Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Method · Uncertainty

Decide how much a small difference is worth before you measure it

There is no universal number of repetitions. There is a procedure: name the variance sources, estimate them in a pilot, size the study against the smallest difference that would change a decision, and report uncertainty at the level the claim is made.

The situation

How many repetitions and which uncertainty summaries are needed before we treat a small performance difference as meaningful?

What you leave with

Written sampling assumptions, a repeated-trial protocol sized from a pilot, and uncertainty summaries — intervals, rank stability and sensitivity runs — reported alongside every headline number.

For: Research sponsor; evaluation engineer

Why small differences need a protocol

A single run of a small item set produces a number that looks precise. My self-audit of an LLM evaluation — eight open model variants, 293 persisted raw outputs — showed how much that appearance can hide:

  • Identical calls did not reliably return identical structure: mean node-set Jaccard ranged from 0.39 to 0.96 across models, and 72% of prompt-model cells were never node-set-perfect.
  • Under a joint cluster bootstrap over prompts, the two least reproducible models held rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each.
  • Two equally defensible rules for merging repeated campaigns changed four of eight rows and moved the study-wide headline by 7 percentage points.
  • Four of the eight endpoints were withdrawn within ten weeks, so the study as specified can no longer be run.

The same pattern appears at smaller scale in my bounded-autonomy harness: in two pilots of eight tasks each, nominally identical configurations produced task success of 0.25 versus 0.12, and the paper treats its tuning results as diagnostic rather than established for that reason. Both are preprints, and both are about specific settings; the point carried over here is procedural, not a general rate.

Assumptions

The protocol assumes items are a reasonable sample of the inputs the claim is about, that items are independent of one another (if they come in families, the family becomes the resampling unit), and that the systems do not change during the measurement window. Each assumption is written into the report so a reader can see where it might fail.

Illustrative example — not a client engagement. An evaluation engineer sees a 2.5-point gain from a new system prompt on 80 items, one run each. A pilot of 30 items with five repeats shows within-item variation comparable to the gain. The protocol fixes 300 paired items with three repeats, resampling by item, and a two-point threshold. The main run gives a 1.1-point gain whose interval spans zero; under the alternative merge rule the gain is 0.6. The report says the prompts are not distinguished on this evidence.

What this can and cannot establish

It can establish whether a difference or a rank position is stable against the variance sources the study modelled, and how sensitive the conclusion is to defensible analysis choices. It cannot correct for an unrepresentative item set, for contamination (see leakage and contamination checks), or for changes to a system after the measurement date. Fixing the design before results exist is covered in pre-specified evaluation plans.

Useful if

  • Two models, prompts or retrieval settings differ by a few points and someone wants to ship the winner.
  • Your evaluation runs each item once, or averages repeats without saying how they were merged.
  • You report a ranked table and need to know which positions in it the evidence actually supports.

Not the right fit if

  • The difference is large relative to any plausible noise and the decision is cheap to reverse; a pilot is enough.
  • The item set is not representative of the use you care about. More repeats make a biased estimate more precise, not more correct — fix the dataset first.
  • You want a single significance test to approve a launch. This method reports uncertainty; it does not convert it into a yes.

What this produces

  • Sampling assumptions. What population the items stand for, which sources of variation are modelled (items, repeated sampling, seeds, configuration, endpoint drift) and which are not.
  • Pilot variance estimates. Between-item and within-item variation from a small pilot, used to choose item and repeat counts.
  • Repeated-trial protocol. Item count, repeats per item, pairing across systems, resampling unit, merge rule and decision rule, fixed before the main run.
  • Uncertainty summaries. Paired differences with cluster-bootstrap intervals, per-system rank stability, and per-cell provenance linking every number to its raw runs.
  • Executed sensitivity comparison. The main result recomputed under the alternative merge rule, seed or configuration that a reasonable reviewer would have chosen.

What you need before starting

  • The claim and its threshold. The comparison being made and the smallest difference that would change the decision.
  • An item set you can defend. Items drawn from the population of interest, or an honest statement that they are not (see evaluation dataset design).
  • Run access and a budget. Enough calls or compute for a pilot plus the main study, and the ability to persist raw outputs.

Protocol

  1. 1
    State the claim and the minimum meaningful difference. For example: prompt B improves task success over prompt A by at least two points on support tickets. Anything smaller would not justify the migration.
  2. 2
    Map variance sources. Item sampling, within-item stochasticity from repeated calls, seeds and configuration, and changes to a hosted endpoint over the measurement window. Decide which the study will model.
  3. 3
    Pilot. Run a small subset with several repeats per item per system to estimate between-item and within-item variance, and to check whether identical calls even return identical outputs.
  4. 4
    Fix the design. Choose items and repeats so the expected interval on the paired difference is narrower than the minimum meaningful difference. Pair systems on the same items. Set the resampling unit to the item (all repeats of an item stay together) and write down the merge rule for repeated campaigns.
  5. 5
    Run and persist. Store every raw output with its item, system, repeat index, seed, endpoint version and measurement date. Aggregates are computed from these, never stored instead of them.
  6. 6
    Summarise and apply the decision rule. Report the paired difference with its cluster-bootstrap interval; for rankings, the fraction of bootstrap replicates in which each system holds its rank; and the result under the pre-declared alternative merge rule. Call the difference meaningful only if it clears the agreed threshold with the agreed interval; otherwise report the systems as not distinguished on this evidence.

Limits and unfavourable results

  • Intervals describe only the variance sources that were modelled. An unrepresentative item set, an unmodelled configuration choice or a later endpoint change can move results outside them.
  • Bootstrap intervals on small item sets are themselves unstable; with very few items the honest summary is often a range of plausible outcomes, not a sharp interval.
  • Repeatability is not accuracy. A system can give the same wrong answer every time.
  • Hosted endpoints change and are withdrawn. A protocol fixes what was measured and when; it cannot guarantee the measurement can be repeated.

'Not distinguished on this evidence' is a valid and common outcome. It is reported with the interval and the item count that would be needed to distinguish the systems, if that is affordable.

Evidence behind this page

Questions

How many repeated trials does an LLM benchmark need?

It depends on how much results vary between items and between repeated calls on the same item, and on the smallest difference you care about. Estimate both variances in a pilot, then choose items and repeats so the interval on the paired difference is narrower than that threshold. More items usually help more than more repeats of the same items.

Which confidence interval should an AI evaluation report?

One that resamples at the unit the claim generalises over — usually the item or prompt — keeping all repeats of an item together. A naive bootstrap over individual runs treats correlated repeats as independent and produces intervals that are too narrow.

Is a ranked leaderboard table enough?

Not on its own. Report how often each system keeps its rank under resampling. In my self-audit of eight models, the bottom two held rank in 99% and 86% of replicates but the middle four only in 27% to 48%, so the table was reliable about the worst model and not about the best.

What should be reported besides a mean and an interval?

Raw per-run outputs, per-cell provenance, the merge rule for repeated runs with the result under a reasonable alternative, and the measurement date with endpoint versions.

See also

Prepare an evaluation brief

Send the comparison, the size of the difference you are seeing and how the current numbers were produced. I will say whether the evidence can support the claim and what a sized study would need.

Last reviewed 2026-10-07.