Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Method · Pre-specification

Agree the success criteria before anyone sees a result

Most evaluation choices are defensible in isolation — which metric, which runs to merge, which items to drop. The problem is making them after the numbers are visible. A pre-specified plan fixes them in advance and records every later change, so readers can tell confirmatory results from exploratory ones.

The situation

What should be agreed before collecting results to reduce cherry-picking and post-hoc changes to success criteria?

What you leave with

A versioned, timestamped plan — research question, primary outcomes, exclusions, analysis and decision rules — and an amendment log recording every change with its reason and whether results had been seen.

For: Research sponsor; evaluator

What an unfixed choice can do

In my self-audit of an LLM evaluation, two equally defensible rules for merging repeated measurement campaigns changed four of eight rows in the results table and moved the study-wide headline by 7 percentage points. Neither rule was wrong. But if the rule were chosen after seeing which one produced the more interesting headline, the 7 points would be an artefact of the choice. The same audit found four of eight endpoints withdrawn within ten weeks — exactly the kind of event a plan should decide how to handle before it happens.

Two other preprints show the practice in use and the reason for it:

  • Principle-Bench (workshop poster) authored its paraphrase, adversarial and boundary perturbations under a pre-registered rubric, so the perturbations could not be tuned towards whichever judge was being tested.
  • The bounded-autonomy harness reports small pilots in which configuration variance was large enough to mask tuning effects, and states plainly that they do not establish reliable adapter effects. Declaring a pilot diagnostic in advance keeps it from being presented as confirmatory later.

Why this matters more when someone is paying

A sponsor rarely asks for a particular result, but both sides know which result is preferred. A plan agreed before results exist removes the need to negotiate success criteria afterwards, when the pressure is greatest. It works alongside the publication and independence terms in publication, IP and independence, and the uncertainty design in uncertainty and repeated trials.

Assumptions

The plan assumes enough is known from a pilot to set sensible item counts and thresholds, that both parties can access the timestamped version, and that confirmatory data has not been examined before the freeze. Where that last assumption fails — results already partly seen — the plan says so and the claim is weakened accordingly.

Illustrative example — not a client engagement. A sponsor wants to know whether a new retrieval setting reduces unsupported answers. The plan fixes unsupported-answer rate on 500 sealed questions as the primary outcome, scored by a frozen rubric, with a three-point minimum improvement. Mid-study, one model endpoint is retired; the plan’s missing-data rule says the comparison continues on the remaining systems and the change is logged. The primary outcome improves by 1.8 points. A secondary metric improves by six. The report leads with the 1.8 points and says the pre-specified threshold was not met.

What this can and cannot establish

It can establish that reported analyses were the planned ones, which changes were made later and why, and which findings are confirmatory. It cannot validate the measurement itself, prove nothing was examined before the freeze without third-party timestamping, or rescue a study whose question was poorly chosen.

Useful if

  • A sponsor and an evaluator need to agree what counts as success before the sponsor's preferred answer is known.
  • An evaluation has several plausible metrics, merge rules or item filters, and the headline depends on which is chosen.
  • You want a report whose readers can see which analyses were planned and which were added later.

Not the right fit if

  • The work is genuinely exploratory and no conclusion will be acted on. Label it exploratory instead; pre-specifying a fishing trip adds paperwork, not rigour.
  • You need the plan registered with a formal external registry. I can draft a plan suitable for one, but the default here is a timestamped, shared, versioned document.
  • The goal is to guarantee a favourable result. A pre-specified plan makes an unfavourable one harder to hide.

What this produces

  • Versioned research question. The question, the decision it informs and the systems compared, with a version number, date and content hash.
  • Primary and secondary outcomes. Exact metric definitions and computation, with one or few primary outcomes; everything else labelled secondary or exploratory.
  • Inclusion, exclusion and missing-data rules. Which items, runs and systems count, what happens to failed calls, timeouts and withdrawn endpoints, and when the study stops.
  • Analysis and decision rules. Comparisons, uncertainty method, merge rule for repeated runs, the minimum meaningful difference and the rule that turns results into a conclusion.
  • Amendment log and deviations section. Every change after freezing, dated, with its reason and whether results had been seen, and a report section showing pre-specified and amended analyses side by side.

What you need before starting

  • A decision owner on each side. Someone for the sponsor and someone for the evaluation who can both sign off the plan.
  • Enough pilot information. A pilot or prior run to know the item counts, failure rates and variance the plan must accommodate. Pilot data is labelled and excluded from the confirmatory analysis.
  • A place to timestamp. A version-controlled repository, a shared document with history, or a registry — somewhere both parties can later verify what was frozen and when.

Protocol

  1. 1
    Write the question and the decision. One paragraph: what is being compared, under what conditions, and which outcome would change what decision. Version it as v1.0.
  2. 2
    Define primary outcomes exactly. Give the metric formula, unit of analysis and aggregation. If a judge model or rubric scores outputs, freeze its prompt or rubric version too.
  3. 3
    Fix inclusion, exclusion and missing-data rules. State which items and runs count, how failures and timeouts are scored, and what happens if a hosted endpoint changes or is withdrawn mid-study.
  4. 4
    Fix the analysis. Comparisons, uncertainty method and resampling unit, the merge rule for repeated campaigns, a named alternative to run as a sensitivity check, and the decision rule with its threshold.
  5. 5
    Freeze and share. Commit the plan, record the hash and date, and send it to the sponsor before confirmatory runs. Where possible, keep held-out results unscored until the freeze.
  6. 6
    Log amendments and report against the plan. Any change after the freeze gets a dated entry: what changed, why, and whether any results had been seen; nothing is edited silently. The report presents the pre-specified analysis first, amended and exploratory analyses after it, clearly labelled, with the amendment log attached.

Limits and unfavourable results

  • Pre-specification shows that the reported analysis was the planned one. It does not make the plan good: a poorly chosen metric stays poorly chosen.
  • A self-held timestamp is weaker than third-party registration. It deters quiet changes between collaborators but cannot prove nothing was run before the freeze.
  • Strict plans can be wrong about the world — an endpoint disappears, a metric turns out to be undefined. Amendments handle this, but every amendment weakens the confirmatory claim a little.
  • Exploratory findings are not worthless; they are hypotheses for the next pre-specified study, and are reported as such.

If the pre-specified primary outcome does not support the hypothesis, that is the headline, even if a secondary or post-hoc analysis looks better. The better-looking analysis is reported as exploratory.

Evidence behind this page

Questions

What is a pre-registration plan for an AI evaluation?

A document, fixed and timestamped before confirmatory results exist, that sets the research question, primary outcomes, inclusion and exclusion rules, analysis method and decision rule, with a log for any later amendment. It separates the analyses that test a claim from the ones that suggest new claims.

How do you prevent cherry-picking in model evaluation?

Fix the choices that can move the headline — metric, item filter, merge rule for repeated runs, judge prompt, decision threshold — before results are visible, run the main alternative as a declared sensitivity check, and report any change with its reason and whether results had been seen.

What should a sponsor and evaluator agree before collecting results?

The question and decision, the primary outcome, what happens to failed runs and withdrawn endpoints, the decision rule, and what will be published if the result is unfavourable. The last point belongs in the plan, not in a later negotiation.

Can the plan change after it is frozen?

Yes, through the amendment log. Each change is dated and justified, the report says whether results had been seen, and the original analysis is still shown where it can be computed.

See also

Prepare an evaluation brief

Send the question, the systems being compared and the outcome each party hopes for. A first draft of the plan is usually possible from a non-confidential description.

Last reviewed 2026-10-07.