Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Free resource · Reproducibility

Could another team rerun this?

A practical worksheet for the week before you release a paper, benchmark or evaluation. Tick what already exists, see which gaps would block an independent rerun, and decide whether to fix them, declare them or hold the release.

The situation

Can another team rerun our study, and which missing artefacts would block that effort?

What you leave with

A checklist summary separating the essential gaps that block an independent rerun from improvements, ready to annotate with evidence links and turn into release decisions.

For: Research author; engineering evaluator

Checklist or method?

This page and reproducibility packages answer different questions. The method page is about how to build a package that someone else can rerun, with a worked example. This page is about whether a package you already have is ready: a short worksheet, run against a release candidate, that tells you which gaps would stop a stranger. If most essentials are unticked, go to the method page first.

What “essential” means here

An item is marked essential when, without it, an outside team would have to guess something that changes the result: which commit produced the table, which seeds, which data split, which hardware, which model version on which date. The other items make a rerun faster or more convincing, but a careful team could get by without them.

Some of this is learned the hard way. In How Reproducible Are Evaluation Conclusions? I re-audited one of my own evaluations from 293 persisted raw outputs; without those raw outputs, the bootstrap that showed the ranking identified the worst model reliably but not the best could not have been run. Keeping raw per-run outputs, not only means, is essential for that reason.

Applied to my own artefacts

The checklist is easier to trust if it is applied to its author’s releases, gaps included.

ArtefactWhat is in placeKnown gap or declared limit
gpuemu-corpusDrivers per paper, stored seeds, replay scriptsNo tagged release; a citation must name a commit
ebrag-vecdb-2026-paperOne JSON artefact per experiment; runnable Python sliceEach experiment has a script, but reruns call a hosted LLM endpoint, so regenerated outputs can drift; the stored JSON is the reference
on-device-auction-auditResults, calibration data, analysis scriptsSimulator source withheld (hashes recorded); data licensed CC BY-NC 4.0

The third row is the pattern for a partial release: the withheld part is named, a means of verifying it is given, and everything downstream of it is released. A checklist that only rewarded complete releases would push authors to say nothing about what they left out.

Using the checklist

Tick items below as you confirm them; the verdict and missing list update as you go, and your ticks are saved in this browser only. Copy the summary into your release ticket and add, against each line, the link that shows the item exists — or the decision you have made about it.

Checklist

Tick what already exists. Items marked essential block an independent rerun. Saved in this browser only.

Claim and version
Environment
Inputs and data
Scripts and runs
Known limits

    Useful if

    • You are about to release a paper, benchmark or evaluation and want to know what a stranger would trip over.
    • You are reviewing someone else's artefact and need a consistent list of what to look for.
    • You need a quick, shareable record of what is missing for a release ticket or a co-author.

    Not the right fit if

    • You need the full method and a worked example package. That is the reproducibility packages method page.
    • You want assurance that the results are correct. A rerunnable study can still be wrong.
    • Your study cannot release data or code at all and you need a different verification route. The checklist will only tell you that most essentials are missing.

    What this produces

    • A verdict. Whether every essential item is in place, or how many are missing.
    • A missing-items list. Every unticked item, with essentials marked, which you can copy into a ticket or print.
    • A release decision. Made by you from that list: fix now, release with the gap declared, or withhold with a way to verify what was withheld.

    What you need before starting

    • One result to check. A specific table or figure, not the whole paper. Run the checklist once per headline result.
    • The release candidate. The repository, dataset card or supplementary material as it would be published.
    • Someone who did not run the study. Ideally the person ticking the boxes, so that "it's documented" means documented to an outsider.

    Protocol

    1. 1
      Pick the result. Choose one headline table or figure and the claim it supports.
    2. 2
      Tick only what you can point to. An item counts only if you can link the file, commit or section that satisfies it. Paste those links next to each line of the copied summary.
    3. 3
      Treat essentials as blockers. An unticked essential item means a stranger cannot rerun that result. Improvements make a rerun easier but do not block it.
    4. 4
      Decide per gap. Fix it, release with the gap stated in the README, or — for code or data that cannot be released — withhold it with hashes or another means of verification.
    5. 5
      Test with a real rerun. Have someone outside the project regenerate the table on a clean machine from the instructions alone. Only that establishes the result is rerunnable.

    Limits and unfavourable results

    • Ticking every box is not a rerun. The checklist says the pieces appear to be present; only an independent regeneration shows they work together.
    • Reproducibility is not correctness. A study can be fully rerunnable and still measure the wrong thing.
    • Results that depend on hosted models or APIs may not be rerunnable later however complete the package is. The checklist asks for the measurement date for this reason.
    • State is saved in this browser only. Copy or print the summary to keep it.

    Evidence behind this page

    • gpuemu-corpus (GitHub) — Drivers for four papers, stored seeds and replay scripts. No tagged releases, so cite a commit. MIT or Apache-2.0.
    • ebrag-vecdb-2026-paper (GitHub) — Paper source, one JSON artefact per experiment and a runnable Python slice. MIT.
    • on-device-auction-audit (GitHub) — Results, calibration data and analysis scripts; the simulator source is withheld with hashes recorded. Code MIT, data CC BY-NC 4.0.

    Questions

    How is this different from the reproducibility packages method?

    The method page explains how to build a reproducibility package and walks through a complete example. This page is the short worksheet you run against a release candidate to find what is missing. Use the method to build; use the checklist to check.

    Which missing artefacts block a rerun?

    The items marked essential: the exact claim and version, pinned environment, hardware list, data access and licence, scripted splits, stored seeds, a command that regenerates each table, raw per-run outputs, logged deviations and stated limits. Without any one of them, a stranger has to guess.

    Can I release if some code or data has to be withheld?

    Yes, if you say so and give a way to verify what was withheld. My on-device auction study withholds its simulator source but records its hashes and releases the results, calibration data and analysis scripts, so the analysis can be recomputed even though the simulation cannot be rerun from source.

    Does ticking every essential item mean my study is reproducible?

    It means it is ready to be tested. Whether it reproduces is shown by someone else regenerating the result from your package on a clean machine.

    See also

    Prepare an evaluation brief

    If the checklist shows the gaps and you want an outside rerun of the result, send a short brief.

    Last reviewed 2026-10-07.