Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Commissioned study · Replication

Reproduce the result before you rely on it

A replication study re-runs a published AI result — a benchmark score, a paper's headline table, a vendor's evaluation — under the conditions your decision actually depends on, and reports what held, what moved and why.

The situation

Our decision depends on a published benchmark. Can the result be reproduced under the relevant workload and constraints?

What you leave with

Replication protocol, environment lockfile, raw per-run results and a deviation report that separates what reproduced from what did not.

For: R&D lead; technical evaluator

Why replicate before deciding

Published AI results are usually measured once, on one setting, at one point in time. In my own audit of an LLM evaluation, a ranking of eight models identified the worst model reliably but the best model only about two-thirds of the time under a bootstrap over prompts, and two equally defensible ways of merging repeated runs changed four of the eight rows (How Reproducible Are Evaluation Conclusions?). In GPU-kernel benchmarks, a fixed-shape correctness check certified every seeded buggy kernel in a 26-op corpus as correct (The Correctness Illusion).

Neither finding means the original authors were careless. It means a single reported number carries less information than a decision usually needs. A replication tells you how much.

What a replication answers

  1. Does the result reproduce as published? Same setting, same method, re-run.
  2. Does it survive your conditions? Your data, hardware, model version and constraints.
  3. Is the difference meaningful? Compared against the run-to-run uncertainty, not against zero.

Illustrative example — not a client engagement. A team choosing between two retrieval re-rankers relies on a published leaderboard gap of three points. The protocol fixes their query set, five repeats per system and a rule that the gap must exceed the 95% interval of the per-query bootstrap. The faithful re-run reproduces the leaderboard gap; on their queries the gap shrinks inside the interval. The decision memo says the leaderboard does not distinguish the two for this workload, and that latency, not accuracy, should drive the choice.

Commission this if

  • A purchase, architecture or research decision rests on a number someone else measured.
  • The published setting differs from yours — different data, hardware, model version or prompt format — and you need to know whether that matters.
  • You can describe the decision and accept a result that goes either way.

Not the right fit if

  • You need the result confirmed for marketing or a sales deck. A replication that can only succeed is not a replication.
  • The original work has no method description, code or data at all; that is a feasibility or benchmark-design question instead.
  • You want a maintained evaluation pipeline in production. That is delivery work, not a study.

What you receive

  • Replication protocol. The claim being tested, the original setting, the setting you care about, the baselines and the reporting rules, written and agreed before any run.
  • Pinned environment. Lockfile or container image, hardware list, model and API versions, and the measurement date.
  • Raw results. Every run's outputs, not only means, so a third party can recompute the tables.
  • Deviation report. Each difference from the original result, whether it is within run-to-run uncertainty, and which change in conditions plausibly explains it.
  • Decision memo. One page stating what the evidence supports for your decision and what it does not.

What the study needs from you

  • The original claim. The paper, report or benchmark page, and the specific number or comparison you depend on.
  • Your conditions. A description of the workload, data distribution and constraints the result must hold under. Sample data can be shared later under an agreed process.
  • Access. Model endpoints or weights, compute (or a budget for it), and any code the original authors released.
  • A named decision owner. Someone who can say what result would change the decision.

How the work runs

  1. 1
    Scope call. Agree the claim, the decision, conflicts of interest and whether the study is viable with the available artefacts.
  2. 2
    Pre-specified protocol. Fix the comparisons, metrics, number of repeats and the rule for calling a result reproduced — before results exist. Amendments are logged with reasons.
  3. 3
    Faithful re-run. Reproduce the original setting first, as published. If that fails, the gap is reported before any adaptation.
  4. 4
    Your setting. Re-run under your workload and constraints, holding everything else fixed so differences can be attributed.
  5. 5
    Report and review. Deliver the package; walk through it with your team; agree what, if anything, is published.

Limits and unfavourable results

  • A replication tests a result under specified conditions. It does not certify that a system is correct in general.
  • If the original authors withheld code, data or weights, the study can only test what is described; the report says which parts were reconstructed.
  • Hosted model endpoints change and disappear. Results carry a measurement date and may not be re-runnable later.

A failed replication is a valid deliverable. The fee is fixed at scoping and does not change with the outcome, and the report is written the same way whether the result held or not.

Engagement terms

Model
Fixed-scope study with an agreed protocol and acceptance criteria for the deliverables (not for the result).
Who does the work
Dipankar Sarkar personally designs and runs the study. Any specialist help is disclosed and agreed in advance.
Commercial basis
Fixed research scope, quoted after a scoping conversation. The fee does not depend on the result.
Availability
Checked per enquiry.

Conflicts are checked before scoping. See publication, IP and independence.

Evidence behind this page

Questions

What counts as a successful replication?

Whatever rule was agreed in the protocol before any results existed — typically that the replicated value falls within the run-to-run uncertainty of the original, or that the ordering of systems is preserved. Changing the rule after seeing results is logged as an amendment and reported.

Can you replicate a vendor's benchmark claim?

Yes, if the claim is specific enough to test and there is enough description to reconstruct it. If I have built, advised or invested in the vendor or a competitor, that is disclosed before scoping and may mean declining the work.

Will the result be published?

Only if agreed. Publication terms, review windows and what happens to an unfavourable result are settled before the study starts; see publication, IP and independence.

How is this different from benchmark design?

Replication tests an existing claim. Benchmark design creates a new test for a decision that existing benchmarks do not measure. Many studies start as one and need a little of the other.

See also

Commission a scoped study

Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.

Last reviewed 2026-10-07.