Commissioned study · Benchmark design
Design the benchmark around the decision, not the system
Most benchmarks answer someone else's question. A benchmark design study starts from the decision you have to make and builds the test that would change it — tasks, baselines, sampling, metrics and the conditions under which a system counts as failing — before any candidate is run.
The situation
How should we design a benchmark that measures the decision we actually need to make rather than flattering one system?
What you leave with
Pre-specified tasks, baselines, sampling plan, metrics and failure conditions, plus a pilot run showing whether the benchmark can separate the candidates at all.
For: Research manager; platform team
Start from the decision
The usual failure is not a bad metric. It is a benchmark chosen first, with the decision fitted to it afterwards. A team adopts a public leaderboard because it exists, discovers their preferred system ranks well, and the benchmark quietly becomes the justification. A benchmark design study reverses the order: decision, then the difference that matters, then the tasks and metrics that could detect it, then the candidates.
What my own benchmark work taught me
- One number hides the failure that matters. Principle-Bench maps 168 scenarios to two UK FCA principles with paraphrase, adversarial and boundary perturbations, and scores LLM judges on four axes — accuracy, paraphrase robustness, adversarial robustness and calibration (Four-Axis Trustworthiness Benchmark, KDD 2026 workshop poster). Separate axes exist because an accuracy score alone cannot show whether a judge is fragile or miscalibrated.
- Failure conditions should not be averaged away. RegLLM tracks six indicators for regulated agents, including citation validity, escalation correctness and unsafe-action rate, reported separately (Evaluating Bounded Autonomy, preprint).
- Report misses and false alarms together. Of seven test-input strategies on the 26-op gpuemu corpus, boundary-only shape sampling reached 78% bug recall with 0% control false positives; adversarial values reached 99% recall but flagged 94% of correct controls (Test-Input Generation for Tensor Programs, preprint). A benchmark that reported recall alone would pick the noisy strategy.
- Small samples resolve less than they appear to. A small-sample ranking of eight open model variants identified the worst model reliably but not the best (How Reproducible Are Evaluation Conclusions?, preprint). The pilot exists to find this out before the full run.
Illustrative example — not a client engagement. A platform team must decide whether to replace a scripted triage flow with an agent. Public agent benchmarks score general tool use. The designed benchmark samples historical tickets stratified by category, keeps a private held-out portion, uses the scripted flow as baseline, and treats any unauthorised account change as a failure regardless of resolution rate. The pilot shows the two systems differ by less than run-to-run noise on common categories, so the protocol is revised to oversample the rare categories where the decision is actually made.
The method pages on evaluation dataset design and benchmark leakage and contamination cover the task-set side in more depth.
Commission this if
- Public benchmarks measure something adjacent to your decision: a general task where yours is narrow, clean inputs where yours are messy.
- You suspect a benchmark was chosen, or tuned, because one system does well on it.
- You need a test your team can re-run later and that a sceptical reviewer would accept.
Not the right fit if
- You want an evaluation service wired into CI, dashboards and release gates and maintained over time. That is production evaluation engineering.
- You need a leaderboard for a launch announcement.
- The tasks cannot be described, sampled or synthesised without data that cannot be shared in any form.
What you receive
- Decision specification. The decision, the alternatives, and the smallest difference that would change it. Everything else in the benchmark is derived from this.
- Task set and sampling plan. Where tasks come from, how they are stratified, the held-out portion, and the contamination checks run against public training data.
- Baselines. The incumbent and a deliberately simple baseline, so the benchmark is able to say 'no improvement'.
- Metrics and failure conditions. A primary metric, the secondary metrics always reported beside it, and conditions that count as failure regardless of the average score.
- Pilot report and frozen protocol. A small run showing whether the benchmark separates known-good from known-bad cases, then the versioned task set, scoring scripts and protocol.
What the study needs from you
- A decision owner. Someone who can say what result would change the decision and what failure is unacceptable.
- Real or representative inputs. Examples of the inputs the system will face, or a description detailed enough to construct them. Sensitive data can be handled under an agreed process.
- Candidate access. Endpoints, weights or binaries for the systems the benchmark must eventually compare, if a pilot or full run is in scope.
- Compute budget. For the pilot and, if commissioned, the full run.
How the work runs
- 1 Scope call. Agree the decision and conflicts. Sometimes the right answer is that an existing benchmark is adequate and a replication is cheaper; I will say so.
- 2 Decision to metrics. Translate the decision into what must be measured, at what resolution, and which failures are disqualifying.
- 3 Tasks with built-in controls. Include known-correct and known-faulty items so the benchmark's own miss rate and false-alarm rate are measured, not assumed.
- 4 Pilot and sample-size check. Run a small pilot to estimate variance and confirm the benchmark can resolve the difference that matters with an affordable number of items and repeats.
- 5 Freeze and hand over. Version the protocol, items and scripts. A full run on the candidates can follow as part of the same study or as a comparison study.
Limits and unfavourable results
- A benchmark measures the tasks in it. Claims beyond the sampled distribution are stated as untested, not implied.
- Releasing items publicly invites contamination. A private held-out portion is recommended and its handling agreed.
- Human or LLM-judge labels carry their own error. Where a judge is used, the judge is evaluated too and its error reported.
The pilot can show that the benchmark cannot separate the candidates at an affordable sample size, or that the metric you wanted does not track the decision. That is reported as the finding. The fee is fixed at scoping and does not depend on it.
Engagement terms
- Model
- Fixed-scope study with an agreed protocol and acceptance criteria for the deliverables (not for which system wins).
- Who does the work
- Dipankar Sarkar personally designs and runs the study. Any specialist help is disclosed and agreed in advance.
- Commercial basis
- Fixed research scope, quoted after a scoping conversation. The fee does not depend on the result.
- Availability
- Checked per enquiry.
Conflicts are checked before scoping. See publication, IP and independence.
Evidence behind this page
- A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
2026 · Peer-reviewed workshop paper
- Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
2026 · Preprint (not peer reviewed)
- Test-Input Generation for Tensor Programs: What Actually Finds Kernel Bugs
2026 · Preprint (not peer reviewed)
- How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
2026 · Preprint (not peer reviewed)
- gpuemu-corpus (GitHub) — A benchmark with controls built in: 16 correct ops alongside 10 seeded buggy variants, so detection and false-positive rates are both measurable. MIT or Apache-2.0.
- gpuemu-corpus (Hugging Face) — The same 26-op corpus as a dataset with fp64 references. MIT.
Questions
How do you stop a benchmark from flattering one system?
Design from the decision rather than from a candidate; include the incumbent and a simple baseline; build in known-good and known-faulty controls; fix the metrics and failure conditions before any candidate runs; and report every slice, not the ones a system wins. Conflicts with any vendor in scope are disclosed before scoping.
What is the difference between a benchmark design study and an evaluation pipeline?
A study produces a defensible experimental protocol and, optionally, one run of it. A pipeline runs a test continuously and is maintained as software. The protocol from a study can become the specification for a pipeline, but building and operating that pipeline is delivery work.
Can you design a benchmark for LLM-as-judge or agent evaluation?
Yes. The judge or harness has to be evaluated as well as the systems it scores. Principle-Bench, for instance, evaluates LLM judges on accuracy, paraphrase robustness, adversarial robustness and calibration rather than accuracy alone.
Who owns the benchmark afterwards?
Data rights, ownership of the task set and publication terms are agreed before the study starts. See publication, IP and independence for the defaults and what can be negotiated.
See also
Commissioned study
AI Benchmark Replication Studies
Re-run a published benchmark or paper result under your workload, with raw results and a deviation report.
Commissioned study
Independent Model and Retrieval Comparison Studies
Run shortlisted models or retrieval methods on one workload; report uncertainty, input-regime effects and costs.
Method
Evaluation Dataset Design and Provenance
Dataset specification, sampling design with controls, per-item provenance ledger, exclusion log and access restrictions.
Method
Benchmark Leakage and Contamination Checks
Overlap audit, written contamination hypotheses and a sensitivity analysis that tests whether the conclusion survives.
Method
Pre-Specified Evaluation Plans
Versioned research question, primary outcomes, exclusion and analysis rules, and an amendment log, fixed before results.
Commission a scoped study
Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.
Last reviewed 2026-10-07.