Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Free resource · Evaluation brief

Write the brief before you meet the evaluator

Most evaluation conversations spend their first hour working out what the decision is. This builder takes you through the questions in the order that matters, keeps the draft in your browser, and produces a plain markdown brief you can keep, share internally or send.

The situation

How can we turn an uncertain technical claim into a structured experiment brief before meeting an evaluator?

What you leave with

An ungated markdown brief covering the claim, the decision it informs, the conditions that matter, the baseline, criteria stated in advance and the cost of being wrong.

For: Technical evaluator; research sponsor

Why the decision comes first

An evaluation is only useful if it can change something. Starting from the decision — what you will do if the claim holds, what you will do if it fails, and who decides — fixes what the test has to distinguish. It also exposes the case where no plausible result would change anything, in which case the evaluation is not worth commissioning.

Why conditions and baseline are separate fields

Published results are measured somewhere else. My RAG conflict-detection audit found that which detector is preferable depends on the input distribution, and The Correctness Illusion found that fixed-shape checks used by kernel benchmarks certify buggy kernels as correct. In both cases, how and where the test was run decided the answer. Writing your conditions down — and naming what the claim must beat — is how the brief keeps the test about your situation rather than the original author’s.

Why the criteria are written before any result

Results usually allow more than one reasonable reading. In How Reproducible Are Evaluation Conclusions?, a ranking of eight open model variants identified the worst model reliably but not the best. Had the criterion been chosen after seeing the table, either story could have been told. The builder asks for the difference you would care about and the rule for calling it while nobody yet knows which way it goes.

Why the cost of being wrong matters

The two errors are rarely equal. Adopting a model that turns out worse may be cheap to reverse; missing a regression in a regulated workflow may not be. The relative cost sets how strict the criterion should be, how many repeats are worth paying for, and whether a cheaper, smaller study is enough.

Illustrative example — not a client engagement. A team is offered a cheaper hosted model. Their brief states the decision (“switch if quality holds”), the conditions (their own 400 historical tickets, current latency limit), the baseline (the current model on the same tickets), the criterion (no more than a two-point drop in graded accuracy, judged on the bootstrap interval) and the error costs (a false switch means customer-facing regressions; a missed switch means one more quarter at the higher price). The asymmetry tells them to set the bar strictly and accept a slower decision.

Using the builder

Fill in the fields below in any order; the brief assembles as you type and is saved in this browser only. When you are done you can copy it, download it as markdown, or choose Send this brief, which opens the contact form with the brief pre-filled. Nothing is sent until you submit that form.

Build your brief

Nothing is sent until you choose to. Your draft is saved in this browser only.

One sentence. "Model X answers our support queries as accurately as the current system at half the cost."

Paper, vendor benchmark, internal pilot, a colleague's result.

What will you do differently if the claim holds, and if it fails? Who decides?

Your inputs, data distribution, hardware, latency or cost limits. A result outside these conditions is not useful to you.

What the claim should be compared against: current system, a simpler method, a published number.

State it before any results exist. Include the size of difference you would care about.

What happens if you act on a false positive? On a false negative?

Data, code, model endpoints, compute. Describe it; do not paste confidential material.

Deadline, budget range if known, and whether results may be published.

Vendors, investments or relationships connected to the outcome.

Your brief

 

Useful if

  • You have a claim — from a paper, a vendor, an internal pilot — and an upcoming decision that depends on it.
  • You want your own team to agree what result would change the decision before anyone runs anything.
  • You are approaching evaluators, including me, and want to compare them against the same brief.

Not the right fit if

  • The decision is already made and you need evidence to support it. A brief written backwards from the answer is not an evaluation plan.
  • You want a number today. The brief frames a test; it does not run one.
  • You would need to paste confidential material to fill it in. Describe what exists instead.

What this produces

  • A markdown brief. Copy it, download it as a .md file, or hand it to the contact form. It is yours either way and is not tied to working with me.
  • A visible gap list. Sections you leave empty are omitted from the output, so what is missing from the brief is what your team has not decided yet.

What you need before starting

  • The claim and its source. The sentence, table or slide you are being asked to rely on.
  • The decision owner. The person who can say what result would change the decision. If nobody can, start there.
  • A rough picture of your conditions. Inputs, data, hardware, cost or latency limits. Approximate is fine.

Protocol

  1. 1
    State the decision. Write what you will do if the claim holds and if it fails, and who decides. Then rewrite the claim as one testable sentence under those terms.
  2. 2
    Fix the conditions and the baseline. Name the workload the result must hold under and what it should beat: your current system, a simpler method or a published number.
  3. 3
    State the criteria before any results exist. The difference you would care about and the rule for calling it. It is the easiest field to leave vague.
  4. 4
    Price the errors. Write down what acting on a false positive costs, and what missing a true positive costs.
  5. 5
    Add access, timing and interests. Describe the data, models and compute available, the deadline and publication position, and any interests connected to the outcome.
  6. 6
    Review, then decide what to do with it. Read the generated brief with the decision owner. Keep it, circulate it, or send it.

Limits and unfavourable results

  • A brief establishes nothing about the claim. It makes a fair test possible; it is not the test.
  • Stating criteria in advance limits post-hoc reinterpretation but does not remove it. Any later change should still be logged with its reason.
  • Drafts live in this browser's local storage. Clearing site data, using a private window or switching devices loses them, so download a copy if it matters.

Evidence behind this page

Questions

Is anything sent when I fill in the builder?

No. The draft is stored only in your browser. Copy and download stay on your machine. Send this brief opens the contact form with the brief filled in, and nothing reaches me until you submit that form.

What makes a claim testable?

It names the system, the comparison, the conditions and a measurable outcome. "Model X is better" is not testable; "Model X answers our support queries as accurately as the current system at half the cost" is, once accuracy and the query set are defined.

Why state the criteria before seeing results?

Because many results can be read more than one defensible way. In my own self-audit of an LLM evaluation, a small-sample ranking identified the worst model reliably but not the best. A criterion chosen after seeing such a table can make either reading look like the finding; one written first cannot.

Can I use the brief with a different evaluator?

Yes. It is plain markdown with no gate, licence or sign-up. A good evaluator of any kind should be able to work from it.

See also

Prepare an evaluation brief

Prefer to talk it through first? Send whatever you have. A half-finished brief is a perfectly good starting point.

Last reviewed 2026-10-07.