Free resource · Evaluation brief
Write the brief before you meet the evaluator
Most evaluation conversations spend their first hour working out what the decision is. This builder takes you through the questions in the order that matters, keeps the draft in your browser, and produces a plain markdown brief you can keep, share internally or send.
The situation
How can we turn an uncertain technical claim into a structured experiment brief before meeting an evaluator?
What you leave with
An ungated markdown brief covering the claim, the decision it informs, the conditions that matter, the baseline, criteria stated in advance and the cost of being wrong.
For: Technical evaluator; research sponsor
Why the decision comes first
An evaluation is only useful if it can change something. Starting from the decision — what you will do if the claim holds, what you will do if it fails, and who decides — fixes what the test has to distinguish. It also exposes the case where no plausible result would change anything, in which case the evaluation is not worth commissioning.
Why conditions and baseline are separate fields
Published results are measured somewhere else. My RAG conflict-detection audit found that which detector is preferable depends on the input distribution, and The Correctness Illusion found that fixed-shape checks used by kernel benchmarks certify buggy kernels as correct. In both cases, how and where the test was run decided the answer. Writing your conditions down — and naming what the claim must beat — is how the brief keeps the test about your situation rather than the original author’s.
Why the criteria are written before any result
Results usually allow more than one reasonable reading. In How Reproducible Are Evaluation Conclusions?, a ranking of eight open model variants identified the worst model reliably but not the best. Had the criterion been chosen after seeing the table, either story could have been told. The builder asks for the difference you would care about and the rule for calling it while nobody yet knows which way it goes.
Why the cost of being wrong matters
The two errors are rarely equal. Adopting a model that turns out worse may be cheap to reverse; missing a regression in a regulated workflow may not be. The relative cost sets how strict the criterion should be, how many repeats are worth paying for, and whether a cheaper, smaller study is enough.
Illustrative example — not a client engagement. A team is offered a cheaper hosted model. Their brief states the decision (“switch if quality holds”), the conditions (their own 400 historical tickets, current latency limit), the baseline (the current model on the same tickets), the criterion (no more than a two-point drop in graded accuracy, judged on the bootstrap interval) and the error costs (a false switch means customer-facing regressions; a missed switch means one more quarter at the higher price). The asymmetry tells them to set the bar strictly and accept a slower decision.
Using the builder
Fill in the fields below in any order; the brief assembles as you type and is saved in this browser only. When you are done you can copy it, download it as markdown, or choose Send this brief, which opens the contact form with the brief pre-filled. Nothing is sent until you submit that form.
Build your brief
Nothing is sent until you choose to. Your draft is saved in this browser only.
Your brief
Useful if
- You have a claim — from a paper, a vendor, an internal pilot — and an upcoming decision that depends on it.
- You want your own team to agree what result would change the decision before anyone runs anything.
- You are approaching evaluators, including me, and want to compare them against the same brief.
Not the right fit if
- The decision is already made and you need evidence to support it. A brief written backwards from the answer is not an evaluation plan.
- You want a number today. The brief frames a test; it does not run one.
- You would need to paste confidential material to fill it in. Describe what exists instead.
What this produces
- A markdown brief. Copy it, download it as a .md file, or hand it to the contact form. It is yours either way and is not tied to working with me.
- A visible gap list. Sections you leave empty are omitted from the output, so what is missing from the brief is what your team has not decided yet.
What you need before starting
- The claim and its source. The sentence, table or slide you are being asked to rely on.
- The decision owner. The person who can say what result would change the decision. If nobody can, start there.
- A rough picture of your conditions. Inputs, data, hardware, cost or latency limits. Approximate is fine.
Protocol
- 1 State the decision. Write what you will do if the claim holds and if it fails, and who decides. Then rewrite the claim as one testable sentence under those terms.
- 2 Fix the conditions and the baseline. Name the workload the result must hold under and what it should beat: your current system, a simpler method or a published number.
- 3 State the criteria before any results exist. The difference you would care about and the rule for calling it. It is the easiest field to leave vague.
- 4 Price the errors. Write down what acting on a false positive costs, and what missing a true positive costs.
- 5 Add access, timing and interests. Describe the data, models and compute available, the deadline and publication position, and any interests connected to the outcome.
- 6 Review, then decide what to do with it. Read the generated brief with the decision owner. Keep it, circulate it, or send it.
Limits and unfavourable results
- A brief establishes nothing about the claim. It makes a fair test possible; it is not the test.
- Stating criteria in advance limits post-hoc reinterpretation but does not remove it. Any later change should still be logged with its reason.
- Drafts live in this browser's local storage. Clearing site data, using a private window or switching devices loses them, so download a copy if it matters.
Evidence behind this page
- How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
2026 · Preprint (not peer reviewed)
- An Input-Regime Audit of Conflict Detection for Retrieval-Augmented Generation
2026 · Manuscript under submission
- The Correctness Illusion in LLM-Generated GPU Kernels
2026 · Preprint (not peer reviewed) · cited by 6 (Google Scholar, October 2026)
Questions
Is anything sent when I fill in the builder?
No. The draft is stored only in your browser. Copy and download stay on your machine. Send this brief opens the contact form with the brief filled in, and nothing reaches me until you submit that form.
What makes a claim testable?
It names the system, the comparison, the conditions and a measurable outcome. "Model X is better" is not testable; "Model X answers our support queries as accurately as the current system at half the cost" is, once accuracy and the query set are defined.
Why state the criteria before seeing results?
Because many results can be read more than one defensible way. In my own self-audit of an LLM evaluation, a small-sample ranking identified the worst model reliably but not the best. A criterion chosen after seeing such a table can make either reading look like the finding; one written first cannot.
Can I use the brief with a different evaluator?
Yes. It is plain markdown with no gate, licence or sign-up. A good evaluator of any kind should be able to work from it.
See also
Method
Pre-Specified Evaluation Plans
Versioned research question, primary outcomes, exclusion and analysis rules, and an amendment log, fixed before results.
Sponsored research
Write a Sponsored AI Research Brief
Structure a sponsored AI study around a claim, a decision, a baseline, realistic access and stop criteria.
Commissioned study
AI Benchmark Replication Studies
Re-run a published benchmark or paper result under your workload, with raw results and a deviation report.
Free resource
Reproducibility Readiness Checklist
Check which missing artefacts would stop another team rerunning your study, and decide what to fix before release.
Prepare an evaluation brief
Prefer to talk it through first? Send whatever you have. A half-finished brief is a perfectly good starting point.
Last reviewed 2026-10-07.