Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Research direction · Efficient inference

Judging what a smaller model actually gives up

Distillation, quantisation and other compression methods promise most of the quality at a fraction of the cost. Whether that holds depends on which tasks were measured, how cost was counted and whether the comparison was matched. This page sets out an evaluation approach and the open questions. I have not published results in this area.

The situation

When a model is compressed or distilled, which quality and resource measurements are needed to judge the trade-off?

What you leave with

Matched evaluation protocol, resource accounting and task-specific degradation analysis, designed before any numbers are produced.

For: ML systems researcher; R&D sponsor

Status: an open research direction

I have not published work on distillation or model compression. This page exists because the evaluation problems are close to ones I have studied, and because I would like to study them properly. Everything below is either general background, stated cautiously, or a proposed approach. It is not a result.

Background

Distillation trains a smaller model to imitate a larger one; quantisation and pruning reduce the cost of an existing model. All three trade quality for cost, and the published evidence for a given trade-off often comes from aggregate benchmark scores measured under conditions that differ between the two models. Resource savings are also conditional: a speed-up at one batch size and sequence length may shrink or vanish at another.

Adjacent evidence from my own work

Two of my preprints bear on how such a comparison should be run, though neither studies compression:

  • How Reproducible Are Evaluation Conclusions? found that a small-sample ranking of eight models identified the worst model reliably but not the best, and that two defensible ways of merging repeated runs changed four of eight rows. A claim that a smaller model “matches” a larger one is a ranking claim and needs the same scrutiny.
  • Operator-Aware Mixed-Precision Tolerance Calibration found that hand-picked numerical tolerances were far looser than those derived from the kernels’ own error, with calibration raising bug-detection recall from 73.2% to 82.4%. Lower-precision variants of a model raise a similar question: what deviation should count as acceptable, and who chose the threshold?

Open questions

  1. Where does degradation concentrate? Aggregate scores can hide large per-task losses. Which task properties predict them is, to my knowledge, not well characterised for most deployment settings.
  2. What is the right unit of cost? Per-token latency, cost per correct answer and energy per task can rank the same pair of models differently.
  3. How stable is the trade-off? If repeated runs of the original model already vary, part of an apparent degradation may be noise. The comparison has to be made against that spread.
  4. Does the cost of producing the smaller model change the answer? For low-volume uses it may.

Illustrative example — not a client engagement. A research group distils a model for classification and summarisation and reports a two-point average drop. A matched protocol with five repeats per model shows the classification tasks within run-to-run spread but a much larger drop on long-document summarisation, and a latency gain that halves at the longest sequence lengths. The trade-off statement supports the smaller model for classification only.

Useful if

  • You have a compressed or distilled model and an aggregate score that looks close to the original, and you need to know where it is not.
  • You are planning a compression experiment and want the comparison designed before results exist.
  • You are interested in this as an open research question and accept that the outcome may be negative.

Not the right fit if

  • You want an existing published result on distillation from me. There is none yet.
  • You need a compressed model trained, deployed or served. That is delivery work.
  • You need a vendor's efficiency claim endorsed rather than tested.

What this produces

  • Matched protocol. Same prompts, decoding settings, seeds, hardware and software stack for the original and compressed model, fixed in writing before runs.
  • Resource accounting. Latency at stated percentiles, throughput, peak memory and, where measurable, energy, each at stated batch size and sequence length, plus the one-off cost of producing the smaller model.
  • Task-specific degradation analysis. Per-task and per-slice quality deltas with uncertainty from repeated runs, so a small average loss cannot hide a large loss on one task.
  • Trade-off statement. Which tasks and operating points the evidence supports using the smaller model for, and which it does not.

What you need before starting

  • Both models. The reference model and the compressed or distilled variant, with access sufficient to control decoding and hardware.
  • The decision tasks. The tasks the smaller model would actually serve, not only a general benchmark suite.
  • Measurement hardware. The hardware class the decision concerns, since resource results do not transfer cleanly between classes.

Protocol

  1. 1
    Write the comparison down first. Tasks, metrics, operating points, number of repeats and what difference counts as material, before running anything.
  2. 2
    Hold everything else fixed. Change only the model. Any unavoidable difference, such as a precision change, is recorded and reported.
  3. 3
    Measure quality per task with repeats. Report run-to-run spread and per-task deltas, not only the aggregate.
  4. 4
    Measure resources at the operating point. Warm-up excluded, percentiles reported, configuration recorded alongside every number.
  5. 5
    State the trade-off with its limits. Including slices where the evidence is too thin to conclude either way.

Limits and unfavourable results

  • This is a research direction. No paper of mine reports distillation or compression results, and none is cited here as if it did.
  • Resource measurements depend on hardware, batch size, sequence length and software versions, and may not transfer to another stack.
  • A matched comparison can show where a smaller model degrades on the measured tasks; it cannot show it is safe to substitute on unmeasured ones.

Evidence behind this page

Questions

How should a distilled model be evaluated against the original?

With a matched protocol (same prompts, decoding, seeds and hardware), per-task quality deltas with run-to-run uncertainty, and resource measurements at the operating point that matters. An aggregate benchmark average alone is not enough to judge the trade-off.

Which resource measurements matter for efficient inference?

Latency at stated percentiles, throughput, peak memory and, where it can be measured, energy, each reported with batch size, sequence length, hardware and software versions. The cost of producing the smaller model belongs in the accounting too.

Have you published results on distillation?

No. This page describes an evaluation approach and open questions. The cited papers are about evaluation method — ranking stability and numerical tolerances — not about distillation.

What makes small-model benchmarks misleading?

Commonly: unmatched settings between the models compared, single runs, and averages that mix tasks where the smaller model is fine with tasks where it degrades sharply. These are general cautions, not findings of mine about specific models.

See also

Commission a scoped study

Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.

Last reviewed 2026-10-07.