Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Commissioned study · Feasibility

Find out whether it can work before you build it

A feasibility study turns 'could this approach work for us?' into the smallest experiment that could answer no — with the success threshold, stop criteria and budget fixed before it runs — and ends in an evidence-led recommendation to proceed, change course or stop.

The situation

We need to know whether an approach is technically viable before committing to implementation. What experiment would answer that?

What you leave with

Bounded feasibility protocol, stop criteria agreed in advance and an evidence-led proceed-or-stop report that names what remains unknown.

For: R&D lead; product research sponsor

A demo is not a feasibility result

Most “is this feasible?” questions arrive with a demo attached. A demo shows that an approach can work on the inputs someone chose. Feasibility asks whether it works on the inputs you will actually face, within the constraints you actually have, better than the alternative you already have. The difference is usually where the risk lives.

Designing the smallest experiment that could say no

  1. Name the decision. “Build the on-device version” or “replace the rules engine” — something concrete that changes with the answer.
  2. List the load-bearing assumptions. For an approach to be viable, several things must all hold: accuracy on hard cases, latency within budget, behaviour under stale or missing data, cost at expected volume.
  3. Set a threshold for each, against a baseline. Written before running, with the decision owner.
  4. Test the riskiest first. If the assumption most likely to fail is also cheap to test, the study may end in its first stage.
  5. Stop when the answer is clear. Stop criteria are part of the protocol; see pre-specified evaluation plans.

What feasibility answers look like in my own work

  • A partial yes. Can cheap static metrics stand in for running kernels when detecting regressions? Pairing static PTX metrics with measured runtime on five GPU classes showed structural bugs are visible in the static signal, but semantic bugs that swap a constant compile to identical PTX (Static PTX Metrics, preprint). Useful as a filter; not viable as the only check. That is the shape of a good feasibility answer — it says where the approach works and where it does not.
  • A clear yes on one mechanism. When an LLM-generated constraint program is unsatisfiable, does returning a minimal unsatisfiable core instead of an error message help? It localised the fault and cut fabricated solutions from 79% to 7% (Minimal-Core-Guided Repair, IJCAI-ECAI 2026 workshop poster).
  • A cost that changes the design. What happens when privacy moves ML-mediated auction decisions on device? An on-device simulation showed large overspend from stale budget information and an incentive misalignment in payment units (On-Device Auction Misalignment, preprint). The simulator source is withheld, with hashes recorded in the artefact repository — a limitation stated rather than hidden.

Illustrative example — not a client engagement. A product team wants to move a recommendation model on device for privacy reasons. The load-bearing assumptions are accuracy after compression, latency on the slowest supported phone, and behaviour when server-side signals are hours stale. The protocol tests staleness first, because it is cheapest to simulate and most likely to fail. Stage one shows accuracy degrades past the agreed threshold once signals are more than a day old; the report recommends proceeding only with a defined refresh interval, and states that compression was not tested.

Commission this if

  • An implementation commitment depends on an approach nobody has tested in your setting.
  • The evidence so far is a demo or a handful of examples, and you need to know how it behaves on the inputs it will actually see.
  • You would rather spend a bounded amount learning that it fails than discover it halfway through implementation.

Not the right fit if

  • The approach is known to work and the question is how to build it. That is delivery work.
  • You are choosing between several approaches that each work. That is a comparison study.
  • The decision has been taken and the study is meant to justify it.

What you receive

  • Viability question and threshold. The measurable condition under which the approach counts as viable for your decision, set with the decision owner before anything runs.
  • Bounded protocol. The smallest experiment that could falsify viability: data, baseline, metrics and the order in which assumptions are tested.
  • Stop criteria. Conditions, written in advance, under which the study ends early because the answer is already clear.
  • Experiment artefacts. Code, configurations, seeds and raw results, so the experiment can be re-run or extended.
  • Proceed-or-stop report. The recommendation, the evidence for it, the assumptions that were not tested, and what a next phase would need to establish.

What the study needs from you

  • The decision and its owner. What will be built, or not, depending on the answer — and who decides.
  • Constraints. Latency, cost, privacy, hardware or regulatory limits the approach must meet to be viable at all.
  • Data or a credible proxy. Representative inputs, or enough description to construct a proxy. The report states which results rest on proxies.
  • A baseline. The current approach, or the simplest alternative, so 'viable' means better than something rather than better than nothing.

How the work runs

  1. 1
    Frame. Agree the viability threshold, the decision it informs and any conflicts. Some questions turn out to need a comparison or a replication instead; I will say so.
  2. 2
    Order assumptions by risk. List what has to be true for the approach to work, then test the assumption most likely to fail first, at the lowest cost.
  3. 3
    Pre-specify. Fix thresholds, stop criteria, repeats and the budget for each stage before results exist. Amendments are logged with reasons.
  4. 4
    Run in stages. Check stop criteria at each stage boundary and review with the decision owner before spending on the next.
  5. 5
    Report. Proceed, change course or stop — with the evidence, its uncertainty and what was out of scope.

Limits and unfavourable results

  • Feasibility in a bounded experiment is not production readiness. Scale, integration and operations are untested unless explicitly scoped.
  • A 'proceed' means the riskiest tested assumptions held. It does not mean every assumption did.
  • Results on proxy data can mislead. The report separates what was measured on real inputs from what was measured on proxies.

Stop is a complete and useful answer, and often the cheapest one you will ever get. The fee is fixed at scoping and does not depend on whether the recommendation is to proceed or to stop; the report is written the same way either way.

Engagement terms

Model
Fixed-scope staged study with pre-agreed stop criteria and acceptance criteria for the deliverables (not for the recommendation).
Who does the work
Dipankar Sarkar personally designs and runs the study. Any specialist help is disclosed and agreed in advance.
Commercial basis
Fixed research scope, quoted after a scoping conversation. The fee does not depend on the result.
Availability
Checked per enquiry.

Conflicts are checked before scoping. See publication, IP and independence.

Evidence behind this page

  • on-device-auction-audit (GitHub) — Results, calibration data and analysis scripts for the on-device auction study. Code MIT, data CC BY-NC 4.0; simulator source withheld with hashes recorded.
  • gpuemu-corpus (GitHub) — Includes the driver for the static-PTX study, with stored seeds and replay scripts. MIT or Apache-2.0.

Questions

What experiment would tell us whether an AI approach is viable?

The smallest one that could show it is not. Write down what must be true for the approach to work, set a threshold for each against a baseline, and test the assumption most likely to fail first. If that one survives, move to the next. A demo tests whether something can work once; a feasibility study tests where it stops working.

How long does a feasibility study take?

It depends on the number of assumptions and how expensive each is to test. The protocol fixes a budget per stage, and the quote follows the scoping conversation. Stop criteria mean the study can end early when the answer is already clear.

What happens if the recommendation is to stop?

You receive the same package as for a proceed: the evidence, the uncertainty, what was not tested, and what a different approach would need to show. The fee is unchanged.

Is a feasibility study the same as building a prototype?

No. A prototype is built to show the approach can work on a favourable path. A feasibility study is designed to find the conditions under which it fails, and its code is an experimental artefact, not a first version of the product.

See also

Commission a scoped study

Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.

Last reviewed 2026-10-07.