Expected Value of Information: How to Decide Whether a Study Is Worth Running (2026)
If no realistic result would change what you do, the value of the information is zero and the correct sample size is zero. How to run the value-of-information test before you plan a study.
Answer first: the question that decides whether a study is worth running is not "how important is this?" or "can we afford it?" It is "what would we do differently depending on the answer?" If no realistic result changes the action, the value of the information is zero and the correct sample size is zero, no matter how large the decision is. Health economics has a formal name and a formula for this - the expected value of perfect information - and it is the cleanest available answer to the question every research team is asked and almost none can answer: which studies should we not run?
The test that comes before sample size
Most research planning starts one step too late. Teams argue about how many interviews are enough, what confidence level to target, whether the sample is representative. Those are all questions about how precisely to answer. They assume the answer is worth having.
The prior question is whether the answer can change anything. Formally:
Value of information = (probability your current belief is wrong) x (cost of acting on the wrong belief)
Both terms have to be non-zero. A question you are almost certainly right about has low value even if the stakes are enormous. A question you are genuinely uncertain about has low value if being wrong costs nothing, or if the decision is already fixed. The product of a small number and a large number is still small, and the product of anything and zero is zero.
This is the whole idea, and it is a deliberately uncomfortable one, because it means some of the most interesting questions your team can ask are worth exactly nothing to answer.
Where this comes from: the expected cost of uncertainty
The formal machinery is called value of information (VOI) analysis, and it is standard practice in health technology assessment, where a national body must decide both whether to fund a treatment and whether to fund more research about it. The framing translates directly to product work.
The clearest statement of the logic comes from the Centre for Health Economics at the University of York, in the pilot study Karl Claxton and colleagues produced to support research recommendations for NICE:
"Decisions based on existing information will be uncertain, and there will always be a chance that the wrong decision will be made. If the wrong decision is made, there will be costs in terms of health benefit and resources forgone. Therefore, the expected cost of uncertainty is determined jointly by the probability that a decision based on existing information will be wrong and the consequences of a wrong decision."
That expected cost of uncertainty has a second name:
"The expected costs of uncertainty can be interpreted as the expected value of perfect information (EVPI), since perfect information can eliminate the possibility of making the wrong decision."
And then the sentence that makes it operational, because it converts a philosophical point into a budget line:
"this is also the maximum that the health care system should be willing to pay for additional evidence to inform this decision in the future, and it places an upper bound on the value of conducting further research."
The decision rule that falls out is blunt: "If the costs of investigation exceed the EVPI, then the proposed research will not be cost-effective." The York team put it in portfolio terms - population EVPI "can be used to rule out research recommendations which will not be worth while."
The empirical result of that pilot is the part worth remembering. Across six technology appraisals, the value of further research ranged from GBP 2.8 million to GBP 865 million - a factor of roughly 300 between the cheapest and most valuable question in the same portfolio. And the analysis contradicted the experts: in some cases "the analysis indicated that the original research recommendations should not be regarded as a priority," with liquid-based cytology screening the worked example at the low end, against GBP 865 million for clopidogrel in stroke patients.
Expert committees had recommended further research on both. The arithmetic said one of those recommendations was not worth making. That is the failure mode this method exists to catch, and there is no reason to think product research committees are better calibrated than NICE appraisal committees.
The four inputs you actually need
You do not need a Bayesian model to run this. You need four numbers you can estimate in a fifteen-minute conversation with whoever owns the decision.
| Input | The question to ask | Typical failure |
|---|---|---|
| The decision | What action does this study feed, and who takes it? | "It's for general understanding" - which means no decision, so no value |
| The options | What are the two or three things we might actually do? | Only one option is real; the others are decoration |
| The switch point | What result would make us pick a different option? | Nobody can name one, at any value |
| The cost of being wrong | If we pick wrong, what does it cost over the life of the decision? | Stated as "a lot" rather than a number with a denominator and a time horizon |
Write those four lines down before you write a single question. If line three comes back empty, you have found a zero-value study, and the honest recommendation is to skip it.
A worked example
A team wants to research a pricing change: move the mid tier from EUR 29 to EUR 39.
- Decision: ship the new price on 1 October, or hold at EUR 29.
- Options: two. There is no third.
- Switch point: the team will hold if more than 20% of current mid-tier accounts say they would downgrade or churn at the new price.
- Cost of being wrong: the mid tier holds roughly 4,000 accounts. A 10-point error in the churn response is about 400 accounts at EUR 29-39 per month, so somewhere between EUR 140,000 and EUR 190,000 of annual recurring revenue.
Now the arithmetic. The team's current belief is that about 12% would leave, and they are genuinely unsure - they would not be shocked by 5% or by 25%. The switch point at 20% sits inside that range. That is exactly the condition under which information has value: the plausible range of answers straddles the switch point. A study costing a few hundred euros against a five- or six-figure exposure clears the bar by two orders of magnitude. Run it.
Change one thing and the answer inverts. Suppose the team's belief is 3% with high confidence, and the switch point is still 20%. Now no realistic result crosses the line. The study is interesting and worthless in the same breath. Either accept the decision as made, or - the better move - go find the question where the range does straddle a switch point. In this case that is probably not "will they churn" but "which accounts will churn and what would have kept them," which feeds a different decision with a different switch point.
The three shapes of a zero-value study
| Shape | What it looks like | The tell |
|---|---|---|
| No switch point | Every result leads to the same action | Ask "what would you do if the answer were the opposite?" and get a shrug |
| Already decided | Contracts signed, launch dated, headcount moved | The study's deadline is after the decision's deadline |
| Reversible at low cost | A two-way door you can walk back in a week | The rollback costs less than the study |
The third one is the least obvious and the most common source of wasted research. Research is a substitute for the ability to change your mind cheaply. When the change is genuinely cheap to reverse, shipping is the experiment, and the study is a slower, weaker version of the thing you could just do. This is why the value-of-information test and a good rollback plan are substitutes for one another, and why teams with strong feature flags should rationally do less pre-launch research and more post-launch measurement than teams without them.
The second shape is a different animal, and it deserves its own treatment - a study commissioned after the decision is settled is not a mistake so much as a signal about the organisation. We take that apart in evidence or ammunition.
What the newest evidence says: precision is not where the value is
There is a fresh and unusually direct piece of evidence here. In February 2026, Alberto Abadie (MIT), Guido Imbens (Stanford), Anish Agarwal (Columbia) and colleagues at Amazon published an empirical framework for estimating the value of evidence-based decision making, applied to the Upworthy Research Archive - 4,857 online experiments comparing headline-and-image packages.
Two of their counterfactuals should change how research budgets are set.
First, precision buys surprisingly little. Halving the estimation variance - roughly a 29% reduction in standard errors, which in practice means doubling your sample - increased the value of evidence-based decision making by 8.33%, from 1.4057 to 1.5228.
Second, the spread of your ideas buys a great deal. Increasing the heterogeneity of true effects by 50% raised the value by 34.32%; halving it cut the value by 44.22%.
Read those two together. Running bigger studies on the same portfolio of ideas produces single-digit gains. Generating a more varied portfolio of ideas - some genuinely bold, some genuinely different - produces gains several times larger. If your research programme is entirely devoted to measuring the same class of small changes more precisely, the arithmetic says you are optimising the wrong term. This is the strongest available argument for cheap, fast, high-variance discovery work over expensive confirmatory work, and it comes out of a portfolio of nearly five thousand real experiments.
The same paper adds a third finding worth flagging to anyone who ships on p-values: decision rules that require statistical significance "can leave substantial value unrealized and, in some cases, generate negative expected value." In their data, a 5% significance rule delivered 1.0287 against 1.4057 for a value-based rule - a reduction of about 27%, rising to roughly 30% of attainable value discarded under their nonparametric estimates. Significance is a rule about error control, not about value, and the two come apart.
What changes when a study costs hours instead of weeks
The value-of-information test has a moving part most write-ups ignore: the cost side. Research is worth doing when its value exceeds its cost, so anything that lowers the cost widens the set of decisions that deserve evidence.
This is where AI-moderated research changes the arithmetic rather than just the workflow. A traditional moderated study - recruit, schedule, moderate, transcribe, code, synthesise - carries enough cost that it can only be justified for decisions with large exposure. Platforms like Koji run the interviews automatically, in voice or text, with the AI generating its own follow-up questions in the moment, and return an analysed report without a moderator. A study that used to cost several weeks and several thousand euros can cost a day and a fraction of that.
The consequence is not "do more research." It is that the threshold moves, and a whole class of mid-sized decisions - the ones previously resolved by the loudest opinion in the room because research could not be justified - now clear the bar. Koji starts free with 10 credits and no card, interviews run as low as €1 per qualified interview, and conversations that fall below the quality gate are never charged, so the cost side of the calculation is a number you can look up rather than estimate.
Two Koji features map directly onto the four inputs above:
- The research brief. Koji's AI consultant turns a plain-language goal into a structured brief with problem framing, target participant and a typed question plan. The brief is where you write the decision and the switch point down, which is the step teams skip. See how to write a research brief.
- Structured questions. Six types -
open_ended,scale,single_choice,multiple_choice,rankingandyes_no- let you put the switch point in the instrument itself. "Would you downgrade at EUR 39?" is ayes_no; "how likely, 0-10" is ascale; the reason is anopen_endedthe AI probes. A switch point defined on a typed question can be read straight off the report instead of argued about. See structured questions.
How this differs from two neighbouring methods
Two related articles in this library answer questions that sound similar and are not. Worth being precise, because using the wrong one wastes the effort.
| Method | The question it answers | When to use it |
|---|---|---|
| Value of information (this article) | Should this study exist at all? | Before any planning, per decision |
| The research risk model | Given that we are studying it, how much rigour does it deserve? | After VOI clears, to size the study |
| The re-research audit | Do we already own this answer? | Before both, as a cheap first filter |
The order is: do we already know it, would knowing it change anything, and only then how carefully should we find out. Most teams run only the third check.
Common mistakes
- Treating importance as value. The most important question in the company can have zero information value if the decision is settled or if no answer changes the action. Importance sets the cost of being wrong, which is only one of the two terms.
- Setting the switch point after seeing the data. A threshold chosen after the fact is not a threshold, it is a rationalisation. Write it in the brief. This is the same discipline that blind analysis enforces on the analysis side.
- Pricing the study but not the delay. A four-week study on a decision that must be made in three weeks has zero value regardless of quality, because the decision will be taken without it.
- Forgetting that value falls as certainty rises. The second study on the same question is worth much less than the first, because it can only move a belief that has already moved. This is the formal reason re-research is expensive.
- Assuming the range straddles the switch point. It usually does not. Ask people for their honest range first; if the whole range sits on one side of the line, you have your answer without fieldwork.
Frequently asked questions
What is the expected value of perfect information in plain terms?
It is the most you should ever be willing to pay to remove all uncertainty from a decision. It is calculated as the probability that your current belief is wrong multiplied by the cost of acting on a wrong belief. Because perfect information is the best information can possibly be, that number is a ceiling: any real study, which gives partial information, is worth less. If a proposed study costs more than the ceiling, it cannot be worth running.
How do I estimate the cost of being wrong without a financial model?
Use the exposure, not the profit. Count the units affected - accounts, users, deals - over the period the decision holds, and multiply by a per-unit value you already track. You are not producing an accounting figure, you are producing an order of magnitude, and orders of magnitude are enough to separate a study worth hundreds from one worth hundreds of thousands. If two people on the team estimate independently and land within a factor of three, that is precise enough to act on.
Does this mean we should only research decisions with big numbers attached?
No, and this is the most common misreading. A small decision where you are genuinely uncertain and the plausible answers straddle your switch point can easily be worth more research than a huge decision where you are confident or already committed. The test is the product of the two terms, not the size of either one. Small, frequent, reversible decisions are usually better served by shipping and measuring than by studying.
What if the stakeholder cannot name a result that would change their mind?
Then you have learned something more valuable than the study would have told you, and you should say so plainly and early. Sometimes the honest answer is that the decision is made and the request is for support rather than evidence - see evidence or ammunition for how to tell the difference and what to do about it. Sometimes it means the decision has not been framed sharply enough yet, and the right next step is a framing conversation rather than fieldwork.
How does this relate to sample size and statistical power?
Value of information decides whether to run the study; power decides how big it has to be to detect the effect you care about. They answer different questions in sequence, and running the power calculation first is how teams end up precisely measuring things that do not matter. Once value of information clears, statistical power and minimum detectable effect tells you what size instrument the switch point requires.
Does cheaper research make this test unnecessary?
It makes it more important, not less. When studies are expensive, the budget itself filters out low-value work, badly but automatically. When a study costs a day and a few credits, that filter disappears and teams can generate large volumes of research that answers nothing, which then has to be read, synthesised and stored. The value-of-information test is what replaces the discipline that cost used to impose for free.
Related Resources
- The Research Risk Model - how much rigour a question deserves once you have decided to study it
- The Re-Research Audit - the cheapest first filter: do you already own the answer?
- Evidence or Ammunition - what a study is for when no result could change the decision
- Structured Questions Guide - the six question types, and how to put a switch point in the instrument
- Statistical Power and Minimum Detectable Effect - sizing the study once it clears the value test
- How to Write a Research Brief - where the decision and the switch point get written down
Related Articles
Blind Analysis: How to Analyze Research Before You Know the Answer
Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.
How to Write a Research Brief: Templates, Examples, and AI-Assisted Generation
A step-by-step guide to writing an effective user research brief. Covers the 7 essential components, participant targeting, methodology selection, and how Koji's AI generates briefs automatically from a plain-language goal.
The Base Rate for a Product Bet: What Experiment Portfolios Say About How Often a Feature Works
Roughly one product idea in three improves the metric it was built for. The published portfolio numbers, how to build your own reference class, and why idea variance beats sample size.
The Re-Research Audit: How Much of Your Budget Buys an Answer You Already Own
Count how many of your last twenty studies answered a question you already owned. The protocol, the four causes, and where the duty belongs.
Evidence or Ammunition: What a Study Is For When No Result Could Change the Decision
A study commissioned after the decision is settled is not waste, it is a signal. How to tell whether you are being asked for evidence, input or support, before fieldwork starts.
The Research Risk Model: How to Decide Which Questions Deserve a Rigorous Study
Rigor is a residual, not an input. Borrow the audit risk model to allocate research effort by what is left over after inherent risk and your existing controls, instead of by how important the question feels.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.