The Base Rate for a Product Bet: What Experiment Portfolios Say About How Often a Feature Works
Roughly one product idea in three improves the metric it was built for. The published portfolio numbers, how to build your own reference class, and why idea variance beats sample size.
Answer first: across every large experiment portfolio that has published its numbers, roughly one product idea in three improves the metric it was designed to improve. Microsoft's own figure is "only about one-third"; a third come back flat and a third actively hurt. That number is the base rate for a product bet, and it is the single most useful piece of context you can put next to a new idea - because your team's confidence in the idea contains no information about how often ideas like it work. The outside view starts with somebody else's data and ends with your own.
The numbers, from teams that measured
The strongest evidence comes from organisations that ran enough controlled experiments to know their own hit rate and were willing to publish it. Ron Kohavi and colleagues reported Microsoft's in Online Experimentation at Microsoft, and the sentence is worth reading exactly as written:
"Evaluating well-designed and executed experiments that were designed to improve a key metric, only about one-third were successful at improving the key metric!"
The distribution behind that headline is the part teams find hardest to accept. In the same body of work, summarising the experience of Microsoft's experimentation platform team:
- 1/3 of ideas were positive and statistically significant
- 1/3 of ideas were flat, with no statistically significant difference
- 1/3 of ideas were negative and statistically significant
Kohavi adds that at Bing "the success rate is lower," and that "the low success rate has been documented many times across multiple companies." The same write-up notes that at Amazon it is common practice to evaluate every new feature, and the success rate there "is smaller than most people think."
The published figures from elsewhere line up.
| Source | Domain | Reported rate |
|---|---|---|
| Microsoft (Kohavi et al.) | Software features, controlled experiments | About 1/3 improve the target metric |
| Bing (same) | Search | Lower than 1/3 |
| Netflix, cited in Moran, Do It Wrong Quickly (2007) | Web product | Considers 90% of what they try to be wrong |
| QualPro (Holland & Cochran, 2005) | 150,000 business improvement ideas over 22 years | 75% had no impact on performance or hurt it |
| Avinash Kaushik, Experimentation and Testing primer (2006) | Web analytics | "80% of the time you/we are wrong about what a customer wants" |
| Regis Hadiaris, Quicken Loans, cited in Moran (2008) | Financial services testing | Correct about the outcome of a test "about 33% of the time" after five years |
Two things about that table. First, the estimates were produced independently, in different industries, by people with no incentive to make their organisations look bad, and they converge on the same order of magnitude. Second, the QualPro figure is offline - 150,000 ideas tested through multivariable experiments over twenty-two years - which means this is not a quirk of web products or of a particular metric definition.
Kohavi's own summary of the implication: "It is humbling to see how bad experts are at estimating the value of features (us included)."
What these numbers do not say
Being precise here matters, because the base rate is easy to misuse in both directions.
They are rates for deliberate, metric-targeted changes that survived internal review and got built. They are not a claim that two-thirds of product work is worthless. Bug fixes, compliance work, infrastructure, table-stakes features and platform investments are not in this population - they were never justified by "this will move metric X."
They are also rates for specific metrics over specific windows. An idea that is flat on the target metric may be genuinely good on a metric nobody measured, or good over a horizon longer than the experiment. Kohavi's own advice for the flat third is to stop the launch anyway, on the grounds that every deployment carries cost - but that is a judgement about deployment cost, not proof the idea was bad.
And they come from organisations with mature experimentation practice, which cuts both ways. Their ideas may be better-filtered than yours, or the filter may already have removed the obvious winners before anything reaches an experiment.
None of that rescues the optimistic reading. Whatever your adjustments, the honest prior for "this feature will move the metric we built it for" is closer to one in three than to the certainty in the room when it was proposed.
Why the base rate is the number you do not have
Every prioritisation framework in wide use asks a team to score confidence. ICE asks for it directly. RICE puts it in the numerator. Weighted scoring buries it in the weights. In every case the number comes from the inside view: the team looks at the specifics of this idea and forms a judgement.
The outside view asks a different question. Not "how good is this idea?" but "of the ideas that looked like this one at this stage, what fraction worked?" Those two questions produce systematically different answers, and the outside view is usually closer, because the inside view has access to all the reasons this idea is compelling and none of the reasons the last twelve compelling ideas did not work.
The practical consequence is arithmetic. A team scoring confidence at 8/10 across a quarterly roadmap of twelve items is implicitly forecasting roughly ten successes. The published base rate forecasts four. Nothing in the roadmap review will surface that gap, because confidence is scored per item and the base rate is a property of the portfolio.
Put the base rate at the top of the scoring sheet and the scores change without any argument about individual items.
Building your own reference class
Borrowed base rates are a starting point. Your own are better, because they reflect your product, your users and your team's particular way of being wrong. Constructing them takes an afternoon and one rule.
The rule: the reference class must be defined by what you knew at the time of the bet, not by what happened. A class of "features we shipped" is contaminated - it excludes everything killed in review, which is exactly the population you want to reason about when deciding whether to build.
A workable protocol:
- Choose the decision point. The moment an idea entered the roadmap is usually right, because that is the decision the base rate has to inform.
- List every candidate that reached it over the last four to six quarters, including those later cancelled. Twenty to forty items is enough.
- Record the target metric each was meant to move, as stated at the time. If it was not stated, mark the item "no target" - the size of that group is itself a finding, and it is usually large.
- Score each outcome as improved, flat, or hurt, using the measurement you actually have. Where nothing was measured, score it "unknown" rather than guessing. A high unknown rate means you cannot compute a base rate yet, and that is the finding.
- Compute the three fractions and, if you have enough items, break them out by category - new surface area, optimisation of an existing flow, pricing or packaging, onboarding. These usually differ sharply, and the differences are actionable in a way the pooled number is not.
Two cautions. Your own class is small, so treat it as an adjustment to the published numbers rather than a replacement - a team of thirty items and a 40% hit rate has not established that it beats the industry. And if the unknown group is large, the visible ideas skew toward the ones somebody bothered to measure, which is usually the ones that worked - the same survivorship mechanism that makes any wins-only record look better than reality.
The spread of your ideas matters more than the precision of your tests
There is a striking recent result on where the value in an experimentation programme actually sits. In February 2026, Alberto Abadie (MIT), Guido Imbens (Stanford), Anish Agarwal (Columbia) and colleagues at Amazon published an empirical framework for the value of evidence-based decision making, applied to the Upworthy Research Archive - 4,857 online experiments testing headline-and-image packages.
They ran two counterfactuals on that portfolio:
| Change | Effect on the value of evidence-based decision making |
|---|---|
| Halve estimation variance (about a 29% cut in standard errors - roughly, double your sample) | +8.33% (1.4057 to 1.5228) |
| Increase the heterogeneity of true effects by 50% (a more varied idea portfolio) | +34.32% (1.4057 to 1.8882) |
| Halve the heterogeneity of true effects (a more uniform idea portfolio) | -44.22% (1.4057 to 0.7841) |
The ordering is the finding. Measuring the same ideas more precisely produces single-digit gains. Generating a more varied set of ideas produces gains three to five times larger. A programme whose entire improvement effort goes into bigger samples and tighter confidence intervals is optimising the smaller term.
This is what makes the base rate actionable rather than merely deflating. A one-in-three hit rate is not an argument for building less. It is an argument for more variance in what you try - because the returns come from the tail, and a portfolio of small, safe, similar bets has no tail to find. The same paper also reports that requiring statistical significance before acting discards roughly 27% to 30% of attainable value relative to a value-based decision rule, which is the same lesson from the analysis side: rules tuned to avoid being embarrassed by a false positive are not tuned to capture value.
What interviews add that a base rate cannot
A base rate tells you how often ideas like this one work. It cannot tell you why the two-thirds fail, and that is where the improvable part lives.
This is the natural division of labour between quantitative portfolio data and qualitative research, and it is worth being concrete about it:
- The base rate sets the prior. Expect one in three, plan the roadmap accordingly, and stop being surprised.
- Interviews move the prior. An idea that survives contact with twenty customers who can describe the problem in their own words, unprompted, is not the same bet as one that survived a workshop. It belongs in a different reference class - and if you have been recording your outcomes as described above, you can eventually show by how much.
- Interviews explain the flat third. The most expensive outcome in the table is not the negative third, which you detect and roll back. It is the flat third, where the feature works as designed and nobody cares. That gap between "built correctly" and "wanted" is a research question, not a measurement question.
Practically, this is where a fast interview platform changes what is affordable. Traditional discovery - recruit, schedule, moderate, transcribe, code - costs enough that most roadmap items never get tested against a customer before they are built, which is precisely why the base rate is what it is. Koji runs the interviews automatically in voice or text, generates its own follow-up questions from what each participant says, and returns an analysed report, so twenty conversations fit inside the window between "idea proposed" and "idea scheduled" rather than after it.
Two mechanics matter for reference-class work specifically:
- Structured questions make outcomes comparable across studies. Six types -
open_ended,scale,single_choice,multiple_choice,rankingandyes_no- mean the same pre-build question can be asked identically across every idea in a quarter. Ascaleon problem severity asked the same way twelve times is a reference class; twelve differently worded workshops are not. See structured questions. - The record persists. Reference-class construction fails most often because nobody wrote down what was expected. Studies that stay linked to the idea they tested make the retrospective an afternoon's work rather than an archaeology project - see the roadmap evidence audit.
How this differs from calibration scoring
A neighbouring article, calibration scoring for research teams, covers proper scoring rules - how to grade a probabilistic claim once reality settles it, and how to decompose a track record into calibration and resolution. It is the right tool for measuring whether your team's stated probabilities are any good.
This article supplies the other half: the prior those probabilities should start from. Calibration tells you whether your 70% claims come true 70% of the time. The base rate tells you that, absent specific evidence, 70% was probably the wrong number to state in the first place. Use the base rate to set the forecast; use the scoring rule to find out whether you are getting better at adjusting away from it.
Common mistakes
- Defining the reference class by outcome. "Features we shipped" excludes everything killed in review. Define it at the decision point, before selection happened.
- Using the base rate to argue against building. The correct response to a one-in-three hit rate is more variance and cheaper tests, not fewer bets. The Abadie result quantifies why.
- Treating a small internal sample as decisive. Thirty items and a 40% hit rate is a hint, not a refutation of the published figures. Adjust, do not replace.
- Ignoring the flat third. It is the largest source of wasted engineering and the only one qualitative research can attack directly.
- Scoring confidence per item and never checking the portfolio. Twelve items at 8/10 confidence is a forecast of ten wins. Write that forecast down where the roadmap review can see it.
- Assuming your ideas are better than the reference class without evidence. Every organisation in that table believed the same thing before it measured. Kohavi's note that "many people dismissed" the statistics when Microsoft first shared them internally is the standard reaction.
Frequently asked questions
What is the base rate for a product feature succeeding?
Across published experiment portfolios, roughly one third of deliberate, metric-targeted changes improve the metric they were built to improve. Microsoft's figure is "only about one-third," split into approximately a third positive, a third flat and a third negative. Bing reports a lower rate, Netflix has been described as considering 90% of what it tries to be wrong, and QualPro found across 150,000 tested ideas that 75% had no impact or hurt performance. One in three is the reasonable default prior.
Does this mean two-thirds of product work is wasted?
No. The population is deliberate changes intended to move a specific metric, not all product work - bug fixes, compliance, infrastructure and table-stakes features are excluded. It also does not mean the flat third had no value, only that it did not move the metric it targeted within the measurement window. What it does mean is that the confidence teams express when proposing ideas is far higher than the historical hit rate justifies.
How do I build a reference class for my own product?
List every idea that reached your roadmap decision point over the last four to six quarters, including cancelled ones, record the target metric each was meant to move as stated at the time, and score each as improved, flat, hurt or unknown. Compute the three fractions and break them out by category if you have enough items. The critical rule is to define the class by what was known at the decision point rather than by what shipped, because a shipped-only list has already been filtered by the outcome you are trying to predict.
Should we run bigger experiments to get more reliable answers?
The evidence says that is the smaller lever. Abadie and colleagues, working with 4,857 online experiments, found that halving estimation variance raised the value of evidence-based decision making by about 8%, while increasing the spread of true effects across the portfolio by half raised it by about 34%. Bigger samples on the same ideas produce modest gains; a more varied set of ideas produces much larger ones. Invest in generating different bets before investing in measuring the same ones more precisely.
How does the base rate change how we prioritise?
Put it at the top of the scoring sheet. Frameworks like ICE and RICE ask for a confidence score derived from the specifics of each idea, which reliably produces a portfolio forecast far above the historical rate. Comparing the implied number of successes across the whole roadmap against a one-in-three prior surfaces the gap without requiring anyone to argue about individual items, which is what makes it a usable intervention rather than a demoralising one.
Can customer interviews improve the hit rate?
They can move an idea into a better reference class, and they are the only tool that explains the flat third - features that were built correctly and that nobody wanted. An idea that survives twenty conversations in which customers describe the problem unprompted is a different bet from one that survived a workshop. The way to prove this for your own team is to record, in advance, which ideas were interview-backed and compare the outcome rates once you have enough items.
Related Resources
- Structured Questions Guide - the six question types, and why identical typed questions make a reference class possible
- Calibration Scoring for Research Teams - grading the forecasts this article supplies the prior for
- Expected Value of Information - deciding which of those bets deserves a study at all
- ICE Prioritization Framework - where the inside-view confidence score enters, and where the base rate belongs
- Survivorship Bias in Customer Research - why an outcome-defined reference class flatters itself
- The Roadmap Evidence Audit - keeping the record that makes a reference class cheap to build
Related Articles
ICE Prioritization Framework: Score Impact, Confidence & Ease
A complete guide to the ICE prioritization framework — how to score ideas by Impact, Confidence, and Ease, run an ICE session, and use customer research to defend your Confidence scores.
Outcome Bias: Why You Cannot Grade a Decision by Its Result (2026)
Decision quality is never observed; only outcomes are, and an outcome is decision quality plus luck. Why this has no technical fix, and what you can audit instead.
Calibration Scoring for Research Teams: How to Find Out If Your Insights Were Actually Right (2026)
Research is graded on process and almost never on outcome. Forecasting tournaments solved this with proper scoring rules. Here is how to score a research team on whether its claims came true.
RICE Prioritization Framework: How to Score and Rank Product Ideas
Master the RICE scoring framework (Reach, Impact, Confidence, Effort) for product prioritization. Includes the formula, worked examples, free template, and how customer research transforms Confidence scores.
The Roadmap Evidence Audit: Trace Every Claim on Your Roadmap Back to a Source
Take your next ten roadmap items and ask five questions of each. Most teams find their highest-cost commitment is backed by nothing but confident retelling. The audit takes an hour and needs no new tooling.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.
Expected Value of Information: How to Decide Whether a Study Is Worth Running (2026)
If no realistic result would change what you do, the value of the information is zero and the correct sample size is zero. How to run the value-of-information test before you plan a study.