{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-07T14:26:24.974Z"},"content":[{"type":"documentation","id":"f7d17834-5de1-4865-bd6d-58af0bb671dc","slug":"regression-to-the-mean-research","title":"Regression to the Mean: Why Your Fix Looks Like It Worked (2026)","url":"https://www.koji.so/docs/regression-to-the-mean-research","summary":"Regression to the mean is the tendency for extreme measurements to be followed by less extreme ones. Expected bounce-back with no intervention equals (1 - r) x d, where r is test-retest reliability and d is the baseline deviation. Because organisations always select intervention targets on extremes, the bias is structurally coupled to good prioritisation. Fixes: randomise within the extreme group, evaluate on an independent measure, use two baselines, add a control series, and collect mechanism evidence at the moment of selection because explanations do not regress.","content":"Regression to the mean is the tendency for an extreme measurement to be followed by one closer to the average, purely because the first measurement contained luck as well as signal. **If you selected something because its number was bad, its next number will usually be better whether or not you did anything.** That is the single most common reason a product fix, an account-rescue programme, or a churn intervention appears to work when it did not.\n\nThe size of the effect is not mysterious. It is arithmetic, and you can compute it in advance. If a measure has test-retest reliability *r*, and a unit sits *d* points away from the average, then the expected movement back toward the average on the next measurement is **(1 - r) x d**. Nothing else is required: no theory of your product, no assumption about your customers. Just the reliability of your instrument and how extreme your selection was.\n\nThis guide covers what regression to the mean is, the formula that makes it predictable, the five places it hides in product and customer research, and the study designs that separate a real improvement from a statistical bounce-back.\n\n## Where the idea comes from\n\nFrancis Galton named the phenomenon in 1886, in a paper to the Anthropological Institute titled \"Regression Towards Mediocrity in Hereditary Stature.\" Working from 928 adult children born to 205 sets of parents, and taking 68.25 inches as what he called the level of mediocrity, Galton found that the height deviation of a child was on average about two-thirds of the height deviation of the mid-parents. Tall parents had children who were tall but less tall; short parents had children who were short but less short. Crucially, Galton observed that the regression was **directly proportional to the parental deviation** - the further from average you start, the further back you come.\n\nThat proportionality is the whole idea, and it is why regression to the mean is predictable rather than merely annoying. It is also why the phenomenon is completely general. Galton was measuring inheritance, but the same arithmetic governs blood pressure readings, exam scores, NPS by account, weekly activation rates by cohort, and support-ticket volume by team.\n\nThe modern methodological treatment most often cited by practitioners is Barnett, van der Pols and Dobson, \"Regression to the mean: what it is and how to deal with it,\" published in the *International Journal of Epidemiology* in 2005 and cited more than 1,900 times. Their central practical warning is the one product teams most need: the effect **grows with measurement error, and grows again when follow-up is examined only on a sub-sample that was selected using the baseline value.** Both conditions describe standard product practice almost perfectly.\n\n## The formula, and what it implies about your metrics\n\nTake any measure you record twice on the same unit. Let *r* be the correlation between the two measurements in the absence of any intervention - the test-retest reliability of your instrument. Let *d* be how far the unit sat from the population average at baseline.\n\nThe expected second measurement is the average plus *r* x *d*. So the expected apparent improvement, with no intervention at all, is:\n\n**Expected bounce-back = (1 - r) x d**\n\nTwo things follow immediately, and both are uncomfortable.\n\n**First, low-reliability metrics manufacture large fake improvements.** A single-item score collected from one respondent per account, a weekly rate computed on 40 sessions, a satisfaction question asked once after a support contact - these carry a great deal of noise. The lower the reliability, the larger the share of any deficit that will evaporate on its own.\n\n**Second, the more aggressively you triage, the worse the problem gets.** Selecting the bottom decile rather than the bottom half makes *d* larger, which makes the fake improvement larger in exact proportion.\n\nHere is what that looks like for an account sitting 20 points below your average on a 0-100 measure:\n\n| Test-retest reliability (r) | Expected bounce-back with zero intervention | Share of a 20-point deficit that closes by itself |\n| --- | --- | --- |\n| 0.90 (well-constructed multi-item scale, large n) | 2.0 points | 10% |\n| 0.70 (typical multi-item survey scale) | 6.0 points | 30% |\n| 0.50 (single-item score, small sample per unit) | 10.0 points | 50% |\n| 0.30 (single item, one respondent, short window) | 14.0 points | 70% |\n\nRead the bottom row carefully. If you pick your worst-scoring accounts on a noisy single-item measure, run a rescue programme, and re-measure, **you should expect roughly 70% of the gap to close on its own.** A programme that closes 60% of the gap has, on this evidence, made things slightly worse.\n\nMost teams cannot fill in the first column for a single metric on their dashboard. That is the real finding. You cannot evaluate an intervention against a null you have never computed.\n\n## The reliability ledger\n\nThe fix at the measurement layer is a small, dull artefact that almost nobody maintains: a **reliability ledger**. One row per recurring metric, with its test-retest correlation over the interval you actually re-measure at, and the sample size per unit that estimate assumes.\n\nYou get *r* the boring way - re-measure a set of units over the normal interval while doing nothing to them, and correlate. A holdout that receives no intervention is the cleanest source. Once the ledger exists, every proposal to act on an extreme group carries an automatic, pre-computed prediction of how much improvement is free. That number becomes the bar the intervention has to clear.\n\nThe ledger also settles arguments about metric design. A measure with *r* = 0.45 is not merely imprecise; it is a measure on which half of every triage-driven improvement will be arithmetic. That is a stronger argument for a better instrument than any appeal to rigour.\n\n## The triage asymmetry: why this is structural, not rare\n\nRegression to the mean is usually taught as a curiosity. In product organisations it is closer to a permanent condition, because of a coupling nobody designs on purpose:\n\n**Organisations always select intervention targets on extremes.** The lowest-NPS accounts get the save play. The worst-converting step gets the redesign. The cohort with the steepest drop-off gets the research. The rep with the weakest numbers gets the coaching. This is good management - it is what prioritisation *means*.\n\nAnd it guarantees the bias. Which produces the line worth putting on a wall:\n\n**The better your prioritisation process, the more extreme your selection, and the larger your regression-to-the-mean bias.**\n\nTeams that prioritise badly - spreading effort evenly - are relatively protected. Teams with disciplined, data-driven triage have engineered the exact condition Barnett and colleagues warn about: follow-up examined on a sub-sample selected using the baseline value. Rigour in prioritisation buys you contamination in evaluation, unless you plan for it.\n\n## The five places it hides in product and customer research\n\n**1. Save plays and at-risk account programmes.** Accounts enter the programme because a health score, usage number, or survey response hit a low. Most such lows are partly transient - a bad month, a departed champion, a billing dispute, one grumpy respondent. Some of those resolve on their own. The programme takes credit.\n\n**2. Post-incident satisfaction.** You survey after a bad outage, scores crater, you ship reliability work, scores recover. Some of the recovery is the work. Some is that the outage week was, by construction, the worst week.\n\n**3. Redesigning the worst-performing screen.** The funnel step with the biggest drop-off gets rebuilt. Drop-off improves. But the step was chosen because it was extreme in a period-specific measurement, and traffic mix, seasonality and sampling all move week to week.\n\n**4. Coaching and enablement evaluated on the bottom quartile.** This is Kahneman and Tversky's flight-school example, transplanted. In \"On the psychology of prediction\" (1973) they described instructors who had concluded that criticism worked and praise backfired, because a cadet praised for an excellent manoeuvre usually did worse next time and a cadet criticised for a poor one usually did better. Both patterns are what regression to the mean predicts when performance mixes skill with luck. Kahneman famously made the point concrete by having instructors throw coins at a chalk target with their backs turned: the best throwers got worse, the worst got better, and no feedback had been given at all. **The instructors had been systematically rewarded for a false belief - because praise really was reliably followed by decline.**\n\n**5. Any pre/post comparison on a cohort defined by the pre measurement.** This is the general case that contains the other four. The moment your inclusion rule references the baseline value of the outcome, regression to the mean is in your estimate.\n\n## The selection-on-extremes audit\n\nYou do not need to be a statistician to run the check. Ask three questions about any claimed improvement.\n\n**Question 1: What selected this unit into the study?** If the answer references the outcome measure itself - \"we took the accounts with the lowest scores,\" \"we picked the step with the highest drop-off\" - regression to the mean is present. If the units were selected on something independent of the outcome (all accounts renewing in Q3, everyone in a randomly assigned holdout, every user who hit a feature flag), it is not.\n\n**Question 2: How noisy is the selecting measure?** Look it up in the reliability ledger. If you cannot, that is your answer: the honest position is that you do not know how much of the improvement is real.\n\n**Question 3: How extreme was the cut?** Bottom decile regresses harder than bottom half, in direct proportion to the deviation.\n\nThree answers give you the predicted free improvement. Compare it with the observed improvement. Only the difference is available for your intervention to claim.\n\n## The four designs that get you a real answer\n\n**Randomise within the extreme group.** You do not have to abandon triage. Take the bottom decile - all of it - and randomise half into the intervention. Both halves regress by the same expected amount, so the difference between them is clean. This is the single highest-value change most teams can make, and it costs nothing but the discipline of holding some accounts back.\n\n**Select on one measurement, evaluate on another.** Regression to the mean attaches to the specific measurement used for selection. If you triage on last month's health score but evaluate on an independent measure - retention, expansion, an instrument fielded fresh - you break the coupling. Do not triage and evaluate on the same number.\n\n**Use two baselines.** Take the average of two pre-period measurements rather than one. Averaging cuts the noise component, which raises effective reliability and shrinks (1 - r). It will not eliminate the bias, but it is cheap and it moves the number.\n\n**Add a control series.** Where randomisation is impossible, compare the treated group against an untreated comparison group measured over the same period, or model the pre-existing trend explicitly and test for a break at the intervention point. Both approaches are covered in detail in our guide to [quasi-experimental design](/docs/quasi-experimental-design-guide), which is the right next read if you cannot randomise.\n\n## The qualitative escape hatch\n\nHere is the part most statistical treatments miss, and it is the reason a research platform belongs in this conversation at all.\n\n**Regression to the mean is a property of a number. A mechanism does not regress.**\n\nWhen an account score drops from 72 to 41, the number 41 is a draw from a distribution and it will drift back. But if you talk to eleven of those accounts at the moment of the drop and nine of them independently describe the same unresolved invoicing error, you are no longer holding a number. You are holding a causal account, and causal accounts do not bounce back on their own - they persist until someone fixes the invoicing.\n\nThis gives you a genuinely different antidote, and one that works even when a control group is politically impossible:\n\n- **Ask why at the moment of selection, not after the intervention.** The explanation you collect at the low point is not subject to regression. The score is.\n- **Report explanations as coverage, not as rates.** \"Nine of eleven at-risk accounts named the same billing failure unprompted\" is a finding. \"At-risk NPS improved 14 points\" is, on its own, mostly arithmetic.\n- **Check whether the named mechanism was actually removed.** If nine accounts blamed invoicing and invoicing is unchanged, a recovered score is a warning, not a win.\n- **Re-interview the units that improved.** If they cannot name anything you did, they regressed.\n\nThe reason this has historically been advice nobody takes is timing. Regression to the mean is worst exactly when you must move fast - the account is at risk *now*, the funnel broke *this week* - and traditional qualitative research takes two to three weeks to recruit, schedule, moderate and analyse. By the time you have explanations, the number has already bounced and the decision has been made.\n\n## The modern approach with AI-moderated research\n\nAn AI-native research platform changes the timing constraint, which is the constraint that actually binds.\n\nWith Koji, the moment a cohort trips a threshold you can field a study to that exact cohort and have analysed findings the same day rather than the same month. Interviews run in parallel and around the clock, so 30 conversations take about as long as one. That is what makes \"ask why at the moment of selection\" a real workflow rather than a counsel of perfection.\n\nThree capabilities matter specifically for the regression-to-the-mean problem:\n\n**Structured questions alongside open-ended conversation.** Koji supports six structured question types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - in the same session. That means you can capture the number and the mechanism in one instrument: a `scale` question gives you the comparable metric, and the AI moderator's follow-up on the same topic gives you the explanation that does not regress. Legacy survey tools like SurveyMonkey or Typeform give you the number without the follow-up; traditional moderated interviews give you the follow-up at ten times the cost per participant. See our [structured questions guide](/docs/structured-questions-guide) for how to combine the six types in one study.\n\n**Consistent moderation across arms.** If you are running the randomise-within-the-extreme-group design, human moderators introduce a second problem: the moderator who interviews your treatment group is not the moderator who interviewed your control group three weeks earlier. An AI moderator asks the same core questions the same way in both arms, which removes moderator variance as a competing explanation. That is covered further in our guide to [interviewer bias](/docs/interviewer-bias).\n\n**Repeatable studies for measuring reliability.** Building a reliability ledger requires re-fielding the same instrument to untouched units on a schedule. Cloning a study and re-running it is a two-minute operation, which turns *r* from a theoretical quantity into something you can actually look up.\n\n**The honest framing:** Koji does not make regression to the mean go away. Nothing does - it is arithmetic. What an AI-native platform changes is the cost of collecting the one kind of evidence that is immune to it, fast enough to be collected at the moment of selection instead of after the bounce.\n\n## Quick reference\n\n| Situation | Regression to the mean present? | What to do |\n| --- | --- | --- |\n| Randomised treatment and control | No bias in the difference | Compare arms; both regress equally |\n| Bottom-decile accounts, no control | Yes, large | Randomise within the decile, or use two baselines plus a control series |\n| Selected on Q1 score, evaluated on Q3 retention | Largely broken | Confirm the two measures are genuinely independent |\n| Post-incident satisfaction recovery | Yes | Compare against unaffected customers over the same window |\n| Whole population measured before and after a launch | Not from selection | Watch for external events and trend instead |\n\n## Frequently asked questions\n\n### Is regression to the mean the same as a placebo effect?\n\nNo, and conflating them causes real errors. A placebo effect is a genuine response by the participant to being treated. Regression to the mean happens with no participant, no treatment and no awareness - it would occur if you measured rocks. They can and often do operate together, which is one reason uncontrolled before-and-after evaluations overstate effects so consistently.\n\n### How do I know the test-retest reliability of my metric?\n\nMeasure the same units twice over your normal re-measurement interval while doing nothing to them, then correlate the two sets of values. An untreated holdout gives you this for free. If you have never done it, assume reliability is lower than you would like, especially for single-item measures collected from few respondents per unit - and treat any uncontrolled improvement in an extreme group as unproven rather than as evidence.\n\n### Does a bigger sample fix regression to the mean?\n\nNo. This is the most common misunderstanding. A larger sample makes your *estimate* of the bias more precise; it does not shrink the bias. If you select the bottom decile of 100,000 accounts instead of the bottom decile of 500, you get a beautifully precise measurement of an effect that is still substantially arithmetic. Only design changes - randomisation, independent evaluation measures, control series, multiple baselines - reduce it.\n\n### If we cannot run a control group, is the evaluation worthless?\n\nNot worthless, but it cannot support a causal claim on its own. Two things rescue it. First, model the pre-existing trend and test for a break at the intervention point rather than comparing two points. Second, collect mechanism evidence: ask the selected units why the number moved, at the moment it moved, and check whether the cause they name was actually removed. A recovered metric plus an unremoved cause is a regression story.\n\n### Does regression to the mean affect qualitative research too?\n\nThe numbers inside qualitative studies regress like any others - a satisfaction rating collected in an interview is still a noisy measurement. The explanations do not. A described mechanism, a workflow a participant walks you through, a reason given for a decision - these are not draws from a distribution around a mean, which is exactly why they are the most robust evidence you can collect about a group selected on an extreme.\n\n### How does this relate to statistical significance?\n\nThey answer different questions and you need both. Significance asks whether an observed difference is larger than sampling noise would produce. Regression to the mean is a *systematic* effect, not sampling noise - it will happily produce differences that are large, highly significant, and entirely artefactual. A statistically significant improvement in a group selected on its baseline extreme is exactly what the arithmetic predicts. See [statistical significance in survey research](/docs/statistical-significance-survey-research) for the complementary picture.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - combine scale, ranking, and open-ended questions in one AI-moderated study\n- [Quasi-Experimental Design](/docs/quasi-experimental-design-guide) - how to measure impact when randomisation is impossible\n- [Statistical Significance in Survey Research](/docs/statistical-significance-survey-research) - what significance does and does not tell you\n- [Statistical Power and Minimum Detectable Effect](/docs/statistical-power-minimum-detectable-effect) - sizing a study to detect the change you care about\n- [Internal Benchmarks and Percentile Norms](/docs/internal-benchmarks-percentile-norms) - building the norm bank that makes deviations interpretable\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - the selection problem on the other end of the distribution\n- [Reliability vs Validity in Research](/docs/reliability-vs-validity-research) - where the r in the formula comes from\n- [Panel Conditioning](/docs/panel-conditioning-repeat-participants) - what repeated measurement does to the people being measured\n\n---\n\n**Try it on a real cohort.** Koji gives you 10 free interview credits - enough to field a mechanism study against your at-risk cohort this week and find out whether the number is going to bounce back on its own.","category":"Research Methods","lastModified":"2026-08-07T03:19:57.691773+00:00","metaTitle":"Regression to the Mean: Why Your Fix Looks Like It Worked (2026)","metaDescription":"Regression to the mean makes noise look like success. Learn the (1-r) x d formula, the five product-research traps it hides in, and the designs that separate a real win from a bounce-back.","keywords":["regression to the mean","regression to the mean example","statistical regression","regression fallacy","pre post study bias","test retest reliability","at risk account programme evaluation","why did our metric bounce back"],"aiSummary":"Regression to the mean is the tendency for extreme measurements to be followed by less extreme ones. Expected bounce-back with no intervention equals (1 - r) x d, where r is test-retest reliability and d is the baseline deviation. Because organisations always select intervention targets on extremes, the bias is structurally coupled to good prioritisation. Fixes: randomise within the extreme group, evaluate on an independent measure, use two baselines, add a control series, and collect mechanism evidence at the moment of selection because explanations do not regress.","aiPrerequisites":["Basic familiarity with product or customer metrics","Understanding of averages and correlation"],"aiLearningOutcomes":["Compute the expected bounce-back for any metric using (1 - r) x d","Run the three-question selection-on-extremes audit on any claimed improvement","Choose between randomising within the extreme group, independent evaluation measures, two baselines, and control series","Build a reliability ledger for recurring metrics","Use mechanism evidence from interviews as the one form of evidence immune to regression"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}