{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-12T14:06:35.813Z"},"content":[{"type":"documentation","id":"21e6f40d-b84a-4be2-b615-6606053459e0","slug":"case-control-research-churn-lost-deals","title":"Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)","url":"https://www.koji.so/docs/case-control-research-churn-lost-deals","summary":"A churn or win-loss study that interviews only the customers who left is a case-control study missing its control group, which makes every finding unfalsifiable. The guide covers control selection and matching using Doll and Hill 1950 as the model protocol, the named failure modes (Berkson admission bias, recall bias, confounding by indication), the base-rate arithmetic that makes a 90 percent accurate churn flag wrong about two times in three, and the limits of what any outcome-sampled design can establish.","content":"\n\n**Answer first: when you interview thirty customers who churned and look for the reason, you are running a case-control study - a design that starts from an outcome and works backwards to find its causes. Epidemiology has been running these since 1950 and has documented exactly how they fail. The single biggest failure is the one product teams make by default: having no control group. Without a comparison group of customers who did not churn, every reason your churned customers give is unfalsifiable, because the customers who stayed would very likely have given you the same reasons.**\n\nThis guide covers the design properly - what it can establish, what it structurally cannot, and how to run one that survives scrutiny.\n\n## What a case-control study actually is\n\nMost research designs are **prospective**: you take a group, do something or observe them, and wait to see what happens. A case-control study runs the other way. You start with people who already have the outcome (the **cases**), assemble a comparable group who do not (the **controls**), and compare their histories to find exposures that differ.\n\nThis is the only feasible design when the outcome is rare, expensive, or slow. You cannot randomly assign customers to churn. You cannot wait three years to see which enterprise deals are lost. So you sample on the outcome and reason backwards.\n\nProduct research does this constantly and almost never names it:\n\n| What the team calls it | Cases | Controls (usually missing) |\n| --- | --- | --- |\n| Churn interviews | Customers who cancelled | Customers who renewed |\n| Win-loss analysis | Deals lost | Deals won at the same stage |\n| Support-ticket mining | Users who reported a problem | Users in the same flow who did not |\n| Post-incident user research | Users who hit the failure | Users on the same version who did not |\n| \"Why did power users become power users?\" | Power users | Signups from the same cohort who did not |\n\nIn every row, the right-hand column is the study. The left-hand column on its own is a collection of anecdotes about people who share an outcome.\n\n## The 1950 study that invented the design, and what it got right\n\nThe founding modern case-control study is Doll and Hill, **\"Smoking and carcinoma of the lung: preliminary report\", *British Medical Journal* 2(4682):739-748, published 30 September 1950**. It is worth reading not for the conclusion, which everyone knows, but for the protocol, which is a masterclass in the thing product teams skip.\n\nDoll and Hill interviewed **649 men and 60 women with carcinoma of the lung**. Among the men, **0.3 percent were non-smokers**; among the women, **31.7 percent**. Those are the cases. The design work was all in the controls.\n\nFour almoners **engaged wholly on research** conducted every interview using, in the paper's own words, **\"a set questionary\"** - one standardised instrument, applied identically to both arms. For each lung-carcinoma patient, the almoners were instructed to interview **a patient of the same sex, within the same five-year age group, and in the same hospital at or about the same time.** Where more than one suitable patient was available, the choice fell on the first in the ward list the ward sister considered fit for interview.\n\nRead that matching rule again, because it is doing four separate jobs. Same sex and age group removes two confounders by construction. Same hospital means the controls are drawn from the **same source population** as the cases - the same catchment, the same referral patterns, the same social composition. Same time removes seasonal and period effects. And a fixed selection procedure removes the interviewer's discretion about who becomes a control.\n\nThe authors were also explicit about the threat they could not fully close. On the risk that patients who escaped notification were systematically different, they reasoned that this was unlikely to bias the inquiry because the points of interest were either unknown or known only in broad outline to the staff doing the notifying. And where the matching rule had to be relaxed - at two specialist hospitals where a same-hospital control was not always available - **those records were analysed separately.**\n\nThat is the standard. A 1950 study, run with paper forms and four staff, matched its controls on three variables, standardised its instrument across both arms, reasoned explicitly about selection, and segregated the data where its own protocol broke down. A 2026 churn study typically does none of these things.\n\n## Why your churn study is unfalsifiable without controls\n\nHere is the failure in its simplest form. You interview thirty churned customers. Twenty-two mention price. You report that price is the leading driver of churn.\n\nNow ask the question the design was supposed to answer: **what fraction of customers who renewed would also have mentioned price?** If it is also around 70 percent, price explains nothing about churn. It is a fact about your market, not a cause of departure. You have measured the prevalence of a complaint, not its association with an outcome.\n\nThis is not a hypothetical failure mode. Complaints about price, missing features, and confusing navigation are close to universal in every customer population ever sampled, including populations with excellent retention. Any explanation that is equally consistent with both outcomes has no diagnostic value - a point developed at length in our guide to [analysis of competing hypotheses](/docs/analysis-of-competing-hypotheses-research), where the ordering principle is how much a piece of evidence **separates** explanations rather than how strongly it supports the favoured one.\n\nThe remedy is structural, not analytical. You cannot fix a missing control group in the analysis. You have to go and interview the people who stayed, with the same instrument, at the same time.\n\n## How to choose controls, which is the entire design\n\nAlmost every serious criticism of a case-control study is a criticism of its control group.\n\n| Control selection rule | Why it matters | Product research version |\n| --- | --- | --- |\n| Same source population | Controls must come from the population that would have become cases had they had the outcome | Same plan tier, same acquisition channel, same region - not \"whoever answers\" |\n| Matched on known confounders | Removes variables you already know drive the outcome | Match on tenure, company size, contract value, onboarding cohort |\n| Same time window | Removes period effects, releases, pricing changes | Interview stayers in the same weeks as leavers, not months later |\n| Identical instrument | Differences must come from respondents, not from the protocol | The same questions, asked the same way, in both arms |\n| Fixed selection procedure | Removes the researcher discretion that creates bias | Define the sampling rule before you look at the list |\n| Documented relaxations | Broken matching should not silently contaminate the pooled result | Analyse relaxed-match cases separately, as Doll and Hill did |\n\nThe tenure match deserves emphasis because it is the one product teams get wrong most often. If your churned customers averaged 8 months of tenure and your comparison group averaged 30 months, then any difference you find may simply be the difference between new and established customers. You have matched on nothing and confounded everything.\n\n## Berkson bias: when your controls are not a fair comparison\n\nIn 1946 Joseph Berkson published **\"Limitations of the application of fourfold table analysis to hospital data\" (*Biometrics Bulletin* 2(3):47-53)**, showing that when both cases and controls are drawn from a population that has already been filtered - in his case, hospital admission - a spurious association can appear between two conditions that are unrelated in the general population. The filter itself manufactures the correlation.\n\nThis is the sharpest single criticism of Doll and Hill's design: their controls were hospital patients, who are by definition ill. It happens not to have overturned their conclusion, but the objection was structurally valid.\n\nThe product version is everywhere and almost never noticed. **Your controls are usually people who responded to a research invitation.** That is a filter, and it selects for engagement, goodwill, and available time - exactly the traits associated with not churning. Comparing churned customers against a control group of enthusiastic respondents will find differences that are artifacts of the recruitment channel rather than causes of churn.\n\nThis is the mirror image of the problem covered in [survivorship bias](/docs/survivorship-bias-customer-research), which is about the cases you never reach. Berkson bias is about the controls you reach too easily. A study can suffer from both simultaneously.\n\n## Two threats the design carries by construction\n\n**Recall bias.** Because both arms are being asked to remember, and only one arm has experienced a memorable outcome, the two arms remember differently. Someone who cancelled has narrativised the decision; someone who renewed has not thought about it. This is the \"effort after meaning\" effect: people who experienced a notable outcome search their memory harder for causes, so cases can appear to report more exposures than controls purely as an artifact of motivated recall. Our guide to [recall bias](/docs/recall-bias) covers the mitigations, of which shortening the gap between event and report is by far the most effective.\n\n**Confounding by indication.** In medicine this describes the situation where the reason a treatment was given is itself associated with the outcome, so treated patients look worse. The product analogue: customers who were given a discount, assigned a customer success manager, or enrolled in a save programme were selected for those interventions **because they were already at risk**. Comparing them naively makes the intervention look harmful. If your analysis concludes that customers who spoke to support churned more, check whether support contact is a marker of trouble rather than a cause of it.\n\n## Base rates: why your at-risk flag is usually wrong\n\nCase-control logic gives you an association. Turning that into a prediction runs into a separate problem that is arithmetic rather than methodological, and it is the one that most often embarrasses a research team in front of an executive.\n\nSuppose your annual churn rate is 5 percent and you build a churn-risk model that is **90 percent accurate in both directions** - it correctly flags 90 percent of customers who will churn, and correctly clears 90 percent of those who will not. Out of 1,000 customers:\n\n| | Will churn (50) | Will not churn (950) | Total flagged |\n| --- | --- | --- | --- |\n| Model flags at risk | 45 | 95 | 140 |\n| Model clears | 5 | 855 | 860 |\n\nOf the 140 customers the model flags, only 45 actually churn. **The positive predictive value is 32 percent - the flag is wrong about two times out of three**, despite 90 percent accuracy on both arms. Nothing is broken. The base rate is low, so the far larger healthy population contributes more false positives than the small at-risk population contributes true ones.\n\nTwo consequences follow. First, do not evaluate a screener or a risk model on accuracy; evaluate it on positive predictive value at your actual base rate. Second, when you recruit research participants using such a flag, remember that most of them are not the population you think you sampled - which contaminates the case group of your next case-control study.\n\n## What a case-control study cannot tell you\n\nBeing honest about the design's limits is what makes its conclusions credible.\n\n- **It cannot give you incidence.** Because you chose how many cases and how many controls to sample, the ratio between them is an artifact of your recruiting. You can estimate the odds ratio - how much more common an exposure is among cases - but not the churn rate itself.\n- **It cannot establish temporal order on its own.** Did the reduced usage cause the churn, or did the decision to leave cause the reduced usage? Only careful questioning about sequence can separate these, and memory is an unreliable narrator about sequence.\n- **It cannot rule out unmeasured confounders.** Matching handles the confounders you thought of. This limitation is shared with all non-randomised designs; see [quasi-experimental design](/docs/quasi-experimental-design-guide) for the prospective counterpart and the conditions under which it earns a causal reading.\n- **It is not a substitute for an experiment.** It is what you use when an experiment is impossible, which for churn it usually is.\n\n## The modern approach: how Koji helps\n\nThe reason product teams skip the control group is almost never that they think it is unnecessary. It is that it doubles the recruiting and scheduling burden on a study that was already hard to staff. Halving the cost per interview is therefore not a convenience; it is what makes the correct design affordable.\n\n**Both arms, same week, same instrument.** Koji runs AI-moderated interviews in parallel, so fielding a matched control arm alongside the case arm is a scheduling change rather than a second project. That directly restores Doll and Hill's \"at or about the same time\" condition, which is the one most often lost when the comparison group is added later as an afterthought.\n\n**The same questionary, enforced rather than intended.** The 1950 study depended on four trained almoners applying one instrument consistently. An AI moderator applies the identical protocol to interview 3 and interview 300, in both arms, which removes the interviewer-variation threat that makes case and control data non-comparable. It also removes a subtler problem: a human moderator who knows they are talking to a churned customer probes for reasons to leave.\n\n**Structured questions make the two arms directly comparable.** Koji supports six question types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - each carrying a stable question ID from plan through analysis into the report. That stability is what lets you put the case arm and the control arm side by side on the same item: the proportion of leavers who selected a reason in a `multiple_choice` item against the proportion of stayers who selected it, or a `scale` distribution compared across arms. `ranking` is particularly useful here, because forcing both arms to order the same set of factors surfaces differences that open-ended complaint counts hide. See the [structured questions guide](/docs/structured-questions-guide) for how each type is aggregated.\n\n**Reaching the leavers at all.** Churned customers are the hardest group to recruit and the fastest to become unreachable. Asynchronous AI-moderated interviews that a former customer can complete in ten minutes, at any hour, without booking a call with the vendor they just left, materially change response rates on the arm that matters most - see [churned customer interviews](/docs/churned-customer-interviews) for the recruiting approach.\n\n**Automatic thematic analysis across both arms at once.** The comparison you need - which themes are differentially present in cases versus controls - is a cross-tabulation, not a reading exercise. Producing it automatically is the difference between a comparison you actually run and one you intended to run.\n\nWhat Koji does not do is choose your controls. Matching on tenure, plan, and cohort is a design judgement, and it is the judgement the whole study rests on.\n\n## Frequently asked questions\n\n### Do I really need to interview customers who stayed?\n\nYes, and it is the single highest-value change you can make to a churn programme. Without a comparison group you cannot tell a reason for leaving from a universal complaint. If 70 percent of leavers mention price and 70 percent of stayers would too, price explains nothing about churn - and you will not know that until you ask them.\n\n### How many controls should I recruit per case?\n\nOne-to-one matching is the simplest and is what Doll and Hill used. Recruiting two or three controls per case increases statistical precision, with diminishing returns past about four. If controls are much easier to reach than cases, which is usual in churn research, take the extra ratio.\n\n### What should I match controls on?\n\nMatch on the variables you already believe drive the outcome and can measure: tenure, plan tier, company size, acquisition channel, and onboarding cohort. Tenure is the one teams most often forget, and forgetting it means any difference you find may just be the difference between newer and older customers.\n\n### Is win-loss analysis a case-control study?\n\nStructurally, yes - lost deals are the cases and won deals are the controls. Win-loss practice is somewhat ahead of churn practice here because interviewing both winners and losers is a common convention. The gap that remains is matching: comparing lost enterprise deals against won small-business deals confounds deal size with outcome.\n\n### Why is my 90 percent accurate churn model wrong most of the time it fires?\n\nBecause churn is rare. At a 5 percent base rate, a model that is 90 percent accurate on both arms produces roughly 45 true positives and 95 false positives per 1,000 customers, giving a positive predictive value near 32 percent. Judge screeners and risk models on positive predictive value at your real base rate, never on accuracy.\n\n### Can a case-control study prove that something causes churn?\n\nNo design that samples on the outcome can prove causation on its own. A well-matched case-control study with a standardised instrument and a plausible temporal sequence gives you a strong, testable association - which is the right input to an experiment, not a replacement for one.\n\n## Related Resources\n\n- [Churned Customer Interviews](/docs/churned-customer-interviews) - reaching and interviewing the case arm, which is the hard half of the recruiting\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - the mirror problem: the cases you never reach at all\n- [Recall Bias](/docs/recall-bias) - why cases and controls remember differently, and how to shorten the gap\n- [Quasi-Experimental Design](/docs/quasi-experimental-design-guide) - the prospective counterpart, for when you start from an intervention instead of an outcome\n- [Analysis of Competing Hypotheses](/docs/analysis-of-competing-hypotheses-research) - sorting evidence by how much it separates explanations rather than how much it supports one\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types that make two arms directly comparable\n","category":"Research Methods","lastModified":"2026-08-12T03:24:15.27832+00:00","metaTitle":"Case-Control Research for Churn and Win-Loss Studies (2026)","metaDescription":"Every churn interview is a case-control study. Learn control group selection, matching, Berkson bias, confounding by indication, and why a 90 percent accurate churn flag is wrong most of the time it fires.","keywords":["case-control study","control group selection","churn research design","win-loss analysis method","sampling on the outcome","Berkson bias","confounding by indication","base rate positive predictive value","matched controls","retrospective study design"],"aiSummary":"A churn or win-loss study that interviews only the customers who left is a case-control study missing its control group, which makes every finding unfalsifiable. The guide covers control selection and matching using Doll and Hill 1950 as the model protocol, the named failure modes (Berkson admission bias, recall bias, confounding by indication), the base-rate arithmetic that makes a 90 percent accurate churn flag wrong about two times in three, and the limits of what any outcome-sampled design can establish.","aiPrerequisites":["Familiarity with customer churn or win-loss research","Basic understanding of research sampling"],"aiLearningOutcomes":["Recognise when a study is a case-control design","Select and match a control group from the same source population","Identify Berkson bias, recall bias and confounding by indication in product research","Calculate positive predictive value at a realistic base rate","State what a case-control study cannot establish"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}