Heterogeneous Treatment Effects: Nobody Experienced Your Average (2026)
A modest average lift can hide substantial benefit for some, nothing for most, and real harm for a few. How to look for that without manufacturing false findings.
Heterogeneous Treatment Effects: Nobody Experienced Your Average (2026)
Answer first: the number your experiment reports is an average over people the change helped, people it did nothing for, and people it harmed - and no individual user experienced that average. When a treatment effect varies across users, a modest positive result can conceal substantial benefit for a minority, nothing at all for the majority, and real damage to a segment you did not think to look at. Kravitz, Duan and Braslow put it precisely in The Milbank Quarterly: "modest average effects may reflect a mixture of substantial benefits for some, little benefit for many, and harm for a few."
The obvious response - slice the data and find out who benefited - is also the single most reliable way to manufacture a false finding. Both things are true at once, and living with that tension is the actual skill. This guide covers what heterogeneous treatment effects (HTE) are, why the honest advice to stop slicing is incomplete, the discipline that lets you look responsibly, and the research move that finds the harmed subgroup without torturing your data.
Why the average is a construct
Heterogeneity of treatment effect, in Kravitz and colleagues' definition, "is present when the same treatment produces different results in different patients." Swap patients for users and nothing else changes.
An average treatment effect (ATE) is a summary of a distribution. If that distribution is tight - everyone got roughly the same benefit - the average describes reality well and you should ship on it. If the distribution is wide or bimodal, the average is a number that describes the population and no one in it. A +2.1% lift built from +14% for new users and -3% for your most engaged cohort is not a small win. It is two findings with opposite signs, reported as one.
The consequence is not academic. Kravitz and colleagues are blunt about it: "misapplying averages can cause harm, by either giving patients treatments that do not help or denying patients treatments that would help them." In product terms, you ship the change that hurts your power users, or you kill the feature that was transforming the experience for a segment too small to move the mean.
There is a vocabulary worth knowing, because the literature and your analytics tool use different words for the same thing.
| Term | What it means | Product equivalent |
|---|---|---|
| Average treatment effect (ATE) | The mean effect across everyone | The headline lift |
| Conditional average treatment effect (CATE) | The mean effect within a defined subgroup | Effect for enterprise accounts, mobile users, week-one cohort |
| Effect modification / interaction | The effect genuinely differs by a characteristic | The change helps new users and hurts veterans |
| Quantitative interaction | Effect is positive everywhere, but larger in some groups | Everyone benefits, new users benefit more |
| Qualitative interaction | Effect changes sign across groups | Helps one segment, harms another |
The distinction between quantitative and qualitative interaction is the one that should drive your decision. A quantitative interaction is a targeting opportunity: ship broadly, prioritise the segment where the effect is strongest. A qualitative interaction is a product problem: shipping broadly means knowingly harming somebody. Most subgroup discussions never make this distinction, and it is the only one that changes what you do on Monday.
The honest objection: slicing manufactures findings
Anyone who has been burned by subgroup analysis will object at this point, and they are right to. The counter-evidence is strong, and it deserves to be stated at full strength rather than waved past.
The canonical demonstration comes from ISIS-2, a trial of more than 17,000 patients that established aspirin's benefit in acute myocardial infarction. When reviewers pressed the investigators to report which patients benefited, the Oxford team complied in a deliberately pointed way: they subdivided patients by astrological birth sign. Aspirin appeared not to work for patients born under Gemini or Libra. The finding was included precisely to demonstrate that a sufficiently enthusiastic subgroup analysis will find something, and that something can be biologically impossible.
The arithmetic behind that is unforgiving. Wang, Lagakos, Ware, Hunter and Drazen laid it out in The New England Journal of Medicine: "if the null hypothesis is true for each of 10 independent tests for interaction at the 0.05 significance level, the chance of at least one false positive result exceeds 40%." Ten cuts is a modest afternoon in any analytics tool. The same paper cites a trial whose published analysis of calcium plus vitamin D supplementation involved a total of 60 subgroup analyses, where the investigators themselves noted that "Up to three statistically significant interaction tests (P<0.05) would be expected on the basis of chance alone."
Their survey of practice found the reporting standards to match. Of 97 randomised trials reported in the journal over a single year, subgroup analyses were reported for 59 (61%). In about two thirds of those trials it was unclear whether the analyses were prespecified or post hoc, and in more than half it was unclear whether interaction tests had been used at all. Their warning is worth quoting because it forecloses the obvious loophole: "Investigators should avoid the tendency to prespecify many subgroup analyses in the mistaken belief that these analyses are free of the multiplicity problem."
This is the same statistical machinery described in the multiple comparisons problem, and nothing here contradicts it. Slicing data into segments does manufacture findings.
But "do not slice" is a rule about analysis, and heterogeneity is a fact about the world. Suppressing the analysis does not make the effect uniform; it makes the non-uniformity invisible and ships it anyway. The multiple comparisons literature tells you that a subgroup difference you discovered by looking is probably noise. It does not tell you that treatment effects are homogeneous, and it offers no protection to the segment your average just harmed. What you need is not abstinence. It is a way to look that does not generate garbage.
Four rules for looking responsibly
1. Prespecify a small number of effect modifiers, with a reason for each. Not a list of every dimension your analytics tool offers - three to five characteristics for which you can state, in advance and in one sentence, a mechanism by which the effect should differ. "Users who have built keyboard muscle memory should be harmed by a navigation change" is a hypothesis. "Let us look at it by country" is a fishing licence. Prespecification does not exempt you from multiplicity, as Wang and colleagues insist, but it converts a search into a test.
2. Test the interaction, not the two subgroup results. The most common error is comparing significance across groups: the effect was significant in new users and not significant in veterans, therefore it differs. That does not follow - the second group may simply be smaller. The correct test asks directly whether the effect differs, and it needs substantially more sample than the main effect did. As a rough working figure, detecting an interaction of a given size typically requires around four times the sample needed to detect a main effect of that size. Most product experiments are not powered to find heterogeneity at all, which means most subgroup results are underpowered, and underpowered significant findings are the ones most likely to be inflated.
3. Confirm the measure means the same thing in both groups before believing the difference. If your outcome is a survey construct - satisfaction, perceived ease, intent - a difference across segments can be a difference in how the segments interpret the scale rather than a difference in the effect. This is a prerequisite, not a refinement, and it has its own method: see measurement invariance.
The four rules, and what each one is defending against:
| Rule | Failure it prevents | Cheapest way to break it |
|---|---|---|
| Prespecify 3-5 modifiers with a mechanism | Fishing across every available dimension | Opening the segment picker after seeing the result |
| Test the interaction, not two subgroup p-values | Mistaking unequal power for a real difference | "Significant in A, not in B, so it differs" |
| Establish measurement invariance first | Reading a scale-interpretation gap as an effect gap | Comparing survey constructs across languages or roles |
| Check specifically for sign flips | Shipping known harm inside a favourable mean | Reporting only "which group benefited most" |
4. Hunt for qualitative interactions specifically, and treat them as a different class of finding. Sign flips are rarer, more consequential, and less likely to be noise than magnitude differences, because chance variation tends to produce differences in size rather than direction. A prespecified check for whether any segment moved the other way is worth its multiplicity cost in a way that a general sweep for "who benefited most" is not.
Where the average gets locked in
Heterogeneity is not only an analysis problem. It gets built into systems.
An adaptive experiment - see multi-armed bandits versus A/B tests - optimises the pooled mean and reallocates traffic accordingly. If a change helps 80% and harms 20%, the bandit converges on the change, and the harmed 20% are never again served the variant that suited them. The evidence that would have revealed the harm is precisely the evidence the algorithm stopped collecting. The same logic applies to any always-on personalisation or ranking system trained on aggregate engagement: the average becomes the objective, and the objective becomes the product.
This is the capstone of the argument. An experiment that ignores interference measures the wrong comparison; an experiment that adapts on the average measures a moving one; but an experiment that reports only an average is measuring something real and reporting it as if it described a person. The first two are technical errors with technical fixes. The third is a reporting convention, which is why it survives in organisations that have already fixed the other two.
The research move: find the sign flip by asking
Here is the practical resolution of the tension. The reason slicing manufactures findings is that you are searching a large space of possible subgroups with a weak instrument and no prior. The fix is to bring a prior - and priors about who a change hurts do not come from the data.
They come from users.
A short qualitative study run alongside or immediately after an experiment does something no amount of slicing can: it generates candidate effect modifiers with mechanisms attached, which you can then prespecify and test properly on the next cycle. Instead of forty cuts hoping one clears p < 0.05, you get three hypotheses that someone can explain.
The questions that surface heterogeneity are not the questions most post-launch surveys ask:
- "Did anything you used to do quickly get slower?" Regression for power users is the most commonly missed qualitative interaction, because power users are a minority of accounts and a majority of retained revenue.
- "Was there anything you used to rely on that is gone or harder to find?" Finds the segment whose workflow depended on a detail nobody documented.
- "Who on your team is most affected by this, and are they affected differently from you?" In B2B, heterogeneity often runs across roles inside one account, which no account-level cut can ever see.
- "If you could go back to the old version for one thing, what would it be?" The sharpest single question for detecting a sign flip, because it asks for a direction rather than a rating.
The users who answer the last question with something specific and immediate are your harmed subgroup. You have found them by description rather than by cut, and now you can define the segment, prespecify it, and test whether the interaction is real - which is exactly the discipline the multiple comparisons literature asks for.
The modern approach: heterogeneity at a sample size that can find it
The reason teams settle for the average is cost. Finding an interaction needs several times the sample of a main effect, and qualitative work at that scale used to be impossible - recruiting, scheduling, moderating and coding a hundred interviews across four segments is a quarter of work, and the decision will not wait a quarter.
Koji removes that constraint. AI-moderated interviews run in parallel, by voice or text, with as many participants as the decision warrants, and thematic analysis is produced automatically rather than by a researcher with a spreadsheet. A study that would have taken six weeks returns in days, which means it can run before the rollout decision instead of after the complaints.
Two capabilities matter specifically for heterogeneity work.
Sample size that supports segment-level reading. A 20-interview study can describe an experience. It cannot compare four segments. Running 150 AI-moderated interviews costs a fraction of 20 human-moderated ones, and segment-level comparison becomes possible rather than aspirational.
Structured questions that make segments quantitative in the same conversation as the mechanism. Koji supports six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - inside one interview, which is what lets a single study both define a segment and measure within it:
- single_choice and multiple_choice capture the segmenting characteristics you prespecified - role, tenure, primary workflow - as clean typed fields rather than free text you have to code later.
- scale gives you a within-segment distribution, so you can see a bimodal response that a mean would flatten.
- yes_no gives a per-segment rate: "Did anything get harder?" answered by 150 people, split by tenure, is an interaction estimate with a mechanism attached.
- ranking exposes sign flips directly, because a segment that ranks the old version above the new one on a specific task has told you the direction of its effect.
- open_ended with AI follow-up produces the why, which is the part that turns a subgroup difference into a prespecifiable hypothesis rather than a coincidence.
Because the structured responses carry types, the report aggregates and cross-tabs them without manual coding, so "68% of weekly-or-more users report a task that got slower, against 11% of monthly users" is available while the rollout decision is still open. Legacy survey tools can capture the segment field, but they cannot probe the answer that explains it - the question set is frozen before anyone has spoken.
You do not need a PhD in causal inference to avoid harming a segment. You need to ask a large enough group of the right people what changed, and to be disciplined about which of their answers you then test.
Frequently asked questions
What are heterogeneous treatment effects?
Heterogeneous treatment effects (HTE) occur when the same change produces different results for different users. Kravitz, Duan and Braslow define it as present "when the same treatment produces different results in different patients," and warn that "modest average effects may reflect a mixture of substantial benefits for some, little benefit for many, and harm for a few." When effects are heterogeneous, the average treatment effect summarises a distribution rather than describing any individual, and shipping on the average can mean knowingly harming a segment.
If subgroup analysis is so unreliable, why do it at all?
Because unreliability of the analysis does not make the underlying effects uniform. The multiple comparisons critique is correct: with 10 independent interaction tests at the 0.05 level, the chance of at least one false positive exceeds 40%, and the ISIS-2 trial famously demonstrated the point by showing aspirin apparently failing for patients born under Gemini or Libra. But declining to look does not protect the harmed segment - it only makes the harm invisible. The resolution is to prespecify a small number of effect modifiers with stated mechanisms, test interactions rather than comparing subgroup significance, and generate candidate segments from qualitative research rather than from data mining.
What is the difference between a quantitative and a qualitative interaction?
A quantitative interaction means the effect points the same direction everywhere but varies in size - everyone benefits, some more than others. A qualitative interaction means the effect changes sign across groups: it helps one segment and harms another. The distinction drives the decision. A quantitative interaction is a targeting opportunity, so ship broadly and prioritise where the effect is strongest. A qualitative interaction is a product problem, because shipping broadly means accepting known harm to a segment.
How much sample do I need to detect heterogeneity?
Substantially more than for the main effect. As a working rule, detecting an interaction of a given size requires roughly four times the sample needed to detect a main effect of that size. This is why most product experiments cannot find heterogeneity even when it exists, and why underpowered subgroup findings that do reach significance tend to be inflated. If your experiment was sized for the headline number, it was not sized to answer who benefited.
Should I prespecify subgroup analyses?
Yes, but prespecification is not an exemption from the statistics. Wang and colleagues in the New England Journal of Medicine warn that "Investigators should avoid the tendency to prespecify many subgroup analyses in the mistaken belief that these analyses are free of the multiplicity problem." Prespecify a small number - three to five - for which you can state a mechanism in one sentence, and account for multiplicity anyway. Their survey found subgroup analyses reported in 61% of trials, with roughly two thirds not making clear whether the analyses were prespecified or post hoc.
How does qualitative research help with heterogeneous treatment effects?
It supplies the priors that make responsible subgroup analysis possible. Data mining searches a large space of possible segments with no prior and produces false positives; interviews produce a small number of candidate effect modifiers with mechanisms attached, which you can then prespecify and test properly. Questions such as "did anything you used to do quickly get slower?" and "if you could go back for one thing, what would it be?" identify harmed users by description rather than by cut. Koji makes this practical by running AI-moderated interviews at a sample size that supports segment-level comparison and returning coded, typed results in days.
Related Resources
- The Multiple Comparisons Problem: Why Slicing Data Into Segments Manufactures Findings - the statistical case against the analysis this guide disciplines rather than dismisses
- Measurement Invariance: Why You Cannot Compare Scores Across Segments - the prerequisite before believing any cross-segment difference
- Multi-Armed Bandits vs A/B Tests - how adaptive allocation locks in an average and starves the harmed subgroup of data
- Interference Between Users - the other way an experiment measures something other than what you think
- Customer Segmentation Research: How to Build Segments That Actually Drive Decisions - defining segments that are worth prespecifying
- Structured Questions Guide - the six question types that make segment-level comparison quantitative
Related Articles
Customer Segmentation Research: How to Build Segments That Actually Drive Decisions
How to use qualitative interviews — rather than demographic surveys — to build behavioral and motivational customer segments that product, marketing, and sales teams actually use.
Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)
Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.
The Multiple Comparisons Problem: Why Slicing Data Into Segments Manufactures Findings (2026)
Test 20 segments at the 5 percent threshold and you have a 64 percent chance of finding at least one difference that is not there. Learn how to count the tests you actually ran, when to control the family-wise error rate versus the false discovery rate, and why a correction cannot rescue a bad prior.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.