Why the Top and Bottom Segments in Your Report Are Both the Smallest Ones (2026)
Rank your segments by score and the top and bottom of the list fill up with your smallest segments. This is arithmetic, not bad luck, and no multiplicity correction touches it.
Sort your segments by satisfaction score and look at the top three. Then look at their sample sizes. In most research readouts, the segments at the top of the ranking and the segments at the bottom of the ranking are the same kind of segment: the small ones.
This is not a bias you can correct by being careful. It is arithmetic. Ranking segments by their observed averages is, in part, ranking them by how few people answered - and the smaller your segments, the more of the ranking is sample size rather than signal.
The answer, stated first
An observed segment average is the true value plus noise, and the size of the noise depends on n. Small segments have noisy averages, so they land far from the middle in both directions. When you sort and take the top of the list, you are selecting on that noise, and the selection preferentially picks exactly the segments whose noise is largest.
The consequence: your best-performing segment is disproportionately likely to be one of your smallest, and it is likely to be closer to average next quarter - not because anything changed, but because it was never actually at the top.
The same is true of the bottom of the ranking. The worst segment is also disproportionately small. Teams notice this less often because a bad number prompts investigation rather than celebration, but it is the identical effect with the sign flipped.
The demonstration that settles it
Here is the cleanest way to see that this is arithmetic and not bad luck. Build a set of segments in which every segment has exactly the same true score, vary only their sizes, and rank them.
Simulating 40 segments that all share a true mean of 4.10 on a 1-5 scale, with a within-segment standard deviation of 1.2 and sizes ranging from 6 to 200 (average size 49), then repeatedly ranking them by observed mean:
- Average size of the top four segments by observed mean: 16
- Average size of all segments: 49
The top of the leaderboard is populated by segments about a third the typical size, in a world where no segment is better than any other. Every "insight" in that ranking is an artefact of the denominator. Re-run it with a different random seed and the names change; the pattern does not.
The published version, with real populations
Gelman and Price demonstrated the same effect on U.S. county disease rates, in a paper whose title is the conclusion: all maps of parameter estimates are misleading. Working with a model constructed so that county parameters were "distributed randomly, with no spatial correlation" - that is, with no real geographic pattern at all to find - they report that when counties with very high observed rates are highlighted on a map, "almost all of the highlighted counties are low-population counties."
Their quantification is the number worth remembering. In their cancer-rate model the average county population is 80,000. The expected average population of the counties that get highlighted as the top 10 per cent is 16,000 when highlighting is based on raw observed rates. One fifth the typical size, selected out of a population with no real spatial structure whatsoever.
This is why maps of disease rates by county reliably light up the sparsely populated interior of the United States, and why the same counties often appear on maps of the lowest rates. It is also, precisely, why your segment leaderboard looks the way it does.
Why this is not the multiple comparisons problem
The resemblance is close enough to cause confusion, so it is worth separating them carefully.
The multiple comparisons problem is about testing: run twenty significance tests at a 5 per cent threshold and you expect one false positive. The fix is a correction to the threshold, or control of the false discovery rate.
This is different in three ways that matter:
- There is no test. You sorted a column. No p-value was computed, no threshold was crossed, and no correction has anywhere to attach.
- It survives a correct model. Gelman and Price are explicit that they count a pattern as an artefact "if it occurs even when inferences are based on the correct statistical model". Getting the statistics right does not remove it.
- The direction is predictable. A multiple comparisons false positive could be any segment. This one specifically favours the small ones, every time.
You can apply every multiplicity correction in the literature and still hand your executive team a ranking whose top is selected on sample size. The two problems compound; neither is the other's fix.
It is also distinct from regression to the mean, which is the same statistical fact observed across time - you pick an extreme at one point and it moves back at the next. Here there is no second measurement. The ranking is wrong the first time you look at it.
Where this shows up in product and research work
- Segment leaderboards in dashboards. Any table sorted by a rate with a variable denominator.
- Top requested feature by segment. Small segments produce extreme request percentages.
- Account-level health scores. The accounts flagged as healthiest and least healthy are both disproportionately the ones with the fewest logged interactions.
- Win rates by source. A channel with 6 opportunities shows a 67 per cent win rate and goes into the board deck.
- NPS by cohort. NPS is unusually vulnerable because it discards the middle and differences two proportions, which inflates the variance relative to a mean.
- Experiment results sliced after the fact. The subgroup with the biggest lift is usually the smallest subgroup.
- Vendor or competitor comparisons built on review counts. The highest-rated product on a review site is frequently the one with nine reviews.
The common structure is a rate or an average, a denominator that varies across rows, and a sort.
What to do instead
1. Never sort on a raw estimate alone
Sorting is the operation that does the damage. If the table must be sorted, sort on a credibility-weighted estimate rather than the raw one - see credibility weighting for small segment estimates for the blend. This does not make the ranking correct, but it removes the crude version of the artefact.
2. Understand that shrinking has its own bias, in the opposite direction
This is the part that catches people who have already learned lesson one. Gelman and Price continue: in their model, the expected average population of the highlighted counties is 16,000 under raw ranking but 190,000 under posterior-mean ranking - against a true average of 80,000. Shrinkage does not land you on the truth. It overshoots, and the top of a shrunk ranking becomes disproportionately populated by your largest segments, because those are the ones whose estimates were allowed to stay extreme.
In my 40-segment simulation the same reversal appears in milder form: ranking by the shrunk estimate lifts the average size of the top four from 16 to 39, against a true average of 49 - much better, and still not neutral.
So there is no point estimate you can rank that is free of the effect. That result deserves its own article, and it has one: when shrinkage hides the one segment that actually changed.
3. Show the uncertainty, and let it do the work
Put a confidence interval on every segment and display it. A ranking in which the top eight intervals all overlap communicates the truth immediately, and non-technical stakeholders read it correctly without being taught anything. This is the single highest-value change most teams can make to a segment report.
4. Rank on the probability of being extreme, not on the estimate
If you genuinely need a shortlist, compute for each segment the posterior probability that its true value exceeds some threshold you care about, and rank on that. It is still not fully neutral with respect to sample size, but it is honest about what it is claiming, and it forces you to name the threshold - which is usually a productive argument.
5. Equalise n at design time
The cleanest fix happens before any data exist. If you know which segments you intend to compare, recruit to comparable sample sizes rather than taking whatever the population mix delivers. A deliberately balanced design makes the whole problem mostly disappear, because the artefact is driven by variation in n across rows. When every row has the same denominator, sorting is fair.
This used to be impractical. Filling a quota of 40 interviews in six segments meant six recruitment pushes, six scheduling cycles, and a month. With an AI-moderated platform like Koji it is a quota setting, because no segment costs more moderator time than any other.
How Koji makes the balanced design affordable
The reason segment sizes vary so wildly in most research is not that anyone wanted them to. It is that traditional research fills whatever slots it can get, and the segments that are hard to reach stay small forever - so those segments are permanently the noisiest rows in the table and permanently the ones topping and bottoming the leaderboard.
Koji attacks this at the root. Because interviews are AI-moderated and asynchronous, there is no moderator calendar to fill and no per-interview marginal cost in analyst time. Running 40 conversations in your smallest segment costs roughly the same effort as running 40 in your largest, which means balanced n is a design choice rather than a budget negotiation. That single change removes most of the artefact described in this article, because the artefact lives entirely in the variation of sample sizes across rows.
Koji's structured questions make the segment definitions clean at the point of collection. Using single_choice or multiple_choice questions as screeners assigns each respondent to a segment during the conversation itself, so you are not reconstructing segments later from mismatched CRM fields. A scale question gives you the per-respondent numeric values needed to compute a real confidence interval per segment rather than a bare mean, ranking questions expose priority differences that a single average hides, and yes_no questions give clean incidence rates with a known denominator.
Koji's real-time reporting also lets you see the size imbalance while the study is still open. If enterprise has 9 responses and mid-market has 84, you can push more interviews into enterprise that week instead of discovering the problem in analysis, when the only remaining options are a caveat or a cutoff.
And when a segment does look genuinely different, the open_ended responses and the AI's automatic follow-up questions give you the mechanism. A ranking tells you which row is highest. The transcripts tell you whether there is a reason - and a real reason, articulated by the people in that segment, is far better evidence that a difference is real than its position in a sorted table.
Frequently asked questions
Does this mean segment analysis is useless?
No. It means sorted tables of raw averages are a bad interface for segment analysis. The underlying differences can be perfectly real and worth acting on. What you cannot do is read the ordering off a leaderboard and treat the top row as the finding. Use intervals, use credibility weighting, and prefer balanced designs.
My top segment has 40 responses, not 6. Am I safe?
Safer, but the question is relative, not absolute. What matters is how the sizes vary across the rows you are comparing. Forty is small if the other segments have 4,000 and large if they have 25. If every segment is around 40, sorting is close to fair.
Is this the same as Simpson's paradox?
No. Simpson's paradox is about a relationship reversing when you aggregate or disaggregate across a confounder - a composition effect, treated in mix shift. The effect here needs no confounder at all and appears when segments are identical in every respect except size.
Would a bigger overall sample fix it?
Only if the extra sample is distributed so that the segment sizes become more equal. Doubling every segment reduces the noise but leaves the relative pattern intact, so small segments still occupy the extremes. Doubling only the small segments fixes it properly.
How do I explain this to a stakeholder who wants a ranked list?
Show them the ranking with intervals attached, and point out how many of the intervals overlap. Then show the sample sizes next to the ranks. In practice the sample-size column does more persuasive work than any explanation of variance, because the pattern is visible immediately once someone looks for it.
Does this affect qualitative research too?
Yes, in a looser form. If you run six interviews in one segment and thirty in another, the six-person segment will produce both your most extreme-sounding quotes and your most unrepresentative themes, simply because a small sample has a wider spread of possible compositions. Koji makes this easy to avoid by lowering the cost of bringing the thin segment up to parity.
Related Resources
- Structured questions guide - the six question types and when to use each
- Credibility weighting for small segment estimates - the blend that partly fixes this
- When shrinkage hides the one segment that actually changed - why the fix has its own bias
- The multiple comparisons problem - the related but separate testing problem
- Regression to the mean - the same fact observed over time
- Mix shift and composition effects - when segment composition drives the total
Related Articles
Credibility Weighting: How Much of a Small Segment Score Should You Believe? (2026)
An eight-person segment scoring 4.6 against a 4.1 average is neither 4.6 nor unusable. Credibility weighting gives you the exact weight to apply, using a formula actuaries have relied on since 1918.
When Shrinkage Hides the One Segment That Actually Changed (2026)
Every small-segment correction assumes your segments are draws from one population. When a segment genuinely breaks away, the correction pulls it back toward a mean it no longer belongs to, silently.
Mix Shift: Why Your Score Fell When Every Segment Improved (2026)
Your headline metric can fall while every segment inside it improves. Learn the Kitagawa decomposition that splits a metric change into rate and composition components, and how to act on it.
The Multiple Comparisons Problem: Why Slicing Data Into Segments Manufactures Findings (2026)
Test 20 segments at the 5 percent threshold and you have a 64 percent chance of finding at least one difference that is not there. Learn how to count the tests you actually ran, when to control the family-wise error rate versus the false discovery rate, and why a correction cannot rescue a bad prior.
Regression to the Mean: Why Your Fix Looks Like It Worked (2026)
Regression to the mean makes ordinary noise look like a successful intervention. Learn the formula that predicts how much of your improvement is arithmetic, the five product-research traps it hides in, and the designs that separate a real win from a bounce-back.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.