The Composite Sample Problem: Why Your Aggregate Score Hides the Account on Fire (2026)
Combining many responses into one number gives you an unbiased mean and destroys the between-unit variance. Here is what compositing costs you, and the design that keeps the mean and the outlier.
Answer first: When you average many customers into a single satisfaction number, you have performed the operation that environmental labs call compositing: combining many samples and analysing the combination once. Compositing is a legitimate and cheap way to estimate a mean. It also destroys all information about variation between the units you combined, which means the aggregate score is structurally incapable of showing you the one account that is about to churn. The fix is not to abandon the average. It is to composite in several groups instead of one.
What compositing is
In environmental sampling, when you need to know the average concentration of a contaminant across a field, you have a choice. You can collect thirty samples and pay for thirty laboratory analyses, or you can collect thirty samples, mix them, and pay for one.
Pacific Northwest National Laboratory, whose Visual Sample Plan software is used to design these programmes, describes the operation directly: "Composite sampling requires collecting multiple samples from the environmental population, combining them, and then analyzing the combined sample." The motive is equally direct: "The purpose of composite sampling is to reduce the number of analyses that must be performed, thereby reducing costs."
That is a good trade in many situations, and it is the trade every aggregate customer metric makes, whether that metric comes out of a survey tool or out of a Koji report. A monthly satisfaction score is a composite. So is an average rating, a mean task completion time, and a single NPS figure for a whole product.
The cost, named by the people who use the method
The same source is candid about what you give up: "A negative consequence of compositing is the loss of spatial information for the individual increments."
In soil that means you know the field averages 40 parts per million and you no longer know which corner of it was the source. In customer feedback it means you know satisfaction is 8.8 and you no longer know which account supplied which part of that number.
A note on rigour: the stronger claim you sometimes see, that compositing eliminates your ability to estimate variance at all, does not appear in that source and is not quoted here as though it did. It does not need to be borrowed, because it follows from arithmetic you can check yourself, which is what the next section does.
The arithmetic of what gets hidden
Take 40 accounts. Thirty-nine of them are genuinely happy and would rate you 9 out of 10. One is in serious trouble and would rate you 1.
Composite them into a single number:
(39 x 9 + 1 x 1) / 40 = 352 / 40 = 8.8
Now consider a second, entirely different population: 40 accounts that every one rate you 8.8. Its composite is also 8.8.
The two numbers are identical and the situations could not be less alike. One of them contains a customer who is leaving. The aggregate does not merely under-weight that customer, which would be a question of sensitivity. It cannot represent them at all, because a single measurement has no room to carry a distribution. Averaging 39 nines and a one is not a lossy summary of the account in trouble; it is the removal of it.
This is why "our score is healthy" and "we lost a major account last quarter" are not contradictory statements. They are compatible, and the metric was never able to warn you.
Now composite in groups instead
Here is the part teams miss, and it costs almost nothing. Keep the same 40 accounts and the same total sample. Instead of one composite of 40, run four composites of 10.
Assume the troubled account falls in group three:
| Group | Composition | Composite score |
|---|---|---|
| 1 | 10 accounts at 9 | 9.0 |
| 2 | 10 accounts at 9 | 9.0 |
| 3 | 9 accounts at 9, one at 1 | 8.2 |
| 4 | 10 accounts at 9 | 9.0 |
Group three reads 8.2 against 9.0 everywhere else. You have not identified the account, but you now know that one quarter of your base contains something that the other three quarters do not, and you know exactly which tenth of your customers to go and look at. You went from one analysis to four. You sampled precisely the same 40 people.
The single composite reported 8.8 with no range at all. The grouped design reports a mean you can still compute and a spread of 8.2 to 9.0 that tells you where to dig. The information was not missing from your customers. It was destroyed by the decision to mix them.
Why this is not the same as slicing a dashboard
An obvious objection: modern analytics lets me break the score down by segment whenever I like, so nothing is lost.
That is true when you have retained the individual measurements, and it is precisely false when you have not. The distinction is whether the mixing happened before or after measurement.
- Aggregated after measurement. You hold 40 individual scores and compute an average for display. Nothing is lost. Any slice you think of later is available.
- Composited before measurement. You asked a question in a way that only ever produced one pooled number, or you discarded the unit-level records, or the unit-level records exist but are not linked to anything you can segment by. The slice is not available at any price, because the data to support it was never separable.
Most feedback programmes are closer to the second case than teams realise, usually for a mundane reason: the responses are anonymous, or not joined to the account, or the open-text answers were summarised into themes without preserving which respondent said what. Once that happens the aggregate is a true composite and a later request to "break this down by account tier" cannot be honoured.
Choosing your compositing groups
If you are going to composite, and cost usually means you are, the grouping is the whole design. Two rules carry most of the value.
Group by the axis along which you expect trouble to concentrate. Compositing averages away variation within a group and preserves variation between groups. So put your groups on the dimension where a problem would cluster: plan tier, onboarding cohort, region, the team that owns the account. Grouping at random preserves nothing useful.
Make the groups small enough that one bad unit still moves the number. In the example above, one troubled account in ten moved the group from 9.0 to 8.2, which is visible. The same account in a group of 100 would have moved it to 8.92, which is not. The smaller the group, the rarer the problem you can still see. This is the same detection logic covered in Nobody Mentioned It, approached from the aggregation side rather than the sampling side.
There is a floor on how small you should go when you publish the results, for privacy rather than statistical reasons, and k-Anonymity for Segment Reporting covers where that floor sits.
Common mistakes
- Treating a low variance report as good news. If your process composites, low reported variance may mean your instrument cannot see variance. Check which.
- Reporting a composite mean with a confidence interval. A single pooled measurement gives you no within-group replication, so an interval computed from it describes measurement noise, not the spread across customers.
- Collecting anonymously by default. Anonymity is often the right call, but it is also the decision that converts a separable sample into a permanent composite. Make it deliberately.
- Compositing across genuinely different populations. Mixing enterprise and self-serve into one number produces an average of two materials, which is the heterogeneity problem in Why a Better Analysis Cannot Rescue a Bad Sample.
- Assuming the outlier will show up somewhere else. It will show up in churn, later, as a surprise. A Koji study run per cohort surfaces it while it is still a conversation.
How Koji avoids the composite trap
Compositing exists because analysis is expensive. In a laboratory the cost is per assay; in customer research the cost has always been per conversation and per hour of synthesis. That cost is exactly why teams reach for one pooled number: thirty separate conversations, each read and coded individually, is a week of work, and one aggregate survey result is an afternoon.
Koji removes the reason to composite in the first place. Every interview is conducted individually by the AI interviewer and analysed individually, so the unit-level record always exists and stays attached to the participant. The report gives you the aggregate view, and the aggregate is never the only thing you have: you can open any single conversation behind any theme and read what that person actually said, in their words, with the AI follow-up questions that established context.
Three specifics matter here:
- Structured questions keep quantitative answers separable. Koji's six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) capture typed values per participant, so a scale question aggregates into a distribution rather than collapsing to a mean. A distribution shows you the account at 1 out of 10. A mean cannot.
- Themes stay traceable to their source. Automatic analysis groups findings into themes, and each theme carries the supporting quotes and the conversations they came from, so synthesis does not become the step that composites your data.
- Running several small studies is free. Because there is no moderator to schedule, grouping your base into four cohorts of ten and running four studies costs what one study of forty costs. The grouped design stops being a luxury.
The aggregate number is still useful and Koji still gives it to you. It simply stops being the only thing that survived.
Frequently asked questions
Is compositing always a bad idea?
Not at all. Compositing is a sound method for estimating a mean at low cost, and that is genuinely what you want for some questions. It becomes a problem when the mean is used to answer a question about outliers, variation or risk, because those are exactly the quantities the operation removes.
Can I recover the variance from a composite afterwards?
No. That is the defining property, and it is worth being blunt about because teams spend real effort trying. Once units are physically or statistically merged into a single measurement, the between-unit differences are not attenuated or hidden, they are absent from the result. Only the design decision made beforehand determines whether you have them.
How many groups should I composite into?
Enough that a single problem unit still moves its group's number by an amount you would notice. A practical starting point is groups of about ten, which in the worked example moved one group from 9.0 to 8.2. Then align the group boundaries with the dimension along which you expect problems to cluster rather than splitting at random.
Does a distribution chart solve this?
Yes, provided the underlying answers were captured per participant. A distribution is precisely the object that a composite destroys, so showing one proves the data was never composited. If a tool can render a real distribution rather than a mean with error bars, your data is still separable.
Is this different from Simpson's paradox?
Related but distinct. Simpson's paradox is about an aggregate relationship reversing when you condition on a subgroup, which requires that you can still see the subgroups. The composite problem is upstream of that: it is about aggregation that leaves no subgroups to condition on, so the paradox could not even be detected.
What if my feedback is anonymous for good reasons?
Anonymity and separability are not actually in conflict, which is the useful thing to know. You do not need to know who someone is to keep their answers as a distinct record; you only need a stable anonymous identifier and the segment attributes you will later want to group by. Collecting plan tier and cohort without collecting identity preserves most of the analytical value.
Related Resources
- Structured Questions Guide - the six question types, and why typed answers aggregate into distributions
- Nobody Mentioned It - the detection side of the same arithmetic
- Why a Better Analysis Cannot Rescue a Bad Sample - the heterogeneity term underneath the pooling decision
- k-Anonymity for Segment Reporting - how small a group can be before you should not publish it
- The Moving Average on Your Dashboard Is Hiding the Week That Mattered - the same loss along the time axis
- Heterogeneous Treatment Effects - why nobody experienced your average
Related Articles
Heterogeneous Treatment Effects: Nobody Experienced Your Average (2026)
A modest average lift can hide substantial benefit for some, nothing for most, and real harm for a few. How to look for that without manufacturing false findings.
k-Anonymity for Segment Reporting: How Small Is Too Small to Publish? (2026)
The rule for minimum base size in research reporting, stated exactly: every visible combination of attributes must be shared by at least k respondents - and why generalisation beats suppression.
The Moving Average on Your Dashboard Is Hiding the Week That Mattered (2026)
Smoothing conserves the area under an event and destroys its height - and every alert threshold you own is a height. The arithmetic of what a rolling average deletes, how late it reports, and why it cannot give you a number for now.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Activating Research Insights: Turn Findings Into Product Decisions
A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.
How to Analyze Open-Ended Survey Responses with AI (2026 Guide)
Stop manually coding free-text survey responses. Learn how AI analyzes open-ended answers at scale — surfacing themes, sentiment, and quotes in minutes, plus why an AI interview captures 10x more depth than any survey can.