{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-30T18:53:06.945Z"},"content":[{"type":"documentation","id":"4771080b-b794-4454-8d5b-fc8e9f2eff2f","slug":"composite-sample-aggregate-feedback-variance","title":"The Composite Sample Problem: Why Your Aggregate Score Hides the Account on Fire (2026)","url":"https://www.koji.so/docs/composite-sample-aggregate-feedback-variance","summary":"Averaging customers into one metric is compositing: combining units and measuring once. It yields a cheap mean and removes all between-unit variance, so the aggregate cannot show an outlier account. Worked example: 39 accounts at 9 plus one at 1 composites to 8.8, identical to 40 accounts all at 8.8. Compositing in four groups of ten instead surfaces the affected group at 8.2 versus 9.0.","content":"**Answer first:** When you average many customers into a single satisfaction number, you have performed the operation that environmental labs call compositing: combining many samples and analysing the combination once. Compositing is a legitimate and cheap way to estimate a mean. It also destroys all information about variation between the units you combined, which means the aggregate score is structurally incapable of showing you the one account that is about to churn. The fix is not to abandon the average. It is to composite in several groups instead of one.\n\n## What compositing is\n\nIn environmental sampling, when you need to know the average concentration of a contaminant across a field, you have a choice. You can collect thirty samples and pay for thirty laboratory analyses, or you can collect thirty samples, mix them, and pay for one.\n\nPacific Northwest National Laboratory, whose Visual Sample Plan software is used to design these programmes, describes the operation directly: \"Composite sampling requires collecting multiple samples from the environmental population, combining them, and then analyzing the combined sample.\" The motive is equally direct: \"The purpose of composite sampling is to reduce the number of analyses that must be performed, thereby reducing costs.\"\n\nThat is a good trade in many situations, and it is the trade every aggregate customer metric makes, whether that metric comes out of a survey tool or out of a Koji report. A monthly satisfaction score is a composite. So is an average rating, a mean task completion time, and a single NPS figure for a whole product.\n\n### The cost, named by the people who use the method\n\nThe same source is candid about what you give up: \"A negative consequence of compositing is the loss of spatial information for the individual increments.\"\n\nIn soil that means you know the field averages 40 parts per million and you no longer know which corner of it was the source. In customer feedback it means you know satisfaction is 8.8 and you no longer know which account supplied which part of that number.\n\nA note on rigour: the stronger claim you sometimes see, that compositing eliminates your ability to estimate variance at all, does not appear in that source and is not quoted here as though it did. It does not need to be borrowed, because it follows from arithmetic you can check yourself, which is what the next section does.\n\n## The arithmetic of what gets hidden\n\nTake 40 accounts. Thirty-nine of them are genuinely happy and would rate you 9 out of 10. One is in serious trouble and would rate you 1.\n\nComposite them into a single number:\n\n(39 x 9 + 1 x 1) / 40 = 352 / 40 = **8.8**\n\nNow consider a second, entirely different population: 40 accounts that every one rate you 8.8. Its composite is also **8.8**.\n\n**The two numbers are identical and the situations could not be less alike.** One of them contains a customer who is leaving. The aggregate does not merely under-weight that customer, which would be a question of sensitivity. It cannot represent them at all, because a single measurement has no room to carry a distribution. Averaging 39 nines and a one is not a lossy summary of the account in trouble; it is the removal of it.\n\nThis is why \"our score is healthy\" and \"we lost a major account last quarter\" are not contradictory statements. They are compatible, and the metric was never able to warn you.\n\n### Now composite in groups instead\n\nHere is the part teams miss, and it costs almost nothing. Keep the same 40 accounts and the same total sample. Instead of one composite of 40, run four composites of 10.\n\nAssume the troubled account falls in group three:\n\n| Group | Composition | Composite score |\n| --- | --- | --- |\n| 1 | 10 accounts at 9 | 9.0 |\n| 2 | 10 accounts at 9 | 9.0 |\n| 3 | 9 accounts at 9, one at 1 | 8.2 |\n| 4 | 10 accounts at 9 | 9.0 |\n\nGroup three reads 8.2 against 9.0 everywhere else. You have not identified the account, but you now know that one quarter of your base contains something that the other three quarters do not, and you know exactly which tenth of your customers to go and look at. You went from one analysis to four. You sampled precisely the same 40 people.\n\nThe single composite reported 8.8 with no range at all. The grouped design reports a mean you can still compute and a spread of 8.2 to 9.0 that tells you where to dig. **The information was not missing from your customers. It was destroyed by the decision to mix them.**\n\n## Why this is not the same as slicing a dashboard\n\nAn obvious objection: modern analytics lets me break the score down by segment whenever I like, so nothing is lost.\n\nThat is true when you have retained the individual measurements, and it is precisely false when you have not. The distinction is whether the mixing happened before or after measurement.\n\n- **Aggregated after measurement.** You hold 40 individual scores and compute an average for display. Nothing is lost. Any slice you think of later is available.\n- **Composited before measurement.** You asked a question in a way that only ever produced one pooled number, or you discarded the unit-level records, or the unit-level records exist but are not linked to anything you can segment by. The slice is not available at any price, because the data to support it was never separable.\n\nMost feedback programmes are closer to the second case than teams realise, usually for a mundane reason: the responses are anonymous, or not joined to the account, or the open-text answers were summarised into themes without preserving which respondent said what. Once that happens the aggregate is a true composite and a later request to \"break this down by account tier\" cannot be honoured.\n\n## Choosing your compositing groups\n\nIf you are going to composite, and cost usually means you are, the grouping is the whole design. Two rules carry most of the value.\n\n**Group by the axis along which you expect trouble to concentrate.** Compositing averages away variation *within* a group and preserves variation *between* groups. So put your groups on the dimension where a problem would cluster: plan tier, onboarding cohort, region, the team that owns the account. Grouping at random preserves nothing useful.\n\n**Make the groups small enough that one bad unit still moves the number.** In the example above, one troubled account in ten moved the group from 9.0 to 8.2, which is visible. The same account in a group of 100 would have moved it to 8.92, which is not. The smaller the group, the rarer the problem you can still see. This is the same detection logic covered in [Nobody Mentioned It](/docs/detection-floor-nobody-mentioned-interviews), approached from the aggregation side rather than the sampling side.\n\nThere is a floor on how small you should go when you publish the results, for privacy rather than statistical reasons, and [k-Anonymity for Segment Reporting](/docs/k-anonymity-segment-reporting-minimum-base-size) covers where that floor sits.\n\n## Common mistakes\n\n- **Treating a low variance report as good news.** If your process composites, low reported variance may mean your instrument cannot see variance. Check which.\n- **Reporting a composite mean with a confidence interval.** A single pooled measurement gives you no within-group replication, so an interval computed from it describes measurement noise, not the spread across customers.\n- **Collecting anonymously by default.** Anonymity is often the right call, but it is also the decision that converts a separable sample into a permanent composite. Make it deliberately.\n- **Compositing across genuinely different populations.** Mixing enterprise and self-serve into one number produces an average of two materials, which is the heterogeneity problem in [Why a Better Analysis Cannot Rescue a Bad Sample](/docs/sampling-error-irreducible-interview-analysis).\n- **Assuming the outlier will show up somewhere else.** It will show up in churn, later, as a surprise. A Koji study run per cohort surfaces it while it is still a conversation.\n\n## How Koji avoids the composite trap\n\nCompositing exists because analysis is expensive. In a laboratory the cost is per assay; in customer research the cost has always been per conversation and per hour of synthesis. That cost is exactly why teams reach for one pooled number: thirty separate conversations, each read and coded individually, is a week of work, and one aggregate survey result is an afternoon.\n\nKoji removes the reason to composite in the first place. Every interview is conducted individually by the AI interviewer and analysed individually, so the unit-level record always exists and stays attached to the participant. The report gives you the aggregate view, and the aggregate is never the only thing you have: you can open any single conversation behind any theme and read what that person actually said, in their words, with the AI follow-up questions that established context.\n\nThree specifics matter here:\n\n- **Structured questions keep quantitative answers separable.** Koji's six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) capture typed values per participant, so a scale question aggregates into a distribution rather than collapsing to a mean. A distribution shows you the account at 1 out of 10. A mean cannot.\n- **Themes stay traceable to their source.** Automatic analysis groups findings into themes, and each theme carries the supporting quotes and the conversations they came from, so synthesis does not become the step that composites your data.\n- **Running several small studies is free.** Because there is no moderator to schedule, grouping your base into four cohorts of ten and running four studies costs what one study of forty costs. The grouped design stops being a luxury.\n\nThe aggregate number is still useful and Koji still gives it to you. It simply stops being the only thing that survived.\n\n## Frequently asked questions\n\n### Is compositing always a bad idea?\n\nNot at all. Compositing is a sound method for estimating a mean at low cost, and that is genuinely what you want for some questions. It becomes a problem when the mean is used to answer a question about outliers, variation or risk, because those are exactly the quantities the operation removes.\n\n### Can I recover the variance from a composite afterwards?\n\nNo. That is the defining property, and it is worth being blunt about because teams spend real effort trying. Once units are physically or statistically merged into a single measurement, the between-unit differences are not attenuated or hidden, they are absent from the result. Only the design decision made beforehand determines whether you have them.\n\n### How many groups should I composite into?\n\nEnough that a single problem unit still moves its group's number by an amount you would notice. A practical starting point is groups of about ten, which in the worked example moved one group from 9.0 to 8.2. Then align the group boundaries with the dimension along which you expect problems to cluster rather than splitting at random.\n\n### Does a distribution chart solve this?\n\nYes, provided the underlying answers were captured per participant. A distribution is precisely the object that a composite destroys, so showing one proves the data was never composited. If a tool can render a real distribution rather than a mean with error bars, your data is still separable.\n\n### Is this different from Simpson's paradox?\n\nRelated but distinct. Simpson's paradox is about an aggregate relationship reversing when you condition on a subgroup, which requires that you can still see the subgroups. The composite problem is upstream of that: it is about aggregation that leaves no subgroups to condition on, so the paradox could not even be detected.\n\n### What if my feedback is anonymous for good reasons?\n\nAnonymity and separability are not actually in conflict, which is the useful thing to know. You do not need to know who someone is to keep their answers as a distinct record; you only need a stable anonymous identifier and the segment attributes you will later want to group by. Collecting plan tier and cohort without collecting identity preserves most of the analytical value.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and why typed answers aggregate into distributions\n- [Nobody Mentioned It](/docs/detection-floor-nobody-mentioned-interviews) - the detection side of the same arithmetic\n- [Why a Better Analysis Cannot Rescue a Bad Sample](/docs/sampling-error-irreducible-interview-analysis) - the heterogeneity term underneath the pooling decision\n- [k-Anonymity for Segment Reporting](/docs/k-anonymity-segment-reporting-minimum-base-size) - how small a group can be before you should not publish it\n- [The Moving Average on Your Dashboard Is Hiding the Week That Mattered](/docs/moving-average-hides-events-research) - the same loss along the time axis\n- [Heterogeneous Treatment Effects](/docs/heterogeneous-treatment-effects-research) - why nobody experienced your average\n","category":"Analysis & Synthesis","lastModified":"2026-09-29T03:36:26.140063+00:00","metaTitle":"The Composite Sample Problem: Why Your Aggregate Score Hides the Account on Fire (2026)","metaDescription":"Your aggregate score is a composite sample: it estimates the mean and destroys the variance that would show you the account about to churn.","keywords":["composite sampling feedback","aggregate satisfaction score problem","why averaging customer feedback hides outliers","between-unit variance","segment feedback reporting"],"aiSummary":"Averaging customers into one metric is compositing: combining units and measuring once. It yields a cheap mean and removes all between-unit variance, so the aggregate cannot show an outlier account. Worked example: 39 accounts at 9 plus one at 1 composites to 8.8, identical to 40 accounts all at 8.8. Compositing in four groups of ten instead surfaces the affected group at 8.2 versus 9.0.","aiPrerequisites":["Familiarity with averages and variance","Basic understanding of customer feedback metrics"],"aiLearningOutcomes":["Recognise when a metric is a composite sample","Explain why variance cannot be recovered after compositing","Design grouped compositing that preserves outlier detection","Distinguish aggregation before measurement from aggregation after it"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}