{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-30T18:56:08.932Z"},"content":[{"type":"documentation","id":"e3e84d43-92f9-42a9-872d-6e2ca6f7f107","slug":"sampling-error-irreducible-interview-analysis","title":"Why a Better Analysis Cannot Rescue a Bad Sample (2026)","url":"https://www.koji.so/docs/sampling-error-irreducible-interview-analysis","summary":"Sampling error contains a component (the fundamental error in Gy sampling theory) that survives a perfect procedure. Its variance scales with grain size cubed over sample mass, so halving population heterogeneity is algebraically equal to an eightfold sample increase. Analysis quality cannot reduce it; stratification can.","content":"**Answer first:** There is a component of sampling error that survives a flawless procedure. It is set by how lumpy your population is and by how much of that population you actually sampled, and no analytical care downstream can reduce it. In particulate sampling theory this is the fundamental error, and its variance scales with the cube of the grain size divided by the sample mass. Translated into research: making your sample more homogeneous is worth roughly eight times as much as doubling the number of interviews. Most teams pull the expensive lever.\n\n## The error that survives a perfect procedure\n\nAnalytical chemists who sample crushed ore, grain shipments or contaminated soil have a problem researchers will recognise. The material is not uniform. Take a scoop from one side of the pile and a scoop from the other and you get two different answers, and neither is the answer for the pile.\n\nPierre Gy formalised this in the 1950s, and the discipline that grew out of it separates sampling error into components you can attack and one you cannot. *Chemical Engineering* states the uncomfortable part plainly: \"Since there is no such thing as two identical samples, a perfectly extracted sample (random sample) will always be inflicted by a residual error, called the fundamental error (FE).\"\n\nRead that carefully, because the word doing the work is *perfectly*. The other error terms in the model, the ones caused by a badly shaped scoop or a tool that fails to take everything it delimits or a sample that segregates in transit, are all procedural. Do the procedure correctly and they shrink until they are much smaller than the fundamental term, at least for particles above roughly 100 microns, which is the condition the source attaches to that claim. The fundamental error is what remains when you have done everything right.\n\nThis is a different claim from the ones research methodology usually makes. The familiar warnings are about procedure: do not lead the witness, do not recruit only from your happiest cohort, do not let one loud participant set the agenda. Those are real and they are fixable. The fundamental error is not a mistake at all. It is a property of the material.\n\n### Why this matters for interview research\n\nA customer base is a heterogeneous material in exactly the technical sense. Your users are not interchangeable particles. Some of them are a two person startup using one feature; some are a 4,000 seat enterprise with a compliance workflow nobody on your team has seen. When you interview twelve people you have taken a scoop, and the question is not whether your scoop was honest. It is how much the pile varies and how big your scoop was relative to it.\n\n## The two terms you can actually move\n\nGy's model for the fundamental error gives you a variance, and in that variance two quantities are under your control.\n\nThe first is sample mass. *Chemical Engineering* states it directly: \"The fundamental sampling error decreases in proportion to the square root of the sample mass\". The same source separately notes that the variance of the fundamental error decreases as the sample size increases. If the standard deviation falls with the square root of mass, that is the same square root law that governs every other sample size calculation you have ever run.\n\nThe second is the grain of the material. Here the source is specific about how sharply it bites: the variance of the fundamental error is a strong function of the coarse end of the particle size distribution, the ninety fifth percentile, as dictated by a diameter cubed term. The largest particles dominate. Not the average particle. The biggest ones.\n\n### The exchange rate, and it is not close\n\nPut the two together and you get an exchange rate between the two levers. Variance is proportional to grain size cubed and inversely proportional to mass.\n\nHalve the grain size. The cubed term becomes one eighth of what it was, so the variance falls to one eighth.\n\nNow try to achieve that same one eighth by adding sample instead. Variance is inversely proportional to mass, so you need eight times the mass.\n\n| Move | Effect on variance | Effect on standard deviation |\n| --- | --- | --- |\n| Double the sample | Falls to 1/2 | Falls to 0.7071, a 29 percent improvement |\n| Quadruple the sample | Falls to 1/4 | Falls to 1/2 |\n| Halve the grain size | Falls to 1/8 | Falls to about 1/2.83 |\n| Eightfold the sample | Falls to 1/8 | Falls to about 1/2.83 |\n\nThe last two rows are algebraically identical. **Halving the grain of your material buys exactly what an eightfold increase in sample buys.** That is not an approximation or a rule of thumb, it falls straight out of the exponents.\n\n## What grain size means in a customer sample\n\nThe analogy is only useful if the translation is honest, so here it is precisely. Grain size is not variance. It is the size of the largest discrete chunk of difference in your material.\n\nA population of 5,000 SMB users who all do broadly the same job with your product is fine grained. It varies, but it varies smoothly. A population of 4,800 SMB users plus 200 enterprise accounts with a fundamentally different workflow is coarse grained, and the enterprise accounts are the ninety fifth percentile particles that dominate the variance term. Whether one of them lands in your twelve interviews changes your answer more than anything else about the study.\n\nThat is the practical content of the diameter cubed term: **your uncertainty is set by your most unlike-the-others users, not by your typical ones.**\n\n### The lever this hands you\n\nYou cannot make your customers more alike. You can, however, stop treating them as one material. Sampling the enterprise accounts and the SMB accounts as two separate populations, each with its own sample and its own reported number, reduces the grain within each one. In sampling theory this is stratification, and the exchange rate above is why it is not a nicety. It is the cheap lever.\n\nConcretely, if you have been running a single monthly study of 20 interviews across a mixed base, splitting it into two studies of 10 within homogeneous strata can leave you better off on precision than tripling the single study would have. The arithmetic says halving the grain is worth eight times the mass; even a partial reduction in grain beats a doubling.\n\n## Why more interviews is the expensive lever\n\nThe square root law is unforgiving and every researcher has felt it without naming it. Going from 10 interviews to 20 improves your standard error by 29 percent. Going from 10 to 40 halves it. Going from 10 to 1,000 improves it tenfold, and nobody is running 1,000 interviews to answer a roadmap question. Even where a platform like Koji makes large samples affordable, the square root law still governs what they buy.\n\nThis is why the perennial question of how many interviews are enough has no satisfying answer when the sample is heterogeneous. The honest answer is that at a given grain, past a certain point additional interviews buy very little, and the money is better spent on making the strata tighter. [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) works the sample size question directly; this article is about the term that sits underneath it and that adding participants cannot touch.\n\n## The claim that this disproves\n\nThere is a common and comfortable belief that a rigorous analysis can compensate for a thin or lopsided sample. Careful coding, a second coder, a well specified codebook, a thoughtful synthesis workshop. All of those are worth doing, and all of them operate downstream of the fundamental error.\n\nNo analysis reads information that is not in the material you collected. If your twelve interviews happened to contain zero enterprise accounts, no coding rubric recovers the enterprise perspective, and a very rigorous analysis of those twelve will produce a confidently wrong answer with excellent provenance. The rigour is real; it is applied to the wrong stage. This is also why Koji keeps every reported theme tied to the conversation it came from: good provenance is a property of the analysis, and it cannot certify the sample underneath it. If you want the error budget of a combined metric rather than a single sample, [Error Propagation in Research Metrics](/docs/error-propagation-derived-research-metrics) covers what happens to uncertainty when you start combining numbers.\n\n## Common mistakes\n\n- **Treating precision as an analysis problem.** More careful coding of the same transcripts cannot reduce sampling error. It reduces a different error.\n- **Reporting one number for a coarse population.** A single satisfaction figure across a base with two genuinely different user types is an average of two materials.\n- **Sizing on the typical user.** The variance is driven by the outliers at the coarse end of the distribution, so size for whether you can see them at all.\n- **Assuming a bigger sample fixes a biased one.** It does not. A larger sample from the wrong frame is a more confident wrong answer, which is why [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) matters independently of n.\n- **Stratifying after the fact.** Slicing a mixed sample into segments post hoc does not reduce the grain you sampled at; it just produces small segments. [Why the Top and Bottom Segments in Your Report Are Both the Smallest Ones](/docs/segment-ranking-sample-size-artifact) shows the artifact this creates.\n\n## How Koji changes which lever is cheap\n\nThe exchange rate above is a fact about arithmetic, not about tooling. What tooling changes is the *price* of each lever, and that is what decides which one a team actually pulls.\n\nStratifying properly means running several smaller studies instead of one big one, each with its own screener and its own recruitment. Under a traditional model that is the expensive option, because every study carries a fixed cost of a researcher's calendar: scheduling, moderating, transcribing and synthesising each cohort separately. Faced with three studies of 10 or one study of 30, most teams correctly conclude that the single study is affordable and the stratified design is not, and then they absorb the grain.\n\nKoji inverts that price. Because the AI interviewer runs conversations asynchronously and in parallel, a study of 10 costs no moderator time, and three studies of 10 cost no more moderator time than one. Each cohort gets its own study with its own screening criteria, and each produces its own report rather than being averaged into a pooled figure. The cheap lever becomes the cheap option in practice as well as in theory.\n\nTwo further pieces of this matter for the grain question specifically. Koji's structured questions give you six typed question formats (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), so the same quantitative question can be asked identically across strata and compared without a coding pass in between, which is what makes stratum-level numbers commensurable at all. And because the AI asks its own follow-up questions, a thin stratum still yields depth: 10 well probed conversations in a homogeneous cohort carry more usable signal than 30 shallow ones spread across a mixed base.\n\nKoji does not abolish the fundamental error. Nothing does. It makes the one move that provably reduces it affordable at the sample sizes product teams actually work at.\n\n## Frequently asked questions\n\n### Does this mean sample size does not matter?\n\nNo. Sample size matters and it is one of only two levers in the model. The point is that it is the weaker of the two: variance falls in proportion to mass, while it falls with the cube of the grain. Both are worth pulling. If you can only pull one, pull the grain.\n\n### Is this the same thing as stratified sampling?\n\nStratification is the practical technique this model recommends, so they point the same way, but the argument is different. Standard treatments of stratified sampling justify it by wanting representation of each group. The sampling theory argument is quantitative and stronger: reducing within-stratum heterogeneity attacks a cubed term while adding participants attacks a linear one.\n\n### How do I know how coarse my population is?\n\nLook for discrete clusters rather than smooth spread. If you can name a group of users whose workflow is categorically different from the rest, that group is your coarse particle, and the relevant question is whether your sample size can reliably include it. Interviewing a handful from each named group and comparing is the cheapest diagnostic.\n\n### Can a better AI analysis reduce sampling error?\n\nNo, and this is the clean dividing line. Analysis quality governs how accurately you read the material you collected. Sampling error governs how well that material represents the population. A perfect analysis of an unrepresentative sample is precisely as wrong as a sloppy one, just harder to argue with.\n\n### Where does this leave qualitative research that never claimed to be representative?\n\nIn good shape, provided the claim matches the design. Exploratory interviews that surface possibilities do not need a representativeness claim. The trouble starts when a finding from an unstratified dozen conversations gets reported as a property of the user base. The fix is to state what the sample can support, not to abandon small samples.\n\n### Does the diameter cubed relationship transfer exactly to people?\n\nNo, and it should not be asserted as though it does. The cubed term comes from the geometry of particles, where mass scales with volume. Customers have no such geometry. What transfers is the structural lesson, which does not depend on the exponent: heterogeneity enters the variance far more aggressively than sample size does, and the extremes dominate. Treat the eightfold figure as an illustration of that asymmetry within the physical model, not as a measured constant for human populations.\n\n## Related Resources\n\n- [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) - the sample size question this article sits underneath\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that make stratum-level numbers comparable\n- [Error Propagation in Research Metrics](/docs/error-propagation-derived-research-metrics) - what happens to uncertainty when you combine numbers\n- [Heterogeneous Treatment Effects](/docs/heterogeneous-treatment-effects-research) - why nobody experienced your average\n- [Quota Sampling Guide](/docs/quota-sampling-guide) - the practical mechanics of sampling within strata\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - why a bigger sample does not fix a biased frame\n- [The List Experiment](/docs/list-experiment-item-count-research) - a design that fixes misreporting and still cannot fix the sample\n","category":"Research Methods","lastModified":"2026-09-30T03:40:15.729055+00:00","metaTitle":"Why a Better Analysis Cannot Rescue a Bad Sample (2026)","metaDescription":"Sampling error has a floor no analysis can cross. Learn the two levers that move it and why narrowing your sample beats adding interviews.","keywords":["sampling error qualitative research","fundamental sampling error","heterogeneous sample interviews","stratified sampling research","sample size vs sample quality"],"aiSummary":"Sampling error contains a component (the fundamental error in Gy sampling theory) that survives a perfect procedure. Its variance scales with grain size cubed over sample mass, so halving population heterogeneity is algebraically equal to an eightfold sample increase. Analysis quality cannot reduce it; stratification can.","aiPrerequisites":["Basic familiarity with sample size reasoning","Understanding of variance and standard deviation"],"aiLearningOutcomes":["Distinguish sampling error from analysis error","Compute the exchange rate between sample size and heterogeneity","Identify the coarse-grained subgroups that dominate your uncertainty","Design stratified studies that attack the dominant variance term"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}