Nothing Was Measured Twice: Why the Only Error Bar You Can Compute Is the Smallest One
Your margin of error covers sampling and nothing else, because sampling is the only step of a typical study that gets repeated. The Type A and Type B distinction, the definitional floor, and the cheapest ways to buy back replication.
Short answer: the margin of error on your research report is not the error on your research report. It is the only component you were able to calculate, and you were able to calculate it because sampling is the one step in your study that got repeated. Every other error source - the question wording, the mode, the interviewer, the coder, the analyst, the sampling frame - happened exactly once, so it has no statistical estimate, and in practice gets treated as zero. The result is a published uncertainty that is systematically the smallest of the uncertainties present. This guide explains the Type A / Type B distinction that makes the problem precise, and gives you the cheapest ways to buy back the replication you are missing.
This is the capstone of the measurement-quality series. The earlier guides assume you can eventually check your work. This one is about what happens when nothing was measured twice.
Type A and Type B: the distinction nobody in product research uses
The Guide to the Expression of Uncertainty in Measurement (GUM, JCGM 100:2008) is the international standard for stating how wrong a number might be. Its central organising idea is a two-way classification of uncertainty components, based purely on how you evaluated them:
- Type A evaluation: the uncertainty is calculated from a series of repeated observations. It is the familiar statistically estimated standard deviation.
- Type B evaluation: the uncertainty is evaluated by any other means - a calibration certificate, a manufacturer specification, published data, prior measurements, or informed judgement about the physics of the situation.
The GUM is careful that this is a classification of methods, not of error types: "The purpose of the Type A and Type B classification is to indicate the two different ways of evaluating uncertainty components and is for convenience of discussion only; the classification is not meant to indicate that there is any difference in the nature of the components resulting from the two types of evaluation."
Now apply that to a normal product research study.
| Component of the study | Did it happen more than once? | Evaluation available | What most teams report |
|---|---|---|---|
| Sampling from the frame | Yes - every respondent is a draw | Type A | Reported, usually as a margin of error |
| Question wording | No - one wording was chosen | Type B only | Nothing |
| Mode (voice, text, email) | No - usually one mode | Type B only | Nothing |
| Interviewer or moderator | No - or unrecorded | Type B only | Nothing |
| Coding of open responses | No - coded once | Type B only | Nothing |
| Analytic choices | No - one analyst, one path | Type B only | Nothing |
| Sampling frame definition | No - one frame | Type B only | Nothing |
One row has a number attached to it. Six do not. And a Type B component that nobody evaluates does not get carried as an unknown - it gets carried as zero, because that is what leaving it out of the arithmetic means.
So the reported uncertainty is not merely incomplete. It is biased in a knowable direction: it is the smallest component, presented alone.
Why "we did not measure it" is not the same as "it is small"
The GUM is blunt about the limit of any uncertainty statement: "an unrecognized systematic effect cannot be taken into account in the evaluation of the uncertainty of the result of a measurement but contributes to its error."
Two consequences follow, and they are the reason this is a capstone rather than a technique.
First, the missing components are usually larger than the one you have. The reason to suspect this is not ideology - it is that whenever anyone does replicate one of those steps, the spread they find is substantial. Ask the same question two ways and the answers move. Give the same dataset to independent analysts and they reach different conclusions, which is documented in the many-analysts problem. Let people choose their own mode and the distribution shifts, as covered in mode effects. Each of these is a Type B component that becomes Type A the moment somebody bothers to replicate it - and when they do, it rarely comes back small.
Second, even your one Type A number is shakier than it looks. The GUM computes the uncertainty of an estimated standard deviation itself, and the result is sobering: the relative uncertainty in your estimate of the standard deviation is approximately 1 divided by the square root of 2(n-1). For n = 10 observations that is 24 percent; for n = 50 it is 10 percent. The GUM calls this "surprisingly large" and draws the conclusion directly: "Type A evaluations of standard uncertainty are not necessarily more reliable than Type B evaluations, and ... in many practical measurement situations where the number of observations is limited, the components obtained from Type B evaluations may be better known than the components obtained from Type A evaluations."
That inverts the usual instinct. The calculated number is not automatically the trustworthy one. A well-reasoned judgement about how much question wording moves your metric can be more reliable than a standard deviation estimated from a handful of observations.
The floor you cannot replicate past
There is one component that no amount of replication will reduce, and it is worth naming because teams spend money attacking it with the wrong tool.
The International Vocabulary of Metrology (VIM, JCGM 200:2012, entry 2.27) defines definitional uncertainty as the "component of measurement uncertainty resulting from the finite amount of detail in the definition of a measurand", and adds the crucial note: "Definitional uncertainty is the practical minimum measurement uncertainty achievable in any measurement of a given measurand."
In other words, the vagueness of what you are measuring sets a floor. Below that floor, buying more sample, more replicates or a better instrument purchases nothing at all.
"How satisfied are our customers" has enormous definitional uncertainty. Satisfied with what, over what period, weighted how, among which customers? Two competent teams reading that phrase will measure genuinely different quantities, and no methodological rigour applied downstream can close the gap. The VIM adds a second note with a sharp practical edge: "Any change in the descriptive detail leads to another definitional uncertainty." Tightening the definition does not remove the floor - it moves you to a different one, which may be lower.
The full treatment of choosing and defending a definition is construct validity. The point here is narrower and arithmetic: if your definitional uncertainty is 8 points on a 100-point scale, an extra 400 respondents is money set on fire.
Buying replication cheaply
The fix is not more sample. It is deliberately measuring something twice so that a Type B component becomes Type A. Here is the menu, ordered by cost.
| What you replicate | How to buy it | What it converts to Type A | Typical cost |
|---|---|---|---|
| Coding of open responses | Independently re-code a random 20% slice, compare | Coder / analyst variance | Low |
| Question wording | Ask two phrasings of the same construct, split across respondents | Wording variance | Low - a split-ballot design |
| Respondent stability | Re-ask a short subset a week later | Repeatability of the respondent | Low |
| Mode | Run part of the fieldwork in voice and part in text | Mode variance | Medium |
| Whole-study execution | A second team runs the same brief blind | Between-team variance | Medium |
| Sampling frame | Draw from a second panel or source | Frame variance | Medium to high |
Three notes on using this table.
Replicate a slice, not the study. You do not need to double your fieldwork. A 20% double-coded slice gives you a usable estimate of coding variance, and a split-ballot on two phrasings costs nothing but design time - it is the same technique described in split-ballot experiments, used here to estimate an uncertainty rather than to prove a wording effect.
Rotate what you replicate. Pick a different component each quarter rather than trying to nail all six at once. After a year you have a Type A estimate for most of your error budget and a defensible Type B judgement for the rest.
Write the Type B components down even when you cannot estimate them. A line in the report reading "mode variance not evaluated; a prior study put it near 4 points" is enormously more useful than silence, because silence is read as zero. This is the practice the GUM calls an uncertainty budget, and the reason total survey error is worth pairing with this guide - that framework tells you which components exist and how to spend across them; this one tells you which of them you are actually able to quantify.
What to put on the slide
A defensible uncertainty statement for a product research finding has four parts, and none of them require a statistician:
- The point estimate, at a precision the instrument can support.
- The Type A interval, labelled for what it is: "sampling error only".
- A short list of the Type B components you did not evaluate, named.
- Any component you did replicate this cycle, with the number it produced.
The change that follows from this is behavioural, not mathematical. Once "sampling error only" appears next to the interval, nobody in the room can treat that interval as the total. The argument shifts from "is this real" to "what else could move it", which is the argument worth having.
The GUM ends its framework section on exactly this note, and it applies verbatim to research: the guide "cannot substitute for critical thinking, intellectual honesty and professional skill ... The quality and utility of the uncertainty quoted for the result of a measurement therefore ultimately depend on the understanding, critical analysis, and integrity of those who contribute to the assignment of its value."
How Koji changes the economics
Every row of that replication table used to be expensive, which is the honest reason research teams report sampling error alone. The cost structure of AI-moderated research is different in exactly the places that matter.
- Re-running a study is cheap. The single biggest barrier to replication is that a second wave meant a second round of recruiting and moderation. Re-running a Koji study against the same brief is not a new project, which is what makes a repeatability check realistic instead of theoretical.
- Structured questions make the split-ballot design trivial. All six types -
open_ended,scale,single_choice,multiple_choice,rankingandyes_no- carry stable IDs, so running two phrasings of the samescaleitem across split samples produces two directly comparable distributions rather than two datasets someone has to reconcile by hand. - Voice and text are the same study. Because Koji runs both voice and text interviews, splitting fieldwork across modes to estimate mode variance is a configuration choice rather than a second procurement.
- Coding variance is inspectable. Every theme links back to the exact transcript message that produced it, so an independent re-code of a 20% slice is a genuine second measurement and a disagreement can be traced to a specific passage rather than argued abstractly.
- The interviewer is one fewer unestimated component. In traditional research each moderator is an unrecorded, unreplicated Type B source. A single version-pinned AI interviewer does not eliminate that variance, but it makes it a known, measurable property rather than an invisible one - which is the subject of the AI interviewer house effect.
SurveyMonkey, Typeform and Qualtrics will all print a margin of error for you. None of them makes measuring anything twice cheap, and that - not the statistics - is why the smallest error bar is the only one the industry publishes.
Frequently asked questions
What is the difference between Type A and Type B uncertainty?
Type A uncertainty is evaluated by statistical analysis of a series of repeated observations - it is a standard deviation you calculate from data. Type B uncertainty is evaluated by any other means: prior studies, published data, specifications, or informed judgement. The GUM is explicit that this classifies the evaluation method, not the nature of the error, and that neither type is inherently more reliable. In product research, sampling error is almost the only Type A component available, because sampling is the only step of a typical study that gets repeated.
Why is the margin of error not the total error of a survey?
Because it only covers sampling variability - the fact that you talked to these people rather than those people. It says nothing about the effect of how the question was worded, which mode was used, who conducted the interview, how open responses were coded, or which analytic path was taken. Each of those happened once in your study, so no statistical estimate exists for them, and leaving them out of the arithmetic is mathematically identical to assuming they are zero. They are not zero, and when anyone bothers to replicate one of them the spread is usually substantial.
How reliable is a standard deviation estimated from a small sample?
Much less reliable than most people assume. The relative uncertainty of an estimated standard deviation is roughly 1 divided by the square root of 2(n-1), which works out to about 24 percent at n = 10 and about 10 percent at n = 50. The GUM describes this as surprisingly large and concludes that with limited observations, a well-reasoned Type B judgement may actually be better known than a calculated Type A value. The practical implication is that a small-sample error bar deserves no special deference just because it was computed rather than reasoned.
What is definitional uncertainty and why does it matter?
Definitional uncertainty is the component of measurement uncertainty that comes from the finite amount of detail in the definition of what you are measuring. The VIM notes that it is the practical minimum uncertainty achievable in any measurement of that quantity - it is a floor. If "customer satisfaction" is defined loosely enough that two competent teams would measure different things, no amount of sample size, replication or instrument quality will close that gap. Sharpening the definition moves you to a different floor rather than removing it, so the definitional work has to happen before the methodological work.
What is the cheapest way to start estimating non-sampling error?
Independently re-code a random 20 percent slice of your open responses and compare the results. It costs a fraction of a study, it needs no new fieldwork, and it converts coder variance from an unestimated Type B component into a Type A number you can put on a slide. The second cheapest is a split-ballot on two phrasings of your most important item, which costs only design time. Rotate through a different component each quarter rather than attempting to quantify everything at once.
How should I report uncertainty if I only have sampling error?
Report the interval and label it precisely as "sampling error only", then list by name the components you did not evaluate. This is more honest and more useful than either omitting the interval or presenting it as the total, because an unlabelled interval is read as the whole uncertainty and an unnamed component is read as zero. Naming the unevaluated sources moves the conversation from whether the finding is real to what else could move it, which is the discussion that actually improves the next study.
Related Resources
- Structured Questions Guide - the six question types and stable IDs that make split-ballot replication designs cheap to run.
- Total Survey Error - the full taxonomy of error components and how to allocate a fixed budget across them.
- Same Data, Different Answers: The Many-Analysts Problem - what happens when the analytic step actually does get replicated.
- Split-Ballot Experiments - the design that turns question wording from an unestimated component into a measured one.
- Construct Validity - the definitional work that sets the floor no replication can lower.
- Interlaboratory Comparison for Research Teams - buying the between-team component that no internal check can reveal.
Related Articles
Construct Validity: How to Tell Whether You Are Measuring the Thing You Named (2026)
Construct validity is the question of whether your engagement score measures engagement. A guide to operationalization, convergent and discriminant evidence, jingle-jangle fallacies, and the discriminant test that kills most product metrics.
Interlaboratory Comparison for Research Teams: How to Find Out If Your Numbers Are Off
You cannot detect your own systematic bias from inside your own process. How to run the research equivalent of a proficiency testing scheme, including assigned values, z and zeta scores, and what to do when a round fails.
Same Data, Different Answers: The Many-Analysts Problem in Product Research
When 73 teams analyzed identical data to test one hypothesis, over 95 percent of the variance in their results was unexplained. Your analysis is one draw from a distribution you never see.
Measurement System Analysis: How Much of Your Segment Difference Is the Instrument? (2026)
How to separate real variation between customers from variation created by measuring them. The intraclass correlation, the four classes of monitor, probable error, and how to run an honest R&R study on a research metric.
Mode Effects: When Letting People Choose Voice or Text Changes the Answer
Pew randomly assigned 3,003 people to phone or web and got answers that differed by up to 18 points on identical questions. Here is what that means when your respondents pick their own mode.
Split-Ballot Experiments: How Much of Your Number Is the Question?
Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)
Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.