{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-20T13:03:33.372Z"},"content":[{"type":"documentation","id":"1377abaf-940b-48c7-9f8e-3bf7a4345694","slug":"measurement-system-analysis-research-metrics","title":"Measurement System Analysis: How Much of Your Segment Difference Is the Instrument? (2026)","url":"https://www.koji.so/docs/measurement-system-analysis-research-metrics","summary":"Measurement System Analysis separates the variance in a research metric into real between-customer variation and variation created by the measurement process. The key statistic is the intraclass correlation: the share of observed variance that is real. Above 0.80 the instrument attenuates real signal by less than 10 percent (a First Class Monitor); below 0.20 more than 55 percent is attenuated and improvement cannot be tracked at all. Measurement error usually makes real differences look smaller, so a noisy instrument is dangerous for declaring that two segments are equivalent. Increasing sample size does not fix measurement error. Probable error (0.675 times the measurement-error standard deviation) sets how many digits are worth reporting.","content":"**Short answer:** every number your research produces is the sum of two things - real variation between the people you measured, and variation created by the act of measuring. Measurement System Analysis (MSA) is the discipline that separates them. The single most useful output is the *intraclass correlation*: the share of your observed variance that is real. Above 0.80, your instrument passes real differences through almost intact. Below 0.20, more than half the real signal is attenuated away and no amount of extra sample will bring it back. Most product teams have never computed this number for any metric they ship, which is why segment comparisons get argued about for months without resolution.\n\nManufacturing solved this problem in the 1960s and product research never imported the solution. This guide does the import.\n\n## The variance you report is two variances added together\n\nWhen you run a satisfaction study and find that enterprise customers score 7.4 and SMB customers score 6.8, you have observed a 0.6-point gap. You want to treat that gap as a fact about your customers. It is not. It is a fact about your customers *plus* your instrument.\n\nThe arithmetic is simple and it is the whole foundation of the field:\n\n```\nobserved variance = product variance + measurement variance\nsigma_x^2 = sigma_p^2 + sigma_e^2\n```\n\nThe ratio of the real part to the whole is the *intraclass correlation coefficient*, a statistic that goes back to Ronald Fisher in 1921:\n\n```\nintraclass correlation (rho) = sigma_p^2 / sigma_x^2\n```\n\nIf rho is 0.90, ninety percent of the spread you are looking at is real and ten percent is your instrument talking. If rho is 0.15, you are mostly reading your own noise back to yourself and calling it a customer insight.\n\nThe consequence that matters is *attenuation*. A measurement system with meaningful error does not just add fuzz around the true difference - it systematically shrinks the difference you observe. Real gaps look smaller than they are. This is why so many product teams conclude that \"the segments are basically the same\" and ship a one-size-fits-all experience: the segments were different, and the instrument flattened them.\n\n## The four classes of monitor\n\nThe most practical framework here comes from Donald J. Wheeler, whose 2006 ASQ/ASA Fall Technical Conference paper *An Honest Gauge R&R Study* is freely available and is the clearest treatment of the subject anywhere. Wheeler argues that the widely used automotive-industry guidelines (the AIAG categories of Good, Marginal and Unacceptable) are \"excessively conservative\" - they effectively demand an intraclass correlation of 0.99 or better to call a measurement system good, and condemn almost everything else. He replaces them with four classes that describe what a measurement system can actually *do*.\n\n| Class | Intraclass correlation | Attenuation of real signal | What you can still do with it |\n|---|---|---|---|\n| First Class Monitor | 1.00 to 0.80 | Less than 10% | Detect a three-standard-error shift more than 99% of the time using the single-point rule |\n| Second Class Monitor | 0.80 to 0.50 | 10% to 30% | Detect the same shift more than 88% of the time using the single-point rule |\n| Third Class Monitor | 0.50 to 0.20 | 30% to 55% | Detect the same shift more than 91% of the time, but only with the full set of run rules |\n| Fourth Class Monitor | Below 0.20 | More than 55% | Detection \"rapidly vanishes\"; unable to track improvement at all |\n\nTwo things in that table are worth sitting with.\n\nFirst, a Second Class Monitor is *usable*. A measurement system that is losing 20% of your real signal is not a scandal - it is a normal working instrument, and Wheeler's point is that condemning it wastes money that would be better spent elsewhere. The reason to compute the number is not to pass an audit. It is so you know how much of a difference you have to see before you believe it.\n\nSecond, the Fourth Class boundary is where the honest answer becomes \"stop\". Below rho = 0.20, more than 55% of any real change is attenuated away, and you cannot track whether an improvement worked. Teams in this position typically respond by collecting more responses. More sample tightens the confidence interval around a number that is still more than half instrument. It does not help.\n\n## What counts as an \"operator\" in product research\n\nIn a factory, a gauge R&R study measures the same parts repeatedly, with several different operators, and decomposes the variance into part-to-part, repeatability (same operator, same part, different trial) and reproducibility (different operators, same part). The vocabulary maps onto research more cleanly than most people expect.\n\n| Metrology term | The research equivalent | Where the variance comes from |\n|---|---|---|\n| Part | The customer, account, or session being measured | The thing you actually care about |\n| Repeatability | The same respondent, asked the same question, twice | Momentary state, attention, recall instability |\n| Reproducibility (operator) | A different interviewer, coder, analyst, or question wording | The asker, not the answerer |\n| Gauge | The question set, scale, and coding scheme | The instrument itself |\n| Measurement increment | The number of digits you report | Resolution: reporting 7.42 when the probable error is 0.9 |\n\nThe International Vocabulary of Metrology (VIM, JCGM 200:2012) makes the distinction crisp. A *repeatability condition of measurement* is one that \"includes the same measurement procedure, same operators, same measuring system, same operating conditions and same location, and replicate measurements on the same or similar objects over a short period of time\". Change the operator and you are no longer measuring repeatability - you are measuring reproducibility, which is nearly always the larger term.\n\nIn research, the \"operator\" is usually invisible. Nobody records which interviewer ran which session, or which analyst coded which transcript, so the operator component never gets estimated and is silently assumed to be zero. It is not zero. The variance a single interviewer introduces has its own dedicated treatment in [the AI interviewer house effect](/docs/ai-interviewer-house-effect); what MSA adds is the arithmetic that tells you whether that variance is large *relative to the differences you want to act on*.\n\n## Running an honest R&R study on a research metric\n\nYou do not need a psychometrics team. You need to measure some of the same things twice, on purpose.\n\n1. **Pick the metric and the decision.** \"Enterprise vs SMB satisfaction, used to decide whether to build a separate enterprise onboarding flow.\" A metric with no decision attached does not need an R&R study, it needs deleting.\n2. **Choose 10 to 20 units that span the real range.** Not a random sample - a deliberate spread, from your happiest accounts to your angriest. The R&R study needs real part-to-part variation to compare against, and a sample of near-identical units will make any instrument look terrible.\n3. **Measure each unit at least twice, under repeatability conditions.** Same question, same mode, short interval.\n4. **Vary one operator dimension.** Two interviewers, or two coders, or two phrasings of the same question. One dimension per study; you can run more later.\n5. **Decompose the variance.** A two-way ANOVA gives you the part, repeatability and reproducibility components directly. The intraclass correlation is the part component divided by the total.\n6. **Classify the monitor.** Read the class off the table above and write it down next to the metric in your documentation.\n7. **Compute the probable error and fix your reporting precision.** See the next section.\n\nWheeler's honest procedure runs to thirteen steps; the seven above are the version that survives contact with a product team, and they get you the number that changes decisions.\n\n## Probable error: stop reporting digits your instrument cannot resolve\n\nThe *probable error* is defined as 0.675 times the standard deviation of pure measurement error - it is the median amount by which any single measurement will be wrong. Half your measurements err by less than this; half err by more.\n\nIt gives you a rule with immediate practical bite. The smallest useful measurement increment is 0.2 probable errors and the largest is 2 probable errors. Report more precision than that and the extra digits are decoration.\n\nWork an example. Suppose you re-ask a 0-10 satisfaction question a week apart and the standard deviation of the differences implies a measurement-error standard deviation of about 1.3 points. The probable error is 0.675 x 1.3 = 0.88 points. Your useful reporting increment sits between 0.18 and 1.76 points. So \"satisfaction is 7.4, up from 7.2\" is not a finding. It is a rounding artefact presented as a trend, and the whole disagreement it will cause in the next review is manufactured.\n\nThis one calculation, applied to the three or four numbers your organisation argues about most, retires more bad meetings than any dashboard redesign.\n\n## What this changes about segment comparisons\n\nThe most common serious error in product research is comparing two groups whose observed difference is smaller than the measurement error of the instrument, and then reasoning about *why* they differ.\n\nBefore you explain a gap, check that the gap survives your instrument. Three questions, in order:\n\n- **Is the observed gap larger than one probable error?** If not, you have nothing to explain.\n- **What is the intraclass correlation of this metric?** If it is below 0.50, the true gap is meaningfully larger than the one you measured, and any effect size you quote is an underestimate.\n- **Does the instrument mean the same thing to both groups?** This is a different failure from noise, and it has its own test - see [measurement invariance](/docs/measurement-invariance-comparing-groups). A metric can have a superb intraclass correlation and still be uncomparable across segments.\n\nNote the asymmetry, because it is the least intuitive part of the whole subject: measurement error usually makes real differences look *smaller*, not larger. A noisy instrument is a conservative one for detecting differences and a dangerous one for declaring equivalence. \"We tested it and the segments were the same\" is the claim most likely to be an artefact of a Third or Fourth Class monitor.\n\n## How Koji makes the R&R study cheap\n\nThe reason almost nobody runs measurement system analysis on research metrics is not ignorance. It is that measuring the same thing twice, with two different askers, has historically meant twice the recruiting, twice the moderator time and twice the analysis. The economics never worked.\n\nAI-moderated interviews change the arithmetic, because the expensive human is no longer in the loop:\n\n- **The operator is version-pinned and identical.** Every Koji interview is run by the same AI interviewer, which removes the largest uncontrolled reproducibility component in traditional research - the human moderator having a good or bad day. What remains is measurable rather than mysterious.\n- **[Structured questions](/docs/structured-questions-guide) give you a stable gauge.** All six types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` - carry stable question IDs from the interview plan through to the report, so the same item can be compared across waves and across studies without hand-matching. A `scale` question produces a numeric distribution you can actually decompose; a `ranking` question produces average positions; `single_choice` and `multiple_choice` produce frequency distributions; `yes_no` produces a proportion. The `open_ended` type is where the AI follow-up probing happens, and its codes are what you double-code in a reproducibility check.\n- **Re-asking is nearly free.** A repeatability wave that would have cost a week of moderator time is a re-run of the same study. This is what makes step 3 above realistic rather than aspirational.\n- **Coding reproducibility is testable.** Because every theme links back to the exact transcript message that produced it, a second pass over the same transcripts is a genuine reproducibility check rather than an exercise in trusting the summary. Compare that with a traditional survey stack, where the coding step happens in a spreadsheet and leaves no trace at all.\n- **The quality gate removes one variance source before you start.** Koji scores conversations for quality and only conversations meeting the bar consume credits, which strips out a class of low-effort responses that would otherwise land squarely in your measurement-error term.\n\nTraditional survey tools - SurveyMonkey, Typeform, Qualtrics - will happily give you a mean to two decimal places. None of them will tell you how many of those decimals are real. That is the gap this analysis fills.\n\n## Frequently asked questions\n\n### What is measurement system analysis in user research?\n\nMeasurement system analysis is the practice of estimating how much of the variation in a research metric comes from real differences between the people or accounts measured, and how much comes from the measurement process itself - the question wording, the interviewer, the coder, the mode. The headline output is the intraclass correlation, the share of observed variance that is real. It is standard practice in manufacturing metrology and almost unknown in product research, which is why so many segment comparisons are irreproducible.\n\n### What is a good intraclass correlation for a research metric?\n\nAbove 0.80 the instrument is a First Class Monitor and attenuates real signal by less than 10%. Between 0.80 and 0.50 it is a Second Class Monitor, losing 10% to 30% - still perfectly usable if you know it. Between 0.50 and 0.20 you need the full set of run rules to detect changes. Below 0.20 more than 55% of any real signal is attenuated and you cannot track improvement at all. The automotive AIAG guidelines are far stricter, effectively demanding 0.99, and Wheeler argues they are excessively conservative and condemn measurement systems that would still do useful work.\n\n### How do I run a gage R&R study on a survey or interview metric?\n\nPick 10 to 20 units that span the real range, measure each at least twice under the same conditions, then vary exactly one operator dimension - two interviewers, two coders, or two phrasings. Decompose the variance with a two-way ANOVA into part, repeatability and reproducibility components, and divide the part component by the total to get the intraclass correlation. The critical design choice is step one: your units must have genuine spread, because the study compares measurement error against real variation and near-identical units will make any instrument look broken.\n\n### Does more sample size fix measurement error?\n\nNo, and this is the most expensive misunderstanding in the area. Increasing your sample size shrinks the standard error of the mean - the uncertainty about *where the average sits*. It does nothing to the measurement variance of each individual reading, so it does not reduce attenuation and does not improve your ability to resolve real differences between units. If your intraclass correlation is 0.15, doubling your sample gives you a tighter estimate of a mostly-noise number. Fix the instrument first, then buy sample.\n\n### What is probable error and how do I use it?\n\nProbable error is 0.675 times the standard deviation of pure measurement error, and it is the median amount by which a single measurement will be wrong. Its practical use is setting reporting precision: the smallest useful measurement increment is 0.2 probable errors and the largest is 2 probable errors. If your probable error on a 0-10 scale is 0.88 points, then reporting a move from 7.2 to 7.4 is reporting noise with a decimal point attached. Computing this once for your three most-argued-about metrics is the highest-return hour available in research operations.\n\n### Is measurement error the same as measurement invariance?\n\nThey are different failures and they need different tests. Measurement error is random noise that attenuates real differences and is diagnosed by measuring the same thing twice. Measurement invariance is about whether a scale means the same thing to two different groups - whether a 7 from an enterprise buyer and a 7 from an SMB user represent the same underlying quantity. An instrument can be extremely precise and still be non-invariant, in which case the comparison is invalid no matter how much data you collect. Test both before you explain a segment gap.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and how stable question IDs let you compare the same item across waves.\n- [The AI Interviewer House Effect](/docs/ai-interviewer-house-effect) - what happens to the reproducibility component when the interviewer count drops to one.\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - the separate test you need before comparing scores across segments, languages, or time.\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - the coding-agreement version of the reproducibility component.\n- [Regression to the Mean](/docs/regression-to-the-mean-research) - why an unreliable instrument guarantees that extreme groups bounce back without any intervention.\n- [Reliability vs Validity](/docs/reliability-vs-validity-research) - the conceptual frame that sits above all of this, and why a precise instrument can still be measuring the wrong thing.\n","category":"Research Methods","lastModified":"2026-08-20T03:25:55.667399+00:00","metaTitle":"Measurement System Analysis for Research Metrics (2026 Guide)","metaDescription":"Separate real customer variation from measurement noise: intraclass correlation, the four classes of monitor, probable error, and how to run a gage R&R on a research metric.","keywords":["measurement system analysis","gage r&r user research","intraclass correlation research metric","measurement error segment differences","probable error reporting precision","research metric noise"],"aiSummary":"Measurement System Analysis separates the variance in a research metric into real between-customer variation and variation created by the measurement process. The key statistic is the intraclass correlation: the share of observed variance that is real. Above 0.80 the instrument attenuates real signal by less than 10 percent (a First Class Monitor); below 0.20 more than 55 percent is attenuated and improvement cannot be tracked at all. Measurement error usually makes real differences look smaller, so a noisy instrument is dangerous for declaring that two segments are equivalent. Increasing sample size does not fix measurement error. Probable error (0.675 times the measurement-error standard deviation) sets how many digits are worth reporting."}],"pagination":{"total":1,"returned":1,"offset":0}}