{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-05T09:56:25.872Z"},"content":[{"type":"documentation","id":"73842203-f390-4d92-afd5-dbd0b6409db4","slug":"statistical-power-minimum-detectable-effect","title":"Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)","url":"https://www.koji.so/docs/statistical-power-minimum-detectable-effect","summary":"Minimum detectable effect (MDE) is the smallest true difference a study can reliably identify, given sample size, confidence level and power. For two independent waves of equal size at 95% confidence and 80% power, MDE is approximately 2.02 times the single-wave margin of error, because comparing two noisy waves multiplies the standard error by the square root of 2 and power requires 2.80 standard errors rather than 1.96. Practical figures: 400 responses per wave gives a ±4.9-point margin of error but a 9.9-point MDE for percentages, a 0.50-point MDE on a 0-10 scale with SD 2.5, and a 14.8-point MDE for NPS at 40% promoters and 20% detractors. Detecting a 5-point percentage change needs about 1,570 responses per wave. Power is destroyed by segmentation (cell n, not total n), multiple comparisons (20 tests at 5% yields a 64% chance of a false positive), and independent rather than repeated-measures waves; a wave-to-wave correlation of 0.5 cuts MDE by about 29%. Moving from 80% to 90% power costs about 34% more sample.","content":"**Short answer: the smallest change your survey can reliably detect is about *twice* its margin of error. A 400-response wave has a ±4.9-point margin of error, which teams read as \"we can spot a 5-point move\" — but the minimum detectable effect for a wave-over-wave comparison at that sample size is 9.9 points. Anything smaller is invisible, and running the study anyway means paying for an answer you cannot get.**\n\nThis is the most expensive misunderstanding in survey research, and it is almost always discovered too late: after the tracker has run for four quarters, when someone asks why the number keeps bouncing around and nobody can say whether the product changes did anything.\n\nThree related concepts get confused. This guide separates them, gives you the numbers, and shows you what to do when the sample you can afford is not the sample you need.\n\n## The three questions, and which one you are actually asking\n\n| Question | Concept | Covered in |\n|---|---|---|\n| How precise is this single number? | Margin of error | [Margin of Error in Surveys](/docs/survey-margin-of-error-guide) |\n| Is this observed difference real, or noise? | Statistical significance | [Statistical Significance in Survey Research](/docs/statistical-significance-survey-research) |\n| **How big would a change have to be before I could see it at all?** | **Statistical power and minimum detectable effect** | This guide |\n\nMargin of error and significance are both **backward-looking**: you have the data, and you are describing it. Power and MDE are **forward-looking**: they tell you, before you spend anything, whether the study is capable of answering the question. Skipping this step is how teams end up with an underpowered study — one that had almost no chance of detecting the effect it was commissioned to find, and which therefore produces a \"no significant difference\" result that means nothing at all.\n\n## Statistical power in one paragraph\n\n**Power** is the probability that your study finds a real effect, given that the effect exists. Convention is 80% power — meaning that if the change you care about is genuinely there, you have an 80% chance of detecting it and a 20% chance of missing it. Turning that around: an 80%-powered study still fails one time in five even when it is right about the world.\n\n**Minimum detectable effect (MDE)** is the flip side. Fix your sample size, your confidence level and your power, and the arithmetic hands you the smallest true difference the study can reliably pick up. Everything below the MDE is beneath the study's resolution.\n\nFor the standard case — 95% confidence, two-sided, 80% power — the multiplier is **2.80** (the sum of 1.96 and 0.84, the z-values for those two thresholds).\n\n## The rule that fixes most of the confusion\n\nFor two independent waves of equal size:\n\n> **MDE ≈ 2 × margin of error.**\n\nThe precise ratio is 2.02, and it comes from two compounding penalties. First, comparing two waves means both are noisy, which multiplies the standard error by √2. Second, power costs more than confidence does: you need 2.80 standard errors rather than the 1.96 used in a margin of error. Multiply those together and you get roughly double.\n\nSo when a stakeholder points at ±5 and says the study can see a 5-point move, the honest answer is: it can see a 10-point move.\n\n## MDE tables you can use directly\n\n**Percentages (two waves, equal n per wave, 95% confidence, 80% power, worst case at 50%)**\n\n| Responses per wave | Margin of error (one wave) | Minimum detectable change |\n|---|---|---|\n| 100 | ±9.8 pts | 19.8 pts |\n| 200 | ±6.9 pts | 14.0 pts |\n| 400 | ±4.9 pts | 9.9 pts |\n| 600 | ±4.0 pts | 8.1 pts |\n| 1,000 | ±3.1 pts | 6.3 pts |\n| 2,000 | ±2.2 pts | 4.4 pts |\n| 4,000 | ±1.5 pts | 3.1 pts |\n\nTo go the other way: to detect a 5-point change in a percentage you need roughly **1,570 responses per wave**. To detect 3 points, about 4,350.\n\n**0–10 rating scales (assuming a standard deviation of 2.5, typical for satisfaction items)**\n\n| Responses per wave | Minimum detectable change in the mean |\n|---|---|\n| 100 | 0.99 points |\n| 200 | 0.70 points |\n| 400 | 0.50 points |\n| 1,000 | 0.31 points |\n| 2,000 | 0.22 points |\n\n**Net Promoter Score (assuming 40% promoters, 20% detractors, so NPS = +20)**\n\nNPS is noisier than people expect, because it is a difference between two proportions and inherits the variance of both.\n\n| Responses per wave | Margin of error on NPS | Minimum detectable NPS change |\n|---|---|---|\n| 200 | ±10.4 | 20.9 points |\n| 400 | ±7.3 | 14.8 points |\n| 1,000 | ±4.6 | 9.4 points |\n| 2,000 | ±3.3 | 6.6 points |\n| 4,000 | ±2.3 | 4.7 points |\n\nRead that middle row again. **With 400 responses per wave — a very typical tracker — the smallest NPS movement you can reliably detect is about 15 points.** Almost every quarterly NPS conversation in the industry is about movements far smaller than that. See [NPS Benchmarks by Industry](/docs/nps-benchmarks-by-industry-2026) for what real movements look like, and [Brand Tracking Studies](/docs/brand-tracking-study-guide) for wave design.\n\n## Four things that quietly destroy your power\n\n**1. Segmentation.** Power is driven by the n in each cell, not the total. Split 1,000 responses across five segments and each segment has 200 — an MDE of 14 points, not 6.3. If the study exists to compare segments, size for the *smallest* segment you intend to report.\n\n**2. Multiple comparisons.** Test 20 segments at the 5% threshold and the probability of at least one false positive is 64%, not 5%. Trackers that slice by region, plan, tenure and platform every quarter are manufacturing significant-looking noise. Decide your comparisons in advance and correct for the rest.\n\n**3. Independent waves instead of the same people.** Re-surveying the same panel makes the two waves correlated, which cancels part of the noise. At a wave-to-wave correlation of 0.5 the MDE drops by roughly 29% — the same benefit as doubling your sample, for free. This is the single cheapest power upgrade available to a tracker.\n\n**4. Raising power without raising sample.** Moving from 80% to 90% power requires about **34% more sample**. If someone wants more certainty, that is the price.\n\n## What to do when you cannot afford the sample\n\nMost teams read the NPS table, discover they need 2,000 responses per wave to see a 7-point move, and conclude the research is impossible. It isn't — the design is just wrong for the question.\n\n**Stop trying to detect small changes in a summary metric.** A 5-point NPS shift is not a finding; it is a number that will move again next quarter. Size your tracker for the movements that would actually change a decision — usually 10 points or more — and stop reporting anything below the MDE as if it were signal.\n\n**Move the question from \"did it move\" to \"why\".** Detecting a 3-point satisfaction change requires thousands of responses. Understanding *what changed for customers* requires depth, not volume. Forty conversations that probe why churned customers left will change more decisions than a tracker sized to prove that satisfaction fell by 0.2.\n\n**Use a longitudinal panel.** Same people, repeated waves. It buys the equivalent of double the sample.\n\n**Reduce the variance instead of increasing the n.** Tighter question wording, fewer response categories misused as scales, and cleaner sampling all shrink the standard deviation, and MDE scales directly with it. See [How to Write Unbiased Survey Questions](/docs/survey-question-wording-guide) and [Survey Data Quality](/docs/survey-data-quality-guide) — every low-effort response you filter out is variance you did not have to pay for.\n\n**Pre-register the effect you care about.** Write down, before fielding: \"we will act if the metric moves by X\". If X is below your MDE, redesign the study now rather than explaining an inconclusive result later.\n\n## How Koji changes this arithmetic\n\nThe power problem is really a cost problem: statistical resolution scales with the square of your sample, so every halving of the MDE costs four times the sample. Traditional research responds by buying more panel — the most expensive possible answer.\n\nKoji attacks it from the other side.\n\n**Every interview yields more per respondent.** A Koji study is a conversation, not a form. The AI interviewer asks the structured question, then probes the answer — so a single participant produces both a countable data point and the reasoning behind it. That reasoning is what lets a team act on a movement that is too small to be statistically certain, because they know the mechanism rather than just the delta.\n\n**Structured questions make the quantitative side rigorous.** Koji supports six question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — and aggregates them automatically into distributions and charts, so your MDE arithmetic applies to real structured data rather than to hand-coded open text. See [Structured Questions Guide](/docs/structured-questions-guide).\n\n**Consistent administration lowers variance.** Human moderators drift: they rephrase, they prompt unevenly, they skip items when a session runs long. That drift is variance, and variance is MDE. An AI interviewer asks every participant the same question the same way, every time, at any hour — which tightens the distribution and improves your effective power without a single extra response.\n\n**Cost per participant is low enough to size properly.** Koji charges 1 credit for a text interview and 3 for a voice interview, and only conversations that pass the platform's quality bar consume credits at all — so the sample you need to hit your MDE is achievable rather than theoretical, and low-effort responses do not eat your budget or inflate your variance.\n\nThe practical pattern that works: run a properly sized structured study to establish whether the number moved, and let the same participants' open-ended answers explain why. One study, both halves of the question — which is the thing a survey tool and an interview platform used separately can never quite deliver.\n\n## Frequently asked questions\n\n**What is the difference between margin of error and minimum detectable effect?**\nMargin of error describes the precision of a single estimate at one point in time. Minimum detectable effect describes the smallest *difference* between two estimates that your study could reliably identify. For two independent waves of equal size, the MDE is roughly twice the margin of error — so a ±4.9-point margin of error corresponds to a 9.9-point detectable change.\n\n**How many responses do I need to detect a 5-point change?**\nFor a percentage measured at around 50%, about 1,570 responses per wave, at 95% confidence and 80% power. For a 3-point change, about 4,350 per wave. Sample requirements grow with the square of the precision you want, which is why chasing small movements gets expensive so fast.\n\n**Why is NPS so hard to move statistically?**\nBecause NPS is the difference between two proportions, it carries the sampling variance of both promoters and detractors. With 400 responses per wave, the minimum detectable NPS change is about 15 points — far larger than the quarter-to-quarter movements most teams discuss in review meetings.\n\n**What does 80% power actually mean?**\nIt means that if the effect you care about genuinely exists at the size you specified, you have an 80% chance of detecting it in this study and a 20% chance of missing it. Raising power to 90% requires roughly 34% more sample.\n\n**Does an inconclusive result mean there was no change?**\nNo, and this is the most common misreading. A non-significant result in an underpowered study means the study could not tell — not that nothing happened. That is precisely why the MDE should be calculated before fielding: it converts \"we found nothing\" into the far more useful \"we could not have found anything smaller than X\".\n\n**Do these calculations apply to qualitative interviews?**\nNot directly. Power analysis assumes you are estimating a number from a sample. Qualitative studies are sized by information coverage rather than statistical power — see [How Many User Interviews Do You Need?](/docs/how-many-user-interviews). Where the two meet is a platform like Koji that runs structured questions and open-ended probing in the same conversation: the structured half obeys the arithmetic on this page, and the qualitative half explains the result.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that produce analysable quantitative data\n- [Margin of Error in Surveys](/docs/survey-margin-of-error-guide) — precision for a single estimate\n- [Statistical Significance in Survey Research](/docs/statistical-significance-survey-research) — testing a difference you have already observed\n- [Survey Sample Size: How Many Responses Do You Really Need?](/docs/survey-sample-size-guide) — sizing from the precision side\n- [Brand Tracking Studies](/docs/brand-tracking-study-guide) — wave design and measuring change over time\n- [NPS Benchmarks by Industry 2026](/docs/nps-benchmarks-by-industry-2026) — what a meaningful NPS movement looks like\n- [How Many User Interviews Do You Need?](/docs/how-many-user-interviews) — the qualitative counterpart to sample size\n\n*Want depth and numbers from the same study? [Start free with 10 credits](https://www.koji.so) — text interviews cost 1 credit, and only quality conversations consume them.*","category":"Research Methods","lastModified":"2026-08-04T03:25:22.474534+00:00","metaTitle":"Statistical Power & Minimum Detectable Effect: Can Your Survey See the Change? (2026)","metaDescription":"MDE is roughly twice your margin of error. Tables for percentages, rating scales and NPS, plus what to do when you cannot afford the sample you need.","keywords":["minimum detectable effect survey","statistical power survey","how big a change can my survey detect","underpowered survey","wave over wave significance","nps sample size","mde calculation"],"aiSummary":"Minimum detectable effect (MDE) is the smallest true difference a study can reliably identify, given sample size, confidence level and power. For two independent waves of equal size at 95% confidence and 80% power, MDE is approximately 2.02 times the single-wave margin of error, because comparing two noisy waves multiplies the standard error by the square root of 2 and power requires 2.80 standard errors rather than 1.96. Practical figures: 400 responses per wave gives a ±4.9-point margin of error but a 9.9-point MDE for percentages, a 0.50-point MDE on a 0-10 scale with SD 2.5, and a 14.8-point MDE for NPS at 40% promoters and 20% detractors. Detecting a 5-point percentage change needs about 1,570 responses per wave. Power is destroyed by segmentation (cell n, not total n), multiple comparisons (20 tests at 5% yields a 64% chance of a false positive), and independent rather than repeated-measures waves; a wave-to-wave correlation of 0.5 cuts MDE by about 29%. Moving from 80% to 90% power costs about 34% more sample.","aiPrerequisites":["Familiarity with margin of error and confidence levels","A metric you intend to track over time","A defined decision threshold — how big a change would change your action"],"aiLearningOutcomes":["Distinguish margin of error, statistical significance and minimum detectable effect","Apply the rule that MDE is roughly twice the margin of error","Read MDE off tables for percentages, rating scales and NPS","Identify the four design choices that silently destroy statistical power","Redesign a study when the required sample is unaffordable, instead of running it underpowered"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}