Back to docs
Research Methods

Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)

Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.

Short answer: the smallest change your survey can reliably detect is about twice its margin of error. A 400-response wave has a ±4.9-point margin of error, which teams read as "we can spot a 5-point move" — but the minimum detectable effect for a wave-over-wave comparison at that sample size is 9.9 points. Anything smaller is invisible, and running the study anyway means paying for an answer you cannot get.

This is the most expensive misunderstanding in survey research, and it is almost always discovered too late: after the tracker has run for four quarters, when someone asks why the number keeps bouncing around and nobody can say whether the product changes did anything.

Three related concepts get confused. This guide separates them, gives you the numbers, and shows you what to do when the sample you can afford is not the sample you need.

The three questions, and which one you are actually asking

QuestionConceptCovered in
How precise is this single number?Margin of errorMargin of Error in Surveys
Is this observed difference real, or noise?Statistical significanceStatistical Significance in Survey Research
How big would a change have to be before I could see it at all?Statistical power and minimum detectable effectThis guide

Margin of error and significance are both backward-looking: you have the data, and you are describing it. Power and MDE are forward-looking: they tell you, before you spend anything, whether the study is capable of answering the question. Skipping this step is how teams end up with an underpowered study — one that had almost no chance of detecting the effect it was commissioned to find, and which therefore produces a "no significant difference" result that means nothing at all.

Statistical power in one paragraph

Power is the probability that your study finds a real effect, given that the effect exists. Convention is 80% power — meaning that if the change you care about is genuinely there, you have an 80% chance of detecting it and a 20% chance of missing it. Turning that around: an 80%-powered study still fails one time in five even when it is right about the world.

Minimum detectable effect (MDE) is the flip side. Fix your sample size, your confidence level and your power, and the arithmetic hands you the smallest true difference the study can reliably pick up. Everything below the MDE is beneath the study's resolution.

For the standard case — 95% confidence, two-sided, 80% power — the multiplier is 2.80 (the sum of 1.96 and 0.84, the z-values for those two thresholds).

The rule that fixes most of the confusion

For two independent waves of equal size:

MDE ≈ 2 × margin of error.

The precise ratio is 2.02, and it comes from two compounding penalties. First, comparing two waves means both are noisy, which multiplies the standard error by √2. Second, power costs more than confidence does: you need 2.80 standard errors rather than the 1.96 used in a margin of error. Multiply those together and you get roughly double.

So when a stakeholder points at ±5 and says the study can see a 5-point move, the honest answer is: it can see a 10-point move.

MDE tables you can use directly

Percentages (two waves, equal n per wave, 95% confidence, 80% power, worst case at 50%)

Responses per waveMargin of error (one wave)Minimum detectable change
100±9.8 pts19.8 pts
200±6.9 pts14.0 pts
400±4.9 pts9.9 pts
600±4.0 pts8.1 pts
1,000±3.1 pts6.3 pts
2,000±2.2 pts4.4 pts
4,000±1.5 pts3.1 pts

To go the other way: to detect a 5-point change in a percentage you need roughly 1,570 responses per wave. To detect 3 points, about 4,350.

0–10 rating scales (assuming a standard deviation of 2.5, typical for satisfaction items)

Responses per waveMinimum detectable change in the mean
1000.99 points
2000.70 points
4000.50 points
1,0000.31 points
2,0000.22 points

Net Promoter Score (assuming 40% promoters, 20% detractors, so NPS = +20)

NPS is noisier than people expect, because it is a difference between two proportions and inherits the variance of both.

Responses per waveMargin of error on NPSMinimum detectable NPS change
200±10.420.9 points
400±7.314.8 points
1,000±4.69.4 points
2,000±3.36.6 points
4,000±2.34.7 points

Read that middle row again. With 400 responses per wave — a very typical tracker — the smallest NPS movement you can reliably detect is about 15 points. Almost every quarterly NPS conversation in the industry is about movements far smaller than that. See NPS Benchmarks by Industry for what real movements look like, and Brand Tracking Studies for wave design.

Four things that quietly destroy your power

1. Segmentation. Power is driven by the n in each cell, not the total. Split 1,000 responses across five segments and each segment has 200 — an MDE of 14 points, not 6.3. If the study exists to compare segments, size for the smallest segment you intend to report.

2. Multiple comparisons. Test 20 segments at the 5% threshold and the probability of at least one false positive is 64%, not 5%. Trackers that slice by region, plan, tenure and platform every quarter are manufacturing significant-looking noise. Decide your comparisons in advance and correct for the rest.

3. Independent waves instead of the same people. Re-surveying the same panel makes the two waves correlated, which cancels part of the noise. At a wave-to-wave correlation of 0.5 the MDE drops by roughly 29% — the same benefit as doubling your sample, for free. This is the single cheapest power upgrade available to a tracker.

4. Raising power without raising sample. Moving from 80% to 90% power requires about 34% more sample. If someone wants more certainty, that is the price.

What to do when you cannot afford the sample

Most teams read the NPS table, discover they need 2,000 responses per wave to see a 7-point move, and conclude the research is impossible. It isn't — the design is just wrong for the question.

Stop trying to detect small changes in a summary metric. A 5-point NPS shift is not a finding; it is a number that will move again next quarter. Size your tracker for the movements that would actually change a decision — usually 10 points or more — and stop reporting anything below the MDE as if it were signal.

Move the question from "did it move" to "why". Detecting a 3-point satisfaction change requires thousands of responses. Understanding what changed for customers requires depth, not volume. Forty conversations that probe why churned customers left will change more decisions than a tracker sized to prove that satisfaction fell by 0.2.

Use a longitudinal panel. Same people, repeated waves. It buys the equivalent of double the sample.

Reduce the variance instead of increasing the n. Tighter question wording, fewer response categories misused as scales, and cleaner sampling all shrink the standard deviation, and MDE scales directly with it. See How to Write Unbiased Survey Questions and Survey Data Quality — every low-effort response you filter out is variance you did not have to pay for.

Pre-register the effect you care about. Write down, before fielding: "we will act if the metric moves by X". If X is below your MDE, redesign the study now rather than explaining an inconclusive result later.

How Koji changes this arithmetic

The power problem is really a cost problem: statistical resolution scales with the square of your sample, so every halving of the MDE costs four times the sample. Traditional research responds by buying more panel — the most expensive possible answer.

Koji attacks it from the other side.

Every interview yields more per respondent. A Koji study is a conversation, not a form. The AI interviewer asks the structured question, then probes the answer — so a single participant produces both a countable data point and the reasoning behind it. That reasoning is what lets a team act on a movement that is too small to be statistically certain, because they know the mechanism rather than just the delta.

Structured questions make the quantitative side rigorous. Koji supports six question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — and aggregates them automatically into distributions and charts, so your MDE arithmetic applies to real structured data rather than to hand-coded open text. See Structured Questions Guide.

Consistent administration lowers variance. Human moderators drift: they rephrase, they prompt unevenly, they skip items when a session runs long. That drift is variance, and variance is MDE. An AI interviewer asks every participant the same question the same way, every time, at any hour — which tightens the distribution and improves your effective power without a single extra response.

Cost per participant is low enough to size properly. Koji charges 1 credit for a text interview and 3 for a voice interview, and only conversations that pass the platform's quality bar consume credits at all — so the sample you need to hit your MDE is achievable rather than theoretical, and low-effort responses do not eat your budget or inflate your variance.

The practical pattern that works: run a properly sized structured study to establish whether the number moved, and let the same participants' open-ended answers explain why. One study, both halves of the question — which is the thing a survey tool and an interview platform used separately can never quite deliver.

Frequently asked questions

What is the difference between margin of error and minimum detectable effect? Margin of error describes the precision of a single estimate at one point in time. Minimum detectable effect describes the smallest difference between two estimates that your study could reliably identify. For two independent waves of equal size, the MDE is roughly twice the margin of error — so a ±4.9-point margin of error corresponds to a 9.9-point detectable change.

How many responses do I need to detect a 5-point change? For a percentage measured at around 50%, about 1,570 responses per wave, at 95% confidence and 80% power. For a 3-point change, about 4,350 per wave. Sample requirements grow with the square of the precision you want, which is why chasing small movements gets expensive so fast.

Why is NPS so hard to move statistically? Because NPS is the difference between two proportions, it carries the sampling variance of both promoters and detractors. With 400 responses per wave, the minimum detectable NPS change is about 15 points — far larger than the quarter-to-quarter movements most teams discuss in review meetings.

What does 80% power actually mean? It means that if the effect you care about genuinely exists at the size you specified, you have an 80% chance of detecting it in this study and a 20% chance of missing it. Raising power to 90% requires roughly 34% more sample.

Does an inconclusive result mean there was no change? No, and this is the most common misreading. A non-significant result in an underpowered study means the study could not tell — not that nothing happened. That is precisely why the MDE should be calculated before fielding: it converts "we found nothing" into the far more useful "we could not have found anything smaller than X".

Do these calculations apply to qualitative interviews? Not directly. Power analysis assumes you are estimating a number from a sample. Qualitative studies are sized by information coverage rather than statistical power — see How Many User Interviews Do You Need?. Where the two meet is a platform like Koji that runs structured questions and open-ended probing in the same conversation: the structured half obeys the arithmetic on this page, and the qualitative half explains the result.

Related Resources

Want depth and numbers from the same study? Start free with 10 credits — text interviews cost 1 credit, and only quality conversations consume them.

Related Articles

Staged Rollout for AI Features: Shadow Mode, Canary, and Kill Switches (2026)

A research-first guide to staging an AI feature launch. What shadow mode can and cannot measure, what to ask users at each canary ring, how to pre-register rollback thresholds, and why the EU AI Act made the kill switch a legal requirement.

Brand Tracking Studies: How to Measure Brand Health Over Time (2026)

A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.

How Many User Interviews Do You Need? The Sample Size Guide for Qualitative Research

Discover the right number of user interviews for your research. Learn about data saturation, theoretical saturation, and practical frameworks for knowing when you've collected enough qualitative data.

Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms

A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.

NPS Benchmarks 2026: Net Promoter Score by Industry (Complete Reference)

Compare your NPS to 2026 industry benchmarks for SaaS, ecommerce, financial services, healthcare, and more. Includes what counts as "good", scoring math, and how to dig into the "why" behind your score with AI follow-up interviews.

Statistical Significance in Survey Research: A Plain-English Guide (2026)

A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Margin of Error in Surveys: What It Means and How to Calculate It (2026)

A plain-English guide to survey margin of error — the formula, a worked example, what changes it, common misreadings, and why AI-moderated interviews sidestep the breadth-vs-depth trade-off entirely.

Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)

A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.