Confidence Intervals for Small-Sample UX Research: Completion Rates, Ratings and Task Times
How to put confidence intervals around UX metrics from 5 to 20 participants: the adjusted Wald interval for completion rates, t-intervals for rating scales, and log-transformed intervals with the geometric mean for task times, with worked examples.
How Do You Report Confidence Intervals With a Small Sample?
Report every small-sample UX number as a range, not a single value, and pick the interval method that fits the metric:
- Completion rates and other yes/no outcomes: use the adjusted Wald interval.
- Rating scales (SEQ, UMUX-Lite, satisfaction): use a t-based interval around the mean.
- Task times: take logs, build a t-interval, then convert back, and report the geometric mean as the center.
A confidence interval tells readers how far the true value for your whole user population could plausibly sit from what you observed. With 5, 10 or 20 participants that range is wide, and showing it is more honest and more useful than a bare percentage. Lewis and Sauro (2006) put the purpose simply: the computation of confidence intervals "helps by establishing the likely boundaries of measurement."
This guide walks through each method with worked numbers you can check.
Why Small Samples Need Different Math
Most people learn the textbook interval for a proportion: observed rate plus or minus 1.96 standard errors (the Wald interval). It works with large samples. With the samples typical in usability testing, it breaks:
- It can produce impossible bounds below 0% or above 100%.
- When everyone succeeds (10 out of 10), its width collapses to zero, implying you're certain the true rate is 100%. You aren't.
- Its real coverage falls well below the 95% it claims.
Jeff Sauro and James R. Lewis tested this directly. In a 2005 paper for the Human Factors and Ergonomics Society, they compared the Wald, exact, score and adjusted Wald intervals using Monte Carlo simulations, drawing samples of 5, 10 and 15 users from real usability datasets. The adjusted Wald interval gave the best coverage, and their later work recommends it as the default for completion rates and other binary UX metrics.
The adjusted Wald method builds on Agresti and Coull (1998), who showed that a small adjustment to the observed proportion fixes most of the Wald interval's problems.
Method 1: Completion Rates (Adjusted Wald)
Use this for any metric where each participant either did or didn't: task success, conversion in a prototype, "would switch" answers, yes/no questions.
The formula
For a 95% interval, z = 1.96, so z² ≈ 3.84.
- Adjust the proportion: p̂ = (x + z²/2) / (n + z²), where x is the number of successes and n is the sample size. At 95% that is roughly (x + 2) / (n + 4), the well-known "add two successes and two failures" rule.
- Compute the standard error: SE = √( p̂ (1 − p̂) / (n + z²) ).
- Build the interval: p̂ ± z × SE. Clip at 0% and 100%.
Worked example
7 of 10 participants completed the checkout task.
- p̂ = (7 + 1.92) / (10 + 3.84) = 8.92 / 13.84 = 0.645
- SE = √(0.645 × 0.355 / 13.84) = 0.129
- Interval = 0.645 ± 1.96 × 0.129 = 39% to 90%
That matches the figure a UX Magazine article on quantitative UX decision-making reports for the same data: a 95% adjusted Wald interval of about 39% to 90% for 7 of 10 completions. An exact (Clopper-Pearson) interval for the same data runs about 35% to 93%, slightly wider and more conservative.
How sample size changes the picture
The table below applies the same formula (95% confidence) to a few common outcomes. These are our calculations, not published figures:
| Observed | Rate | 95% adjusted Wald interval |
|---|---|---|
| 4 of 5 | 80% | 36% to 98% |
| 5 of 5 | 100% | 51% to 100% |
| 7 of 10 | 70% | 39% to 90% |
| 9 of 10 | 90% | 57% to 100% |
| 10 of 10 | 100% | 68% to 100% |
| 14 of 20 | 70% | 48% to 86% |
| 70 of 100 | 70% | 60% to 78% |
Two lessons stand out. First, a perfect score from five people is consistent with a true success rate as low as about 51%. "Everyone completed it" is not the same as "the task works". Second, getting the interval down to roughly ±10 points takes around 100 participants, which is why small usability studies are better at finding problems than at estimating rates precisely.
Which number do you report as the "rate"?
Report the observed rate (70%) with the interval. Lewis and Sauro's 2006 Journal of Usability Studies paper looked at the best point estimate for small samples and recommends adjustments only for extreme outcomes (near 0% or 100%), where the raw rate overstates certainty. For most results, the observed rate plus the adjusted Wald interval is clear and defensible.
Method 2: Rating Scales (t-Interval)
Use this for the average of a rating: SEQ, UMUX-Lite, SUS, CSAT on a 1–5 scale and similar.
The formula
Mean ± t × (s / √n), where s is the sample standard deviation and t comes from the t-distribution with n − 1 degrees of freedom. For 95% confidence, t is 2.78 at n = 5, 2.26 at n = 10, 2.20 at n = 12 and 2.09 at n = 20. Use t, not 1.96: with small samples, 1.96 makes the interval too narrow.
Worked example
Twelve participants rated a task on the 7-point Single Ease Question: 6, 5, 7, 4, 6, 5, 7, 6, 3, 6, 5, 6.
- Mean = 5.50
- Standard deviation = 1.17
- SE = 1.17 / √12 = 0.34
- Interval = 5.50 ± 2.20 × 0.34 = 4.76 to 6.24
So you can say: "The average ease rating was 5.5 out of 7 (95% CI 4.8 to 6.2)." If a benchmark you care about sits inside that range, you don't yet have evidence that you're above or below it.
Rating data is bounded and often skewed, so treat the interval as an approximation, especially when most answers pile up at one end of the scale.
Method 3: Task Times (Log-Transform and Geometric Mean)
Task times are skewed: most people finish in a minute or two, and a few take far longer. The arithmetic mean gets dragged up by those few, and a plain t-interval can be misleading.
What the research recommends
Sauro and Lewis studied this in a CHI 2010 paper using Monte Carlo simulations on 61 large-sample tasks. They found that for small samples, the geometric mean was a better estimate of the typical task time than the sample median, with 13% less error and 22% less bias. Their practical guidance: use the geometric mean as the center for samples under about 25, and the median for larger samples. MeasuringU also notes that at small sample sizes the median can overstate the middle time by as much as 10%.
The method
- Take the natural log of each time.
- Compute the mean and standard deviation of the logs.
- Build a t-interval on the logs: mean ± t × (s / √n).
- Convert the center and both ends back with eˣ.
Worked example
Ten participants' task times in seconds: 62, 75, 48, 120, 90, 55, 180, 70, 66, 85.
- Arithmetic mean: 85.1 seconds
- Geometric mean: 78.9 seconds
- 95% interval (back-transformed): about 60 to 104 seconds
The geometric mean sits below the arithmetic mean because the 180-second outlier no longer dominates. The interval is asymmetric (wider above than below), which correctly reflects the skew in task-time data.
Common Mistakes
Reporting percentages without the denominator. "80% succeeded" means something very different at n = 5 and n = 500. Always show "4 of 5" or "n = 5" alongside.
Using 1.96 with tiny samples. For means, use the t-value. For proportions, use adjusted Wald rather than the plain Wald.
Treating overlap as "no difference". Two intervals can overlap slightly while the difference between the groups is still statistically meaningful. If you need to compare two designs, test the difference directly.
Mistaking precision for importance. A narrow interval tells you the estimate is precise, not that the result matters. Pair numbers with what participants actually said.
Dropping intervals from executive summaries. The headline is where the range matters most. "Completion: 70% (likely 39–90%)" sets better expectations than "70%".
When Small Samples Are Enough
Wide intervals don't mean small studies are useless. Small qualitative studies are excellent at finding problems: if 3 of 5 people fail at the same step, you've found something worth fixing, whatever the exact rate turns out to be. Intervals matter when you want to estimate or compare: benchmarking, tracking over time, or choosing between designs. For those goals, plan sample sizes around the interval width you can tolerate.
How Koji Helps
Koji is an AI research platform that runs interviews by text or voice and produces a report as responses come in. Several parts of it make interval-friendly UX metrics easier to collect.
Structured questions give you clean numbers. Koji's structured question types include yes/no, single choice, multiple choice, ranking and scale questions with a range and endpoint labels you set (1 to 7 for SEQ or UMUX-Lite, for example). Each answer is stored per question, so the data is ready for the formulas above. See the structured questions guide.
Reports show the distribution, not just the average. For scale questions the report shows the full distribution with the mean and median, and yes/no questions show the split. Participant quotes are linked to the answers, so a low rating comes with the reason.
Export for your own intervals. The Responses tab shows every participant's answers in a grid with a summary row, and exports to CSV with one click. Paste the column into a spreadsheet and apply the adjusted Wald or t-interval formulas directly.
Larger samples without more moderation. The fastest way to narrow an interval is more participants. Because Koji's AI interviewer runs every session, you can move from 10 participants to 50 or 100 without scheduling more moderated calls, and you can recruit through your own share link or through Koji's research panel (launching a panel recruitment needs a paid plan).
Filters out low-effort responses. Each interview gets a quality score from 1 to 5, and only interviews scoring 3 or more enter the report. That keeps click-through responses from adding noise to small samples, where every data point moves the estimate.
The why behind the number. Every scale question can carry AI follow-ups, including one anchored to the participant's own score. A small-sample study still produces a defensible range for the metric and a set of reasons explaining it.
Related Resources
- Structured Questions Guide: collect yes/no, choice and scale data in interviews
- Usability Metrics Guide: task success, time on task and error rate
- Survey Sample Size Guide: plan the sample around the precision you need
- Statistical Significance in Survey Research: testing differences between groups
- Survey Margin of Error Guide: margin of error for larger surveys
- UMUX-Lite Guide: a two-item usability score that pairs well with intervals
Sources
- Sauro, J., & Lewis, J. R. (2005). Estimating Completion Rates from Small Samples Using Binomial Confidence Intervals: Comparisons and Recommendations. Proceedings of the Human Factors and Ergonomics Society Annual Meeting.
- Lewis, J. R., & Sauro, J. (2006). When 100% Really Isn't 100%: Improving the Accuracy of Small-Sample Estimates of Completion Rates. Journal of Usability Studies, 1(3).
- Sauro, J., & Lewis, J. R. (2010). Average Task Times in Usability Tests: What to Report? Proceedings of CHI 2010.
- Agresti, A., & Coull, B. A. (1998). Approximate Is Better than "Exact" for Interval Estimation of Binomial Proportions. The American Statistician.
- "A New Formula for Quantitative UX Decision Making." UX Magazine.
- MeasuringU, "Average Task Times in Usability Tests: What to Report?" (measuringu.com/average-times).
Related Articles
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Margin of Error in Surveys: What It Means and How to Calculate It (2026)
A plain-English guide to survey margin of error — the formula, a worked example, what changes it, common misreadings, and why AI-moderated interviews sidestep the breadth-vs-depth trade-off entirely.
Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)
A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.
UMUX-Lite: The Two-Item Usability Questionnaire (Items, Scoring and Evidence)
How to use UMUX-Lite, the two-item alternative to the System Usability Scale: the exact items, how to score it, what the research says about its reliability and correspondence with SUS, and when to use it.
Usability Metrics: Task Success Rate, Time on Task, and Error Rate Explained
The complete guide to the core usability metrics — task success rate, time on task, and error rate — including industry benchmarks, formulas, sample sizes, and how to capture them automatically with AI-moderated research.