{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-10T11:59:56.218Z"},"content":[{"type":"documentation","id":"44f0cfaf-423e-427f-b34a-fc38f0fbdcff","slug":"confidence-intervals-small-sample-ux","title":"Confidence Intervals for Small-Sample UX Research: Completion Rates, Ratings and Task Times","url":"https://www.koji.so/docs/confidence-intervals-small-sample-ux","summary":"With small UX samples, report each metric as a range. Use the adjusted Wald interval for completion rates and other yes/no outcomes: Sauro and Lewis (2005) found it gave the best coverage in simulations with 5, 10 and 15 users, and 7 of 10 successes gives a 95% interval of about 39% to 90%. Use a t-interval for rating-scale means. For task times, log-transform, build a t-interval and back-transform; Sauro and Lewis (2010) found the geometric mean had 13% less error and 22% less bias than the median for small samples.","content":"## How Do You Report Confidence Intervals With a Small Sample?\n\n**Report every small-sample UX number as a range, not a single value, and pick the interval method that fits the metric:**\n\n- **Completion rates and other yes/no outcomes:** use the **adjusted Wald** interval.\n- **Rating scales** (SEQ, UMUX-Lite, satisfaction): use a **t-based** interval around the mean.\n- **Task times:** take logs, build a t-interval, then convert back, and report the **geometric mean** as the center.\n\nA confidence interval tells readers how far the true value for your whole user population could plausibly sit from what you observed. With 5, 10 or 20 participants that range is wide, and showing it is more honest and more useful than a bare percentage. Lewis and Sauro (2006) put the purpose simply: the computation of confidence intervals \"helps by establishing the likely boundaries of measurement.\"\n\nThis guide walks through each method with worked numbers you can check.\n\n## Why Small Samples Need Different Math\n\nMost people learn the textbook interval for a proportion: observed rate plus or minus 1.96 standard errors (the **Wald** interval). It works with large samples. With the samples typical in usability testing, it breaks:\n\n- It can produce impossible bounds below 0% or above 100%.\n- When everyone succeeds (10 out of 10), its width collapses to zero, implying you're certain the true rate is 100%. You aren't.\n- Its real coverage falls well below the 95% it claims.\n\nJeff Sauro and James R. Lewis tested this directly. In a 2005 paper for the Human Factors and Ergonomics Society, they compared the Wald, exact, score and adjusted Wald intervals using Monte Carlo simulations, drawing samples of **5, 10 and 15 users** from real usability datasets. The **adjusted Wald** interval gave the best coverage, and their later work recommends it as the default for completion rates and other binary UX metrics.\n\nThe adjusted Wald method builds on Agresti and Coull (1998), who showed that a small adjustment to the observed proportion fixes most of the Wald interval's problems.\n\n## Method 1: Completion Rates (Adjusted Wald)\n\nUse this for any metric where each participant either did or didn't: task success, conversion in a prototype, \"would switch\" answers, yes/no questions.\n\n### The formula\n\nFor a 95% interval, z = 1.96, so z² ≈ 3.84.\n\n1. **Adjust the proportion:** p̂ = (x + z²/2) / (n + z²), where x is the number of successes and n is the sample size. At 95% that is roughly (x + 2) / (n + 4), the well-known \"add two successes and two failures\" rule.\n2. **Compute the standard error:** SE = √( p̂ (1 − p̂) / (n + z²) ).\n3. **Build the interval:** p̂ ± z × SE. Clip at 0% and 100%.\n\n### Worked example\n\n7 of 10 participants completed the checkout task.\n\n- p̂ = (7 + 1.92) / (10 + 3.84) = 8.92 / 13.84 = 0.645\n- SE = √(0.645 × 0.355 / 13.84) = 0.129\n- Interval = 0.645 ± 1.96 × 0.129 = **39% to 90%**\n\nThat matches the figure a UX Magazine article on quantitative UX decision-making reports for the same data: a 95% adjusted Wald interval of about 39% to 90% for 7 of 10 completions. An exact (Clopper-Pearson) interval for the same data runs about 35% to 93%, slightly wider and more conservative.\n\n### How sample size changes the picture\n\nThe table below applies the same formula (95% confidence) to a few common outcomes. These are our calculations, not published figures:\n\n| Observed  | Rate | 95% adjusted Wald interval |\n| --------- | ---- | -------------------------- |\n| 4 of 5    | 80%  | 36% to 98%                 |\n| 5 of 5    | 100% | 51% to 100%                |\n| 7 of 10   | 70%  | 39% to 90%                 |\n| 9 of 10   | 90%  | 57% to 100%                |\n| 10 of 10  | 100% | 68% to 100%                |\n| 14 of 20  | 70%  | 48% to 86%                 |\n| 70 of 100 | 70%  | 60% to 78%                 |\n\nTwo lessons stand out. First, **a perfect score from five people is consistent with a true success rate as low as about 51%**. \"Everyone completed it\" is not the same as \"the task works\". Second, getting the interval down to roughly ±10 points takes around 100 participants, which is why small usability studies are better at finding problems than at estimating rates precisely.\n\n### Which number do you report as the \"rate\"?\n\nReport the observed rate (70%) with the interval. Lewis and Sauro's 2006 _Journal of Usability Studies_ paper looked at the best _point_ estimate for small samples and recommends adjustments only for extreme outcomes (near 0% or 100%), where the raw rate overstates certainty. For most results, the observed rate plus the adjusted Wald interval is clear and defensible.\n\n## Method 2: Rating Scales (t-Interval)\n\nUse this for the average of a rating: SEQ, UMUX-Lite, SUS, CSAT on a 1–5 scale and similar.\n\n### The formula\n\nMean ± t × (s / √n), where s is the sample standard deviation and t comes from the t-distribution with n − 1 degrees of freedom. For 95% confidence, t is 2.78 at n = 5, 2.26 at n = 10, 2.20 at n = 12 and 2.09 at n = 20. Use t, not 1.96: with small samples, 1.96 makes the interval too narrow.\n\n### Worked example\n\nTwelve participants rated a task on the 7-point Single Ease Question: 6, 5, 7, 4, 6, 5, 7, 6, 3, 6, 5, 6.\n\n- Mean = 5.50\n- Standard deviation = 1.17\n- SE = 1.17 / √12 = 0.34\n- Interval = 5.50 ± 2.20 × 0.34 = **4.76 to 6.24**\n\nSo you can say: \"The average ease rating was 5.5 out of 7 (95% CI 4.8 to 6.2).\" If a benchmark you care about sits inside that range, you don't yet have evidence that you're above or below it.\n\nRating data is bounded and often skewed, so treat the interval as an approximation, especially when most answers pile up at one end of the scale.\n\n## Method 3: Task Times (Log-Transform and Geometric Mean)\n\nTask times are skewed: most people finish in a minute or two, and a few take far longer. The arithmetic mean gets dragged up by those few, and a plain t-interval can be misleading.\n\n### What the research recommends\n\nSauro and Lewis studied this in a CHI 2010 paper using Monte Carlo simulations on **61 large-sample tasks**. They found that for small samples, the **geometric mean** was a better estimate of the typical task time than the sample median, with **13% less error and 22% less bias**. Their practical guidance: use the geometric mean as the center for samples under about 25, and the median for larger samples. MeasuringU also notes that at small sample sizes the median can overstate the middle time by as much as 10%.\n\n### The method\n\n1. Take the natural log of each time.\n2. Compute the mean and standard deviation of the logs.\n3. Build a t-interval on the logs: mean ± t × (s / √n).\n4. Convert the center and both ends back with eˣ.\n\n### Worked example\n\nTen participants' task times in seconds: 62, 75, 48, 120, 90, 55, 180, 70, 66, 85.\n\n- Arithmetic mean: 85.1 seconds\n- Geometric mean: **78.9 seconds**\n- 95% interval (back-transformed): **about 60 to 104 seconds**\n\nThe geometric mean sits below the arithmetic mean because the 180-second outlier no longer dominates. The interval is asymmetric (wider above than below), which correctly reflects the skew in task-time data.\n\n## Common Mistakes\n\n**Reporting percentages without the denominator.** \"80% succeeded\" means something very different at n = 5 and n = 500. Always show \"4 of 5\" or \"n = 5\" alongside.\n\n**Using 1.96 with tiny samples.** For means, use the t-value. For proportions, use adjusted Wald rather than the plain Wald.\n\n**Treating overlap as \"no difference\".** Two intervals can overlap slightly while the difference between the groups is still statistically meaningful. If you need to compare two designs, test the difference directly.\n\n**Mistaking precision for importance.** A narrow interval tells you the estimate is precise, not that the result matters. Pair numbers with what participants actually said.\n\n**Dropping intervals from executive summaries.** The headline is where the range matters most. \"Completion: 70% (likely 39–90%)\" sets better expectations than \"70%\".\n\n## When Small Samples Are Enough\n\nWide intervals don't mean small studies are useless. Small qualitative studies are excellent at **finding** problems: if 3 of 5 people fail at the same step, you've found something worth fixing, whatever the exact rate turns out to be. Intervals matter when you want to **estimate** or **compare**: benchmarking, tracking over time, or choosing between designs. For those goals, plan sample sizes around the interval width you can tolerate.\n\n## How Koji Helps\n\nKoji is an AI research platform that runs interviews by text or voice and produces a report as responses come in. Several parts of it make interval-friendly UX metrics easier to collect.\n\n**Structured questions give you clean numbers.** Koji's structured question types include **yes/no**, **single choice**, **multiple choice**, **ranking** and **scale** questions with a range and endpoint labels you set (1 to 7 for SEQ or UMUX-Lite, for example). Each answer is stored per question, so the data is ready for the formulas above. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Reports show the distribution, not just the average.** For scale questions the report shows the full distribution with the mean and median, and yes/no questions show the split. Participant quotes are linked to the answers, so a low rating comes with the reason.\n\n**Export for your own intervals.** The Responses tab shows every participant's answers in a grid with a summary row, and exports to CSV with one click. Paste the column into a spreadsheet and apply the adjusted Wald or t-interval formulas directly.\n\n**Larger samples without more moderation.** The fastest way to narrow an interval is more participants. Because Koji's AI interviewer runs every session, you can move from 10 participants to 50 or 100 without scheduling more moderated calls, and you can recruit through your own share link or through Koji's research panel (launching a panel recruitment needs a paid plan).\n\n**Filters out low-effort responses.** Each interview gets a quality score from 1 to 5, and only interviews scoring 3 or more enter the report. That keeps click-through responses from adding noise to small samples, where every data point moves the estimate.\n\n**The why behind the number.** Every scale question can carry AI follow-ups, including one anchored to the participant's own score. A small-sample study still produces a defensible range for the metric and a set of reasons explaining it.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide): collect yes/no, choice and scale data in interviews\n- [Usability Metrics Guide](/docs/usability-metrics-guide): task success, time on task and error rate\n- [Survey Sample Size Guide](/docs/survey-sample-size-guide): plan the sample around the precision you need\n- [Statistical Significance in Survey Research](/docs/statistical-significance-survey-research): testing differences between groups\n- [Survey Margin of Error Guide](/docs/survey-margin-of-error-guide): margin of error for larger surveys\n- [UMUX-Lite Guide](/docs/umux-lite-guide): a two-item usability score that pairs well with intervals\n\n## Sources\n\n- Sauro, J., & Lewis, J. R. (2005). Estimating Completion Rates from Small Samples Using Binomial Confidence Intervals: Comparisons and Recommendations. _Proceedings of the Human Factors and Ergonomics Society Annual Meeting_.\n- Lewis, J. R., & Sauro, J. (2006). When 100% Really Isn't 100%: Improving the Accuracy of Small-Sample Estimates of Completion Rates. _Journal of Usability Studies_, 1(3).\n- Sauro, J., & Lewis, J. R. (2010). Average Task Times in Usability Tests: What to Report? _Proceedings of CHI 2010_.\n- Agresti, A., & Coull, B. A. (1998). Approximate Is Better than \"Exact\" for Interval Estimation of Binomial Proportions. _The American Statistician_.\n- \"A New Formula for Quantitative UX Decision Making.\" _UX Magazine_.\n- MeasuringU, \"Average Task Times in Usability Tests: What to Report?\" (measuringu.com/average-times).\n","category":"Research Methods","lastModified":"2026-10-10T04:56:50.423609+00:00","metaTitle":"Confidence Intervals for Small Samples in UX Research","metaDescription":"Use adjusted Wald for completion rates, t-intervals for ratings and the geometric mean for task times. Formulas and worked examples for 5 to 20 participants.","keywords":["confidence interval small sample","adjusted Wald interval","completion rate confidence interval","usability statistics","geometric mean task time","UX metrics sample size"],"aiSummary":"With small UX samples, report each metric as a range. Use the adjusted Wald interval for completion rates and other yes/no outcomes: Sauro and Lewis (2005) found it gave the best coverage in simulations with 5, 10 and 15 users, and 7 of 10 successes gives a 95% interval of about 39% to 90%. Use a t-interval for rating-scale means. For task times, log-transform, build a t-interval and back-transform; Sauro and Lewis (2010) found the geometric mean had 13% less error and 22% less bias than the median for small samples.","aiPrerequisites":["Basic statistics (mean, standard deviation)","Familiarity with usability metrics"],"aiLearningOutcomes":["Calculate an adjusted Wald confidence interval for a completion rate","Build a t-based confidence interval for a rating-scale mean","Estimate typical task time with the geometric mean and a log-transformed interval","Avoid common mistakes when reporting small-sample UX metrics"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}