The List Experiment: Measure a Behavior Without Ever Asking About It (2026)
A list experiment asks people how many statements are true, never which ones. The difference between two group averages is your prevalence estimate, and no respondent is ever on record.
A list experiment estimates how many of your users do something sensitive without ever putting the question to anyone. Respondents see a short list of statements and report only how many are true - not which. One group gets an extra statement. The difference between the two group averages is your answer.
The short answer
Randomized response protects the respondent by adding noise to their answer. A list experiment protects them by never collecting the answer at all.
You split your sample in two. The control group sees four innocuous statements and reports a count from 0 to 4. The treatment group sees the same four plus the sensitive one, and reports a count from 0 to 5. Since the groups are randomly assigned, they should average the same count on the four shared items. Any excess in the treatment group has to come from the fifth statement.
If the control group averages 1.80 and the treatment group averages 2.13, then 0.33 - 33% - is your prevalence estimate. No respondent ever said yes to anything in particular.
The technique is also called the unmatched count technique or the item count technique. It was introduced by D. Raghavarao and Walter T. Federer in 1979.
Why counting beats confessing
The respondent is never on record
This is the structural difference from every mitigation that relies on a promise. With a direct question, a truthful yes exists somewhere in a database, and the respondent knows it. With a list experiment, the most incriminating thing any individual ever produces is the number 3. That number is compatible with many combinations of true statements, so it commits them to nothing.
The technique is described as a way to improve, through anonymity, the number of true answers to possibly embarrassing or self-incriminating questions - and the anonymity here is a property of the response format rather than of your data handling.
What the design assumes
One assumption does all the work: that the control group would have given the same average count, were it not for the critical question. Randomization is what buys you that assumption, and it is why the two groups must be assigned at random rather than split by convenience, timing, or segment.
If you let people self-select into groups, or you field the two versions a week apart, the assumption fails silently and the difference you measure is partly a difference between the groups rather than the effect of the extra item.
A worked example
You want to know what share of your users have shared a paid seat with someone outside their organization. Asking directly gets you a number you do not believe.
You field two versions to 500 people each.
| Group | Items | Mean count | n |
|---|---|---|---|
| Control | 4 innocuous | 1.80 | 500 |
| Treatment | Same 4 + sensitive | 2.13 | 500 |
Estimate: 2.13 - 1.80 = 0.33, so 33% of users have shared a seat.
Reading the uncertainty honestly
The point estimate is not the finding. With those group sizes and realistic response spread, the standard error of the difference is about 6.33 points, so the 95% confidence interval runs roughly 33% plus or minus 12.4 points - from about 21% to about 45%.
That is a wide interval, and reporting 33% without it would be misleading. What the study supports is a statement like "somewhere between a fifth and a half of users have done this, and it is certainly not rare" - which, for a decision about whether to build seat-sharing detection, is usually enough.
The precision cost, stated plainly
The method is very simple to use but yields only the number of people bearing the property of interest, and it leads to a larger sampling error than direct questions. Here is the size of that penalty:
| Design | Total n | Standard error |
|---|---|---|
| Direct question at 33% | 1,000 | 1.49 points |
| List experiment | 1,000 (500 per arm) | 6.33 points |
About 4.26 times less precise from the same 1,000 people. You are spending sample to buy deniability, exactly as in randomized response, and for the same reason.
Designing the list
Most failed list experiments fail here rather than in the analysis.
Avoid the ceiling
If a respondent in the treatment group finds all five statements true, answering 5 tells you everything. The protection collapses for exactly the people you most wanted to protect - and they can see that it has, so they under-report.
The fix: make sure at least one control item is rare enough that almost nobody can hit the ceiling. Include one statement that is true for perhaps 5% of people.
Avoid the floor
The mirror problem. A respondent who answers 0 has denied everything, including the sensitive item. Include one statement that is true for nearly everyone, so that 0 is an implausible answer and nobody lands there.
Choose items that do not move together
The variance of your estimate depends on how much the control counts vary. Four items that all tend to be true or all tend to be false together produce a wide spread and a noisy estimate. Items that are negatively correlated - where being true on one makes another less likely - tighten the distribution and buy you precision for free.
Keep the control items genuinely boring
Every control item must be non-sensitive. If one of your four filler statements is itself mildly embarrassing, the treatment and control groups are both distorted and the difference no longer isolates the item you care about.
Field both versions at the same time
Same window, same recruitment source, same instrument. The assumption that the two groups are otherwise identical is the only thing standing between you and an uninterpretable number.
What you get and what you give up
You get a prevalence estimate and nothing else
There is no individual-level variable to cross-tabulate. If you want the rate among enterprise users specifically, you have to run the whole two-arm design within that segment, at full sample size. Segment comparisons multiply your recruiting requirement fast.
You cannot follow up
The most frustrating limitation. A respondent has told you a number, so there is no thread to pull - you cannot ask why, or when, or what would change it. Pair the list experiment with separate qualitative work if you need the reasoning; see projective techniques for the qualitative route to the same material.
It does not fix who showed up
Like every question-design fix, this one corrects for misreporting, not for nonresponse bias. If seat-sharers avoid your study, an unbiased estimator applied to a biased sample still gives you the wrong number.
Running a list experiment with Koji
Three things make this practical on an AI-native platform that were awkward before.
Random assignment has to be clean and invisible. Koji assigns respondents to the control or treatment list without the respondent ever seeing that two versions exist, which protects the design from the single most common contamination - people comparing notes.
The response format is a count, not prose. Koji's structured questions cover six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and a scale or single_choice item captures the count as a discrete value that flows straight into the group means. Nothing needs hand-coding, and the AI interviewer is instructed not to probe the count, which is the one place probing would break the method.
The sample requirement is the real barrier. A two-arm design at 500 per arm is 1,000 conversations. That is a budget conversation with a traditional panel and a routine study when Koji runs the interviews in parallel and analyzes them automatically. The design is from 1979; what changed is the cost of fielding it.
Koji will not stop you from writing a ceiling-prone list. That judgment stays with you.
A working procedure
- Write the sensitive statement first, in the exact words you want estimated.
- Build four control items: one rare, one near-universal, two ordinary and mutually unrelated.
- Pilot the control list alone on 50 people and check the mean sits comfortably between 1 and 3, away from both ends.
- Field both arms simultaneously at your target sample.
- Compute the difference in means and its confidence interval. Report both.
- State the estimate as a range in the readout, never as a single number.
Frequently asked questions
How many control items should a list experiment use?
Four is the standard choice. Three leaves too little room to avoid ceiling and floor effects; five or more increases the counting burden and adds variance without adding protection.
Can a list experiment give a negative prevalence estimate?
Yes, when true prevalence is low and sampling noise runs the wrong way, the treatment mean can come in below the control mean. Report the estimate and its interval as they came out rather than truncating at zero, and treat it as evidence the behavior is rare.
Is a list experiment better than randomized response?
They solve the same problem differently. A list experiment is easier for respondents to understand because there is no device and no instruction to follow, while randomized response gives a cleaner individual-level guarantee. If comprehension is your worry, choose the list experiment.
How do I compare prevalence across segments?
Run the full two-arm design inside each segment. You cannot slice a list experiment after the fact, because no individual-level value exists to slice.
What sample size does a list experiment need?
Budget roughly four times a direct question at the same precision. In the worked example, 1,000 respondents produced a standard error of 6.33 points against 1.49 for a direct question on the same sample.
Can Koji field both arms of a list experiment?
Yes. Koji handles the random assignment, captures the count as a structured question, and keeps the two versions from ever meeting. You supply the item list and the judgment about ceiling and floor effects.
Related Resources
- Structured Questions in AI Interviews - the six question types, and the count capture this design needs
- Randomized Response - the other way to buy deniability, with an individual-level guarantee
- Social Desirability Bias - why the direct question was failing in the first place
- Projective Techniques - the qualitative counterpart when you need the reasoning, not the rate
- Nonresponse Bias - the error no question-design fix can reach
- Stated vs Revealed Preferences - the wider gap between what people say and what they do
- Why a Better Analysis Cannot Rescue a Bad Sample - the limit on all of this
Related Articles
Nonresponse Bias: How Missing Respondents Skew Your Data
Nonresponse bias occurs when the people who do not answer your survey differ systematically from those who do. Learn why a low response rate is not the same as bias, how to detect it, and how to reduce it.
Projective Techniques in Market Research: The Complete Guide
A practitioner's guide to projective techniques — word association, sentence completion, collage, personification and more. Learn when to use them, real examples, and how AI moderation runs them at scale.
Why a Better Analysis Cannot Rescue a Bad Sample (2026)
Sampling error has a floor that no amount of analytical care can cross. Here is the model that sets the floor, the two levers that move it, and why narrowing who you talk to beats running more interviews.
Social Desirability Bias: What It Is and How to Eliminate It in Research
Social desirability bias makes people tell you what sounds good instead of what is true. Learn what causes it, why it quietly wrecks product decisions, and the seven evidence-based ways to reduce it — including why AI-moderated interviews get more honest answers.
Stated vs. Revealed Preferences: Why Customers Say One Thing and Do Another (2026)
Customers routinely say one thing and do another — the say-do gap. This guide explains stated vs. revealed preferences, why the gap exists, what the data shows about its size, and how to design research that gets past what people claim to what they actually do.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.