{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-01T04:39:12.811Z"},"content":[{"type":"documentation","id":"a40f487c-cd65-4288-8808-5db3ad7c6b4b","slug":"list-experiment-item-count-research","title":"The List Experiment: Measure a Behavior Without Ever Asking About It (2026)","url":"https://www.koji.so/docs/list-experiment-item-count-research","summary":"A list experiment, also called the unmatched count or item count technique, estimates the prevalence of a sensitive behavior by asking respondents how many statements on a list are true rather than which ones. A randomly assigned control group sees four innocuous items and a treatment group sees the same four plus the sensitive item; the difference in mean counts is the prevalence estimate. It was introduced by D. Raghavarao and Walter T. Federer in 1979. With control mean 1.80 and treatment mean 2.13 at 500 per arm, the estimate is 33% with a standard error near 6.33 points, about 4.26 times less precise than a direct question on the same 1,000 people. Design must avoid ceiling and floor effects and use items that are not positively correlated.","content":"A list experiment estimates how many of your users do something sensitive without ever putting the question to anyone. Respondents see a short list of statements and report only how many are true - not which. One group gets an extra statement. The difference between the two group averages is your answer.\n\n## The short answer\n\nRandomized response protects the respondent by adding noise to their answer. A list experiment protects them by never collecting the answer at all.\n\nYou split your sample in two. The control group sees four innocuous statements and reports a count from 0 to 4. The treatment group sees the same four plus the sensitive one, and reports a count from 0 to 5. Since the groups are randomly assigned, they should average the same count on the four shared items. Any excess in the treatment group has to come from the fifth statement.\n\nIf the control group averages 1.80 and the treatment group averages 2.13, then 0.33 - 33% - is your prevalence estimate. No respondent ever said yes to anything in particular.\n\nThe technique is also called the unmatched count technique or the item count technique. It was introduced by D. Raghavarao and Walter T. Federer in 1979.\n\n## Why counting beats confessing\n\n### The respondent is never on record\n\nThis is the structural difference from every mitigation that relies on a promise. With a direct question, a truthful yes exists somewhere in a database, and the respondent knows it. With a list experiment, the most incriminating thing any individual ever produces is the number 3. That number is compatible with many combinations of true statements, so it commits them to nothing.\n\nThe technique is described as a way to improve, through anonymity, the number of true answers to possibly embarrassing or self-incriminating questions - and the anonymity here is a property of the response format rather than of your data handling.\n\n### What the design assumes\n\nOne assumption does all the work: that the control group would have given the same average count, were it not for the critical question. Randomization is what buys you that assumption, and it is why the two groups must be assigned at random rather than split by convenience, timing, or segment.\n\nIf you let people self-select into groups, or you field the two versions a week apart, the assumption fails silently and the difference you measure is partly a difference between the groups rather than the effect of the extra item.\n\n## A worked example\n\nYou want to know what share of your users have shared a paid seat with someone outside their organization. Asking directly gets you a number you do not believe.\n\nYou field two versions to 500 people each.\n\n| Group | Items | Mean count | n |\n| --- | --- | --- | --- |\n| Control | 4 innocuous | 1.80 | 500 |\n| Treatment | Same 4 + sensitive | 2.13 | 500 |\n\nEstimate: 2.13 - 1.80 = 0.33, so 33% of users have shared a seat.\n\n### Reading the uncertainty honestly\n\nThe point estimate is not the finding. With those group sizes and realistic response spread, the standard error of the difference is about 6.33 points, so the 95% confidence interval runs roughly 33% plus or minus 12.4 points - from about 21% to about 45%.\n\nThat is a wide interval, and reporting 33% without it would be misleading. What the study supports is a statement like \"somewhere between a fifth and a half of users have done this, and it is certainly not rare\" - which, for a decision about whether to build seat-sharing detection, is usually enough.\n\n### The precision cost, stated plainly\n\nThe method is very simple to use but yields only the number of people bearing the property of interest, and it leads to a larger sampling error than direct questions. Here is the size of that penalty:\n\n| Design | Total n | Standard error |\n| --- | --- | --- |\n| Direct question at 33% | 1,000 | 1.49 points |\n| List experiment | 1,000 (500 per arm) | 6.33 points |\n\nAbout 4.26 times less precise from the same 1,000 people. You are spending sample to buy deniability, exactly as in [randomized response](/docs/randomized-response-technique-research), and for the same reason.\n\n## Designing the list\n\nMost failed list experiments fail here rather than in the analysis.\n\n### Avoid the ceiling\n\nIf a respondent in the treatment group finds all five statements true, answering 5 tells you everything. The protection collapses for exactly the people you most wanted to protect - and they can see that it has, so they under-report.\n\nThe fix: make sure at least one control item is rare enough that almost nobody can hit the ceiling. Include one statement that is true for perhaps 5% of people.\n\n### Avoid the floor\n\nThe mirror problem. A respondent who answers 0 has denied everything, including the sensitive item. Include one statement that is true for nearly everyone, so that 0 is an implausible answer and nobody lands there.\n\n### Choose items that do not move together\n\nThe variance of your estimate depends on how much the control counts vary. Four items that all tend to be true or all tend to be false together produce a wide spread and a noisy estimate. Items that are negatively correlated - where being true on one makes another less likely - tighten the distribution and buy you precision for free.\n\n### Keep the control items genuinely boring\n\nEvery control item must be non-sensitive. If one of your four filler statements is itself mildly embarrassing, the treatment and control groups are both distorted and the difference no longer isolates the item you care about.\n\n### Field both versions at the same time\n\nSame window, same recruitment source, same instrument. The assumption that the two groups are otherwise identical is the only thing standing between you and an uninterpretable number.\n\n## What you get and what you give up\n\n### You get a prevalence estimate and nothing else\n\nThere is no individual-level variable to cross-tabulate. If you want the rate among enterprise users specifically, you have to run the whole two-arm design within that segment, at full sample size. Segment comparisons multiply your recruiting requirement fast.\n\n### You cannot follow up\n\nThe most frustrating limitation. A respondent has told you a number, so there is no thread to pull - you cannot ask why, or when, or what would change it. Pair the list experiment with separate qualitative work if you need the reasoning; see [projective techniques](/docs/projective-techniques) for the qualitative route to the same material.\n\n### It does not fix who showed up\n\nLike every question-design fix, this one corrects for misreporting, not for [nonresponse bias](/docs/nonresponse-bias). If seat-sharers avoid your study, an unbiased estimator applied to a biased sample still gives you the wrong number.\n\n## Running a list experiment with Koji\n\nThree things make this practical on an AI-native platform that were awkward before.\n\nRandom assignment has to be clean and invisible. Koji assigns respondents to the control or treatment list without the respondent ever seeing that two versions exist, which protects the design from the single most common contamination - people comparing notes.\n\nThe response format is a count, not prose. Koji's structured questions cover six types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - and a `scale` or `single_choice` item captures the count as a discrete value that flows straight into the group means. Nothing needs hand-coding, and the AI interviewer is instructed not to probe the count, which is the one place probing would break the method.\n\nThe sample requirement is the real barrier. A two-arm design at 500 per arm is 1,000 conversations. That is a budget conversation with a traditional panel and a routine study when Koji runs the interviews in parallel and analyzes them automatically. The design is from 1979; what changed is the cost of fielding it.\n\nKoji will not stop you from writing a ceiling-prone list. That judgment stays with you.\n\n## A working procedure\n\n1. Write the sensitive statement first, in the exact words you want estimated.\n2. Build four control items: one rare, one near-universal, two ordinary and mutually unrelated.\n3. Pilot the control list alone on 50 people and check the mean sits comfortably between 1 and 3, away from both ends.\n4. Field both arms simultaneously at your target sample.\n5. Compute the difference in means and its confidence interval. Report both.\n6. State the estimate as a range in the readout, never as a single number.\n\n## Frequently asked questions\n\n### How many control items should a list experiment use?\n\nFour is the standard choice. Three leaves too little room to avoid ceiling and floor effects; five or more increases the counting burden and adds variance without adding protection.\n\n### Can a list experiment give a negative prevalence estimate?\n\nYes, when true prevalence is low and sampling noise runs the wrong way, the treatment mean can come in below the control mean. Report the estimate and its interval as they came out rather than truncating at zero, and treat it as evidence the behavior is rare.\n\n### Is a list experiment better than randomized response?\n\nThey solve the same problem differently. A list experiment is easier for respondents to understand because there is no device and no instruction to follow, while randomized response gives a cleaner individual-level guarantee. If comprehension is your worry, choose the list experiment.\n\n### How do I compare prevalence across segments?\n\nRun the full two-arm design inside each segment. You cannot slice a list experiment after the fact, because no individual-level value exists to slice.\n\n### What sample size does a list experiment need?\n\nBudget roughly four times a direct question at the same precision. In the worked example, 1,000 respondents produced a standard error of 6.33 points against 1.49 for a direct question on the same sample.\n\n### Can Koji field both arms of a list experiment?\n\nYes. Koji handles the random assignment, captures the count as a structured question, and keeps the two versions from ever meeting. You supply the item list and the judgment about ceiling and floor effects.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and the count capture this design needs\n- [Randomized Response](/docs/randomized-response-technique-research) - the other way to buy deniability, with an individual-level guarantee\n- [Social Desirability Bias](/docs/social-desirability-bias) - why the direct question was failing in the first place\n- [Projective Techniques](/docs/projective-techniques) - the qualitative counterpart when you need the reasoning, not the rate\n- [Nonresponse Bias](/docs/nonresponse-bias) - the error no question-design fix can reach\n- [Stated vs Revealed Preferences](/docs/stated-vs-revealed-preferences) - the wider gap between what people say and what they do\n- [Why a Better Analysis Cannot Rescue a Bad Sample](/docs/sampling-error-irreducible-interview-analysis) - the limit on all of this\n","category":"Research Methods","lastModified":"2026-09-30T03:35:55.371613+00:00","metaTitle":"The List Experiment: Item Count Technique for Research","metaDescription":"A list experiment asks how many statements are true, never which. The difference in group means gives prevalence, plus design rules.","keywords":["list experiment","item count technique","unmatched count technique","indirect questioning","sensitive question survey design","prevalence estimation research"],"aiSummary":"A list experiment, also called the unmatched count or item count technique, estimates the prevalence of a sensitive behavior by asking respondents how many statements on a list are true rather than which ones. A randomly assigned control group sees four innocuous items and a treatment group sees the same four plus the sensitive item; the difference in mean counts is the prevalence estimate. It was introduced by D. Raghavarao and Walter T. Federer in 1979. With control mean 1.80 and treatment mean 2.13 at 500 per arm, the estimate is 33% with a standard error near 6.33 points, about 4.26 times less precise than a direct question on the same 1,000 people. Design must avoid ceiling and floor effects and use items that are not positively correlated.","aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}