Randomized Response: How to Measure a Behavior Nobody Will Admit To (2026)
Randomized response adds noise to each answer on purpose, then subtracts it from the total. You get an honest population estimate and no individual ever answers the sensitive question.
Randomized response lets you estimate how many of your users do something they will not admit to, without any individual ever answering the sensitive question. The respondent privately flips a coin, and the coin sometimes answers for them. You cannot tell who was telling the truth. You can still recover the true percentage, because you know exactly how much noise you added.
The short answer
Ask a sensitive question directly and you measure two things at once: how many people do the thing, and how many people are willing to say so. Those numbers are not the same, and nothing in your data separates them.
Randomized response separates them by design. You deliberately corrupt every individual answer with a known amount of random noise. Because the noise is known, you can subtract it from the aggregate. Because it is random and private, no one can tell which answers it touched. The respondent gets genuine deniability; you get an unbiased estimate of the population figure.
The price is precision. In the worked example below, the same 400 respondents produce a standard error of 4.77 percentage points instead of 2.29. You would need roughly 1,734 people to match what a direct question achieves with 400 - if the direct question were honest, which is the whole reason you are not using it.
Why the direct question fails
The three things a direct question actually measures
When you ask "have you ever shared your login with someone outside your team?", a "no" can mean three different things: they have not done it, they have done it and will not say so, or they have done it and do not consider it sharing. A direct question collapses all three into one number and gives you no way to pull them apart.
The first two are the dangerous pair. If willingness to admit correlates with the behavior itself - and for policy violations, workarounds, and anything involving embarrassment, it usually does - then your undercount is not random. It is largest exactly where the behavior is most common.
Anonymity is a promise; randomization is a mechanism
Most teams reach for a promise. They add a line saying responses are anonymous and assume the problem is handled. A promise requires the respondent to trust you, your vendor, your database, and your future self.
Randomization requires none of that. The respondent does not have to believe you, because the protection is arithmetic rather than policy. Even an adversary holding the complete raw dataset cannot determine any individual's true answer. That is a materially stronger guarantee than a privacy notice, and it is the reason the technique survives in contexts where trust in the researcher is low.
This distinction matters for how you describe the study. Our guide to anonymizing customer interview data covers protecting data after collection. Randomized response protects it before collection, which is the only point at which the respondent is deciding whether to lie.
How randomized response works
The forced-response design
The cleanest version to run in practice is the forced-response design:
- The respondent privately flips a coin, without showing anyone.
- If it comes up heads, they answer "yes" regardless of the truth.
- If it comes up tails, they answer the sensitive question honestly.
Every "yes" is now ambiguous. It might be a coin. It might be a confession. Nobody, including you, can tell which - and crucially, the respondent knows nobody can tell, which is what makes honesty on the tails branch rational.
The arithmetic
Let p be the true proportion who do the thing. Half the respondents are forced to say yes. The other half answer honestly, and p of them say yes. So the observed proportion of yes answers is:
Observed yes = 0.5 + 0.5 x p
Rearranged, the estimate is:
p = 2 x (observed yes) - 1
That is the entire method. One line of arithmetic, and it is exact rather than approximate.
A worked example with real numbers
You run the study with 400 respondents. 260 of them say yes.
| Quantity | Value |
|---|---|
| Respondents | 400 |
| Observed yes answers | 260 |
| Observed proportion | 0.65 |
| Estimated true prevalence | 2 x 0.65 - 1 = 0.30 |
So 30% of your users do the thing. Sanity-check it forward: of 400 people, about 200 flip heads and are forced to say yes. The other 200 answer honestly, and 30% of them - 60 people - say yes. Total: 260. The arithmetic closes.
Notice what you never learned: which of those 260 were confessing. That information does not exist in the dataset. It was destroyed on purpose, before it reached you.
Where the design came from
The technique was introduced by the statistician S. L. Warner in 1965 and modified by B. G. Greenberg and coauthors in 1969. Warner's original design used a different randomizing rule, in which the device decides whether the respondent answers the sensitive question or its negation. With the die-based version of that design, 75 yes answers out of 100 respondents implies a true prevalence of 12.5% - a reminder that the observed number and the estimate can be very far apart, and that reading the raw response rate as if it were the answer would be badly wrong.
What it costs: the variance tax
This is the part most write-ups skip, and it is the part that decides whether you should run one.
The same study, two standard errors
You bought deniability with noise, and noise widens your confidence interval. Using the worked example above:
| Design | n | Standard error |
|---|---|---|
| Direct question (if honest) | 400 | 2.29 points |
| Randomized response | 400 | 4.77 points |
The randomized estimate is about 2.08 times less precise from the same number of people.
How many more people you need
To get randomized response down to the precision a direct question gives you at 400 respondents, you need roughly 1,734 respondents - about 4.3 times the sample. That multiplier is not a flaw in the method; it is the exchange rate between privacy and precision, and it is knowable in advance.
When the tax is worth paying
Run the comparison honestly. A direct question gives you a tight confidence interval around a biased number. Randomized response gives you a wide confidence interval around an unbiased one. A precise wrong answer is worse than an imprecise right one whenever the bias is larger than the width of the interval - which, for genuinely sensitive behavior, it usually is.
The decision rule: if you expect under-reporting to shift the estimate by more than a few points, pay the variance tax. If the topic is only mildly awkward, a well-worded direct question and a good privacy statement will serve you better. Our article on social desirability bias covers the lighter-touch options.
Designing one that works
Choose a randomizing device the respondent controls
The respondent must perform the randomization themselves, privately, with something they can see. A coin, a die, the last digit of their phone number. If your platform generates the random number, the guarantee evaporates - you could have logged it, and a suspicious respondent will assume you did.
Write the instruction so it survives a distracted reader
Comprehension is the single biggest failure mode. A respondent who does not understand the rule will default to answering honestly, or to answering no, and either one biases your estimate. Keep the instruction to three short steps, state explicitly that you cannot tell which branch they took, and say why that is mathematically true rather than merely promised.
Verify comprehension before you trust the estimate
Add a check: after the instruction, ask what they should do if the coin shows heads. This is a perfect use for a single_choice structured question, and respondents who fail it should be analyzed separately. If a large share fails, your estimate is not salvageable and you need to rewrite the instruction rather than adjust the numbers.
Decide the analysis before you collect
Because p = 2 x (observed yes) - 1 can return a negative number when the true prevalence is near zero and sampling noise runs against you, agree in advance how you will report that. The honest answer is to report the estimate and its confidence interval as they come out, including an out-of-range point estimate, rather than quietly truncating to zero.
What randomized response cannot do
No individual-level data, by construction
You cannot cross-tabulate the sensitive item against anything measured at the individual level, because you do not have individual values. You can compare group-level estimates by running the design separately within each segment - which multiplies your sample requirement again. Check measurement invariance before you treat two segment estimates as comparable.
It cannot fix a bad sample
Randomized response corrects for lying. It does nothing about who showed up. If the people most likely to do the thing are also least likely to join your study, you still have nonresponse bias sitting underneath an otherwise unbiased estimator.
Non-compliance still biases you
Some respondents will ignore the coin and simply answer no, because they do not trust the design or cannot be bothered. This pushes your estimate downward, and no arithmetic recovers it. Treat your result as a lower bound on the true prevalence, and say so in the report.
Running this with Koji
Randomized response is unusually well suited to an AI-moderated interview, for three reasons.
The instruction has to be delivered identically to every respondent, and it has to be understood. A static survey page cannot tell whether someone read it. Koji's AI interviewer delivers the setup the same way every time, and because it is conversational, it can confirm the respondent has understood the rule before the sensitive item is asked - then move on without ever seeing the coin.
The design needs a clean binary capture. Koji's structured questions give you six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and the yes_no type records the randomized answer as a discrete value your report can total directly, rather than as prose someone has to code by hand.
The variance tax means you need a bigger sample than you are used to. This is where an AI-moderated approach changes the arithmetic of the decision: 1,700 conversations is an impossible ask for a moderator-led study and a routine one when Koji runs every interview in parallel. The method has existed since 1965; what has changed is that the sample size it demands is now affordable.
One honest limitation: Koji cannot verify that the respondent actually flipped the coin. No platform can. That is inherent to the design, not to the tool.
A working procedure
- Confirm the behavior is genuinely sensitive. If people would answer it directly, do not pay the variance tax.
- Fix your target precision, then compute the sample you need at roughly 4 times the direct-question figure.
- Write the three-step instruction and a comprehension check.
- Pilot with 30 people and read their comprehension-check answers before scaling.
- Field the study, keeping the wording frozen for the whole wave.
- Apply p = 2 x (observed yes) - 1, report the confidence interval, and label the result a lower bound.
Frequently asked questions
Does randomized response actually produce higher estimates than direct questioning?
For genuinely sensitive behavior, that is the expected pattern, and it is the main evidence the design is working. If your randomized estimate comes back lower than a direct question on the same topic, suspect non-compliance or a misunderstood instruction rather than a real finding.
Can I use randomized response for anything other than yes or no questions?
The forced-response design described here is binary. Extensions exist for quantitative items, but they cost even more precision and are harder to explain to respondents. If you need a distribution rather than a proportion, a list experiment is usually the better tool.
How do I explain the method without making people more suspicious?
Lead with the guarantee rather than the mechanism. Tell them the point of the coin is that you will never know their answer, then give the three steps. Respondents accept the design readily when the benefit to them is stated first.
What sample size do I need for randomized response?
Budget roughly four times what you would need for a direct question at the same precision. In the worked example, matching a direct question's precision at 400 respondents took about 1,734.
Is this legal and ethical to run on customers?
It is more protective than a standard sensitive question, because you never collect the individual answer. Normal consent and data-handling rules still apply, and you should state plainly that the study is designed so that individual responses cannot be recovered.
Can Koji run a randomized response study?
Yes. Build the instruction and comprehension check as structured questions, capture the randomized item as a yes_no question, and let Koji run the volume the design requires. The coin flip stays with the respondent, which is exactly where it belongs.
Related Resources
- Structured Questions in AI Interviews - the six question types, including the yes_no capture this design depends on
- Social Desirability Bias - the problem randomized response solves, and the lighter-touch alternatives
- Projective Techniques - the qualitative route to the same unspoken material
- Nonresponse Bias - the error randomized response does not fix
- Anonymizing Customer Interview Data - protecting answers after collection rather than before
- Total Survey Error - where to spend a fixed budget across competing sources of error
- Measurement Invariance - before comparing two segment estimates
Related Articles
Anonymizing Customer Interview Data: A Practical Guide for Privacy-Safe Research
Five operational techniques for handling PII in AI customer interviews — from intake-time anonymization to stakeholder-safe quote sharing — without sacrificing research signal.
Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)
Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.
Nonresponse Bias: How Missing Respondents Skew Your Data
Nonresponse bias occurs when the people who do not answer your survey differ systematically from those who do. Learn why a low response rate is not the same as bias, how to detect it, and how to reduce it.
Projective Techniques in Market Research: The Complete Guide
A practitioner's guide to projective techniques — word association, sentence completion, collage, personification and more. Learn when to use them, real examples, and how AI moderation runs them at scale.
Social Desirability Bias: What It Is and How to Eliminate It in Research
Social desirability bias makes people tell you what sounds good instead of what is true. Learn what causes it, why it quietly wrecks product decisions, and the seven evidence-based ways to reduce it — including why AI-moderated interviews get more honest answers.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)
Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.