Scoring Answers You Cannot Verify: Incentive-Compatible Research (2026)
How to score honest answers in customer research when no ground truth exists, using peer prediction, Bayesian truth serum and the surprisingly popular rule.
Short answer: you usually cannot check whether a customer told you the truth about what they do, what they would pay, or how often they hit a problem. You can still score the answer, by using everyone else's answers as the yardstick. Ask each participant two things, what they believe and what they think other people will say, then give weight to the answers that turn out to be more common than the group collectively predicted. Drazen Prelec published this idea as a Bayesian truth serum in Science in 2004, and the same arithmetic drives the surprisingly popular algorithm, which reduced errors by 21.3 percent against simple majority voting.
The unusual part is that nobody gets caught out. There is no behavioural log to cross-reference, no trick question and no attention trap. The method works on the structure of the answers themselves, which is why it survives in exactly the situations where every other quality check fails.
Most research questions have no answer key
Research operations spends a lot of effort on checks that quietly assume a ground truth exists somewhere. Attention checks assume you know the correct response. Screener validation assumes you can confirm a job title. Behavioural triangulation assumes the behaviour is instrumented.
Now list the questions that actually drive roadmap decisions:
- How often do you hit this problem in a normal week?
- What are you doing today instead of using our product?
- Would you pay 40 dollars a month for this?
- Which of the five things you just described actually blocks you?
None of these has an answer key. The behaviour happens outside your product, the purchase never happened, and the importance ranking exists only in the participant's head. Economists call this class of problem information elicitation without verification. A 2024 review on arXiv, Mechanisms for belief elicitation without ground truth, defines the task as "eliciting truthful information from multiple individuals when such information cannot be verified" and counts over 25 published mechanisms that attempt it.
The standard research response to unverifiable questions is to reduce the incentive to misreport: anonymity, neutral wording, indirect questioning. Those are real and they work, and social desirability bias and questionnaire design cover them properly. But every one of them is defensive. They lower the pressure to shade an answer without ever telling you which answers were shaded.
Ask two questions instead of one
Mechanism design takes a different route. Rather than making dishonesty unattractive, it makes honesty the highest-scoring strategy, and it uses a second question as the lever.
Every participant answers twice:
- Their own answer. What do you think, or what do you actually do?
- Their prediction about everyone else. What share of other participants will give each answer?
The second question looks redundant. It is the entire mechanism. Prelec's scoring rule combines two components:
- A prediction score, which rewards a participant for being accurate about the distribution of other people's answers.
- An information score, which rewards a participant whose own answer turns out to be more common than the group collectively predicted it would be.
The information score is the interesting half. It pays you for holding a view that the crowd systematically underestimates, and under reasonable assumptions about how people form beliefs, the only way to maximise it reliably is to report what you actually think.
Why the surprisingly popular answer wins
Prelec, Seung and McCoy published the cleanest version in Nature in 2017. Their decision rule is one sentence: "select the answer that is more popular than people predict."
The logic is asymmetric knowledge. Somebody holding a well-informed minority view usually understands why most people disagree, so their prediction of the group correctly allows for the popular wrong answer. Somebody holding the popular wrong answer cannot imagine the minority view being widespread, so they under-predict it. Averaged over the sample, the better-informed answer is systematically under-predicted relative to the support it actually has.
Here is what that looks like on a question a product team would really ask. One hundred participants, asked whether they keep a manual spreadsheet workaround alongside your product, and then asked what share of others will say yes.
| Quantity | Value |
|---|---|
| Said yes | 40 of 100 |
| Said no | 60 of 100 |
| Average predicted share saying yes | 25 percent |
| Gap between actual and predicted | plus 15 points |
| Plain majority verdict | No |
| Surprisingly popular verdict | Yes |
A majority vote closes this question as most customers do not maintain workarounds, and the roadmap moves on. The surprisingly popular rule reaches the opposite conclusion, because the 40 who admit the workaround correctly anticipated that most people would not, while the 60 who said no assumed their own behaviour was typical. The minority was carrying the information.
That failure mode is exactly what the Nature paper accuses voting of. Democratic methods, the authors write, "are biased for shallow, lowest common denominator information, at the expense of novel or specialized knowledge that is not widely shared." In product research, novel and specialised knowledge is the entire point of talking to customers at all.
The size of the improvement is measurable. Across their test questions, MIT reported that "the surprisingly popular algorithm reduced errors by 21.3 percent compared to simple majority votes, and by 24.2 percent compared to basic confidence-weighted votes." Note what the second number rules out: asking people how confident they are does not rescue majority voting. Confidence and information are different things, which is also why a loud stakeholder is not a well-informed one.
Does paying for honesty change what people admit?
Yes, and the best evidence comes from researchers studying themselves. John, Loewenstein and Prelec surveyed over 2,000 psychologists about their own questionable research practices in Psychological Science in 2012, comparing plain anonymous self-report against self-report backed by incentives for truth telling.
Their finding: "The impact of truth-telling incentives on self-admissions of questionable research practices was positive, and this impact was greater for practices that respondents judged to be less defensible." They concluded that "some questionable practices may constitute the prevailing research norm."
Read that second clause again, because it is the part that matters operationally. The incentive bought the most additional truth on the items people were most ashamed of. That is the opposite of the usual pattern in survey methodology, where the hardest questions are where methods fail worst. Here the harder the question, the more the mechanism was worth.
One honest caveat: that study asked academics about their own research conduct. Your participants are customers describing product behaviour, a different population and a different kind of embarrassment. Treat the direction of the effect as the transferable result, not its magnitude.
Where these mechanisms break down
The arXiv review is refreshingly blunt about the state of the field. Although many of these mechanisms guarantee truthfulness in theory, it reports that "empirical evidence regarding the effects of mechanisms on truth-telling is limited and generally weak", and identifies why: "most mechanisms are very complex and cannot be easily conveyed to research subjects."
Three limits worth planning around:
- Sample size. The group prediction has to mean something. Below roughly 20 participants the predicted distribution is mostly noise, and the gap you are reading is sampling error. This is a different constraint from the one governing how many interviews are enough for discovery.
- Comprehension. A participant who does not understand the scoring rule cannot game it, but also cannot be moved by it. If you cannot explain the incentive in two sentences, you are running an ordinary survey with extra steps.
- Shared misconception. If your whole sample believes the same wrong thing, no scoring rule recovers the truth, because the mechanism only ever compares participants against each other. That limitation is structural, and it is the same one that governs group accuracy.
How Koji handles this
The two-question format is administratively annoying and that is the main reason teams do not run it. Every participant answers twice, the second answer is a distribution rather than a value, and somebody has to compute the gap. Koji removes all three costs.
- Structured questions carry the prediction naturally. Koji supports six question types: open_ended, scale, single_choice, multiple_choice, ranking and yes_no. The own-answer half is typically single_choice or yes_no, and the prediction half is a scale question asking for a percentage, so both halves are first-class structured data rather than prose somebody has to re-read.
- The AI consultant is customizable, so the instruction to follow every belief question with its matching prediction question is written once into the brief and then applied identically across hundreds of interviews. No moderator forgets it on interview 40.
- Real-time reporting computes the gap for you. Because scale and choice answers are aggregated automatically, the actual share and the average predicted share are both available as the interviews land, so the surprisingly popular answer surfaces without an analyst exporting anything.
- AI-moderated and voice interviews keep the sample size reachable. The mechanism needs tens of participants to work at all, which is a real barrier when each interview costs an hour of a researcher's day. Running them concurrently is what makes a 20-plus sample routine rather than a project.
- Answers stay traceable to the transcript, so when a minority answer turns out to be the surprisingly popular one, you can read the eight people who gave it instead of arguing about whether the number is real.
The pattern to notice is that Koji is not detecting lies. It is making a second question cheap enough that you can afford to ask it, which is the only thing standing between most teams and a mechanism published in Science over twenty years ago.
Common mistakes
- Treating the prediction question as a warm-up. It is not rapport-building, it is half the instrument. Dropping it leaves you with an ordinary opinion poll.
- Reading the gap as a magnitude. A plus 15 point gap tells you which answer to trust, not that 15 percent more people do the thing. The gap is a selection signal, not an estimate.
- Running it on questions that do have an answer key. If the behaviour is instrumented, go and look at the instrument. Koji reports are most useful here on the questions your telemetry structurally cannot see.
- Assuming a majority is a finding. The whole contribution of this literature is that the popular answer is biased toward what is easy to know.
- Confusing this with an attention check. Attention checks remove participants who are not reading. This scores participants who are.
Frequently asked questions
Can I use this if I only have 10 interviews?
Not reliably. The mechanism compares each answer against the group's predicted distribution, and with 10 participants that predicted distribution carries too much sampling error to trust a 10 or 15 point gap. Use it as a qualitative prompt at that size, by all means, but do not let a small gap overturn a decision. Around 20 participants is a sensible floor, and 50 or more makes the gap stable.
Is this the same as an attention check?
No, and the two do opposite jobs. An attention check identifies people who are not engaging, so you can remove them. An incentive-compatible scoring rule assumes everyone is engaging and works out whose answer carries information the rest of the sample lacks. You can and should run both: screen first, then score.
Do I have to pay participants for it to work?
Payment is the cleanest way to make the incentive real, but the mechanism also works as a design discipline without money attached. Simply asking for the prediction changes what you learn, because you get the actual distribution and the believed distribution, and the gap between them is informative even when nobody is scored on it. Koji makes the unpaid version nearly free, since the second question is just another structured question in the brief.
What if everyone in my sample believes the same wrong thing?
Then this will not save you, and no scoring rule built on peer comparison can. The mechanism measures answers against other answers, so a misconception shared by the whole sample is invisible to it. That is a real structural limit, not a tuning problem, and the defence is sampling for genuine variety rather than scoring harder.
Does Koji support the two-question format?
Yes, directly. Pair the belief question with a prediction question in the same study: the own-answer half is usually yes_no, single_choice or ranking, and the prediction half is a scale question capturing a percentage. Koji's AI consultant asks both consistently across every interview, and real-time reporting aggregates both halves so the actual and predicted shares sit side by side.
How is this different from just asking better questions?
Better question wording lowers the pressure to misreport, which is genuinely valuable and covered in questionnaire design. It still leaves you unable to tell which answers were shaded. This approach does not try to prevent misreporting at all. It changes the payoff structure so that reporting honestly is the best available strategy, and it gives you a computed signal about which answers to weight.
Related Resources
- Structured Questions: The Complete Guide - the six question types, and how to pair a belief question with a prediction question
- Social Desirability Bias - the defensive half of the problem, and seven evidence-based reductions
- Stated vs Revealed Preferences - why say-do gaps appear even when nobody is lying
- The Mom Test - question wording that avoids inviting a flattering answer
- Why a More Accurate Reviewer Can Make Your Panel Worse - what happens when the people you are comparing stop being independent
- Attention Check Questions - screening out disengagement before you start scoring
Related Articles
Attention Check Questions: How to Catch Low-Effort Survey Responses Without Annoying Real Participants
Attention check questions catch inattentive, low-effort, and fraudulent survey responses. Learn the main types, how many to use, the pitfalls, and why a conversational AI interview reduces the need for them in the first place.
The Mom Test: How to Ask Customer Interview Questions That Get Honest Answers
A complete guide to the Mom Test methodology by Rob Fitzpatrick—covering the three core rules, good vs. bad interview questions, avoiding confirmation bias, and how AI scales honest customer discovery conversations.
Questionnaire Design: The Complete Guide to Writing Questions That Get Honest Answers
A research-backed guide to questionnaire design — defining your constructs, writing unbiased questions, choosing response scales, ordering for flow, pre-testing, and avoiding the biases that quietly ruin your data.
Social Desirability Bias: What It Is and How to Eliminate It in Research
Social desirability bias makes people tell you what sounds good instead of what is true. Learn what causes it, why it quietly wrecks product decisions, and the seven evidence-based ways to reduce it — including why AI-moderated interviews get more honest answers.
Stated vs. Revealed Preferences: Why Customers Say One Thing and Do Another (2026)
Customers routinely say one thing and do another — the say-do gap. This guide explains stated vs. revealed preferences, why the gap exists, what the data shows about its size, and how to design research that gets past what people claim to what they actually do.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.