Short answer: yes — and the research says they often tell a machine more than they tell a person. Across three decades of survey-methodology work, moving an interview from a human interviewer to an automated one consistently increases disclosure, especially on sensitive or embarrassing topics. The open question in 2026 is not whether people will open up to an AI interviewer. It is whether the AI is a good enough listener to follow up on what they say.
This post separates those two questions, because almost every buyer conversation about AI-moderated research confuses them.
The objection, stated fairly
If you are evaluating AI customer interview tools, someone on your team has already said some version of this:
"People will not be honest with a bot. You lose rapport. You lose the human connection that makes interviews work."
It is a reasonable instinct. It is also, on the disclosure question specifically, the opposite of what the evidence shows — and it has been the opposite since well before large language models existed.
What thirty years of mode-effect research actually found
Survey methodologists have a name for this: mode effects. Change the mode of administration — paper, computer, telephone, face-to-face — and answers to the same question change in predictable ways.
Weisband and Kiesler (1996) ran the foundational meta-analysis, published at CHI, covering studies from 1969 to 1994. Across 39 studies and 100 measures, computer administration increased self-disclosure relative to paper (a modest overall effect, Cohen's d of about 0.20). The moderators matter more than the headline: effects were larger when the comparison was a face-to-face interview, larger when the form asked for sensitive information, and larger among medical and psychiatric patients. In other words, the machine advantage grows exactly as the stakes of embarrassment grow.
Lind, Schober, Conrad and Reichert (2013), in Public Opinion Quarterly, took up the obvious follow-up question in a paper titled "Why do survey respondents disclose more when computers ask the questions?" — by that point the whether was settled enough that the field had moved on to the why.
Then Lucas, Gratch, King and Morency (2014), in Computers in Human Behavior, ran the cleanest test of the mechanism. Participants in a health-screening interview all talked to the same virtual human. The only thing that changed was what they were told: that the character was controlled by a human operator, or that it was fully automated. Participants who believed they were talking to a computer reported lower fear of self-disclosure, lower impression management, displayed sadness more intensely, and were rated by independent observers as more willing to disclose. The paper has been cited more than 850 times.
Read that design again, because it is the whole argument. The interviewer was identical. The belief about who was listening was the manipulation. And belief alone moved four separate measures of candor.
The candor advantage is not a property of the AI
This is the part most vendors get wrong, including vendors who would like to sell you an AI interviewer.
The disclosure effect does not come from the AI being clever, warm, or well-prompted. It comes from the respondent's sense that no human is evaluating them in real time. Impression management is a response to a perceived audience. Remove the audience and the management drops away.
That has a sharp and slightly inconvenient design consequence:
Anything that reinstates a perceived human evaluator gives the advantage back.
Telling participants that a researcher will listen to every recording, that their account manager will review the transcript, or that their name will appear next to their answers reintroduces exactly the audience the mode effect removed. You can still do those things — sometimes you must, for consent and ethics reasons — but you should do them knowing you are trading candor for something else, and you should say so honestly in your write-up. Our guide to social desirability bias covers the general mechanism; this is that mechanism operating at the level of your consent screen.
Disclosure and elaboration are two different skills
Here is the distinction that clears up most of the confusion in this debate.
| Disclosure | Elaboration | |
|---|---|---|
| The question it answers | Will they say the difficult thing at all? | Will they explain why, four layers down? |
| Historically better | Automated modes | Skilled human interviewers |
| What it depends on | Perceived audience | Real-time comprehension and follow-up |
| Settled by research? | Yes, decades ago | No — this is the live question |
Traditional self-administered surveys bought disclosure by giving up elaboration entirely: a text box cannot ask "what do you mean by that?" Skilled human moderators bought elaboration at the cost of disclosure, plus cost, scheduling, and moderator bias.
LLM-based interviewers are the first mode that can plausibly attempt both at once. That is genuinely new, and it is why the serious 2025 and 2026 research on AI interviewers is almost entirely about follow-up quality rather than candor. Candor was never the hard part.
What the 2025 evidence says about the hard part
The most useful recent assessment is Tirumala, Jain, Leybzon and Buskirk, "Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers", published at COLM 2025. It is a position paper, and refreshingly it does not oversell. Its own summary: field studies suggest AI interviewers already exceed traditional Interactive Voice Response systems for both quantitative and qualitative collection, but real-time transcription error rates, limited emotion detection, and uneven follow-up quality mean fitness for purpose is context-dependent for qualitative work.
The specific findings worth carrying into a buying decision:
- Transcription is better than you fear and worse than the marketing. State-of-the-art English word error rates hover around 5%. But real-time, streaming transcription — which is what a live voice interview requires — runs significantly higher, around 10.9% on average in the work the authors cite. Batch accuracy is not live accuracy.
- Conversational chatbots increase self-disclosure, with a caveat. Rhim and colleagues (2022) found more human-like survey chatbots produced higher self-disclosure — but at the cost of slightly elevated social desirability bias. Humanising the interface a little re-creates a little of the audience effect. This is the trade-off named above, measured.
- Satisficing goes down. Kim and colleagues (2019) found conversational administration reduced satisficing, the habit of picking the easiest valid answer to get the survey over with.
- Voice has its own mode effects. Respondents round numeric answers more in voice than in text (Schober et al., 2015), and long answer lists in audio invite primacy and recency effects (Le Bigot et al., 2013). Voice is not strictly better; it is differently better.
- Recordings are a fraud control. Gomila and colleagues (2017) note that collecting and analysing recordings helps mitigate response fraud — a quality check that text surveys simply cannot run. This matters more every year; see survey data quality.
Separately, Wuttke and colleagues (2024) randomly assigned participants to a conversational interview with either an AI or a human interviewer using identical questionnaires, and measured interviewer adherence to guidelines, response quality, and engagement. Their conclusion was that AI conversational interviewing is viable for producing data comparable to traditional methods, with scalability as the added benefit.
Comparable, at scale. Not magic. That is the honest state of the art, and it is more than enough to change how most teams work.
Where the machine advantage is biggest
Based on the moderators in the literature, candor gains concentrate in situations where a respondent would otherwise be managing your impression of them:
- Criticising your product to your face. Customers soften feedback for humans who built the thing. This is the single most under-appreciated case, because it is not a "sensitive topic" in the clinical sense — it is ordinary politeness, and it quietly corrupts most discovery interviews.
- Admitting confusion. "I did not understand what that button did" is an admission of incompetence in front of a person, and a neutral fact in front of a machine. Critical for onboarding and activation research.
- Money. Willingness to pay, budget authority, what they actually spend today.
- Churn and cancellation reasons. The real reason is often unflattering to the customer, not to you.
- Anything an employer or peer might see. Which is why anonymous employee research is such a strong fit.
- Genuinely sensitive subject matter, where the ethical bar is higher and trauma-informed practice applies regardless of mode.
Where it does not help — and what to do instead
Be straight with your stakeholders about the limits:
- Emotion detection is weak. If your study depends on reading a face or a catch in the voice, the current generation will not do it reliably. The COLM paper says so plainly.
- Live streaming transcription errors compound. Design questions that survive a 10% word error rate: short, unambiguous, one idea per question.
- Long multi-select lists are bad in voice. Move them to structured questions rather than reading twelve options aloud.
- Rapport-dependent longitudinal work with the same participants over months is still a place where a named human researcher earns their keep.
- An AI interviewer will not save a bad discussion guide. Mode does not fix leading questions. The Mom Test still applies.
And one thing that is not a limitation but is frequently confused with one: whether you can trust the AI's analysis is a separate question from whether respondents were candid. We answer that one separately in Can You Trust AI Interviewers?, and the question of whether AI-generated participants are valid — they largely are not — in synthetic users.
How to design an AI interview that earns candor
Six practical rules, drawn directly from the mechanisms above:
- Do not promise a human will review it, unless a human will. Then say so precisely, and expect slightly more guarded answers.
- Say what happens to the recording, once, plainly, up front. Ambiguity about the audience is worse than a clear human audience.
- Decouple identity from response wherever your consent basis allows. Aggregate reporting is a candor feature, not just a compliance one.
- Ask about behaviour, not intention. Mode effects reduce impression management; they do not make people better at predicting themselves.
- Mix conversational depth with structured measurement. Open questions get the elaboration; structured questions get you something countable.
- Pilot with five conversations and read the transcripts yourself. Follow-up quality is the variable that differs most between platforms, and it is visible immediately.
Where Koji fits
Koji is built around exactly this split between disclosure and elaboration.
AI-moderated voice interviews run the conversation without a human in the room, which is the condition the disclosure literature identifies. The AI probes in the moment — asking the follow-up a static survey cannot — so you are not trading elaboration away to get candor.
Structured questions handle the parts of a study that should be countable rather than conversational. Koji supports six types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — so a single study can carry both a numeric backbone and open-ended depth. That is also the clean answer to the voice-mode problem above: put the twelve-option list in a structured question instead of reading it aloud. See structured questions in AI interviews.
Automatic thematic analysis reads every transcript rather than the three your team had time for, and one-click reports turn a completed study into something shareable the same day. No moderator means no moderator bias, and no scheduling: interviews run while you sleep, in parallel, in the respondent's own time.
The practical result is the one that matters to a buying committee: from question to insight in hours rather than weeks, with no research expertise required to run the study — and, according to thirty years of mode-effect research, answers that are more candid than the ones a human moderator would have collected, not less.
If you want to see the mechanics of the interview itself, start with AI-moderated interviews. If you are comparing platforms, Best AI Customer Interview Tools covers the field, and Best Voice Survey Software covers the voice-specific category.
Try it on your own hardest question
Pick the question your team is least comfortable asking customers directly — the pricing one, the "why did you really leave" one, the one where you suspect people are being polite. That is the question where the mode effect is largest, and it is the fastest way to find out whether this holds for your customers.
Koji gives you 10 free credits when you sign up — enough to run real voice interviews and read the transcripts yourself. A text conversation costs 1 credit, a voice conversation 3, and Koji only charges for conversations that clear its quality bar, so a junk response costs you nothing.
Frequently Asked Questions
Do people really tell an AI more than they tell a human interviewer?
On sensitive and self-critical topics, yes. The meta-analytic evidence goes back to Weisband and Kiesler's 1996 review of 39 studies, and the effect is largest precisely where the comparison is a face-to-face interview and where the question is sensitive. Lucas and colleagues (2014) isolated the mechanism: participants told the same virtual interviewer was automated rather than human-operated showed lower fear of disclosure, lower impression management, and were rated by observers as more willing to disclose.
Does that mean AI interviews are always better than human ones?
No. The candor advantage is real and well replicated; the follow-up advantage is not settled. The COLM 2025 assessment by Tirumala and colleagues found AI interviewers already exceed IVR systems but flagged uneven follow-up quality, weak emotion detection, and real-time transcription error rates near 11% as genuine limits for qualitative work. For rapport-dependent longitudinal studies, or research that hinges on reading emotion, a skilled human moderator is still the better tool.
Will making the AI more human-like increase honesty?
Slightly the opposite, on balance. Rhim and colleagues (2022) found more humanised survey chatbots produced higher self-disclosure but also slightly elevated social desirability bias. The more your interviewer feels like a person, the more the respondent starts managing that person's impression. A warm but clearly non-human interviewer is usually the better trade.
Do I have to tell participants they are talking to an AI?
Yes — disclose it, both because it is the ethical baseline and increasingly because regulation requires it. This is not a cost to your data quality. The disclosure literature says the candor benefit comes from participants believing no human is evaluating them, so telling them it is automated is what activates the effect rather than undermining it.
What kinds of questions still go wrong in a voice interview?
Long multi-select lists, precise numeric answers, and anything phrased ambiguously enough that a 10% word error rate could flip the meaning. Respondents round numbers more in voice than in text, and long option lists in audio invite primacy and recency effects. Move those to structured questions and keep the spoken parts conversational.
How do I sanity-check this for my own customers?
Run the same question two ways on comparable samples — once in your existing survey or moderated interview, once as an AI-moderated conversation — and compare the distributions, not just the themes. If the mode effect is operating, you will typically see less flattering answers and fewer non-responses in the automated condition. Treat the divergence as the finding.