Back to docs
Interview Techniques

Face-Threatening Questions: Why Users Will Not Tell You Your Product Is Confusing (2026)

Face threat is a property of the question you wrote, not the participant you recruited. How to score it before you field a study and rewrite the questions that buy a polite, inaccurate answer.

A face-threatening question is any question whose honest answer would make the participant look incompetent, careless, or disloyal. The practical consequence for a research team is not refusal. People almost never refuse these questions. They answer them pleasantly and inaccurately, and the transcript looks perfectly healthy.

The important claim in this guide is that face threat is a property of the question you wrote, not of the participant you recruited. That makes it a design problem with a design fix, available before you field anything, rather than a bias you discover in analysis and apologise for in the readout.

Why a polite answer is worse than no answer

A refusal is visible. It shows up as a blank, and you know to discount it. A face-saving answer is invisible: it is fluent, confident, on topic, and wrong. It enters your themes with the same weight as a true answer, and nothing downstream can separate the two.

This is why face threat is more dangerous than most named biases. It does not degrade your data into obvious noise. It degrades it into plausible signal.

The classic shape in product research is the comprehension question. Ask did you find that screen confusing? and you have asked the participant to volunteer that they could not work out something that, by implication, other people manage. The honest answer costs them something. The dishonest answer costs them nothing and ends an awkward moment. Most people take the cheap option, and the product team concludes the screen is clear.

The three variables that set the size of the threat

The standard model comes from Penelope Brown and Stephen Levinson, whose 1987 book Politeness: Some Universals in Language Usage treats politeness as a rational computation rather than a personality trait. A speaker facing a potentially threatening act weighs three things:

  1. Social distance between the two parties. Strangers and intimates behave differently, and both differ from the awkward middle where a participant is talking to a friendly representative of a company they pay money to.
  2. Relative power. If the asker controls something the answerer wants - an incentive, continued access, a support relationship, a job - the answerer softens.
  3. The imposition of the act itself. Some questions are simply bigger asks than others. What did you have for breakfast and why did you stop paying us are not the same size.

The useful part is that a research team controls two of these directly. You choose the imposition when you word the question, and you influence power and distance when you choose who asks, how the study is introduced, and what the participant believes happens to their answer. None of this requires you to find more honest participants.

Score your guide before you field it. Walk the question list and mark each item high, medium or low on imposition, then mark the whole study high or low on power. Any high-imposition question inside a high-power study is a question you should expect to be answered for politeness rather than accuracy. In a typical discovery guide of twenty questions, three or four will score that way, and they are almost always the four the stakeholders care most about.

The evidence: threat is situational, not personal

Roger Tourangeau and Ting Yan reviewed the survey-methodology literature on this in Psychological Bulletin in 2007 (volume 133, issue 5, pages 859 to 883). Their conclusion is the one that matters for design: "misreporting about sensitive topics is quite common and that it is largely situational." They are explicit about what drives it - "The extent of misreporting depends on whether the respondent has anything embarrassing to report and on design features of the survey" - and about the mechanism, describing a process in which "respondents edit the information they report to avoid embarrassing themselves in the presence of an interviewer or to avoid repercussions from third parties."

Read that carefully, because two separate audiences appear in it. There is the interviewer in the room, and there are third parties who might see the answer later. A question can be threatening because of who is listening now, or because of who the participant thinks will read the transcript. Those need different fixes.

The size of the effect is not small. Timo Gnambs and Kai Kaspar published a meta-analysis in Behavior Research Methods in 2015 (volume 47, issue 4, pages 1237 to 1259) covering 460 effect sizes and a total sample of 125,672 people. Comparing self-administered paper surveys against computerised ones, they found that "computerized surveys led to significantly more reporting of socially undesirable behaviors than comparable surveys administered on paper," and that "This effect was strongest for highly sensitive behaviors and surveys administered individually to respondents."

That last clause is the design lesson. The advantage of a lower-threat setting grows exactly where the stakes are highest. On low-threat questions the administration method barely matters; on the questions you most need a true answer to, it matters a great deal.

Five ways to ask, ranked by what they cost

Brown and Levinson describe a ladder of strategies, and it maps cleanly onto interview guide writing. Ordered from most to least face-threatening:

StrategyWhat it sounds like in researchWhen to use it
Bald on recordWhy did you get this wrong?Almost never in research
Positive politenessLots of people find this part fiddly - how did it go for you?Default for comprehension questions
Negative politenessIf you do not mind me asking, what happened at that step?Sensitive history, money, failure
Off recordWalk me through what you did next.The strongest option for most product research
Do not askInfer it from behaviour insteadWhen the question cannot be de-threatened

The fourth row is the one teams underuse. An off-record question does not ask for an evaluation at all. It asks for a narrative, and lets the difficulty surface as a fact of the story rather than as an admission. Walk me through the last time you tried to export a report costs the participant nothing to answer honestly, and it produces the confusion, the workaround and the abandoned attempt as incidental detail. Was exporting confusing? asks them to file a self-report of incompetence.

This is the same instinct behind the classic advice to ask about past behaviour rather than opinions, and it is why that advice works even when the participant has no reason to lie. Low-threat questions do not require honesty to be accurate.

Rewriting the six shapes that cost the most face

Threatening shapeWhy it costs faceLower-threat rewrite
Did you find this confusing?Asks for an admission of failureWalk me through what you expected to happen here.
Why did you not use the feature?Implies the correct behaviour was to use itTell me about the last time you needed to do this.
Do you understand what this means?A comprehension test with a right answerHow would you explain this to a colleague?
Would you recommend us?Asks for a public commitmentWho have you mentioned us to, if anyone?
Was the price fair?Invites agreement with the party being paidWhat else did you look at before you decided?
Any problems?Requires the participant to raise the complaint unpromptedWhat was the most annoying part of that week?

The rewrites share one move: they replace an evaluation the participant must own with a description they merely report. Notice also that the last rewrite presupposes that something was annoying, which lowers the cost of saying so. Used carelessly that becomes a leading question, so keep the presupposition on the existence of friction and never on its direction or cause.

What lowering face threat does not fix

Be honest about the limits, because overclaiming here produces the opposite problem.

Lowering face threat does not fix sampling. If only your happiest customers answer the invitation, every one of them can answer with perfect candour and your result will still be wrong.

It does not fix recall. A participant who genuinely cannot remember last March will not remember it better because you asked gently.

And it does not eliminate incentive effects. Someone who believes their answers influence whether they keep a discount has a material reason to shade their responses that no amount of careful wording removes. That is a study-design problem: change what the participant has at stake, or accept the ceiling.

Finally, a question can be low-threat and still be badly written. Face is one axis. Clarity, single-barrelled construction and neutral framing are separate axes, and a question has to pass all of them.

How Koji handles this

Koji is built so that the low-threat configuration is the default rather than something you have to remember to switch on.

  • No human is in the room. Koji runs interviews without a moderator, which removes the single largest source of face threat in the Tourangeau and Yan model - the interviewer whose presence the respondent is editing for.
  • Identical wording every time. Koji asks every participant the same question in the same words, so a moderator cannot soften a question for one participant and sharpen it for the next, which is how face threat normally becomes uneven across a sample.
  • AI follow-up that probes narrative, not judgement. When an answer is thin, Koji follows up by asking for the next part of the story rather than asking the participant to rate or justify themselves - the off-record strategy applied automatically.
  • Structured questions where a number is genuinely needed. Koji supports six structured types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so a sensitive magnitude can be collected as a scale or a single_choice item without the participant having to say an uncomfortable sentence out loud. See the structured questions guide for how each type is asked and analysed.
  • Voice or text, participant choice. Some people disclose more in writing than aloud. Koji supports both, and because the wording is fixed the two are comparable.
  • Clear data handling. Koji lets you anonymise transcripts, which addresses the second audience in the research above - the third party the participant imagines reading it later.

The honest caveat: removing the human interviewer removes one audience, not all of them. Participants still know a company is reading. Koji lowers the floor on face threat considerably; it does not take it to zero, and a question that would be humiliating to answer in front of anyone will still be answered carefully.

Frequently asked questions

What is a face-threatening question in user research?

It is a question whose honest answer would make the participant look incompetent, careless or disloyal. Typical examples are comprehension checks such as did you find this confusing, questions about why someone failed to use a feature, and any question that asks a paying customer to criticise the person asking. The defining symptom is that participants answer fluently rather than refusing, so the distortion is invisible in the transcript.

Is a face-threatening question the same as a leading question?

No, and the difference matters because the fixes are different. A leading question points at a particular answer through its wording. A face-threatening question may be perfectly neutral and still be expensive to answer truthfully, because the cost sits in what the answer reveals about the participant rather than in how the question is phrased. A question can be neutral and threatening, leading and harmless, or both at once. See avoiding leading questions for the other axis.

Does an AI interviewer reduce face threat?

It reduces one important component of it. The meta-analytic evidence on self-administered and computerised modes shows more reporting of socially undesirable behaviour when no person is present, with the largest gains on the most sensitive items. Because Koji runs interviews without a moderator, it starts from that lower-threat position by default. It does not remove the participant's awareness that a company will read the transcript, so high-stakes questions still need careful wording.

How do I know which of my questions are face-threatening?

Score each question on imposition, then score the study on power. Mark any question where the honest answer would be an admission of failure, a criticism of the person asking, or a disclosure the participant would not volunteer to a colleague. Anything scoring high on both belongs on the rewrite list. In practice three or four questions in a twenty-question guide will qualify, and they are usually the ones stakeholders most want answered.

Should I just remove every face-threatening question?

No. Removing them means not learning the things that are hardest to learn. The goal is to move each one down the strategy ladder until the honest answer is cheap, usually by converting an evaluation into a narrative. Only drop a question when no rewrite gets the cost low enough and you have a behavioural measure that answers the same thing.

Does anonymity solve the problem?

Partly. Anonymity addresses the third-party audience, which is real, but it does nothing about the interviewer present during the conversation and nothing about the participant's own self-image. Someone can be fully anonymous and still find it unpleasant to say out loud that they could not work out how to cancel. Combine anonymity with lower-imposition wording rather than treating it as a substitute.

Related Resources

Related Articles

How to Avoid Leading Questions in Surveys and Interviews

Leading questions quietly bias your research data. Learn how to spot and rewrite leading, loaded, and double-barreled questions — and how Koji's AI writes neutral questions and probes without steering respondents.

Building Rapport in Research Interviews: How to Make Participants Open Up

Learn proven techniques to build trust and comfort with research participants so they share honest, detailed insights instead of surface-level answers.

Courtesy Bias: Why Respondents Tell You What You Want to Hear

Courtesy bias is when respondents give overly positive, agreeable answers to avoid offending the interviewer or sponsor. Learn where it shows up, the evidence, and how neutral moderation, anonymity, and behavior-based questions fix it.

Demand Characteristics: When Participants Tell You What They Think You Want

Demand characteristics are the cues in a study that let participants guess your hypothesis and change their behavior to fit it. Learn where they come from, how they differ from social desirability and the Hawthorne effect, and how to design research that captures honest behavior.

Epistemic Territory in Interviews: Ask People About What Is Theirs to Know (2026)

Participants answer questions outside their knowledge as fluently as inside it. How to map the epistemic territory of each question and read agreement tokens as evidence.

Mode Effects: When Letting People Choose Voice or Text Changes the Answer

Pew randomly assigned 3,003 people to phone or web and got answers that differed by up to 18 points on identical questions. Here is what that means when your respondents pick their own mode.

Social Desirability Bias: What It Is and How to Eliminate It in Research

Social desirability bias makes people tell you what sounds good instead of what is true. Learn what causes it, why it quietly wrecks product decisions, and the seven evidence-based ways to reduce it — including why AI-moderated interviews get more honest answers.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.