{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-20T17:11:51.868Z"},"content":[{"type":"documentation","id":"281e7b62-cc1a-4d74-b455-68a461448991","slug":"dispreferred-responses-interview-timing","title":"A Fast No and a Slow Yes: What Response Timing Really Tells You (2026)","url":"https://www.koji.so/docs/dispreferred-responses-interview-timing","summary":"Response latency does not predict acceptance versus rejection below about 700 ms. In 195 measured responses, flat rejections were the fastest type (modal gap about -50 ms, 33.3 percent beginning in overlap) and qualified acceptances the slowest (about 500 ms), inverting the folk heuristic at both ends. Delays beyond about 300 ms predict turn format (hedging) rather than direction. Because conversational gaps average 200 to 300 ms while utterance planning needs at least 600 ms, participants begin composing answers before a question ends, so trailing qualifiers land on an answer already in progress.","content":"Every moderator has been taught some version of the same heuristic: a quick answer means enthusiasm, a pause means reluctance, and a long pause means no. The best quantitative study of response timing in conversation says that heuristic is wrong at both ends. **The fastest responses in the corpus are blunt rejections. The slowest are hedged acceptances.** And below roughly 700 milliseconds, timing does not distinguish a yes from a no at all.\n\nIf you are reading latency as confidence -- in a voice interview, in a sales call recording, or in an analysis pipeline that scores hesitation -- you are systematically misreading the two response types you most need to tell apart.\n\n## What the timing data actually shows\n\nKendrick and Torreira examined 195 responding actions drawn from telephone conversation corpora, measuring the gap between the end of a question, request, offer or proposal and the beginning of the response. They then split responses along two independent dimensions: the **action** (an acceptance or a rejection) and the **turn format** (whether the turn carries hedging, qualification and delay markers, or is delivered flat).\n\nThat split produces four types, and their distribution is the whole story:\n\n| Response type | n | Share | Mode of the gap |\n| --- | --- | --- | --- |\n| Flat rejection | 18 | 9.2% | about -50 ms (in overlap) |\n| Normal acceptance | 87 | 44.6% | about 275 ms |\n| Normal rejection | 54 | 27.7% | about 325 ms |\n| Qualified acceptance | 36 | 18.5% | about 500 ms |\n\nRead the mode column top to bottom. The quickest thing anyone does is reject bluntly -- so quickly that 33.3 percent of flat rejections begin in overlap with the question still being asked, against 1.8 percent of normal rejections. The slowest thing anyone does is accept with qualification. A hedged *yeah, I mean, sure, I guess that could work* arrives later than a clean *no*.\n\nThe authors state the headline finding plainly: examining those 195 cases, they \"found that the timing of the most frequent cases did not differ systematically.\" The difference only emerges in the tail. After approximately 700 ms, the proportion of dispreferred responses exceeds that of preferred ones -- which is why they conclude that a gap of 700 ms, \"but not shorter, could allow one to predict the responding action.\"\n\nAn independent experimental study converges on the same number. Roberts and Francis presented scripted request-acceptance sequences with the gap manipulated from 200 to 1200 ms in 100 ms steps and asked listeners to rate how willing the speaker seemed. Between 200 and 700 ms there were no statistically significant differences in those ratings. Between 700 and 800 ms, ratings of willingness dropped significantly. Two different methods, one threshold.\n\n## Why 700 milliseconds, and why not sooner\n\nThe reason short gaps carry no signal is that there is no room in them for anything to happen.\n\nGaps between turns in ordinary conversation average between 200 and 300 ms. Planning even a simple utterance takes at least 600 ms, on the psycholinguistic evidence. Those two numbers are incompatible unless speakers begin planning their response **before the current speaker finishes**, projecting where the turn is going and preparing the answer against that projection.\n\nThis is not a marginal effect. It is the load-bearing fact about how conversation works, and it has a consequence for interviewing that almost nobody acts on:\n\n**The last few seconds of your question land on an answer that is already under construction.**\n\nThink about what moderators habitually put there. Trailing qualifiers. Examples. Reassurance. *...or, you know, whatever comes to mind, there is no right answer, take your time.* All of it arrives after the participant has identified the projectable end of the question and started composing. At best it is ignored. At worst it redirects a half-planned answer, and the participant restarts -- producing exactly the delay the moderator then reads as reluctance.\n\nThe disciplined version is the opposite of the instinct: **put every qualifier before the question stem, then stop talking.** Say *there is no right answer here* first, ask the question second, and let the projectable end of the turn be the actual end of the turn.\n\n## The gap that does carry information\n\nTiming is not useless. It is just informative about something other than what people assume.\n\nKendrick and Torreira's second finding is that small departures from a normal gap -- anything beyond about 300 ms -- decrease the likelihood of an unqualified acceptance and increase the likelihood that the response, whether it turns out to be an acceptance or a rejection, will arrive in a **dispreferred turn format**. The delay predicts hedging, not direction.\n\nThat is a usable signal, and it points at a different action than the folk version. A 600 ms gap does not mean *they are about to say no*. It means *whatever they say next is likely to come wrapped in qualification* -- and the qualification is the part worth probing. The follow-up is not *you seem hesitant, is that a no?* It is a neutral request for the content of the hedge.\n\nThere is a useful analysis lesson buried in the same paper. Stivers and colleagues, in a sample of 219 polar questions, found mean gaps of approximately 75 ms before affirmations and 400 ms before disaffirmations -- a result that reads like strong support for the folk heuristic. Kendrick and Torreira point out that because only mean gap durations were reported, \"one cannot determine whether the timing of the two alternatives differs systematically\" or whether the difference in means is driven by a tail. Their own distributional analysis shows it is the tail. A difference of means between two heavily overlapping distributions is not a classification rule, and treating it as one is a mistake that recurs well beyond conversation timing -- see [can you average Likert scale data](/docs/can-you-average-likert-scale-data) for the same problem in survey analysis.\n\n## What travels and what does not\n\nResponse latency is not a universal constant, but it is close enough to one to be worth calibrating against. Stivers and colleagues sampled ten languages from traditional indigenous communities to major world languages and found, in their words, \"clear evidence for a general avoidance of overlapping talk and a minimization of silence between conversational turns\" in every one. Languages did differ in average gap, but the spread was contained -- differences fell \"within a range of 250 ms from the cross-language mean.\"\n\nThat is a genuinely reassuring result for anyone running multi-market research, and a caution at the same time. A 250 ms baseline shift between two languages is large relative to the 300 ms departure signal and meaningful relative to the 700 ms threshold. **Never compare raw response latencies across languages or markets without re-establishing the baseline within each one.**\n\nTwo further caveats on this whole body of evidence, stated plainly because they bound what you can do with it. First, these corpora are ordinary telephone conversations among people who know each other, not research interviews with a stranger about a product; treat the thresholds as order-of-magnitude calibration rather than cutoffs. Second, the 700 ms threshold is a shift in the *proportion* of dispreferreds, not a classifier -- plenty of acceptances occur after 700 ms, which is exactly what the qualified-acceptance row in the table above is telling you.\n\n## Practical rules for moderated and AI-led interviews\n\n- **Do not score hesitation as doubt.** Any analysis that treats response delay as an enthusiasm or confidence measure will rate hedged agreement below blunt refusal, which is backwards. If you want confidence, ask for it with a scale question.\n- **Front-load your qualifiers.** Reassurance, framing and examples belong before the question stem, because everything after the projectable end of the turn arrives mid-plan.\n- **Stop talking at the end of the question.** The most common moderator error is filling a 500 ms gap that was never trouble, which converts a normal response into a restarted one.\n- **Treat a >300 ms gap as a flag to probe the hedge, not the direction.** Ask for the content of the qualification.\n- **Re-baseline per language and per modality.** Cross-market latency comparisons need a within-market reference.\n- **Do not read timing at all in text interviews.** Typing latency reflects typing speed, device, distraction and message length, and has no established relationship to preference organisation. The findings above are about speech.\n\n## How this works in Koji\n\nKoji runs both voice and text interviews, and the distinction above is the reason to think about which one you are choosing. The [voice versus text guide](/docs/voice-vs-text-interviews) covers the trade in general; timing adds one specific consideration. Voice preserves the response-latency signal but invites over-reading it. Text destroys the signal entirely -- which is not a loss if you were never going to use it correctly, and a real one if your research question is about hesitation.\n\nMore usefully, Koji's AI moderator does not do the thing that corrupts the signal in the first place. It asks the question once, as written, and stops. It does not fill silence out of discomfort, does not add trailing examples when a participant takes a beat, and does not restate the question in a slightly different form three seconds in -- all habits that human moderators develop precisely because a 500 ms gap feels much longer from the asking side than it is. Because the delivery is identical for every participant, the gaps you do observe are comparable across the whole study rather than confounded with moderator behaviour, a point developed in [interviewer variance and moderator drift](/docs/interviewer-variance-moderator-consistency).\n\nWhere hedging genuinely is the finding, capture it as content rather than inferring it from timing. Koji's six structured question types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- let you pair a yes_no or single_choice commitment with a scale measuring certainty and an open_ended probe for the reservation itself, so the qualification lands in the data as text and a number instead of as a pause someone has to interpret. The [structured questions guide](/docs/structured-questions-guide) covers how to combine types this way, and Koji's automatic analysis scores each conversation for quality on a 1-to-5 scale, with only conversations scoring 3 or above consuming a credit.\n\nThe broader point is that the preference for agreement is structural, not a wording problem. Agreement is cheap to produce and disagreement is expensive, which is why [acquiescence bias](/docs/acquiescence-bias) survives every careful rewrite of the question stem. Timing is one visible symptom of that asymmetry. It is not a way to detect it.\n\n## Frequently asked questions\n\n### Does a long pause in an interview mean the participant is about to say no?\n\nNot reliably, and not below about 700 ms. In a study of 195 responses, the timing of the most frequent acceptances and rejections did not differ systematically. Only after roughly 700 ms does the proportion of rejections exceed acceptances, and plenty of qualified acceptances still occur past that point.\n\n### What is the fastest kind of response in conversation?\n\nA flat rejection -- a blunt no with no hedging. Its modal gap is about -50 ms, meaning it typically begins in overlap with the question, and 33.3 percent of flat rejections start before the question finishes. The slowest common response is a qualified acceptance, at about 500 ms.\n\n### If timing does not predict yes or no, what does it predict?\n\nTurn format. Departures beyond about 300 ms from a normal gap reduce the likelihood of an unqualified acceptance and raise the likelihood that the response will arrive hedged or qualified, whichever direction it goes. So a delay is a signal to probe the qualification, not to infer refusal.\n\n### Why do participants start answering before the question ends?\n\nBecause they have to. Gaps between turns average 200 to 300 ms, while planning even a simple utterance takes at least 600 ms. Speakers project where a turn is going and compose against that projection, which means anything you add after the projectable end of your question lands on an answer already in progress.\n\n### Can I compare response times across different markets or languages?\n\nOnly against a within-market baseline. Across ten languages, all showed the same avoidance of overlap and minimisation of silence, but average gaps varied within a range of 250 ms around the cross-language mean -- large enough to swamp the 300 ms signal if you compare raw latencies directly.\n\n### Does response timing work in text-based interviews?\n\nNo. Typing latency reflects typing speed, device, message length and distraction, and has no established relationship to preference organisation. Koji supports both voice and text interviews; if hesitation is part of your research question, use voice, and capture certainty explicitly with a scale question rather than inferring it.\n\n## Related Resources\n\n- [Structured questions guide](/docs/structured-questions-guide) -- combining yes_no, scale and open_ended questions to capture certainty as data instead of inferring it from pauses.\n- [Acquiescence bias](/docs/acquiescence-bias) -- why agreement is structurally cheaper than disagreement, and what actually reduces it.\n- [Voice vs. text interviews](/docs/voice-vs-text-interviews) -- choosing a modality, including what each one preserves and destroys.\n- [Can you average Likert scale data](/docs/can-you-average-likert-scale-data) -- the same trap of reading a difference of means between overlapping distributions as a classification rule.\n- [Handling difficult interview participants](/docs/handling-difficult-interview-participants) -- practical patterns for vague answerers and yes-people.\n- [Interviewer variance and moderator drift](/docs/interviewer-variance-moderator-consistency) -- why consistent delivery is what makes any cross-participant comparison meaningful.\n","category":"Research Methods","lastModified":"2026-09-20T03:24:31.239085+00:00","metaTitle":"Response Timing in Interviews: A Fast No and a Slow Yes","metaDescription":"Pauses do not mean no. What 195 measured responses show about the 700 ms threshold, and why participants plan answers before you finish asking.","keywords":["dispreferred responses interviews","interview silence meaning","response timing interviews","preference organization","why participants agree","moderator pause interview","conversation analysis research"],"aiSummary":"Response latency does not predict acceptance versus rejection below about 700 ms. In 195 measured responses, flat rejections were the fastest type (modal gap about -50 ms, 33.3 percent beginning in overlap) and qualified acceptances the slowest (about 500 ms), inverting the folk heuristic at both ends. Delays beyond about 300 ms predict turn format (hedging) rather than direction. Because conversational gaps average 200 to 300 ms while utterance planning needs at least 600 ms, participants begin composing answers before a question ends, so trailing qualifiers land on an answer already in progress.","aiPrerequisites":["Experience moderating or analysing interview recordings"],"aiLearningOutcomes":["Stop reading response latency as a confidence or enthusiasm signal","Use the 700 ms and 300 ms thresholds correctly and know what each predicts","Front-load qualifiers so they do not land on an answer already being planned","Capture hesitation as data rather than inferring it from pauses"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}