Back to docs
Research Methods

A Fast No and a Slow Yes: What Response Timing Really Tells You (2026)

The fastest responses in conversation are blunt rejections and the slowest are hedged acceptances. Below 700 ms, timing does not distinguish a yes from a no at all.

Every moderator has been taught some version of the same heuristic: a quick answer means enthusiasm, a pause means reluctance, and a long pause means no. The best quantitative study of response timing in conversation says that heuristic is wrong at both ends. The fastest responses in the corpus are blunt rejections. The slowest are hedged acceptances. And below roughly 700 milliseconds, timing does not distinguish a yes from a no at all.

If you are reading latency as confidence -- in a voice interview, in a sales call recording, or in an analysis pipeline that scores hesitation -- you are systematically misreading the two response types you most need to tell apart.

What the timing data actually shows

Kendrick and Torreira examined 195 responding actions drawn from telephone conversation corpora, measuring the gap between the end of a question, request, offer or proposal and the beginning of the response. They then split responses along two independent dimensions: the action (an acceptance or a rejection) and the turn format (whether the turn carries hedging, qualification and delay markers, or is delivered flat).

That split produces four types, and their distribution is the whole story:

Response typenShareMode of the gap
Flat rejection189.2%about -50 ms (in overlap)
Normal acceptance8744.6%about 275 ms
Normal rejection5427.7%about 325 ms
Qualified acceptance3618.5%about 500 ms

Read the mode column top to bottom. The quickest thing anyone does is reject bluntly -- so quickly that 33.3 percent of flat rejections begin in overlap with the question still being asked, against 1.8 percent of normal rejections. The slowest thing anyone does is accept with qualification. A hedged yeah, I mean, sure, I guess that could work arrives later than a clean no.

The authors state the headline finding plainly: examining those 195 cases, they "found that the timing of the most frequent cases did not differ systematically." The difference only emerges in the tail. After approximately 700 ms, the proportion of dispreferred responses exceeds that of preferred ones -- which is why they conclude that a gap of 700 ms, "but not shorter, could allow one to predict the responding action."

An independent experimental study converges on the same number. Roberts and Francis presented scripted request-acceptance sequences with the gap manipulated from 200 to 1200 ms in 100 ms steps and asked listeners to rate how willing the speaker seemed. Between 200 and 700 ms there were no statistically significant differences in those ratings. Between 700 and 800 ms, ratings of willingness dropped significantly. Two different methods, one threshold.

Why 700 milliseconds, and why not sooner

The reason short gaps carry no signal is that there is no room in them for anything to happen.

Gaps between turns in ordinary conversation average between 200 and 300 ms. Planning even a simple utterance takes at least 600 ms, on the psycholinguistic evidence. Those two numbers are incompatible unless speakers begin planning their response before the current speaker finishes, projecting where the turn is going and preparing the answer against that projection.

This is not a marginal effect. It is the load-bearing fact about how conversation works, and it has a consequence for interviewing that almost nobody acts on:

The last few seconds of your question land on an answer that is already under construction.

Think about what moderators habitually put there. Trailing qualifiers. Examples. Reassurance. ...or, you know, whatever comes to mind, there is no right answer, take your time. All of it arrives after the participant has identified the projectable end of the question and started composing. At best it is ignored. At worst it redirects a half-planned answer, and the participant restarts -- producing exactly the delay the moderator then reads as reluctance.

The disciplined version is the opposite of the instinct: put every qualifier before the question stem, then stop talking. Say there is no right answer here first, ask the question second, and let the projectable end of the turn be the actual end of the turn.

The gap that does carry information

Timing is not useless. It is just informative about something other than what people assume.

Kendrick and Torreira's second finding is that small departures from a normal gap -- anything beyond about 300 ms -- decrease the likelihood of an unqualified acceptance and increase the likelihood that the response, whether it turns out to be an acceptance or a rejection, will arrive in a dispreferred turn format. The delay predicts hedging, not direction.

That is a usable signal, and it points at a different action than the folk version. A 600 ms gap does not mean they are about to say no. It means whatever they say next is likely to come wrapped in qualification -- and the qualification is the part worth probing. The follow-up is not you seem hesitant, is that a no? It is a neutral request for the content of the hedge.

There is a useful analysis lesson buried in the same paper. Stivers and colleagues, in a sample of 219 polar questions, found mean gaps of approximately 75 ms before affirmations and 400 ms before disaffirmations -- a result that reads like strong support for the folk heuristic. Kendrick and Torreira point out that because only mean gap durations were reported, "one cannot determine whether the timing of the two alternatives differs systematically" or whether the difference in means is driven by a tail. Their own distributional analysis shows it is the tail. A difference of means between two heavily overlapping distributions is not a classification rule, and treating it as one is a mistake that recurs well beyond conversation timing -- see can you average Likert scale data for the same problem in survey analysis.

What travels and what does not

Response latency is not a universal constant, but it is close enough to one to be worth calibrating against. Stivers and colleagues sampled ten languages from traditional indigenous communities to major world languages and found, in their words, "clear evidence for a general avoidance of overlapping talk and a minimization of silence between conversational turns" in every one. Languages did differ in average gap, but the spread was contained -- differences fell "within a range of 250 ms from the cross-language mean."

That is a genuinely reassuring result for anyone running multi-market research, and a caution at the same time. A 250 ms baseline shift between two languages is large relative to the 300 ms departure signal and meaningful relative to the 700 ms threshold. Never compare raw response latencies across languages or markets without re-establishing the baseline within each one.

Two further caveats on this whole body of evidence, stated plainly because they bound what you can do with it. First, these corpora are ordinary telephone conversations among people who know each other, not research interviews with a stranger about a product; treat the thresholds as order-of-magnitude calibration rather than cutoffs. Second, the 700 ms threshold is a shift in the proportion of dispreferreds, not a classifier -- plenty of acceptances occur after 700 ms, which is exactly what the qualified-acceptance row in the table above is telling you.

Practical rules for moderated and AI-led interviews

  • Do not score hesitation as doubt. Any analysis that treats response delay as an enthusiasm or confidence measure will rate hedged agreement below blunt refusal, which is backwards. If you want confidence, ask for it with a scale question.
  • Front-load your qualifiers. Reassurance, framing and examples belong before the question stem, because everything after the projectable end of the turn arrives mid-plan.
  • Stop talking at the end of the question. The most common moderator error is filling a 500 ms gap that was never trouble, which converts a normal response into a restarted one.
  • Treat a >300 ms gap as a flag to probe the hedge, not the direction. Ask for the content of the qualification.
  • Re-baseline per language and per modality. Cross-market latency comparisons need a within-market reference.
  • Do not read timing at all in text interviews. Typing latency reflects typing speed, device, distraction and message length, and has no established relationship to preference organisation. The findings above are about speech.

How this works in Koji

Koji runs both voice and text interviews, and the distinction above is the reason to think about which one you are choosing. The voice versus text guide covers the trade in general; timing adds one specific consideration. Voice preserves the response-latency signal but invites over-reading it. Text destroys the signal entirely -- which is not a loss if you were never going to use it correctly, and a real one if your research question is about hesitation.

More usefully, Koji's AI moderator does not do the thing that corrupts the signal in the first place. It asks the question once, as written, and stops. It does not fill silence out of discomfort, does not add trailing examples when a participant takes a beat, and does not restate the question in a slightly different form three seconds in -- all habits that human moderators develop precisely because a 500 ms gap feels much longer from the asking side than it is. Because the delivery is identical for every participant, the gaps you do observe are comparable across the whole study rather than confounded with moderator behaviour, a point developed in interviewer variance and moderator drift.

Where hedging genuinely is the finding, capture it as content rather than inferring it from timing. Koji's six structured question types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- let you pair a yes_no or single_choice commitment with a scale measuring certainty and an open_ended probe for the reservation itself, so the qualification lands in the data as text and a number instead of as a pause someone has to interpret. The structured questions guide covers how to combine types this way, and Koji's automatic analysis scores each conversation for quality on a 1-to-5 scale, with only conversations scoring 3 or above consuming a credit.

The broader point is that the preference for agreement is structural, not a wording problem. Agreement is cheap to produce and disagreement is expensive, which is why acquiescence bias survives every careful rewrite of the question stem. Timing is one visible symptom of that asymmetry. It is not a way to detect it.

Frequently asked questions

Does a long pause in an interview mean the participant is about to say no?

Not reliably, and not below about 700 ms. In a study of 195 responses, the timing of the most frequent acceptances and rejections did not differ systematically. Only after roughly 700 ms does the proportion of rejections exceed acceptances, and plenty of qualified acceptances still occur past that point.

What is the fastest kind of response in conversation?

A flat rejection -- a blunt no with no hedging. Its modal gap is about -50 ms, meaning it typically begins in overlap with the question, and 33.3 percent of flat rejections start before the question finishes. The slowest common response is a qualified acceptance, at about 500 ms.

If timing does not predict yes or no, what does it predict?

Turn format. Departures beyond about 300 ms from a normal gap reduce the likelihood of an unqualified acceptance and raise the likelihood that the response will arrive hedged or qualified, whichever direction it goes. So a delay is a signal to probe the qualification, not to infer refusal.

Why do participants start answering before the question ends?

Because they have to. Gaps between turns average 200 to 300 ms, while planning even a simple utterance takes at least 600 ms. Speakers project where a turn is going and compose against that projection, which means anything you add after the projectable end of your question lands on an answer already in progress.

Can I compare response times across different markets or languages?

Only against a within-market baseline. Across ten languages, all showed the same avoidance of overlap and minimisation of silence, but average gaps varied within a range of 250 ms around the cross-language mean -- large enough to swamp the 300 ms signal if you compare raw latencies directly.

Does response timing work in text-based interviews?

No. Typing latency reflects typing speed, device, message length and distraction, and has no established relationship to preference organisation. Koji supports both voice and text interviews; if hesitation is part of your research question, use voice, and capture certainty explicitly with a scale question rather than inferring it.

Related Resources

Related Articles

Acquiescence Bias: Why Respondents Say Yes (and How to Stop It)

Acquiescence bias is the tendency to agree with survey statements regardless of their content. Learn why it happens, how much it distorts data, and how to design questions that measure genuine opinion.

Can You Average Likert Scale Data? What the Evidence Actually Says (2026)

Yes, in most situations, and the tests will behave. But robustness is about p-values, not meaning: the median can freeze while real change happens, and the mean can rank two groups in the opposite order from every other summary.

How to Handle Difficult Participants in User Interviews (2026 Playbook)

A practical guide to the eight most common difficult participant types in user research — over-talkers, vague answerers, performative responders, no-shows, and more — with the language, probes, and AI techniques that get useful answers without making anyone uncomfortable.

Interviewer Variance: Every Participant Answered a Slightly Different Question (2026)

Nothing is malformed, nobody lied, the arithmetic is right -- and the study is still wrong, because the stimulus was not held constant and the transcript shows a constant one.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Voice vs Text Interview: When to Use Each Mode

Choosing between voice and text mode for your AI interview? This guide breaks down response depth, completion rate, audience fit, and cost — plus a decision matrix that tells you which mode wins for each research scenario.