Back to docs
Analysis & Synthesis

Paradata: What Response Time, Hesitation and Drop-Off Tell You About Your Questions

Every interview produces a record of how the answers were produced. Most teams read it to judge respondents. Read it to judge your questions instead, and you get the cheapest instrument improvement available.

Every interview produces two datasets: the answers, and a record of how the answers were produced. The second one is called paradata, and almost every product team throws it away. The teams that do look at it usually look for one thing - bad respondents. That is a legitimate use and a small one. The larger use, and the one this guide is about, is that the same signals indict your questions. A ten-second stall on item 7 is far more likely to mean item 7 is badly written than it is to mean this particular person was distracted, because the stall shows up on item 7 for everybody.

What paradata is

Paradata is, in the standard definition, "auxiliary data collected in a survey that describe the data collection process" (Brady T. West, "Paradata in Survey Research," Survey Practice 4(4), 2011). The term was coined by Mick Couper in a 1998 paper, "Measuring Survey Quality in a CASIC Environment," presented in the Proceedings of the Survey Research Methods Section of the American Statistical Association, and formalised two years later to distinguish paradata, which describe the process, from metadata, which describe the data.

The distinction that trips people up is the one West flags directly: "care should be taken not to confuse paradata with more traditional auxiliary variables." A respondent's company size is not paradata. The fact that they answered the pricing question in four seconds and the security question in fifty-one is.

Paradata familyExamplesWhat it is evidence about
Contact and effortInvitation sends, reminders, attempts before a responseWho is hard to reach, and whether your responders are the easy ones
TimingTotal interview duration, item-level response latency, idle timeWhich items cost effort, and where effort spikes
ProgressBreak-off point, items skipped, sessions resumedWhere the instrument loses people
Response behaviourAnswers changed after selection, backtracking, keystrokesComprehension and option ambiguity
VerbalPauses, hesitation, disfluency, changes in deliveryConceptual misalignment between what you asked and what they heard

The last family used to require a lab. In a voice interview it is a by-product of the recording, which is a genuinely new situation for product research.

The inversion: same signal, different defendant

Most product teams meet paradata through data-quality tooling, where a fast completion is a speeding flag and a straight-line response pattern is a fraud signal. That reading treats the respondent as the defendant. It is a real and necessary discipline, and survey fraud and respondent quality covers it properly.

Survey methodology reads the identical signals with the question in the dock. Both readings are usually available for the same data point, and choosing only the first one is how teams spend three years shipping a question nobody understands.

SignalRespondent-quality readingQuestion-quality readingWhich is more likely
Item answered unusually fastSpeeding, low effortThe item is skimmable, or one option is obviously the expected answerQuestion, if it is fast for most people
Long idle time before answeringDistracted, multitaskingComprehension problem; the item is hard to parse or requires a computationQuestion, if the delay clusters on one item
Break-off at a specific itemLow-commitment participantThat item is the burden cliff, or reads as intrusiveQuestion, almost always
Answer changed after first selectionCareless clickingResponse options are not mutually exclusive, or the stem was misreadQuestion
Hesitation and disfluency in voiceNervousnessThe respondent and the instrument mean different things by a wordQuestion
Slow answer on an attitude itemIndecision, low engagementThe respondent holds a view they believe others do not shareNeither - this is your most valuable respondent

The diagnostic rule that separates the two readings is simple and worth writing on the wall: a signal that concentrates on a person is about the person; a signal that concentrates on an item is about the item. One respondent who answers everything in four seconds is a quality problem. Forty percent of respondents stalling on question 7 is a question problem, and no amount of respondent screening will fix it.

Three findings that make latency worth reading

Response time is a property of the item, not just the person. Yan and Tourangeau's study of web survey response times found that latency is driven by question characteristics - the total number of clauses, the number of words per clause, the number and type of answer categories, and where the question sits in the questionnaire - alongside respondent characteristics such as age, education and internet experience. Because item features move response time, item-level timing is a legitimate instrument diagnostic, not just an attention meter.

Hesitation in speech predicts misunderstanding, not dishonesty. Work on speech survey interfaces by Ehlen, Schober and Conrad models disfluency specifically to predict conceptual misalignment - cases where the respondent and the instrument are using a term differently. Related work by Conrad, Schober and Dijkstra catalogues cues of communication difficulty in telephone interviews. When someone says "well... I guess it depends what you mean by active user," the disfluency is the finding.

Slow answers can mark the opinion you most need. Bassili documented the minority slowness effect: people are measurably slower to express views they believe are not widely shared. In a product context that is the customer who thinks your flagship feature is a waste of time, in a room where everyone else loves it. A pipeline that discards slow responses as low-quality systematically deletes dissent - which is the exact opposite of what a research programme is for.

The question-level paradata review

Run this after every study with more than about 30 responses. It takes fifteen minutes and it is the cheapest instrument improvement available.

What to computeHeuristic thresholdWhat it usually meansAction
Median time per item, rankedAny item over 2x the median of its typeComprehension load or genuine effortRead the item aloud; split it if it has two clauses
Share of respondents below 300 words per minute of reading time on an itemHigh share on a long itemNobody read itShorten to one clause, or convert to a structured type
Break-off rate by item positionAny item with a break-off spikeBurden cliff or perceived intrusionMove it later, make it optional, or ask it conversationally
Answer-change rate per itemAbove about 10 percentOverlapping or unclear optionsRewrite options to be mutually exclusive
Item nonresponse by itemAny item well above the study averageSensitivity or irrelevanceAdd a genuine "not applicable" path
Follow-up depth needed per item (AI-moderated studies)Items that always require a probeThe original question is under-specifiedRewrite the question to ask what the probe asks

That last row is available only in AI-moderated research, and it is the strongest signal in the table. If the AI interviewer has to ask a clarifying follow-up on question 3 in ninety percent of interviews, question 3 is not doing its job - the probe is. Rewrite question 3 to be the probe.

Paradata is data, and data has error

The honest limitation, and West is blunt about it: "The collection of paradata may not be worthwhile if the resulting data are of reduced quality." His review of validation studies found the accuracy and reliability of interviewer observations "can range from quite low (<10%) to relatively high (92%)." Call record data has been found to under-report attempts. Disposition codes get recorded incorrectly. Inter-rater reliability of coded verbal paradata may be low.

Three rules follow:

  1. Never let a paradata signal alone remove a response. Use it to flag, then look at the answer itself.
  2. Timing is contaminated by everything. A respondent on a commute, a slow connection, a phone call - all of it lands in your latency distribution. This is why you compare items within a study, not respondents across studies.
  3. Collect it for a stated purpose. As West puts it, "paradata should be collected for some purpose. The collection and archiving of paradata in the absence of a clearly defined purpose... is a waste of computing system resources." It is also the right stance for participant trust: capture the process signals you will actually use to improve questions, say so in your privacy notice, and do not hoard the rest.

How Koji makes paradata usable

Traditional survey tools give you a completion timestamp and a completion rate. That is enough to know a study went badly and not enough to know which question did it.

Koji produces the process record as a by-product of how the interview works:

  • Per-interview and per-question timing across both voice and text, so you can rank items by effort rather than guessing which one is heavy.
  • Break-off position, so the burden cliff is a location in your guide rather than an aggregate completion percentage. Pair this with survey completion rate to separate "the study is too long" from "question 9 is the problem."
  • Verbal signal in voice interviews. Hesitation, self-correction and "what do you mean by" moments are captured in the transcript rather than lost, which is the family of paradata that used to require a lab.
  • Follow-up depth as a first-class diagnostic. Because the AI interviewer probes vague answers automatically, the number of probes an item needs is itself a measurement of how well the item is written. No static survey tool can produce this number, because a static survey never notices that the answer was vague.
  • Structured questions make items comparable. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and comparing latency within a type is what makes the "2x the median" heuristic meaningful. Comparing a ranking item to a yes_no item tells you nothing; comparing two scale items tells you which scale is broken. See structured questions.
  • A quality gate on the credit ledger. Only conversations that clear a quality score consume a credit, so the respondent-quality reading of paradata is handled for you and you are free to spend your attention on the question-quality reading.

Where this fits

Paradata is the cheapest component of your measurement error to attack, because you already paid to collect it. It sits between two other disciplines: cognitive interviews test questions with a handful of people before launch, and paradata tests the same questions against everybody who answered, continuously, at no extra cost. Cognitive interviewing tells you why an item is hard. Paradata tells you which item to take to a cognitive interview.

Start with one thing on your next study: rank your questions by median response time and read the top three out loud. In most instruments, at least one of them turns out to contain two questions wearing one question mark.

Frequently asked questions

What is paradata in survey and interview research?

Paradata is auxiliary data collected during a study that describes the data collection process rather than the answers themselves - timing, break-off points, contact attempts, answer changes, and in voice research, hesitation and disfluency. The term was coined by Mick Couper in 1998 and is standard in survey methodology, where it is used to monitor data collection and diagnose instrument problems.

How is paradata different from metadata?

Paradata describes the process that produced the data; metadata describes the data itself. The number of seconds a respondent spent on question 4 is paradata. The fact that question 4 is a five-point scale with labelled endpoints is metadata. Both are useful, and confusing them leads teams to file process signals in the schema documentation where nobody looks at them.

Does a fast response mean the respondent was not paying attention?

Sometimes, but the more common explanation is that the question was easy, skimmable, or had an obvious expected answer. The distinguishing test is where the signal concentrates: if one person is fast on everything, that is a respondent-quality issue; if most people are fast on one item, that item is not measuring what you think it is.

Can I use response time to detect low-quality responses?

You can use it as a flag, never as a verdict. Paradata carries its own error - validation studies find reliability of process observations ranging from under 10 percent to over 90 percent depending on the type - so any response flagged by timing should be reviewed on the substance of its answers before being excluded. Removing slow responses is particularly risky, because respondents are measurably slower to state views they believe others do not share.

What is the minority slowness effect?

It is the finding, documented by Bassili, that people take longer to express opinions they believe are not shared by others. For product research this reverses the usual instinct about slow answers: the respondent who takes a long time on "how valuable is this feature to you" may be the one dissenting voice in your sample, and a quality filter that drops slow responses will delete exactly that person.

How do I start using paradata without a methodologist?

Do one thing: after your next study, rank the questions by median response time and by break-off rate, then read the worst three aloud. Rewrite anything that contains two clauses, an undefined term, or an implied right answer. That single loop, repeated study to study, improves an instrument faster than any amount of pre-launch wording debate.

Related Resources

Related Articles

Cognitive Interviews: How to Test Your Survey Questions Before You Launch

A practical guide to cognitive interviewing — the pretesting technique that reveals whether your survey questions and interview guides are understood as intended. Covers think-aloud, verbal probing, sample sizing, and AI-powered approaches.

How Long Should a Survey Be? Ideal Survey Length and Question Count

The data-backed guide to ideal survey length — how many questions to ask, how completion rate drops with each question, the 7-minute abandonment cliff, and why conversational AI interviews beat long static surveys.

Split-Ballot Experiments: How Much of Your Number Is the Question?

Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Completion Rate: How to Stop People Abandoning Your Survey Halfway

How to calculate survey completion rate, why it differs from response rate, what makes people quit mid-survey, and how to fix it — including why conversational formats finish stronger.

Survey Data Quality: How to Detect and Prevent Bad Responses (2026)

The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.

Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)

Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.

Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)

Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.