{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-19T09:53:33.257Z"},"content":[{"type":"documentation","id":"6479e919-19d5-4ee3-abdb-ab8aa0cd429c","slug":"paradata-interview-process-signals","title":"Paradata: What Response Time, Hesitation and Drop-Off Tell You About Your Questions","url":"https://www.koji.so/docs/paradata-interview-process-signals","summary":"Paradata is auxiliary data describing the data collection process rather than the answers: timing, item-level response latency, break-off position, answer changes, and in voice research hesitation and disfluency. The term was coined by Mick Couper in 1998. Product teams typically read these signals to judge respondents for fraud or low effort; survey methodology reads the identical signals to diagnose questions. The distinguishing rule is concentration: a signal concentrated on a person is about the person, a signal concentrated on an item is about the item. Paradata carries its own error, so it should flag rather than exclude. Koji captures per-question timing, break-off position, verbal signal and AI follow-up depth across voice and text, with six structured question types making within-type latency comparison meaningful.","content":"**Every interview produces two datasets: the answers, and a record of how the answers were produced. The second one is called paradata, and almost every product team throws it away.** The teams that do look at it usually look for one thing - bad respondents. That is a legitimate use and a small one. The larger use, and the one this guide is about, is that the same signals indict your questions. A ten-second stall on item 7 is far more likely to mean item 7 is badly written than it is to mean this particular person was distracted, because the stall shows up on item 7 for everybody.\n\n## What paradata is\n\nParadata is, in the standard definition, \"auxiliary data collected in a survey that describe the data collection process\" (Brady T. West, \"Paradata in Survey Research,\" *Survey Practice* 4(4), 2011). The term was coined by Mick Couper in a 1998 paper, \"Measuring Survey Quality in a CASIC Environment,\" presented in the Proceedings of the Survey Research Methods Section of the American Statistical Association, and formalised two years later to distinguish paradata, which describe the process, from metadata, which describe the data.\n\nThe distinction that trips people up is the one West flags directly: \"care should be taken not to confuse paradata with more traditional auxiliary variables.\" A respondent's company size is not paradata. The fact that they answered the pricing question in four seconds and the security question in fifty-one is.\n\n| Paradata family | Examples | What it is evidence about |\n|---|---|---|\n| Contact and effort | Invitation sends, reminders, attempts before a response | Who is hard to reach, and whether your responders are the easy ones |\n| Timing | Total interview duration, item-level response latency, idle time | Which items cost effort, and where effort spikes |\n| Progress | Break-off point, items skipped, sessions resumed | Where the instrument loses people |\n| Response behaviour | Answers changed after selection, backtracking, keystrokes | Comprehension and option ambiguity |\n| Verbal | Pauses, hesitation, disfluency, changes in delivery | Conceptual misalignment between what you asked and what they heard |\n\nThe last family used to require a lab. In a voice interview it is a by-product of the recording, which is a genuinely new situation for product research.\n\n## The inversion: same signal, different defendant\n\nMost product teams meet paradata through data-quality tooling, where a fast completion is a speeding flag and a straight-line response pattern is a fraud signal. That reading treats the respondent as the defendant. It is a real and necessary discipline, and [survey fraud and respondent quality](/docs/survey-fraud-respondent-quality) covers it properly.\n\nSurvey methodology reads the identical signals with the question in the dock. Both readings are usually available for the same data point, and choosing only the first one is how teams spend three years shipping a question nobody understands.\n\n| Signal | Respondent-quality reading | Question-quality reading | Which is more likely |\n|---|---|---|---|\n| Item answered unusually fast | Speeding, low effort | The item is skimmable, or one option is obviously the expected answer | Question, if it is fast for most people |\n| Long idle time before answering | Distracted, multitasking | Comprehension problem; the item is hard to parse or requires a computation | Question, if the delay clusters on one item |\n| Break-off at a specific item | Low-commitment participant | That item is the burden cliff, or reads as intrusive | Question, almost always |\n| Answer changed after first selection | Careless clicking | Response options are not mutually exclusive, or the stem was misread | Question |\n| Hesitation and disfluency in voice | Nervousness | The respondent and the instrument mean different things by a word | Question |\n| Slow answer on an attitude item | Indecision, low engagement | The respondent holds a view they believe others do not share | Neither - this is your most valuable respondent |\n\nThe diagnostic rule that separates the two readings is simple and worth writing on the wall: **a signal that concentrates on a person is about the person; a signal that concentrates on an item is about the item.** One respondent who answers everything in four seconds is a quality problem. Forty percent of respondents stalling on question 7 is a question problem, and no amount of respondent screening will fix it.\n\n## Three findings that make latency worth reading\n\n**Response time is a property of the item, not just the person.** Yan and Tourangeau's study of web survey response times found that latency is driven by question characteristics - the total number of clauses, the number of words per clause, the number and type of answer categories, and where the question sits in the questionnaire - alongside respondent characteristics such as age, education and internet experience. Because item features move response time, item-level timing is a legitimate instrument diagnostic, not just an attention meter.\n\n**Hesitation in speech predicts misunderstanding, not dishonesty.** Work on speech survey interfaces by Ehlen, Schober and Conrad models disfluency specifically to predict conceptual misalignment - cases where the respondent and the instrument are using a term differently. Related work by Conrad, Schober and Dijkstra catalogues cues of communication difficulty in telephone interviews. When someone says \"well... I guess it depends what you mean by active user,\" the disfluency is the finding.\n\n**Slow answers can mark the opinion you most need.** Bassili documented the minority slowness effect: people are measurably slower to express views they believe are not widely shared. In a product context that is the customer who thinks your flagship feature is a waste of time, in a room where everyone else loves it. A pipeline that discards slow responses as low-quality systematically deletes dissent - which is the exact opposite of what a research programme is for.\n\n## The question-level paradata review\n\nRun this after every study with more than about 30 responses. It takes fifteen minutes and it is the cheapest instrument improvement available.\n\n| What to compute | Heuristic threshold | What it usually means | Action |\n|---|---|---|---|\n| Median time per item, ranked | Any item over 2x the median of its type | Comprehension load or genuine effort | Read the item aloud; split it if it has two clauses |\n| Share of respondents below 300 words per minute of reading time on an item | High share on a long item | Nobody read it | Shorten to one clause, or convert to a structured type |\n| Break-off rate by item position | Any item with a break-off spike | Burden cliff or perceived intrusion | Move it later, make it optional, or ask it conversationally |\n| Answer-change rate per item | Above about 10 percent | Overlapping or unclear options | Rewrite options to be mutually exclusive |\n| Item nonresponse by item | Any item well above the study average | Sensitivity or irrelevance | Add a genuine \"not applicable\" path |\n| Follow-up depth needed per item (AI-moderated studies) | Items that always require a probe | The original question is under-specified | Rewrite the question to ask what the probe asks |\n\nThat last row is available only in AI-moderated research, and it is the strongest signal in the table. If the AI interviewer has to ask a clarifying follow-up on question 3 in ninety percent of interviews, question 3 is not doing its job - the probe is. Rewrite question 3 to be the probe.\n\n## Paradata is data, and data has error\n\nThe honest limitation, and West is blunt about it: \"The collection of paradata may not be worthwhile if the resulting data are of reduced quality.\" His review of validation studies found the accuracy and reliability of interviewer observations \"can range from quite low (<10%) to relatively high (92%).\" Call record data has been found to under-report attempts. Disposition codes get recorded incorrectly. Inter-rater reliability of coded verbal paradata may be low.\n\nThree rules follow:\n\n1. **Never let a paradata signal alone remove a response.** Use it to flag, then look at the answer itself.\n2. **Timing is contaminated by everything.** A respondent on a commute, a slow connection, a phone call - all of it lands in your latency distribution. This is why you compare items within a study, not respondents across studies.\n3. **Collect it for a stated purpose.** As West puts it, \"paradata should be collected for some purpose. The collection and archiving of paradata in the absence of a clearly defined purpose... is a waste of computing system resources.\" It is also the right stance for participant trust: capture the process signals you will actually use to improve questions, say so in your privacy notice, and do not hoard the rest.\n\n## How Koji makes paradata usable\n\nTraditional survey tools give you a completion timestamp and a completion rate. That is enough to know a study went badly and not enough to know which question did it.\n\nKoji produces the process record as a by-product of how the interview works:\n\n- **Per-interview and per-question timing** across both voice and text, so you can rank items by effort rather than guessing which one is heavy.\n- **Break-off position**, so the burden cliff is a location in your guide rather than an aggregate completion percentage. Pair this with [survey completion rate](/docs/survey-completion-rate-guide) to separate \"the study is too long\" from \"question 9 is the problem.\"\n- **Verbal signal in voice interviews.** Hesitation, self-correction and \"what do you mean by\" moments are captured in the transcript rather than lost, which is the family of paradata that used to require a lab.\n- **Follow-up depth as a first-class diagnostic.** Because the AI interviewer probes vague answers automatically, the number of probes an item needs is itself a measurement of how well the item is written. No static survey tool can produce this number, because a static survey never notices that the answer was vague.\n- **Structured questions make items comparable.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and comparing latency within a type is what makes the \"2x the median\" heuristic meaningful. Comparing a ranking item to a yes_no item tells you nothing; comparing two scale items tells you which scale is broken. See [structured questions](/docs/structured-questions-guide).\n- **A quality gate on the credit ledger.** Only conversations that clear a quality score consume a credit, so the respondent-quality reading of paradata is handled for you and you are free to spend your attention on the question-quality reading.\n\n## Where this fits\n\nParadata is the cheapest component of your measurement error to attack, because you already paid to collect it. It sits between two other disciplines: [cognitive interviews](/docs/cognitive-interview-guide) test questions with a handful of people before launch, and paradata tests the same questions against everybody who answered, continuously, at no extra cost. Cognitive interviewing tells you why an item is hard. Paradata tells you which item to take to a cognitive interview.\n\nStart with one thing on your next study: rank your questions by median response time and read the top three out loud. In most instruments, at least one of them turns out to contain two questions wearing one question mark.\n\n## Frequently asked questions\n\n### What is paradata in survey and interview research?\n\nParadata is auxiliary data collected during a study that describes the data collection process rather than the answers themselves - timing, break-off points, contact attempts, answer changes, and in voice research, hesitation and disfluency. The term was coined by Mick Couper in 1998 and is standard in survey methodology, where it is used to monitor data collection and diagnose instrument problems.\n\n### How is paradata different from metadata?\n\nParadata describes the process that produced the data; metadata describes the data itself. The number of seconds a respondent spent on question 4 is paradata. The fact that question 4 is a five-point scale with labelled endpoints is metadata. Both are useful, and confusing them leads teams to file process signals in the schema documentation where nobody looks at them.\n\n### Does a fast response mean the respondent was not paying attention?\n\nSometimes, but the more common explanation is that the question was easy, skimmable, or had an obvious expected answer. The distinguishing test is where the signal concentrates: if one person is fast on everything, that is a respondent-quality issue; if most people are fast on one item, that item is not measuring what you think it is.\n\n### Can I use response time to detect low-quality responses?\n\nYou can use it as a flag, never as a verdict. Paradata carries its own error - validation studies find reliability of process observations ranging from under 10 percent to over 90 percent depending on the type - so any response flagged by timing should be reviewed on the substance of its answers before being excluded. Removing slow responses is particularly risky, because respondents are measurably slower to state views they believe others do not share.\n\n### What is the minority slowness effect?\n\nIt is the finding, documented by Bassili, that people take longer to express opinions they believe are not shared by others. For product research this reverses the usual instinct about slow answers: the respondent who takes a long time on \"how valuable is this feature to you\" may be the one dissenting voice in your sample, and a quality filter that drops slow responses will delete exactly that person.\n\n### How do I start using paradata without a methodologist?\n\nDo one thing: after your next study, rank the questions by median response time and by break-off rate, then read the worst three aloud. Rewrite anything that contains two clauses, an undefined term, or an implied right answer. That single loop, repeated study to study, improves an instrument faster than any amount of pre-launch wording debate.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and why comparing latency within a type is what makes it meaningful\n- [Cognitive Interviews](/docs/cognitive-interview-guide) - testing questions with a handful of people before launch\n- [Survey Fraud and Respondent Quality](/docs/survey-fraud-respondent-quality) - the respondent-quality reading of the same signals\n- [Survey Completion Rate](/docs/survey-completion-rate-guide) - separating study length from a single broken question\n- [Survey Data Quality](/docs/survey-data-quality-guide) - detecting and preventing bad responses end to end\n- [Ideal Survey Length](/docs/ideal-survey-length-guide) - how question count drives completion","category":"Analysis & Synthesis","lastModified":"2026-08-17T03:27:04.765031+00:00","metaTitle":"Paradata in Research: Using Response Time and Drop-Off to Fix Your Questions","metaDescription":"Paradata is the process record behind every interview: timing, break-off, answer changes, hesitation. Learn to read it as a question diagnostic, not a respondent verdict.","keywords":["paradata","response latency","survey process data","break-off point","question difficulty","interview metadata","response time survey"],"aiSummary":"Paradata is auxiliary data describing the data collection process rather than the answers: timing, item-level response latency, break-off position, answer changes, and in voice research hesitation and disfluency. The term was coined by Mick Couper in 1998. Product teams typically read these signals to judge respondents for fraud or low effort; survey methodology reads the identical signals to diagnose questions. The distinguishing rule is concentration: a signal concentrated on a person is about the person, a signal concentrated on an item is about the item. Paradata carries its own error, so it should flag rather than exclude. Koji captures per-question timing, break-off position, verbal signal and AI follow-up depth across voice and text, with six structured question types making within-type latency comparison meaningful.","aiPrerequisites":["Familiarity with running surveys or interviews","Basic understanding of data quality checks"],"aiLearningOutcomes":["Define paradata and distinguish it from metadata and auxiliary variables","Apply the concentration rule to decide whether a signal indicts a respondent or a question","Run a question-level paradata review with six standard metrics","Avoid the two common errors: excluding on timing alone, and deleting minority opinion"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}