{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-21T16:49:04.209Z"},"content":[{"type":"documentation","id":"b9a4ee9e-914d-4a12-9dea-f058bb74677f","slug":"speech-disfluency-interview-transcripts","title":"Um, Uh, and the False Start: What Transcript Cleanup Deletes (2026)","url":"https://www.koji.so/docs/speech-disfluency-interview-transcripts","summary":"Filled pauses are signals rather than noise. Clark and Fox Tree (2002) showed uh announces a minor upcoming delay and um a major one. Listeners exploit them: fillers speed recognition of the following word (Fox Tree), improve recall, bias listeners toward new referring expressions (Arnold et al.), and reduce the N400 surprise response (Corley et al.). The decisive control is that a matched cough hampers recall while a filler facilitates it (Fraundorf and Watson), showing fillers work as language rather than as interruption. Disfluencies are roughly 5 to 10 percent of natural conversation. Google Cloud Speech-to-Text maps fillers to silence by default while its medical model retains them. Filler rate aggregated per question across participants is a free index of question difficulty. Koji preserves full transcripts and probes hesitant answers with AI follow-ups.","content":"Every research team eventually decides to clean its transcripts. Strip the *um*s, drop the false starts, tidy the half-finished sentences into readable prose. The result is easier to quote and easier to skim, and it has quietly deleted a layer of data that the speech literature has spent thirty years showing is informative.\n\n**The short version: filled pauses are not failures of speech, they are signals in it -- and listeners demonstrably use them.** If *um* were merely the sound of a stalled speaker, any interruption of similar length would have the same effect on a listener. It does not. That single comparison is the cleanest evidence in this area, and it is covered below.\n\nDisfluency is not rare enough to ignore, either. Estimates in the speech-processing literature put disfluencies at roughly 5 to 10 percent of natural human-to-human conversation, and fillers are the most common type -- as one recent survey of the area puts it, fillers *are disfluencies that occur the most frequently compared to other kinds of disfluencies*.\n\n## What the disfluency record actually shows\n\n### Um and uh are not the same word\n\nClark and Fox Tree's 2002 paper in Cognition made the distinction that organizes everything since: speakers use *uh* to announce an upcoming **minor** delay and *um* to announce a **major** one. On their account these are not involuntary leakage but deliberate announcements, which speakers use to implicate a range of things -- that they are searching for a word, deciding what to say next, intending to keep the floor, or preparing to hand it over.\n\nFor an interviewer this is immediately usable. An *um* before an answer forecasts a longer piece of planning than an *uh* does. A participant who produces *um* and then a fluent forty-second answer was assembling something. A participant who produces *uh* and then stops was not.\n\n### The cough test: why this is not just about time\n\nHere is the comparison that settles whether fillers are noise. Fraundorf and Watson examined what happens to listeners' recall when a speaker produces a filler, and compared it against a matched non-linguistic interruption -- a cough. Fillers **facilitated** recall. Coughs **hampered** recall accuracy.\n\nIf the benefit of *um* were simply that it buys the listener a moment, a cough of the same duration would do the same work. It does the opposite. Whatever a filled pause is doing, it is doing it as part of the language, not as an interruption of it. This is the disfluency equivalent of a negative control, and it is the reason the *clean it up, it is just noise* position does not survive contact with the evidence.\n\n### Listeners run on this signal\n\nThe rest of the picture is consistent. Fox Tree found that fillers helped listeners recognize a following target word faster. Arnold and colleagues found that a filler biases listeners toward expecting a **new** referring expression rather than one already introduced into the conversation -- which is precisely what you would expect if *um* means *something harder is coming*. Corley and colleagues found the effect at the neural level: the N400 response, a standard marker of how surprising a word is, was reduced for unpredictable words when a filler preceded them.\n\nTaken together, a filled pause is a small piece of advance notice, and listeners spend it on comprehension.\n\n## Your transcription layer deletes this by default\n\n### The medical exception gives the game away\n\nGeneral-purpose speech recognition treats fillers as garbage. Google's Cloud Speech-to-Text maps paralinguistic sounds such as fillers to silence. In standard spoken-dialogue pipelines, fillers are discarded as noise before anything downstream sees them.\n\nThe exception is the interesting part: Google's medical speech model **includes** fillers when transcribing. The one domain where a wrong transcript has immediate consequences, and where hesitation around a symptom or a dosage is clinically meaningful, is the domain that keeps them.\n\nThat is a direct verdict on the default setting. It is tuned for producing readable text, not for preserving evidence -- and user research is much closer to the medical case than to the dictation case.\n\n### What you lose, concretely\n\nThree things disappear with the *um*s, and none of them are recoverable later:\n\n- **Difficulty.** You can no longer tell which questions were hard to answer, because the marker of planning effort has been removed from every answer equally.\n- **Certainty.** *It costs, um, about forty dollars* and *It costs forty dollars* become the same sentence in your report. They were not the same answer.\n- **Abandonment.** A false start shows a participant beginning one answer and switching to another. The discarded beginning is often the more honest one, and cleanup deletes the evidence that a switch occurred at all.\n\n## Using filler rate as a difficulty index\n\nThere is a cheap, genuinely useful analysis move here that almost nobody runs.\n\n### Aggregate by question, not by participant\n\nCount filled pauses per hundred words **per question**, averaged across participants. Individual filler rates vary enormously between people, so a single participant's rate tells you mostly about them. Averaged across a study, that individual variation washes out and what remains is a property of the question.\n\nA question that produces a consistently elevated filler rate across a dozen participants is not a question your participants found interesting. It is a question they found hard: ambiguous, memory-dependent, or socially awkward. That is a pre-launch signal about question quality and a post-hoc caveat about the answers you collected, and it is available for free in any transcript nobody sanitized.\n\n## Where this stops and the repair article starts\n\nThree neighbours, deliberately excluded.\n\nWhen a participant stops and asks *what do you mean by onboarding?*, that is **repair** directed at your question, and it has [its own article](/docs/repair-clarification-research-interviews). This one covers the speaker working on their own in-progress utterance.\n\n**Silent** gaps before an answer are a distinct signal from filled ones, with a distinct literature; response timing is handled in [a fast no and a slow yes](/docs/dispreferred-responses-interview-timing). And accent-related transcription error is an accuracy problem rather than a cleanup problem -- see [voice research and transcription accuracy](/docs/voice-research-accents-transcription-accuracy).\n\n## How Koji handles this\n\n- **Transcripts are preserved, not prettified.** Koji keeps what the participant actually said. You can read a clean summary when you want one, but the underlying transcript remains available, so the evidence is still there when a finding needs checking. The [viewing interview transcripts](/docs/viewing-interview-transcripts) guide covers where to find it.\n- **Voice interviews collect the signal at all.** A typed answer has no filled pauses in it. If hesitation, planning effort, and self-correction are relevant to your question -- and for pricing, recall, and anything sensitive they usually are -- the interview has to be spoken. Koji runs voice interviews at survey scale without a moderator, which is the combination that makes this practical rather than aspirational.\n- **AI follow-ups act on hesitation in the moment.** The highest-value response to a heavily hesitant answer is a follow-up question while the participant is still present. Koji's AI interviewer probes thin and uncertain answers automatically, which converts a hedge you would otherwise have to interpret into a statement you can quote.\n- **Structured questions stop you inferring certainty.** When you need a confident number rather than a hesitant one, ask for it directly. Koji supports six question types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- and the [structured questions guide](/docs/structured-questions-guide) covers using a scale question for the quantity and an open_ended question for the reasoning behind it.\n- **Analysis reads the conversation, not a summary of it.** Koji records a confidence level on each extracted answer, so answers that were hesitant or incomplete are visible as such rather than being flattened into the same register as everything else.\n\n## Frequently asked questions\n\n### Should I remove um and uh from interview transcripts?\n\nNot from the transcript of record. Clean a quote for a slide if you need to, but keep the original, because filled pauses carry information about planning effort, certainty, and question difficulty that cannot be reconstructed once deleted.\n\n### What is the difference between um and uh?\n\nClark and Fox Tree's account is that *uh* announces a minor upcoming delay and *um* a major one. In practice an *um* before an answer forecasts a longer piece of planning than an *uh* does, which is a useful cue for whether to wait or to prompt.\n\n### Are filled pauses just a sign of nervousness?\n\nSometimes, but that is not the main thing they do. Listeners use them: fillers speed up recognition of the following word, improve recall, and reduce the neural surprise response to unpredictable words. A matched cough does not produce these benefits and actively hurts recall, which is strong evidence fillers are part of the language rather than an interruption of it.\n\n### Why does my transcription tool delete them?\n\nBecause general-purpose speech recognition optimizes for readable text. Google's Cloud Speech-to-Text maps fillers to silence, while its medical model keeps them -- the domain where hesitation carries consequences is the one that preserves it. Research is closer to the medical case.\n\n### How common are disfluencies in normal speech?\n\nEstimates in the speech-processing literature put disfluencies at roughly 5 to 10 percent of natural conversation, with fillers the most frequent kind. They are not an edge case in your data, which is why removing them changes it.\n\n### Can Koji show me which questions were hardest?\n\nKoji preserves the full transcript, so you can compute filler rate per question across participants, and its analysis records a confidence level per extracted answer. A question with consistently hesitant answers across a study is a question worth rewriting before your next round.\n\n## Related Resources\n\n- [Structured questions in AI interviews](/docs/structured-questions-guide) -- the six question types, and when to ask for a number instead of interpreting a hesitation.\n- [Repair and clarification in research interviews](/docs/repair-clarification-research-interviews) -- when the participant interrupts to ask what you meant.\n- [A fast no and a slow yes](/docs/dispreferred-responses-interview-timing) -- what silent gaps before an answer tell you.\n- [Viewing interview transcripts](/docs/viewing-interview-transcripts) -- where the unedited record lives in Koji.\n- [AI voice interviews](/docs/ai-voice-interviews) -- why a spoken interview collects signal a typed one cannot.\n- [Voice research and transcription accuracy](/docs/voice-research-accents-transcription-accuracy) -- accent handling, and the difference between an error and a cleanup.\n","category":"Research Methods","lastModified":"2026-09-21T03:22:47.909434+00:00","metaTitle":"Speech Disfluency in Interview Transcripts: Um and Uh","metaDescription":"Filled pauses are signal, not noise. What um and uh mean, why listeners use them, and why transcription tools delete them.","keywords":["speech disfluency","filled pause","um and uh","interview transcripts","verbatim transcription","false start","transcript cleanup"],"aiSummary":"Filled pauses are signals rather than noise. Clark and Fox Tree (2002) showed uh announces a minor upcoming delay and um a major one. Listeners exploit them: fillers speed recognition of the following word (Fox Tree), improve recall, bias listeners toward new referring expressions (Arnold et al.), and reduce the N400 surprise response (Corley et al.). The decisive control is that a matched cough hampers recall while a filler facilitates it (Fraundorf and Watson), showing fillers work as language rather than as interruption. Disfluencies are roughly 5 to 10 percent of natural conversation. Google Cloud Speech-to-Text maps fillers to silence by default while its medical model retains them. Filler rate aggregated per question across participants is a free index of question difficulty. Koji preserves full transcripts and probes hesitant answers with AI follow-ups.","aiPrerequisites":["Access to raw or lightly edited interview transcripts","Basic familiarity with qualitative analysis"],"aiLearningOutcomes":["Interpret um and uh as planning signals rather than noise","Explain why cleaned transcripts lose difficulty, certainty and abandonment data","Compute a filler-rate difficulty index per question","Decide when an interview needs to be spoken rather than typed"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}