Back to docs
Research Methods

Mm Hm, Right, Oh: How Your Listening Noises Change the Answer (2026)

The small sounds an interviewer makes while listening are not interchangeable. Learn what mm hm, right, and oh each do to the next thing a participant says.

Most interview training tells you to sound engaged while a participant talks. The conversation-analytic record says something less comfortable: the engaged-sounding tokens are the ones that end the answer, and the nearly contentless one is the only one that reliably buys you more of it.

Mm hm, right, oh, and wow are usually lumped together as backchannels, as though they were a single undifferentiated class of listening noise. They are not. Decades of work on conversational structure treats them as distinct words with distinct jobs, and the jobs point in different directions. If you make the friendly noise at the wrong moment, you do not get a warmer interview. You get a shorter one.

This matters more than it sounds, because these tokens are the part of your interviewing behaviour you are least aware of and least able to standardize across a team.

Four different words, not one class of noise

The research literature separates response tokens into families by what they do to the turn that follows. The distinctions below track Schegloff, Jefferson, Goodwin, and Heritage; Gardner's 2001 monograph remains the most thorough single treatment of the quieter end of the set.

Continuers: mm hm, uh huh

A continuer does one thing: it declines a turn. By producing mm hm, the listener signals that they recognize the speaker has more to say and that they are passing the floor straight back. Schegloff's 1982 account is the standard reference. The token claims nothing about whether you agree, understand, or find the content notable -- and that emptiness is exactly what makes it useful in research. It is the closest thing to a neutral keep going available in speech.

Acknowledgement tokens: yeah, right, okay

These feel warmer, and they behave differently. Work in this tradition, including Drummond and Hopper's analysis of acknowledgement tokens and speakership incipiency, finds that yeah frequently projects incipient speakership -- it is produced by a listener who is moving toward taking the floor, not away from it.

This is the inversion at the centre of the article. The token an interviewer reaches for when they want to seem more supportive than a flat mm hm is a token that, in ordinary conversation, forecasts that the listener is about to start talking. Participants hear that forecast. Many of them wind down to make room for you.

Assessments: wow, great, interesting

An assessment supplies evaluative feedback on what was just said. Goodwin's work on assessments covers the mechanics. In a research interview an assessment is not a neutral encouragement; it is an opinion about the content, delivered by the person who designed the study. Great after one answer and mm hm after the next teaches the participant which kind of answer you are collecting, which is a slow, self-inflicted form of leading that never appears in your interview guide.

Change-of-state tokens: oh

Heritage's 1984 account treats oh as marking that its producer has undergone a change in their "locally current state of knowledge, information, orientation or awareness." It registers the prior turn as news.

Two consequences follow for interviewers. First, oh tells the participant that what they just said was surprising -- useful if you want emphasis, corrosive if you wanted an unbiased account of how normal their behaviour is. Second, and less obviously, marking something as received news tends to close the sequence rather than extend it. You have signalled that the information landed. There is nothing left for the participant to do with it.

The inversion, and what to do about it

The warm tokens are the expensive ones

Rank the four by how supportive they sound and you get roughly: wow > oh > right > mm hm. Rank them by how much additional talk they produce and the order very nearly reverses. The one that sounds least like listening is the one that most reliably keeps a participant going, because it is the only one that commits to nothing.

The practical rule is narrow and easy to apply: use continuers while you want more, and save everything else for when you have deliberately decided to intervene. An oh is a tool for marking a moment you want emphasized. A right is a tool for taking the floor back from a participant who has drifted. Both are legitimate. Neither is neutral, and neither should be produced automatically.

The consistency problem you cannot see

The deeper issue is variance. Nobody writes their response tokens into an interview guide, nobody reviews them in a transcript, and every moderator on a team has a different distribution of them. Two researchers running the same guide can produce systematically different answer lengths and systematically different emphases without either one asking a different question. That is interviewer variance arriving through a channel no amount of guide review will catch, because the guide is identical in both cases.

Cleaned transcripts make this invisible. Most transcription passes strip listener tokens as noise, which means the one artifact you might audit has had the evidence removed before you see it.

Where this stops and active listening starts

This article is about what four specific words do to the following turn. It is not a general guide to listening well -- for the broader skill set, including reflection, summarizing, and silence as a technique, see active listening techniques, and for the relational side, building rapport.

Two adjacent problems belong to other articles on purpose. The length of a pause before an answer is a separate signal with its own literature and its own article. And when a participant interrupts to ask what you meant, that is repair, not backchanneling.

How Koji handles this

Response tokens are an interviewer-variance problem, and the most direct fix for an interviewer-variance problem is to stop having a different interviewer every time.

  • Every participant gets the same listener. Koji's AI interviewer does not get tired at participant fourteen and start saying great to answers it likes. Whatever backchanneling behaviour the study uses, it is identical across every session, so differences between transcripts are differences between participants rather than between moderators.
  • Text interviews remove the channel entirely. In a text-mode Koji interview there are no listener tokens at all, which eliminates this bias class rather than managing it. Voice interviews keep the natural conversational feel where you want it; the AI voice interviews guide covers choosing between the two.
  • Follow-ups replace assessments. The reason interviewers reach for wow is to encourage elaboration. Koji does that with an actual follow-up question instead -- probing the specific thing that warrants more detail, without telling the participant that their answer was surprising or welcome.
  • Structured questions carry the quantitative load. Where you need a clean, uninfluenced measurement rather than a narrative, Koji's six structured question types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- take the answer directly, with no conversational channel to leak through. The structured questions guide covers mixing them with open conversation.
  • Transcripts keep what happened. Koji preserves the conversation as it occurred rather than tidying it into prose, so a reviewer can check what the interviewer actually did between a participant's sentences.

Frequently asked questions

What is a backchannel in an interview?

It is a short response produced while the other person is still speaking -- mm hm, yeah, oh, wow. The term is convenient but misleading, because it implies a single class of noise. These tokens do measurably different jobs, and choosing between them changes what the participant says next.

Is mm hm better than yeah in a research interview?

For getting more talk, yes. A continuer like mm hm declines the turn and passes the floor back. Yeah often projects incipient speakership -- it forecasts that the listener is about to talk -- and participants routinely wind down in response. Use yeah when you actually want the floor.

Why should I avoid saying oh during an interview?

Because oh registers what was just said as news, and marks it as surprising. That biases the participant toward treating their own behaviour as unusual, and it tends to close the sequence rather than extend it. Save it for moments you have decided to emphasize.

Do these tokens really change the data?

They change answer length and they teach participants which answers you reward, which is a form of leading that never shows up in the interview guide. Because no two moderators distribute these tokens the same way, they also inject interviewer variance that guide review cannot catch.

Should transcripts keep listening tokens?

For anything you plan to audit, yes. Most cleanup passes strip them as noise, which removes the only evidence of what the interviewer did while the participant was talking. Koji keeps the conversation as it happened.

How does Koji avoid this problem?

By making the listener identical in every session. Koji's AI interviewer applies the same behaviour to participant one and participant forty, and text-mode interviews remove the channel altogether. Where an interviewer would use an assessment to fish for elaboration, Koji asks a targeted follow-up question instead.

Related Resources