Back to docs
Research Methods

Why the Loudest Complaint Hides the Real One: Masking in Customer Interviews (2026)

One dominant complaint does not just take up airtime - it raises the threshold for everything quieter, asymmetrically, and your analysis then discards what was buried. A protocol for hearing the masked signal.

A participant with one furious complaint will give you a clean, confident, useless interview. The complaint is real. The problem is that it raises the bar for everything else they might have told you - and the material that falls below that bar does not show up in your transcript as missing. It shows up as absent.

This is structurally identical to auditory masking, and the analogy is worth taking seriously because acoustics has already mapped the shape of the effect: it is asymmetric, it works backwards in time, and the engineering discipline built on top of it deliberately throws away the masked signal on the grounds that nobody will notice.

The short answer

  • A dominant topic suppresses the reporting of quieter ones. The suppression is not symmetric - big masks small, small never masks big.
  • Masking works backwards as well as forwards, so a complaint raised late can bury something said earlier in the same session.
  • Analysis by airtime or mention-count inherits the masking, then launders it into a ranked theme list that looks complete.
  • The fix is not better listening. It is asking about the quiet categories explicitly, so they do not have to compete for the floor.

What masking actually is

In hearing, a sound that is perfectly audible on its own becomes inaudible in the presence of another. The quantity is measured, not metaphorical. The unmasked threshold is "the quietest level of the signal which can be perceived without a masking signal present"; the masked threshold is "the quietest level of the signal perceived when combined with a specific masking noise." A standard worked example takes a tone audible at 10 dB SPL on its own and shows it requiring 26 dB SPL in the presence of a masker - 16 dB of masking. The signal did not change. The threshold moved.

Two properties of the effect matter for interviews.

It is asymmetric. Masking is not a fair fight in which the louder of two signals wins by a proportional margin. Low-frequency maskers "are effective over a wide frequency range" while high-frequency maskers work "over a narrow range of frequencies". The result is called the upward spread of masking, and it "is why an interfering sound masks high frequency signals much better than low frequency signals." A big, broad, dominant signal reaches across and suppresses things unlike itself. A small one does not reach back.

It runs backwards. Temporal masking is not confined to what comes after the masker. Backward masking obscures sounds "immediately preceding the masker", with the effect lasting roughly 20 ms at onset and about 100 ms at offset. The interference travels in both directions in time.

The import, stated honestly

None of the above is a finding about interviews. It is a finding about ears. What transfers is the structure of the mechanism, and structure is the useful part of any import: it tells you what shape of failure to look for and which measurements will be corrupted.

The structural claim is this. In an interview, attention and conversational floor are finite and shared. A topic with high salience for the participant raises the threshold a competing topic must clear before they judge it worth raising. That is not an exotic hypothesis - it is the ordinary consequence of relevance judgements under a limited turn budget. Three consequences follow directly from the acoustic shape:

  • Asymmetry. The dominant complaint suppresses reporting of minor friction. Minor friction never suppresses the dominant complaint. Your theme counts are therefore biased in one direction only, which means they cannot be corrected by averaging across participants who all share the same masker.
  • Backward reach. If the big complaint surfaces at minute 20, it does not only crowd out minutes 21 to 40. It retroactively reframes what the participant said at minute 18, both for them and for whoever reads the transcript. The earlier remark gets recoded as a lesser instance of the later theme.
  • Inaudible loss. Whatever falls below threshold leaves no trace. This is the dangerous part, and it has an exact engineering parallel.

The codec problem

Perceptual audio coding exists because masking is predictable. A psychoacoustic model "provides for high-quality lossy signal compression by describing which parts of a given digital audio signal can be removed or reproduced with reduced quality without significant loss in the perceived quality of the sound." MP3 and AAC do not compress by being clever about redundancy alone. They compute what is masked and discard it, confident that the listener will not hear the absence.

Thematic analysis does the same thing, accidentally, and with the same reassuring result. Themes are ranked by how often and how forcefully they were raised. Material that was masked was raised rarely and weakly, so it ranks low or does not appear. The report that comes out is coherent, confident and audibly complete - for exactly the reason a 128 kbps file sounds complete. The discarded content was selected for being unnoticeable.

A lossy codec at least knows its bitrate. A theme list does not report what it dropped.

Frequency was never the importance axis anyway

There is a tempting response here: if quiet issues are under-counted, count harder. It does not work, and usability research already has the counterexample.

Jakob Nielsen reported on six user interfaces evaluated by heuristic evaluation and found 59 major usability problems against 152 minor ones. His conclusion was that "the lists of usability problems found by heuristic evaluation will tend to be dominated by minor problems". Note that this is the opposite failure from masking - here the trivia crowd out the serious issues by sheer count - and yet it teaches the same lesson. Raw frequency does not track importance in either direction.

That is precisely why Nielsen proposed severity ratings as a supplement. He defines severity as "a combination of three factors: The frequency with which the problem occurs: Is it common or rare? The impact of the problem if it occurs: Will it be easy or difficult for the users to overcome? The persistence of the problem: Is it a one-time problem that users can overcome once they know about it or will users repeatedly be bothered by the problem?" Frequency is one of three inputs, and the only one masking corrupts.

So the defensible position is not "count the quiet things more carefully." It is: stop inferring importance from how much airtime a topic won in a conversation where airtime was contested.

Unmasking: a protocol

  • Do not make quiet topics compete for the floor. Ask about them directly. A question that is asked of everyone cannot be masked, because it does not depend on the participant deciding it was worth raising.
  • Discharge the masker first, deliberately. Let the participant say the big thing in full, acknowledge it, mark it as captured, then explicitly change frame. An unfinished complaint keeps masking; a complaint the participant believes has landed stops competing.
  • Separate elicitation from ranking. Gather candidate issues in one pass, then ask the participant to rank or rate them in another. Ranking forces a distribution across items rather than letting the loudest item absorb the session.
  • Count presence per participant, not mentions. One person mentioning something nine times is one person. Mention counts are the measurement masking corrupts most directly.
  • Watch for the masker you introduced. If your first question named the dominant topic, you installed the masker yourself, and every subsequent silence about anything else is an artefact of your interview guide rather than a finding about your product.

How Koji handles this

Masking is a floor-allocation problem, which makes it unusually tractable for a structured AI interview.

  • Structured questions guarantee coverage. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Every participant is asked the scale, choice and ranking items regardless of how the open conversation went, so a quiet category cannot be masked out of existence by a loud one. See the structured questions guide.
  • Ranking questions force a distribution. A ranking item makes a participant place every option relative to the others, which surfaces the second and third priorities that free conversation buries.
  • AI follow-up probing keeps going after the big answer. Koji AI-moderated and voice interviews continue probing once the dominant complaint is fully expressed, rather than accepting a strong answer as a finished topic the way a rushed human interviewer often does.
  • Consistency across sessions. Because Koji asks the agreed questions the same way every time, a topic that is quiet in one interview and quiet in fifty is genuinely quiet - not a casualty of which interviewer ran which session.
  • Real-time reporting separates the axes. Scale and choice distributions are reported as distributions, so importance can be read from ratings rather than inferred from how often something came up.

The point is not that Koji listens better. It is that a question asked of everyone is immune to the competition for airtime that produces masking in the first place.

Frequently asked questions

What is masking in a customer interview?

It is the suppression of a quieter concern by a dominant one. The quiet concern is not absent from the participant experience; it simply fails to clear the threshold of being worth raising while a bigger issue occupies the conversation. The term is borrowed from auditory masking, where a measurable sound becomes inaudible in the presence of another.

Is this the same as anchoring bias?

No. Anchoring is about a reference value pulling subsequent judgements toward it. Masking is about a strong signal raising the detection threshold for a weak one, so the weak signal is not reported at all. Anchoring distorts the answers you get; masking removes answers you would otherwise have got.

Why does masking make theme counts unreliable?

Because the suppression is asymmetric. A dominant theme reduces the mention count of minor themes, while minor themes never reduce the count of the dominant one. The bias runs in a single direction, so it does not cancel out when you aggregate across participants who share the same dominant concern.

Can I just ask participants what else is bothering them?

It helps, but an open catch-all question is still subject to the same threshold - the participant judges what is worth raising, and the big issue is still setting the bar. Specific questions about named categories work far better, which is why Koji structured questions are asked of every participant rather than left to emerge.

Does masking affect what was said earlier in the interview?

Yes. By analogy with backward temporal masking, a topic raised late can retroactively reduce the weight given to something mentioned earlier, both by the participant and by whoever analyses the transcript. The earlier remark tends to be recoded as a minor instance of the later dominant theme rather than as a separate finding.

How do I tell a genuinely minor issue from a masked one?

Compare how the topic performs when it is asked about directly versus when it has to emerge on its own. An issue that scores low on a direct rating question from everyone is genuinely minor. An issue that is rarely raised spontaneously but rates highly when asked about explicitly was being masked.

Related Resources

Related Articles

Anchoring Bias in Research and Surveys: How the First Number Skews Every Answer

Anchoring bias makes the first number a respondent sees pull every later judgment toward it — distorting pricing research, scale questions, and willingness-to-pay studies. Learn how to design anchors out, including with AI-moderated interviews.

Data Saturation in Qualitative Research: How to Know When You Have Enough

Data saturation is the point at which additional interviews stop producing new information. This guide covers the four types of saturation (theoretical, data, code, meaning), how to recognize and document them, the empirical sample sizes from Hennink and Guest, and how AI-moderated interviews let you reach saturation in days instead of months.

How Many Interviews Are Enough? A Guide to Sample Size

Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.

Probing and Follow-Up Questions: Going Deeper in Research Interviews

Learn the different types of probing questions — clarification, elaboration, and contrast — and when to use each to get richer qualitative data from your participants.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

The Taphonomy of Customer Feedback: Which Complaints Survive to Reach You (2026)

Most customer feedback is destroyed before it reaches you, and the filter has a predictable shape. A taphonomic method for naming the evidence classes your channels systematically lose.

The Complete Guide to Thematic Analysis

Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.