Back to docs
Research Methods

The Dumping Effect: Why the Attributes You Leave Out Change the Scores of the Ones You Keep (2026)

When a perception has no matching scale it lands on the nearest one. The experimental evidence, why it is invisible in the data, and the only structural defense.

A respondent notices something about your product. Your questionnaire offers no place to record it. They do not discard the perception - they put it on the nearest scale you did give them. The score you get back is therefore partly a measurement of the product and partly a measurement of your attribute list, and nothing in the response itself tells you which is which.

This is the dumping effect. It has been documented experimentally since the early 1990s in sensory science, it is routinely designed against in that field, and it is almost never considered in product research - where fixed-form surveys with five to ten attributes are the default instrument.

The answer, up front

When respondents perceive something your scales do not cover, the perception is displaced onto whichever scale is closest, inflating or deflating it. The effect is real, it is measurable, and rehearsing the task does not remove it. The only structural defense is to make sure there is always somewhere for an unanticipated perception to go: an open, unbounded response channel alongside the closed items, and an interviewer who follows up on what appears there. That is a design property of conversational research and an inherent limitation of fixed forms.

The experiment that pinned it down

Clark and Lawless published the definitive demonstration in Chemical Senses (volume 19, issue 6, 1994, pages 583 to 594), under the title Limiting response alternatives in time-intensity scaling: an examination of the halo-dumping effect. The design is simple enough to restate in a sentence: panelists rated a beverage containing a sweetener plus an aromatic flavouring, and a second beverage containing the sweetener alone, while the experimenters varied how many scales they were allowed to use.

The results, in the authors' own words:

  • "The aromatic flavor caused an increase in sweetness intensity and especially so when the panelists were limited to sweetness responses only."
  • "The odor-induced enhancement of sweetness was smaller when panelists were given both flavor and sweetness response options than when the panelists were given only a sweetness scale."

So the measured sweetness of a fixed physical stimulus changed depending on what else the rating form allowed people to say. Nothing about the product moved. The instrument moved.

Two further findings make this harder to dismiss than a typical order effect.

Practice does not fix it. "Prior use of both scales in a previous experimental session did not lessen the halo-dumping enhancement effect." Respondents who had already used the full set of scales in an earlier session still dumped when the set was narrowed again. You cannot train it away, which rules out the comfortable explanation that it is a novice artifact.

It runs in both directions. "In one study, sweetness ratings of sucrose alone were depressed when the additional scale for flavoring was provided, perhaps due to inappropriate partitioning of responses." Adding a scale moved a score that had nothing to do with the added attribute. The distortion is not simply missing attributes inflate their neighbours; changing the attribute list at all can shift scores in either direction.

A recent review by Spence and Di Stefano in Psychonomic Bulletin and Review (volume 33, issue 4, 2026) summarizes the mechanism as "the tendency of participants to dump their feelings and experience onto whatever response scale they have been presented with, no matter whether those scales capture their experience or not."

The effect is also actively managed in current practice rather than treated as a historical curiosity. Jeong, Kwak and Lim, comparing two sensory profiling methods in Foods (volume 13, issue 17, 2024, article 2853), attribute a weak correlation between methods partly to scale coarseness, noting that a disparity in scale granularity "may lead to a dumping effect, potentially limiting the discriminatory power" of the coarser instrument. And Weir and colleagues, in a 2023 study in Physiology and Behavior (volume 271, article 114331), explicitly presented all of their intensity scales in every condition rather than only the relevant ones, stating that they did so to minimize dumping artifacts - a design decision taken purely to protect the measurement.

What this looks like in product research

Nothing about the mechanism is specific to taste. It requires only a perception, a set of scales, and no matching place to put it.

  • You ask about ease of use and visual design. You do not ask about speed. A user who found the feature sluggish has one usable channel for that irritation, and ease-of-use absorbs it. Your redesign then targets the interface, and the interface was never the problem.
  • You ask about the product. You do not ask about support. A customer who waited nine days for a ticket response rates the product lower. Product quality is now carrying a service-desk metric.
  • You ask a battery of feature satisfaction items with no item for price. Value perceptions distribute themselves across the battery, and the whole battery shifts down together in a way that reads like a broad quality problem.
  • You run a concept test with scales for relevance and clarity but none for trust. A concept that felt intrusive scores low on clarity, and you rewrite copy that was already clear.

In every case the arithmetic is untroubled and the conclusion is wrong. The averages are correct, the confidence intervals are honest, and the number is answering a question you did not ask.

This is not the halo effect

The two are easy to conflate and the confusion leads to the wrong fix.

The halo effect is a global impression bleeding into specific judgements: someone who likes the brand rates every attribute higher, including attributes they have never encountered. The cause is an overall evaluation, the direction is uniform, and the countermeasures are separation, order and forcing discrimination - covered in the halo effect in customer research.

The dumping effect is a specific, real perception with no matching scale being displaced onto an adjacent one. The cause is instrument incompleteness, the direction depends on which scale is nearest, and no amount of randomizing, separating or reordering helps, because the missing channel is still missing.

One useful diagnostic: halo predicts that attributes move together; dumping predicts that a particular attribute moves when an unrelated attribute is added to or removed from the form. That second pattern is what Clark and Lawless observed.

You cannot detect it by reading the responses

This is the part worth sitting with, and it is why the effect survives in mature research programs.

An inflated rating is not malformed. It is a plausible number from an attentive respondent who answered the question you asked, in good faith, on the scale you supplied. There is no data-quality signal to catch - no straightlining, no speeding, no contradiction. Reviewing individual responses more carefully cannot surface it, and neither can a larger sample, because the distortion is a property of the instrument, and every respondent is being measured with the same distorted instrument. More responses give you a more precise estimate of the wrong quantity.

The only way to see it is to change the instrument and watch the scores move. Two practical protocols:

The added-attribute test. Field two versions of the same questionnaire on matched samples, identical except that version B adds one attribute. If the shared attributes score differently across versions, the added attribute was being dumped into them. This is a split-ballot design applied to the attribute list rather than to question wording.

The open-channel audit. Add an unstructured anything else you noticed? question and code what comes back. Any theme that appears there with real frequency and has no matching closed item is a dumping candidate, and its likely landing zone is the nearest semantic neighbour on your form.

The second is cheaper, runs on every study, and is the one to institutionalize.

Why this is a structural argument for conversational research

Here is the uncomfortable implication for the standard survey stack. The fix for dumping is a complete attribute list. A complete attribute list requires knowing in advance everything a respondent might notice. If you knew that, you would not be doing exploratory research.

So completeness is unreachable, and the practical target shifts: not enumerate every attribute, but guarantee that an unanticipated perception has somewhere to land other than your nearest scale.

That is exactly what an unbounded response channel provides, and it is where Koji's approach differs structurally rather than cosmetically from a form builder:

  • Every closed question can carry an open follow-up. Pair scale, single_choice, multiple_choice, ranking and yes_no items with open_ended probes so the respondent always has a non-numeric outlet. The six question types and how to combine them are covered in the structured questions guide.
  • The AI interviewer probes what it did not expect. When a participant volunteers something outside the attribute list, Koji follows the thread rather than discarding it. A static form cannot do this - not because of a missing feature, but because there is no one there to notice.
  • Themes surface without being pre-declared. Because Koji analyzes every transcript rather than a sample, a perception with no matching scale shows up as a named theme in the report. That is the open-channel audit above, running automatically on every study.
  • The closed data stays intact. You still get the distributions and the crosstabs from Koji's structured questions. You just also get the thing the distributions were quietly absorbing.
  • It stays affordable enough to do routinely. A text interview costs 1 credit and a voice interview 3, so the open channel is not a luxury reserved for a handful of moderated sessions.

SurveyMonkey, Typeform and Qualtrics all let you bolt an open text box onto the end of a form. The difference is what happens to it: a text box collects an answer nobody follows up on and most teams never code. Koji treats it as the beginning of a conversation, which is the only version of this that actually recovers the dumped perception.

The general principle

Every closed instrument imposes a vocabulary, and every vocabulary is incomplete. The dumping effect is what incompleteness looks like in the data: the answer to a question you did not ask, recorded in the answer to one you did. It is invisible in the response, invisible in the analysis, and visible only when you change the instrument - which means the discipline it demands is not sharper scrutiny of your data, but a standing habit of leaving a door open.

Frequently asked questions

How large is the dumping effect in practice?

It varies with how closely the missing perception resembles the available scales, and no single number generalizes across domains. What the Clark and Lawless work establishes is that it is large enough to change a conclusion: the same physical stimulus produced systematically different intensity ratings depending only on which scales were offered. Treat it as a threat to validity rather than as a correction factor to subtract.

Does adding more attributes solve it?

Partly, and it introduces its own costs. More scales means longer questionnaires, more fatigue and more straightlining, and the same research showed that adding a scale can itself shift an unrelated rating. The better strategy is a short, well-chosen closed set - built the way an attribute lexicon is built - plus a genuine open channel, rather than an ever-expanding grid.

Is an open text box at the end enough?

It is better than nothing and much weaker than a follow-up. A single end-of-survey box gets short, low-effort answers, is frequently skipped, and typically goes uncoded. The value comes from probing at the moment the perception is live, which is what an AI-moderated interview does by default and what open-ended questions in AI interviews covers in detail.

Can dumping affect just-about-right scales too?

Yes, and the consequences are more expensive because JAR data feeds directly into a fix list. If an attribute has no JAR item, dissatisfaction with it lands on a neighbouring item and shows up as an off-target reading you will then try to fix. Because penalty analysis ranks attributes by the liking they cost, a dumped perception can promote the wrong attribute to the top of a roadmap.

How do I check an existing study retrospectively?

Code the open-ended responses you already have and compare the resulting theme list against the closed attribute list. Themes with meaningful frequency and no matching closed item are your dumping candidates. Then check whether the closed attribute nearest each candidate scored unusually low relative to your benchmarks - that is where the perception most likely landed.

Does this apply to voice interviews as well as text?

The mechanism applies to any closed response format in either mode. The protection is the same in both: conversational modes have an unbounded channel by construction, so a perception with no scale still gets spoken aloud, recorded and analyzed rather than silently redistributed into a number.

Related Resources

Related Articles

The Halo Effect in Customer Research: Why One Good Impression Distorts Every Rating

The halo effect makes one positive impression inflate judgments about everything else — corrupting satisfaction scores, brand ratings, and usability tests. Learn where it hides and how structured, AI-moderated research neutralizes it.

Open-Ended Questions in AI Interviews: How Koji Probes Free-Form Answers for Real Depth

Learn how Koji's open_ended question type works in AI interviews — with automatic probing, theme extraction, and verbatim quote capture that goes far beyond what surveys can do.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.

The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)

A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.

A/B Testing vs. User Research: When to Use Each (And When to Use Both)

Understand when A/B testing and qualitative user research each shine, and how to combine them for better product decisions. Includes framework for choosing methods, real case studies, and how AI interviews make mixed methods accessible.