Why the Feature Request That Reaches You Is Bigger Than the Problem (2026)
Relayed feedback does not fade, it amplifies. The bullwhip effect explains why a request arrives louder and lumpier than the experience behind it.
Short answer: customer feedback does not merely fade as it travels through your organisation - it gets louder and lumpier. Every person who relays a complaint adds their own batching, buffering and advocacy, so the variability of what lands in the roadmap meeting is larger than the variability of what customers actually experienced. Supply chain engineering has measured this precise distortion since 1961 and named it the bullwhip effect. The fix is structural: fewer relays, not better dashboards.
The practical consequence is uncomfortable. When a request reaches you as "customers are screaming about the export flow", the screaming may be a property of your reporting chain rather than a property of your customers - and nothing in the request itself tells you which. Only shortening the chain tells you.
What the bullwhip effect is
In supply chains, the bullwhip effect is the phenomenon "where orders to suppliers tend to have a larger variability than sales to buyers". The further you sit from the actual consumer, the wilder the swings you see. The concept "first appeared in Jay Forrester's Industrial Dynamics (1961)", and it got its name at Procter and Gamble decades later.
The magnitude is not subtle. The standard illustration in the supply chain literature is that "a fluctuation in point-of-sale demand of five percent will be interpreted by supply chain participants as a change in demand of up to forty percent" - an eightfold exaggeration produced by nothing except the act of relaying. Note what is being claimed: not that information was lost, but that it was multiplied.
The import, stated honestly
Three distinct things can happen to a complaint between the customer and you, and they are not the same failure.
The first is destruction. Most feedback never arrives at all, and the losses are patterned rather than random. That is a preservation problem, and it is covered in The Taphonomy of Customer Feedback.
The second is inadmissibility. A relayed claim is a claim about a claim, and it cannot be cross-examined. That is an evidence problem, covered in Hearsay in Product Research.
The third - the subject of this article - is amplification. The complaint does arrive, it is genuinely traceable to a real customer, and its magnitude and timing are still wrong, because the relay chain has its own dynamics. This is the failure mode that survives both of the other fixes. You can have perfect preservation and a named, verifiable source and still be reading a forty percent signal generated by a five percent event.
Amplification is also the one of the three that gets worse as your feedback operation gets more professional, because professionalising a feedback pipeline usually means adding layers to it.
The flat-demand proof: the babies were fine
The Procter and Gamble case is the cleanest demonstration in the literature, because the underlying demand was known to be constant. Hau Lee, V. Padmanabhan and Seungjin Whang, writing in MIT Sloan Management Review in April 1997, put it plainly: "While the consumers, in this case, the babies, consumed diapers at a steady rate, the demand order variabilities in the supply chain were amplified as they moved up the supply chain." P&G named the phenomenon the bullwhip effect.
Sit with that. Babies do not have quarters. Nappy consumption is about as stationary a demand signal as commerce offers. Every bit of the volatility P&G saw in its order book was manufactured downstream of the babies, by the chain itself.
Now translate. Your users experience a steady, unremarkable amount of friction with your export flow. By the time that steady friction has been through a support macro, a weekly ticket roundup, a customer success escalation and a quarterly roadmap submission, it can arrive as a crisis or as nothing, and the difference between those two outcomes may have nothing to do with the users. The Koji view is that this is the strongest available argument for talking to customers directly rather than efficiently: efficiency in a relay chain is what creates the distortion.
Why better information does not fix it
The instinctive response is to instrument the chain - more dashboards, better tagging, a single source of truth. The experimental evidence says this is not enough on its own, and the evidence is unusually clean.
The Beer Distribution Game is the canonical teaching simulation for this, studied formally by John Sterman (Management Science, volume 35, issue 3, 1989, pages 321-339, doi:10.1287/mnsc.35.3.321). It "illustrates in a compelling way the effects of poor system understanding and poor communication for even a relatively simple and idealized supply chain". Players reliably produce wild oscillations from steady demand.
Here is the part that matters. Players blame their lack of visibility into real customer orders, and they are wrong: "analysis of the minimum possible score using the optimal strategy under different conditions shows an expected value of perfect information of 0 for the standard game, and simulations that included giving players perfect information still showed poor team performance."
An expected value of perfect information of zero is a strong claim. It says the distortion in that system is not an information deficit. It is a structural property of a chain with delays and independent local decisions. Handing everyone the customer data does not repair it, because each link still reacts locally, and the reactions still compound.
This is the single most actionable finding in this article, because it redirects your effort. If more visibility will not fix an amplifying chain, the only remaining lever is the number of links.
The four relay behaviours that amplify feedback
The supply chain literature attributes the effect to a small set of local behaviours, each of which has an exact analogue in a product feedback pipeline.
Demand forecast updating. In supply chains, forecasting "is applied individually by all members of a supply chain". Each link re-forecasts from the distorted signal it received, so each link's own inference gets layered onto the next one's input. In feedback terms: a support lead infers "this is trending", a CS manager infers "this is a churn risk" from the support lead's inference, and a PM infers "this is our biggest problem" from that. Three inferences stacked on one ticket.
Order batching. "Order batching is the preference of most companies to accumulate demand before ordering." A steady trickle becomes a spike purely by being accumulated. The weekly feedback digest and the quarterly planning cycle are order batching. Ten identical complaints arriving one per week read as background; the same ten arriving in one Monday roundup read as an emergency.
Delay in the circulating signal. The classic worked example is unforgiving: if a retailer sees a permanent drop of ten percent in demand on day one, "he will not place a new order until day 10", "the wholesaler is going to notice the 10% drop at day 10 and will place his order on day 20", and meanwhile the producer "had a surplus of 11% per day, accumulated since day 1". The error does not sit still while the news travels - it accrues.
Exaggeration under scarcity. When capacity is rationed, participants inflate their requests to secure a share. Roadmap slots are rationed capacity. An advocate who knows only three items ship per quarter has a rational incentive to describe their item as existential, and that inflation is indistinguishable, downstream, from severity.
What this changes about how you read a request
Practical moves, in order of how much distortion they remove:
- Count the hops. Before acting on any request, write down the number of humans between the customer and the sentence you just read. Two is common. Four is not rare. This number is the single best predictor of how much you should distrust the stated magnitude.
- Ask for the denominator, not the anecdote. "Several customers" is an amplified quantity. "Nine of the forty-one people we asked" is not. A request that cannot produce a denominator has been through a forecast-updating layer.
- Break the batch. Look at arrival times rather than the digest. A spike in a roundup that corresponds to a flat arrival rate is a batching artefact, and you can see it in the timestamps.
- Treat severity claims and frequency claims separately. Amplification hits magnitude hardest. See Usability Issue Severity Ratings for scoring severity on its own evidence.
- Go direct for anything expensive. For any decision above a threshold you set in advance, do not accept a relayed signal at all. Re-measure at the source.
The last one used to be the expensive option, which is exactly why relay chains grew in the first place. Koji exists to make it the cheap one.
How Koji handles this
The bullwhip literature says the structural fix is to shorten the chain and remove the layers that re-forecast and batch. Koji is built to make a zero-relay chain cheap enough to be the default.
- AI-moderated interviews remove the relay entirely. Instead of reading a summary of a summary, you run interviews with fifty customers this week and read what they said. There is no intermediate link to add its own forecast, because there is no intermediate link.
- Structured questions give you a denominator. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - asked of every participant. Asking all forty-one people a yes_no or scale question about the export flow produces a rate, and a rate cannot be amplified by an advocate. This is the direct antidote to point two above.
- Real-time reporting removes the batch. Results land as interviews complete rather than being accumulated into a weekly digest, which eliminates the order-batching spike at its source.
- Automatic thematic analysis replaces stacked human inference. One analysis pass over the raw transcripts, rather than three humans each re-inferring from the previous summary.
- Voice interviews and customisable AI consultants let you re-measure at the source in days, which is what makes "go direct for anything expensive" a realistic policy rather than an aspiration.
Where a traditional survey tool like SurveyMonkey gives you an instrument you still have to staff, recruit for and analyse - which is precisely why teams fall back on relayed feedback between studies - an AI-native platform collapses the cost of going direct. Teams that adopt AI-assisted research report substantially faster time-to-insight, and the mechanism is not magic: it is the removal of handoffs.
You do not need a PhD in research methods to run this. You need the chain to be short.
Frequently asked questions
What is the bullwhip effect in customer feedback?
It is the tendency for the variability of a relayed complaint to grow at each handoff, so the request that reaches a product team is louder and spikier than the customer experience that caused it. The name comes from supply chains, where orders to suppliers vary more than sales to buyers, and the mechanism transfers directly to any multi-layer feedback pipeline.
Is this the same as the telephone game?
No, and the difference matters. The telephone game is about corruption of content - the message changes. The bullwhip effect is about amplification of magnitude and timing - the message can stay perfectly accurate in substance while its apparent size and urgency multiply. A chain of scrupulously honest relayers still produces a bullwhip.
Does more feedback data fix the amplification?
Usually not. In the Beer Distribution Game the expected value of perfect information is zero for the standard setup, and simulations that gave players perfect information still produced poor performance. Amplification comes from the structure of a delayed, locally-optimising chain, not from a shortage of data. Shortening the chain works; adding dashboards to it mostly does not.
How do I tell an amplified request from a real one?
Ask for a denominator and a timestamp. Real signals survive both questions: someone can tell you how many of how many, and when each instance arrived. Amplified signals collapse into "several customers" and a digest date. Koji makes the denominator available by default, because structured questions are asked of every participant rather than emerging from whoever complained loudest.
Does this mean I should ignore my support team?
No. Support is an excellent detector and a poor quantifier. Use relayed feedback to decide what to investigate and never to decide how big something is. The detection function of a relay chain is valuable and the magnitude function is compromised, and separating those two uses is the whole discipline here.
How many relay layers is too many?
Two is where measurable distortion begins, because that is enough for one link to re-forecast another link's inference. Beyond three, treat all stated magnitudes as unusable and re-measure at the source before committing engineering time.
Related Resources
- Structured Questions in AI Interviews - the six question types, and why asked-of-everyone produces a denominator
- The Taphonomy of Customer Feedback - the complementary failure, where evidence is destroyed rather than amplified
- Hearsay in Product Research - why a relayed claim is not evidence, even when it is accurate
- Proxy Response Bias - when someone answers on the customer's behalf
- The Complete Guide to Thematic Analysis - doing the inference once, over raw data
- Usability Issue Severity Ratings - scoring severity on its own evidence rather than on relayed urgency
Related Articles
Hearsay in Product Research: Why a Relayed Customer Claim Is Not Customer Evidence
Most product decisions rest on relayed claims about what customers want. Borrow the law of evidence's hearsay rule to grade every claim, and promote the ones that matter to first-hand evidence in 48 hours.
Proxy Response Bias: Why the Question, Not the Person, Decides How Wrong a Proxy Is (2026)
Observability and interaction explain over 60% of the gap between self-reports and proxy reports. Proxy error is systematic and directional, which is why adding more proxies makes you more confident and no more correct.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Taphonomy of Customer Feedback: Which Complaints Survive to Reach You (2026)
Most customer feedback is destroyed before it reaches you, and the filter has a predictable shape. A taphonomic method for naming the evidence classes your channels systematically lose.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.
Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)
How to rate the severity of usability problems using Nielsen's 0-4 scale, why single-evaluator ratings are unreliable, how to separate severity from priority, and how to replace guessed frequency estimates with measured data.