{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T17:45:45.900Z"},"content":[{"type":"documentation","id":"58f01177-5762-4480-a3e2-d9c09f92254a","slug":"feedback-volume-reporting-propensity","title":"Feedback Volume Tracks Attention, Not Incidence: How to Read a Complaint Trend","url":"https://www.koji.so/docs/feedback-volume-reporting-propensity","summary":"Every point on a complaint trend is the product of three terms - exposure, incidence and reporting propensity - and only the product is observable. FDA states that \"Many factors can influence whether an event will be reported, such as the time a product has been marketed and publicity about an event\", naming the two best-documented propensity effects: the Weber effect and notoriety bias. Hartnell and Wilson replicated the Weber effect in 2004, but Hoffman and colleagues analysed sixty-two drugs in 2014 and found most showed little evidence for it. That failure to replicate is worse news than confirmation, because a stable propensity curve could have been divided out; an unstable one cannot. A worked case shows reported volume rising 8% while real incidence fell 40%, because a fix lowered incidence while its announcement raised propensity. The fix is a fixed instrument on a chosen sample, which is what a Koji study with reusable structured questions provides.","content":"## The short answer\n\nWhen complaint volume for a feature doubles, the most common explanation is not that the feature got twice as bad. It is that something changed how likely people were to tell you. Attention, publicity, a new in-app prompt, a status page, a support hiring wave and simple age of the feature all move reported volume without touching the underlying failure rate.\n\nThis matters because the trend line is the artifact most teams trust most. A single ticket gets discounted as an anecdote; a rising line gets treated as evidence. It is often weaker evidence than the single ticket, because a line invites you to read a slope that three different mechanisms could have produced. The way out is to measure incidence on a channel you control, which is what a platform like Koji is for.\n\n## The quantity you are actually plotting\n\n### The identity you are actually plotting\n\nEvery point on a complaint trend is a product of three terms:\n\n**reports = exposure x incidence x reporting propensity**\n\nExposure is how many people used the thing. Incidence is the fraction of them who hit the problem. Reporting propensity is the fraction of those who told you. Only the product is observable. Any one of the three can move while the other two sit still, and all three commonly move at once.\n\nThe term you care about is incidence. It is the only one that describes your software. The other two describe your growth and your users, and both are moving constantly - exposure because you are shipping and marketing, propensity because you keep changing how easy and how socially expected it is to complain.\n\n### Why any one factor moving looks like the product changing\n\nA 40% rise in export complaints is equally consistent with all of these:\n\n| What actually happened | Exposure | Incidence | Propensity |\n| --- | --- | --- | --- |\n| The export bug got worse | flat | up 40% | flat |\n| Marketing drove a campaign at power users | up 40% | flat | flat |\n| You added a \"Report a problem\" button next to export | flat | flat | up 40% |\n| A competitor blog post named your export bug | flat | flat | up 40% |\n| You fixed it and announced the fix loudly | flat | down | up more |\n\nThe last row is the one that should worry you, and it gets its own section below. The point of the table is that the trend line is identical in all five cases. Nothing in the shape of the curve distinguishes them. You have to bring outside information.\n\n## The regulator names both mechanisms in its own disclaimer\n\nFDA, which runs the adverse event reporting system formerly called FAERS and now being consolidated as the Adverse Event Monitoring System (AEMS), states the problem in a single sentence on its public dashboard documentation:\n\n> \"Many factors can influence whether an event will be reported, such as the time a product has been marketed and publicity about an event.\"\n\nThose two examples are not casual. They are the two best-documented reporting-propensity effects in safety surveillance, and each has a name. Time on market is the Weber effect. Publicity is notoriety bias. Both have direct product analogues, and neither has anything to do with whether the product changed.\n\n## The Weber effect, and the honest state of the evidence\n\n### What Weber claimed\n\nThe original observation was that reporting of adverse events follows a predictable arc after launch rather than tracking real risk. A replication study by Hartnell and Wilson in Pharmacotherapy in 2004 applied it to the drugs Weber had originally studied, setting an explicit criterion: a reporting pattern counted as demonstrating the effect if \"the highest peak in reports during the first 5 years after product approval occurred during year 2\". Against that criterion their result was unambiguous, though the sample was small: \"All five drugs analyzed in this study demonstrated the Weber effect.\"\n\nThe mechanism is intuitive. At launch, few people are exposed and nobody is looking. Through the first two years, exposure grows, prescribers get curious, and reporting climbs. After that, the product becomes unremarkable, reporting fatigue sets in, and volume declines even as exposure keeps rising.\n\n### What happened when someone checked sixty-two drugs\n\nThen a larger study looked again. Hoffman and colleagues, writing in Drug Safety in 2014, analysed \"Sixty-two drugs approved by the FDA between 2006 and 2010\" against the standard definition, which they stated as the claim that \"AE reporting peaks at the end of the second year after a regulatory authority approves a drug.\"\n\nTheir finding: \"a majority of the drugs showed little evidence for the effect.\"\n\n### Why the failure to replicate is the worse news\n\nThe instinct is to file this as good news. It is the opposite, and this is the most useful thing in the article.\n\nIf the Weber effect were a stable, universal curve, it would be a *correction*. You could divide it out. Reporting propensity would be a known function of time since launch, you would deflate each point by it, and what remained would be an estimate of incidence. A reliable bias is not really a bias; it is a calibration.\n\nWhat the 2014 result says is that the curve is not stable across products. Some show it, most do not, and you cannot know in advance which kind yours is. That makes reporting propensity an *uncorrectable* nuisance rather than a correctable one. The contested literature is a worse situation than a confirmed effect would have been, because it removes the possibility of a standard adjustment.\n\nThe practical consequence: there is no detrending step you can add to your dashboard that converts reported volume into incidence. Anyone who offers you one is assuming a curve that the best available evidence says does not generalise.\n\n## Notoriety: the spike you caused\n\n### Stimulated reporting: the spike you caused\n\nThe second mechanism FDA names is publicity, and in product work you are usually the publisher. Every one of these raises reporting propensity without changing a line of code:\n\n- Adding or moving a feedback entry point, especially into the failure path itself\n- Posting an incident to a status page\n- Shipping a changelog entry that names the area\n- A support agent asking \"has this happened before?\" in a macro\n- A community thread, a competitor comparison post, or a viral complaint\n- Running a survey about the feature, which teaches people the topic is fair game\n\nThe in-product case is the sharpest. Put a \"Report a problem\" link on the export screen and export complaints will rise, immediately and substantially, with zero change in the export failure rate. The channel got cheaper, so more of the existing failures converted into reports. You improved your instrument and it looked like a regression. An invited Koji study avoids this failure mode entirely, because the invitation rather than the user's motivation decides who answers.\n\n### The announcement that cancels its own fix\n\nNow combine the two. You find the export bug, fix it, and publish a note telling affected customers it is fixed.\n\nThe fix lowers incidence. The announcement raises propensity. They push the observable in opposite directions, and the arithmetic is unforgiving. Suppose incidence falls 40%, from 100 affected users per 10,000 to 60, while the announcement lifts reporting propensity from 5% to 9%:\n\n| | Incidence per 10,000 | Propensity | Reports per 10,000 |\n| --- | --- | --- | --- |\n| Before the fix | 100 | 5% | 5.0 |\n| After the fix and the note | 60 | 9% | 5.4 |\n\nReported volume went **up 8%** while the actual problem fell **40%**. A team watching the ticket trend concludes the fix failed and considers rolling it back. A team watching a flat line concludes nothing happened. Both are looking at a large, real improvement.\n\nThis is the same structural trap as a remedy that suppresses the signal it would be judged by, and the only escape is to measure incidence somewhere other than the channel your announcement perturbed.\n\n## Reading a trend without fooling yourself\n\n### Four questions before you believe a complaint trend\n\n1. **Did the denominator move?** Normalise by monthly active users of that specific feature, not by total accounts. A flat per-user rate under a rising raw count is a growth story, not a quality story.\n2. **Did any reporting surface change in the window?** Keep a dated log of feedback entry points, status posts, changelog entries and support macros. This log is cheap to maintain and it is the single highest-value artifact for trend interpretation. Without it you are guessing.\n3. **Did attention change?** Check community threads, review sites and social mentions for the window. A spike that starts the day after a popular post is a notoriety spike.\n4. **Is the issue new?** Newly shipped surfaces are in the rising part of whatever propensity curve they have. Comparing a three-month-old feature to a three-year-old one on raw complaint volume compares their ages as much as their quality.\n\n### What a trend you can trust looks like\n\nA defensible series has three properties the raw inbox count lacks: a fixed denominator, a fixed instrument, and a fixed question. Ask the same structured question, of a comparable sample, on a fixed cadence, and the resulting series moves only when the thing you are measuring moves.\n\nThat is a tracking study, not an inbox, and a repeatable Koji study is the cheapest way to stand one up. The inbox remains the right tool for discovering that something is wrong - it is fast, free and unprompted. It is simply the wrong tool for measuring whether the problem is growing, because two of its three factors are outside your control and one of them is outside your knowledge.\n\nTwo neighbouring failure modes are worth distinguishing from this one. A false trend can also be manufactured by the spacing of your measurements, and a real event can be erased by smoothing. Both of those hold reporting propensity constant and let the measurement distort a true signal. This article is the mirror image: the measurement is fine and the willingness to report is what moved.\n\n## How Koji handles this\n\nKoji addresses this by fixing the two terms the inbox leaves floating.\n\n- **A known, chosen denominator.** You define who is invited, so exposure is a number you set rather than a number you infer. A rate computed on an invited sample is comparable across quarters; a ticket count is not.\n- **A fixed instrument across waves.** Reuse the same structured questions - open_ended, scale, single_choice, multiple_choice, ranking, yes_no - and the question stops being a variable. A scale question asked identically in March and June produces two comparable distributions.\n- **Propensity is flattened by invitation.** Everyone in the sample is asked, so you hear from the quiet majority rather than only from whoever was motivated enough to find your feedback form. That is precisely the term that wrecks inbox trends.\n- **AI follow-ups without interviewer drift.** Koji's AI interviewer probes inconsistent or vague answers the same way in every session, so the depth of the data does not depend on which researcher ran which wave.\n- **Voice or text, unmoderated.** A repeat wave costs no scheduling, which is what makes a genuine cadence realistic instead of aspirational.\n- **Reports build as responses arrive**, so a wave-over-wave comparison is available immediately rather than after a synthesis sprint.\n\nOne practical habit worth adopting regardless of tooling: when you announce a fix, run a short Koji study to the affected cohort in the same week. The study measures incidence on a stable instrument while the announcement is busy inflating your ticket count. You will be able to tell the two apart, which is exactly what the trend line cannot do.\n\n## Frequently asked questions\n\n### Does a rising complaint count ever mean the product got worse?\n\nOften, yes - but the count alone cannot establish it. Rising volume is evidence that something changed in the product of exposure, incidence and reporting propensity. Before attributing it to quality, normalise by usage of that specific feature and check whether any reporting surface or public discussion changed in the same window. If the per-user rate is up and nothing about the channel changed, you have a real signal.\n\n### What is the Weber effect, and should I correct for it?\n\nIt is the observation that adverse event reporting peaks around the end of the second year after a product launches, rather than tracking real risk. You should not correct for it. A 2014 analysis of sixty-two drugs found that most showed little evidence of the effect, which means there is no dependable curve to divide out. Treat time since launch as a known confounder to reason about, not as a coefficient to apply.\n\n### Why did complaints jump when we added a feedback button?\n\nBecause you lowered the cost of reporting. The failures were already happening; more of them now convert into reports. This is called stimulated reporting, and it is the cleanest example of reporting propensity moving on its own. Expect a step change in level at the moment of the change, and never compare volume across that boundary without noting it.\n\n### How do I tell a real regression from a notoriety spike?\n\nCheck the start date against external attention. A notoriety spike usually begins within a day or two of a specific triggering post, arrives with unusually similar wording across reports, and decays over a week or two without a code change. A real regression tends to start at a deploy boundary, correlates with an instrumented error rate, and does not decay on its own.\n\n### Should we stop tracking inbound feedback volume?\n\nNo. Track it as an operations metric, because it drives staffing and response time, and as a discovery signal, because it finds problems you did not know to look for. Just stop treating it as a measure of how common a problem is. Those are different jobs, and one dashboard cannot do both honestly.\n\n### What is the minimum setup for a trend I can defend?\n\nA fixed question set, a defined sampling frame, and a fixed cadence - plus a dated log of every change you make to reporting surfaces. Three of those four are one-time setup. The log is the one people skip, and it is the one that later lets you explain a step change instead of arguing about it.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the fixed instrument that makes a wave-over-wave comparison legitimate\n- [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) - the denominator problem underneath this one\n- [Measurement Interval and False Trends](/docs/measurement-interval-false-trend) - how the spacing of your measurements invents a trend of its own\n- [When a Moving Average Hides an Event](/docs/moving-average-hides-events-research) - the companion case where smoothing erases something real\n- [Common Cause vs Special Cause](/docs/common-cause-special-cause-research-metrics) - deciding whether a move on a chart deserves an investigation\n- [Research Refresh Cadence](/docs/research-refresh-cadence) - how often to re-run a tracking wave","category":"Analysis & Synthesis","lastModified":"2026-09-28T03:47:25.65633+00:00","metaTitle":"Feedback Volume Tracks Attention, Not Incidence","metaDescription":"A complaint trend moves when reporting propensity moves. How to tell a real regression from a notoriety spike or a new feedback button.","keywords":["feedback volume trend","reporting propensity","weber effect reporting","notoriety bias","complaint trend analysis","stimulated reporting"],"aiSummary":"Every point on a complaint trend is the product of three terms - exposure, incidence and reporting propensity - and only the product is observable. FDA states that \"Many factors can influence whether an event will be reported, such as the time a product has been marketed and publicity about an event\", naming the two best-documented propensity effects: the Weber effect and notoriety bias. Hartnell and Wilson replicated the Weber effect in 2004, but Hoffman and colleagues analysed sixty-two drugs in 2014 and found most showed little evidence for it. That failure to replicate is worse news than confirmation, because a stable propensity curve could have been divided out; an unstable one cannot. A worked case shows reported volume rising 8% while real incidence fell 40%, because a fix lowered incidence while its announcement raised propensity. The fix is a fixed instrument on a chosen sample, which is what a Koji study with reusable structured questions provides.","aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}