Back to docs
Analysis & Synthesis

Feedback Volume Tracks Attention, Not Incidence: How to Read a Complaint Trend

A rise or fall in complaint volume is at least as likely to be a change in how willing people are to report as a change in your product. Here is how to tell them apart.

The short answer

When complaint volume for a feature doubles, the most common explanation is not that the feature got twice as bad. It is that something changed how likely people were to tell you. Attention, publicity, a new in-app prompt, a status page, a support hiring wave and simple age of the feature all move reported volume without touching the underlying failure rate.

This matters because the trend line is the artifact most teams trust most. A single ticket gets discounted as an anecdote; a rising line gets treated as evidence. It is often weaker evidence than the single ticket, because a line invites you to read a slope that three different mechanisms could have produced. The way out is to measure incidence on a channel you control, which is what a platform like Koji is for.

The quantity you are actually plotting

The identity you are actually plotting

Every point on a complaint trend is a product of three terms:

reports = exposure x incidence x reporting propensity

Exposure is how many people used the thing. Incidence is the fraction of them who hit the problem. Reporting propensity is the fraction of those who told you. Only the product is observable. Any one of the three can move while the other two sit still, and all three commonly move at once.

The term you care about is incidence. It is the only one that describes your software. The other two describe your growth and your users, and both are moving constantly - exposure because you are shipping and marketing, propensity because you keep changing how easy and how socially expected it is to complain.

Why any one factor moving looks like the product changing

A 40% rise in export complaints is equally consistent with all of these:

What actually happenedExposureIncidencePropensity
The export bug got worseflatup 40%flat
Marketing drove a campaign at power usersup 40%flatflat
You added a "Report a problem" button next to exportflatflatup 40%
A competitor blog post named your export bugflatflatup 40%
You fixed it and announced the fix loudlyflatdownup more

The last row is the one that should worry you, and it gets its own section below. The point of the table is that the trend line is identical in all five cases. Nothing in the shape of the curve distinguishes them. You have to bring outside information.

The regulator names both mechanisms in its own disclaimer

FDA, which runs the adverse event reporting system formerly called FAERS and now being consolidated as the Adverse Event Monitoring System (AEMS), states the problem in a single sentence on its public dashboard documentation:

"Many factors can influence whether an event will be reported, such as the time a product has been marketed and publicity about an event."

Those two examples are not casual. They are the two best-documented reporting-propensity effects in safety surveillance, and each has a name. Time on market is the Weber effect. Publicity is notoriety bias. Both have direct product analogues, and neither has anything to do with whether the product changed.

The Weber effect, and the honest state of the evidence

What Weber claimed

The original observation was that reporting of adverse events follows a predictable arc after launch rather than tracking real risk. A replication study by Hartnell and Wilson in Pharmacotherapy in 2004 applied it to the drugs Weber had originally studied, setting an explicit criterion: a reporting pattern counted as demonstrating the effect if "the highest peak in reports during the first 5 years after product approval occurred during year 2". Against that criterion their result was unambiguous, though the sample was small: "All five drugs analyzed in this study demonstrated the Weber effect."

The mechanism is intuitive. At launch, few people are exposed and nobody is looking. Through the first two years, exposure grows, prescribers get curious, and reporting climbs. After that, the product becomes unremarkable, reporting fatigue sets in, and volume declines even as exposure keeps rising.

What happened when someone checked sixty-two drugs

Then a larger study looked again. Hoffman and colleagues, writing in Drug Safety in 2014, analysed "Sixty-two drugs approved by the FDA between 2006 and 2010" against the standard definition, which they stated as the claim that "AE reporting peaks at the end of the second year after a regulatory authority approves a drug."

Their finding: "a majority of the drugs showed little evidence for the effect."

Why the failure to replicate is the worse news

The instinct is to file this as good news. It is the opposite, and this is the most useful thing in the article.

If the Weber effect were a stable, universal curve, it would be a correction. You could divide it out. Reporting propensity would be a known function of time since launch, you would deflate each point by it, and what remained would be an estimate of incidence. A reliable bias is not really a bias; it is a calibration.

What the 2014 result says is that the curve is not stable across products. Some show it, most do not, and you cannot know in advance which kind yours is. That makes reporting propensity an uncorrectable nuisance rather than a correctable one. The contested literature is a worse situation than a confirmed effect would have been, because it removes the possibility of a standard adjustment.

The practical consequence: there is no detrending step you can add to your dashboard that converts reported volume into incidence. Anyone who offers you one is assuming a curve that the best available evidence says does not generalise.

Notoriety: the spike you caused

Stimulated reporting: the spike you caused

The second mechanism FDA names is publicity, and in product work you are usually the publisher. Every one of these raises reporting propensity without changing a line of code:

  • Adding or moving a feedback entry point, especially into the failure path itself
  • Posting an incident to a status page
  • Shipping a changelog entry that names the area
  • A support agent asking "has this happened before?" in a macro
  • A community thread, a competitor comparison post, or a viral complaint
  • Running a survey about the feature, which teaches people the topic is fair game

The in-product case is the sharpest. Put a "Report a problem" link on the export screen and export complaints will rise, immediately and substantially, with zero change in the export failure rate. The channel got cheaper, so more of the existing failures converted into reports. You improved your instrument and it looked like a regression. An invited Koji study avoids this failure mode entirely, because the invitation rather than the user's motivation decides who answers.

The announcement that cancels its own fix

Now combine the two. You find the export bug, fix it, and publish a note telling affected customers it is fixed.

The fix lowers incidence. The announcement raises propensity. They push the observable in opposite directions, and the arithmetic is unforgiving. Suppose incidence falls 40%, from 100 affected users per 10,000 to 60, while the announcement lifts reporting propensity from 5% to 9%:

Incidence per 10,000PropensityReports per 10,000
Before the fix1005%5.0
After the fix and the note609%5.4

Reported volume went up 8% while the actual problem fell 40%. A team watching the ticket trend concludes the fix failed and considers rolling it back. A team watching a flat line concludes nothing happened. Both are looking at a large, real improvement.

This is the same structural trap as a remedy that suppresses the signal it would be judged by, and the only escape is to measure incidence somewhere other than the channel your announcement perturbed.

Reading a trend without fooling yourself

Four questions before you believe a complaint trend

  1. Did the denominator move? Normalise by monthly active users of that specific feature, not by total accounts. A flat per-user rate under a rising raw count is a growth story, not a quality story.
  2. Did any reporting surface change in the window? Keep a dated log of feedback entry points, status posts, changelog entries and support macros. This log is cheap to maintain and it is the single highest-value artifact for trend interpretation. Without it you are guessing.
  3. Did attention change? Check community threads, review sites and social mentions for the window. A spike that starts the day after a popular post is a notoriety spike.
  4. Is the issue new? Newly shipped surfaces are in the rising part of whatever propensity curve they have. Comparing a three-month-old feature to a three-year-old one on raw complaint volume compares their ages as much as their quality.

What a trend you can trust looks like

A defensible series has three properties the raw inbox count lacks: a fixed denominator, a fixed instrument, and a fixed question. Ask the same structured question, of a comparable sample, on a fixed cadence, and the resulting series moves only when the thing you are measuring moves.

That is a tracking study, not an inbox, and a repeatable Koji study is the cheapest way to stand one up. The inbox remains the right tool for discovering that something is wrong - it is fast, free and unprompted. It is simply the wrong tool for measuring whether the problem is growing, because two of its three factors are outside your control and one of them is outside your knowledge.

Two neighbouring failure modes are worth distinguishing from this one. A false trend can also be manufactured by the spacing of your measurements, and a real event can be erased by smoothing. Both of those hold reporting propensity constant and let the measurement distort a true signal. This article is the mirror image: the measurement is fine and the willingness to report is what moved.

How Koji handles this

Koji addresses this by fixing the two terms the inbox leaves floating.

  • A known, chosen denominator. You define who is invited, so exposure is a number you set rather than a number you infer. A rate computed on an invited sample is comparable across quarters; a ticket count is not.
  • A fixed instrument across waves. Reuse the same structured questions - open_ended, scale, single_choice, multiple_choice, ranking, yes_no - and the question stops being a variable. A scale question asked identically in March and June produces two comparable distributions.
  • Propensity is flattened by invitation. Everyone in the sample is asked, so you hear from the quiet majority rather than only from whoever was motivated enough to find your feedback form. That is precisely the term that wrecks inbox trends.
  • AI follow-ups without interviewer drift. Koji's AI interviewer probes inconsistent or vague answers the same way in every session, so the depth of the data does not depend on which researcher ran which wave.
  • Voice or text, unmoderated. A repeat wave costs no scheduling, which is what makes a genuine cadence realistic instead of aspirational.
  • Reports build as responses arrive, so a wave-over-wave comparison is available immediately rather than after a synthesis sprint.

One practical habit worth adopting regardless of tooling: when you announce a fix, run a short Koji study to the affected cohort in the same week. The study measures incidence on a stable instrument while the announcement is busy inflating your ticket count. You will be able to tell the two apart, which is exactly what the trend line cannot do.

Frequently asked questions

Does a rising complaint count ever mean the product got worse?

Often, yes - but the count alone cannot establish it. Rising volume is evidence that something changed in the product of exposure, incidence and reporting propensity. Before attributing it to quality, normalise by usage of that specific feature and check whether any reporting surface or public discussion changed in the same window. If the per-user rate is up and nothing about the channel changed, you have a real signal.

What is the Weber effect, and should I correct for it?

It is the observation that adverse event reporting peaks around the end of the second year after a product launches, rather than tracking real risk. You should not correct for it. A 2014 analysis of sixty-two drugs found that most showed little evidence of the effect, which means there is no dependable curve to divide out. Treat time since launch as a known confounder to reason about, not as a coefficient to apply.

Why did complaints jump when we added a feedback button?

Because you lowered the cost of reporting. The failures were already happening; more of them now convert into reports. This is called stimulated reporting, and it is the cleanest example of reporting propensity moving on its own. Expect a step change in level at the moment of the change, and never compare volume across that boundary without noting it.

How do I tell a real regression from a notoriety spike?

Check the start date against external attention. A notoriety spike usually begins within a day or two of a specific triggering post, arrives with unusually similar wording across reports, and decays over a week or two without a code change. A real regression tends to start at a deploy boundary, correlates with an instrumented error rate, and does not decay on its own.

Should we stop tracking inbound feedback volume?

No. Track it as an operations metric, because it drives staffing and response time, and as a discovery signal, because it finds problems you did not know to look for. Just stop treating it as a measure of how common a problem is. Those are different jobs, and one dashboard cannot do both honestly.

What is the minimum setup for a trend I can defend?

A fixed question set, a defined sampling frame, and a fixed cadence - plus a dated log of every change you make to reporting surfaces. Three of those four are one-time setup. The log is the one people skip, and it is the one that later lets you explain a step change instead of arguing about it.

Related Resources

Related Articles

Common Cause vs Special Cause: When a Move in Your Research Metric Is Real

Most movement in a research metric is noise, and reacting to it makes the metric worse. How to build a process behaviour chart for NPS, satisfaction or completion rate, and the decision rule that tells you when to investigate.

Why Complaint Counts Cannot Become Rates (And What to Compute Instead)

A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.

Why Your Quarterly Metric Shows a Trend That Is Not There (2026)

Undersampling does not blur a cycle, it counterfeits a different one. How the gap between your measurement waves manufactures smooth trends, flat lines, and reversed directions - and the three-question test that catches it.

The Moving Average on Your Dashboard Is Hiding the Week That Mattered (2026)

Smoothing conserves the area under an event and destroys its height - and every alert threshold you own is a height. The arithmetic of what a rolling average deletes, how late it reports, and why it cannot give you a number for now.

How Long Is User Research Valid? Insight Decay and When to Re-Run a Study

Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.