Back to docs
Analysis & Synthesis

Why Complaint Counts Cannot Become Rates (And What to Compute Instead)

A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.

The short answer

A count of complaints is not a rate and cannot be turned into one. A rate needs a denominator - the number of people who could have reported and did not - and an inbound feedback channel never tells you that number. What it hands you is a numerator of unknown completeness, drawn from a population of unknown size, at a reporting rate that varies by person, by severity and by month.

That is not a reason to ignore your inbox. It is a reason to run arithmetic on it that does not need the missing number. Drug safety regulators have depended on exactly this shape of data for decades, and the method they settled on transfers to product feedback almost unchanged.

The missing denominator, stated precisely

Why "12 complaints about export" is not a number yet

Suppose twelve people wrote in about CSV export this month. Is that a lot? The question has no answer, because "a lot" is a ratio and you have only been given the top half of it. Twelve out of how many people who used export? Out of how many who hit the bug? Out of how many who hit it and noticed?

Every inbound channel produces the same shape. A report exists because a chain of independent things all happened: someone hit the problem, noticed it, decided it was worth the effort, found the channel, and wrote something a human could classify. The count you receive is the product of all five, and you can measure none of them.

The three unknowns behind every complaint count

UnknownWhat it isWhy you cannot recover it
ExposureHow many people actually used the feature in the windowAnalytics can approximate this, but not for the specific path and state that triggered the problem
IncidenceHow many of those people hit the problemOnly instrumentation on the exact failure can tell you, and if you had that you would not need the inbox
Reporting propensityWhat fraction of the people who hit it wrote inUnmeasurable in principle, because the people who did not write in left no record

The third row is the one that ends the argument. Exposure and incidence are hard; reporting propensity is structurally invisible. The people you would need to count are defined by the fact that they generated no data.

What the regulator says about its own database

The clearest statement of this limitation comes from the organisation with the most to lose from it. FDA runs the adverse event reporting system long known as FAERS, which it is now consolidating into the Adverse Event Monitoring System, or AEMS. Its own public documentation is blunt about what the data cannot do:

"AEMS data cannot be used to calculate the incidence of an adverse event or medication error in the U.S. population."

The stated reasons are exactly the three unknowns above. On completeness: "FDA does not receive reports for every adverse event or medication error that occurs with a product." On the size of the slice: adverse event reports "represent a small percentage of total usage numbers of a product." And on what a report even asserts, FDA notes that it "does not require that a causal relationship between a product and event be proven".

That last point deserves its own treatment, and it gets one in the companion article on grading whether a feature actually caused a complaint. For now, note only that a regulator holding millions of reports declines to compute a rate from them. If they will not, a product team with four thousand support tickets should not either.

How incomplete is incomplete? A field that measured it

Product teams rarely know their own reporting rate. Pharmacovigilance researchers went and measured theirs. A systematic review by Hazell and Shakir, published in Drug Safety in 2006, pooled the studies that had tried to quantify it:

"In total, 37 studies using a wide variety of surveillance methods were identified from 12 countries."

"The median under-reporting rate across the 37 studies was 94% (interquartile range 82-98%)."

Read that as a survival rate: roughly six in every hundred reportable events reached the system. The interquartile range matters as much as the median, because it says the rate is not a constant you could look up and divide by.

One important qualifier, because it changes what you may conclude. These studies measured trained clinicians reporting suspected adverse drug reactions under a professional obligation to do so. They are not a measurement of software users filing product feedback, and nobody should quote 94% as the under-reporting rate for a support inbox. What transfers is not the number. What transfers is the demonstrated fact that a spontaneous reporting channel can run at single-digit completeness while still being the primary safety instrument for an entire industry - and that when people finally measured it, the answer was far worse than practitioners had assumed.

Why an uneven reporting rate is worse than a low one

Here is the part that most teams get backwards. A low reporting rate, by itself, is survivable. If exactly 6% of everyone who hit any problem wrote in, then complaint counts would be a constant multiple of real incidence, and the ranking of your issues would be perfectly correct even though every absolute number was wrong by 16x. You could act on the ordering with complete confidence.

The Hazell and Shakir data destroys that comfort, because the rate is not constant. The same review found:

"The median under-reporting rate was lower for 19 studies investigating specific serious/severe ADR-drug combinations but was still high at 85%."

And in five of the ten general practice studies reviewed, under-reporting was heavier for ADRs overall than for the more serious ones - 95% against 80%.

So the multiplier changes with severity. Work through what that does to a ranking. Two issues, A and B. A is mild and hits 10,000 people; B is severe and hits 1,000. At a uniform 6% reporting rate you would receive 600 reports about A and 60 about B, and A correctly outranks B on frequency. Now apply severity-dependent reporting - 5% for mild, 20% for severe - and you receive 500 about A and 200 about B. A still leads, but the gap has collapsed from 10x to 2.5x. Push the severity gradient a little further, as the general-practice numbers suggest it can go, and the order flips outright.

A biased instrument with a known bias is a measuring device. A biased instrument whose bias varies with the thing you are measuring is not. This is why the honest response to inbox data is to stop trying to estimate magnitudes from it, and to switch to a comparison that does not depend on the reporting rate at all.

The arithmetic that does not need a denominator

The proportional reporting ratio, in product terms

The method regulators converged on is disproportionality analysis, and its simplest form is the proportional reporting ratio. Evans, Waller and Davis set out the version still in use in Pharmacoepidemiology and Drug Safety in 2001:

"The proportion of all reactions to a drug which are for a particular medical condition of interest is compared to the same proportion for all drugs in the database, in a 2 x 2 table."

Translate the vocabulary and nothing else changes. For "drug" read feature, surface or release. For "medical condition" read failure mode. You are asking: among reports that mention this feature, what share mention this problem - and is that share higher than among reports that do not mention the feature?

Their signal threshold was deliberately conservative: "3 or more cases, PRR at least 2, chi-squared of at least 4."

A worked example you can run on your own inbox

Take a quarter with 4,000 inbound reports. 200 of them mention CSV export. Within those 200, 60 describe missing or dropped rows. Across the other 3,800 reports, 190 mention missing data.

Mentions data lossDoes notRow total
Mentions export60140200
Does not mention export1903,6103,800
Column total2503,7504,000

The proportion within export reports is 60/200 = 30%. Outside export it is 190/3,800 = 5%. The ratio is 30/5 = 6.0. The chi-squared statistic for this table is about 203. Against the Evans criteria - 60 cases, PRR 6.0, chi-squared 203 - this clears every threshold with room to spare.

The reporting-odds-ratio variant, which uses odds rather than proportions, gives (60/140)/(190/3,610) = 8.1 on the same table. Either statistic points the same way; the choice between them rarely changes a decision at this effect size.

Why the unknown cancels

Now look at what is absent from that calculation. Your total user count never appeared. Neither did the number of people who used export, nor the number who hit the bug silently. Every figure in the table is a count of reports, and both proportions are computed inside the same report set.

Write the reporting rate as r - the unknown fraction of people who hit a problem and wrote in. It sits in the numerator and the denominator of each proportion, and when you divide one proportion by the other it cancels. The quantity you could never measure drops out of the arithmetic entirely.

That is the whole trick, and it is worth stating plainly because it is genuinely non-obvious: you cannot estimate how often something happens, but you can estimate whether it is over-represented, and over-representation is enough to decide what to investigate next.

Three ways this will still mislead you

Competition bias: your own biggest issue deflates everything else

Every PRR is computed against the rest of your inbox, so the inbox is the baseline. Ship a password-reset regression that generates 3,000 tickets and every other issue's proportion falls, because the denominator of the comparison group swelled. Nothing got better; the yardstick moved. Recompute your baselines after any event that changes total volume, and never compare a PRR from one quarter to a PRR from another without checking that the mix behind them is comparable.

Differential reporting: when the cancellation fails

The cancellation of r holds only if the reporting rate is the same for the feature group and the comparison group. Often it is not. If export is used mainly by enterprise accounts who have a named CSM filing tickets on their behalf, their r is several times higher than a self-serve user's, and the PRR inherits that gap as a real-looking signal. Before believing a ratio, ask whether the two groups had equal opportunity and equal motivation to report. This is the same failure that makes complaint trends unreadable over time, which the companion article on reporting propensity takes up in detail.

It is not a risk estimate

A PRR of 6.0 does not mean export users are six times more likely to lose data. It means data-loss language is six times over-represented among export reports. The statistic is about the composition of a report database, not about the world. Teams that forget this end up writing "export is 6x riskier" in a roadmap document, which is a claim the arithmetic cannot support and which the first analyst to check will dismantle.

There is a fourth trap worth naming: a very loud issue can suppress the reporting of a quieter one in the same area, so an absent signal is not evidence of absence. That mechanism is covered in the article on masking in customer interviews.

How Koji handles this

Every limitation above comes from one root cause: nobody chose who would be asked. The fix is not better arithmetic on the inbox - it is running a channel where the denominator is yours by construction. That is what a Koji study is.

  • You know the denominator. You invited 120 people and 48 completed. "38 of 48 hit this" is a rate, with a stated base, and it is defensible in a roadmap review in a way that "60 tickets" never is.
  • Structured questions turn language into counts. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so "did you lose rows on export" becomes a yes_no you can total, severity becomes a scale, and frequency becomes a single_choice. No manual tagging pass, and no arguing about whether two tickets describe the same thing.
  • AI follow-ups close the gap the inbox leaves. When someone says export is broken, Koji's AI interviewer probes for the file size, the destination and the exact step, automatically, on every single response. A ticket gives you one sentence; the study gives you the whole chain.
  • Everyone gets asked, not just the motivated. Reporting propensity is what wrecks inbox arithmetic, and an invited study flattens it: the quiet majority and the vocal minority answer the same questions.
  • Voice or text, with no moderator. Studies run unmoderated in either modality, so getting a real denominator costs scheduling time rather than weeks of it.
  • Analysis and the report are automatic. Results aggregate into a live report as responses land, with distributions for the scale and choice questions and themes plus quotes for the open-ended ones.

The practical sequence is the one regulators use: treat the inbox as a signal generator and a study as the estimator. Disproportionality tells you where to look. Then you go and measure it properly, which takes an afternoon rather than a quarter.

A closing note on tooling: if your tickets live in Zendesk, Koji can trigger an interview straight off a ticket, which turns a single unquantified complaint into a measured one without anyone copying text between tabs.

Frequently asked questions

Can I ever turn support ticket counts into a rate?

Not from the tickets alone. You would need the number of people who hit the problem and did not write in, and that population leaves no trace by definition. What you can do is use the tickets to identify a candidate issue, then measure its real prevalence in a channel with a known base - an invited study, or instrumentation on the specific failure. The ticket count tells you where to point the instrument, not what the instrument would read.

What is a proportional reporting ratio in plain terms?

It is the share of reports about one feature that mention a particular problem, divided by the share of all other reports that mention the same problem. If 30% of export reports mention data loss and 5% of everything else does, the ratio is 6.0. Because both halves are shares of report counts, the unknown reporting rate cancels out, which is why the statistic works on data that cannot support a rate.

Does a high ratio mean the feature caused the problem?

No. It means the problem is over-represented in reports mentioning that feature. Co-occurrence in a report database is consistent with a shared cause, a reporting artifact, or coincidence. Grading whether a specific feature caused a specific complaint is a separate exercise with its own four tests, including removing the feature and putting it back.

How many reports do I need before the comparison means anything?

The conventional floor is three cases, a ratio of at least 2, and a chi-squared of at least 4 - all three together, not any one of them. Below that you are reading noise. In practice, three cases is a very low bar for a consumer product, and most teams should treat it as a threshold for opening an investigation rather than for making a decision.

Why does uneven under-reporting matter more than heavy under-reporting?

Because a constant multiplier preserves your ranking and a varying one destroys it. If everyone under-reported at the same rate, complaint counts would be wrong in magnitude but right in order, and ordering is what you act on. When the rate changes with severity, familiarity or customer segment, the ranking itself becomes an artifact of who bothers to write in.

What is the fastest way to get a real rate instead of a count?

Pick the single issue your disproportionality analysis flagged hardest, write four or five structured questions about it, and send a Koji study to a defined slice of the affected population. You will have a rate with a stated denominator, usually within a day. That number belongs in the roadmap document; the ticket count belongs in the paragraph explaining why you went looking.

Related Resources

Related Articles

The Base Rate Nobody Measured: Why Every Flag in Your Research Stack Has an Unknown Precision (2026)

Every AI tag, sentiment label and risk score is a diagnostic test whose precision depends on a prevalence nobody measured. Accuracy rises as precision collapses. How to audit the unflagged pile and publish a precision footer.

Number Needed to Treat: How Many Users You Must Reach to Keep One (2026)

Every effect in your deck is a rate. None of them is a count of people. Number needed to treat converts a percentage lift into the only figure a roadmap can cost, and the evidence says the persuasive format is the misleading one.

Product Feedback Triage: A Framework for Turning Noise Into a Prioritized Backlog

A practical framework for triaging product feedback at scale — capture, dedupe, tag, route, and validate every request before it ever reaches prioritization. Includes a triage workflow, a severity matrix, and an AI-native approach.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Support Ticket Analysis: How to Mine Customer Service Data for Product Insights

A practical guide to systematically extracting product insights from customer support tickets — covering manual coding workflows, AI-powered thematic analysis, and how to tie ticket themes to business impact.

The Taphonomy of Customer Feedback: Which Complaints Survive to Reach You (2026)

Most customer feedback is destroyed before it reaches you, and the filter has a predictable shape. A taphonomic method for naming the evidence classes your channels systematically lose.