Back to docs
Analysis & Synthesis

Did the Feature Cause the Complaint? A Causality Grading Method for Product Feedback

A complaint that names a feature is a hypothesis about that feature, not evidence about it. Four tests, two scales, and the rechallenge your feature flags already give you for free.

The short answer

A customer who writes "since the new editor shipped, my drafts keep vanishing" has told you two things: drafts vanished, and they believe the editor did it. The first is data. The second is an attribution the customer is not positioned to make, and treating it as data is how teams end up rolling back the wrong release.

Drug safety has spent forty years building the discipline for exactly this: judging whether a single reported event was caused by a specific product, when you cannot run a trial. The method is four questions and a graded verdict, and the strongest of the four tests is one most product teams can run in about ten minutes with a feature flag. A platform like Koji collects the facts those questions need; your flag system supplies the decisive test itself.

What a report actually asserts

What a report actually asserts

The organisations that collect the most reports are the most explicit that a report is not a finding. FDA states plainly that "FDA does not require that a causal relationship between a product and event be proven" before a report enters its system. The openFDA documentation for the same data goes further:

"a causal relationship cannot be established between product and reactions listed in a report."

"The information in these reports has not been scientifically or otherwise verified as to a cause and effect relationship"

"Adverse event reports submitted to FDA do not undergo extensive validation or verification."

A report, in other words, records a temporal coincidence plus a human guess. That guess carries real information - the customer was there and you were not - but it is a starting hypothesis, and the whole point of causality assessment is to grade it rather than to accept or dismiss it.

The four questions that grade a single report

The WHO-UMC system, developed with input from national centres in the WHO Programme for International Drug Monitoring and named for the Uppsala Monitoring Centre, assesses a single case on four criteria: the time relationship between the exposure and the event, whether competing causes can be excluded, what happened when the product was withdrawn, and what happened when it was reintroduced. The last two have names worth learning, because they are the ones with teeth: dechallenge and rechallenge.

1. Temporal relationship

Did the event follow the exposure, with a plausible gap? This sounds trivial and is the test that kills the most hypotheses. Pull the actual timestamps rather than trusting the narrative: the deploy time for the affected cohort, and the first occurrence in logs or in the customer's own account history.

The common failure is a complaint that arrives after a release but describes behaviour that predates it. People notice a problem when something draws their attention to the area, and a release draws attention. If your logs show the same failure at the same rate three weeks before the deploy, the release is exonerated and you have learned something more interesting: the release changed reporting, not behaviour.

2. Competing explanations

What else changed? Candidates in almost every product incident: a concurrent deploy to another service, a dependency or browser update, a data migration, a permissions change, an expiring credential, a seasonal usage shift, or the customer's own environment. In safety assessment this is called excluding concomitant causes, and it is the criterion that separates "possible" from "probable".

Write the competing explanations down explicitly. An undocumented alternative tends to be quietly dropped rather than actually excluded.

3. Dechallenge: what happened when it was removed

Remove the suspected cause. Does the problem stop?

In pharmacovigilance this means stopping the drug. In product terms it means turning the feature off for the affected account, reverting them to the previous version, or moving them out of the experiment arm. A clean dechallenge - problem present, feature removed, problem gone - is substantially stronger evidence than any number of additional reports.

The weakness of a dechallenge alone is that it is confounded by time. Lots of things resolve on their own. If the problem was going to clear up anyway, removing the feature gets undeserved credit.

4. Rechallenge: what happened when it came back

Put it back. Does the problem return?

This is the strongest single-case test that exists, because it is the one that time cannot fake. A problem that disappears on removal and reappears on reintroduction, in the same account, is very difficult to explain any way other than causally. It converts an anecdote into a controlled experiment with a sample size of one - and an n of 1 with a working control beats an n of 50 without one.

In medicine, a deliberate rechallenge is often unethical, which is why the WHO-UMC framework treats it as desirable but not required for a "probable" verdict. In software it is usually trivial and harmless. That asymmetry is the single biggest opportunity in this article and it is examined below.

A graded verdict beats a binary one

The WHO-UMC system does not return yes or no. It returns one of six levels: certain, probable or likely, possible, unlikely, conditional or unclassified, and unassessable or unclassifiable. The last two are not evasions - they are the honest output when the report lacks the information needed to judge it, which is extremely common.

Mapped to product work:

GradeWhat it means for a complaintWhat to do
CertainTiming fits, alternatives excluded, dechallenge clean, rechallenge positiveFix and ship; the investigation is over
ProbableTiming fits, alternatives unlikely, dechallenge clean, no rechallengeFix; run the rechallenge if it is cheap
PossibleTiming fits, alternatives not excludedInvestigate; do not roll anything back yet
UnlikelyTiming implausible or a better explanation existsClose with a note; keep the report for pattern detection
ConditionalNeeds more data before it can be gradedGo get the specific missing fact
UnassessableReport too vague to grade at allAsk the reporter, or accept it is unusable

Why collapsing six grades to two loses what you needed

For regulatory purposes these six are often collapsed into "related" and "not related", with certain, probable and possible mapping to related. Product teams do the same thing by instinct, sorting complaints into real and not-real.

The collapse throws away the two most actionable categories. "Conditional" says there is a specific fact that would settle this - a log line, a browser version, one follow-up question - and naming that fact is a concrete next action. "Unassessable" says the intake form is failing, which is a fixable process problem rather than a fact about the product. Both disappear the moment you force a binary, and what replaces them is an argument about whether a ticket counts.

The scored alternative: Naranjo

Where WHO-UMC relies on expert judgement, the Naranjo algorithm turns the same reasoning into a questionnaire. It is described as "a questionnaire designed by Naranjo et al. for determining the likelihood of whether an adverse drug reaction (ADR) is actually due to the drug rather than the result of other factors."

The ten-item version, and its score bands

Ten items are scored and summed. A total of 9 or more is graded definite; 5 to 8 probable; 1 to 4 possible; and 0 or below doubtful. Two of the ten items are precisely the dechallenge and rechallenge tests. Item three asks: "Did the adverse reaction improve when the drug was discontinued or a specific antagonist was given?" Item four asks whether the reaction reappeared when the product was readministered.

The value of a scored instrument is not that the number is meaningful in itself. It is that two people scoring the same report independently will converge, and that their disagreement localises to a specific item rather than to a general impression. Adapting the ten items to your product once, as a checklist in your issue template, is a half-day of work that pays back permanently.

Your feature flag is a rechallenge apparatus

Your feature flag is a rechallenge apparatus

Here is the observation that should change how your team triages.

The gold standard for establishing single-case causality in medicine is a positive rechallenge, and in medicine it is frequently unavailable because deliberately re-exposing a patient to something that harmed them is unacceptable. Entire methodological literatures exist to work around its absence.

In software, a rechallenge costs one toggle. If you run feature flags, per-account overrides, or percentage rollouts, you already own the apparatus for the strongest causal test in the field. You can turn the feature off for one complaining account, confirm the problem stops, turn it back on, and confirm it returns - with the customer's permission, inside a support conversation, in an afternoon.

Almost no product team does this. The default response to an ambiguous complaint is to wait for more complaints, which is the one move that cannot resolve the ambiguity, because every new report shares the same confounding structure as the first. Fifty uncontrolled anecdotes do not add up to one controlled test.

Why one positive rechallenge beats fifty more complaints

Consider two responses to a single report that the new editor eats drafts.

Wait for volume. Four weeks later you have 50 reports. You still do not know whether the editor is at fault, because all 50 arrived through the same channel, from the same self-selected population, after the same release that drew attention to the area. Volume has raised your confidence that something is wrong while telling you nothing new about what. Worse, the count itself is unreadable as a rate, and it moves with reporting propensity rather than with incidence - two separate problems covered in the companion articles.

Run the challenge. Same afternoon: pull the timestamps (test 1), list what else shipped (test 2), disable the editor for that one account and confirm drafts persist (test 3), re-enable and confirm they vanish again (test 4). You now have a "certain" grade on a single case, and you can fix it before report 2 arrives. Where the four tests need facts the ticket does not contain, a short Koji study sent to the affected cohort collects them in a day.

The second path is faster, cheaper, and produces a stronger conclusion. The only reason the first is more common is that waiting feels like gathering evidence.

Where this sits relative to root cause analysis

Where this sits relative to root cause analysis

This method is upstream of techniques like the 5 Whys. Root cause analysis takes a causal link as given and asks why it exists, drilling from symptom toward mechanism. Causality assessment asks whether the link is there at all.

Running root cause analysis on an ungraded complaint is how teams produce confident, well-documented explanations of things that were never happening. Grade first, then drill. The grade also tells you how much drilling is justified: a "certain" earns a full investigation, a "possible" earns one clarifying question. A single Koji question put to the affected cohort is often enough to settle whether the link is there at all.

How Koji handles this

Causality grading needs specific facts from the person who experienced the problem, and a support thread rarely captures them. This is where an interview beats a ticket.

  • The four tests become structured questions. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so "when did you first notice this" becomes a single_choice on time windows, "did it stop when we switched you back" becomes a yes_no, and severity becomes a scale. The grade falls out of the answers instead of out of a debate.
  • AI follow-ups chase the missing fact. Most reports arrive "unassessable" because one detail is absent. Koji's AI interviewer probes for the file, the step, the browser and the timing automatically, on every response, which is what turns a conditional grade into a gradable one.
  • Ask the whole affected cohort, not the loudest member. Rather than waiting for report 50, send a short study to everyone in the affected segment. You get the temporal pattern across the cohort in a day.
  • Voice or text, unmoderated. No scheduling, so a causality check does not wait for a researcher's calendar.
  • Quality scoring flags thin answers so you know which responses are solid enough to grade and which are not.
  • The report is live, so the timing distribution is visible while you still have the flag in your hand.

Used together with a flag, the sequence is: interview the affected cohort to establish timing and exclude alternatives, then dechallenge and rechallenge one consenting account to settle it. Koji handles the first half; your flag system handles the second.

Frequently asked questions

Is it really safe to reproduce a bug on a customer account on purpose?

Only with explicit, informed consent, on a non-destructive problem, and never where data loss or a financial action is the symptom. For a rendering glitch, a slow load or a confusing flow, most customers are pleased to be asked, because it visibly means someone is working on their issue. Where the symptom is destructive, stop at the dechallenge and accept a probable grade - which is exactly what the WHO-UMC framework does for the same reason.

What if we cannot reproduce the problem at all?

Then the honest grade is conditional or unassessable, and the useful output is the name of the specific missing fact. Write that fact down and go get it - the exact timestamp, the browser build, the file, the account state. A report you cannot grade is not a report you should either fix or dismiss; it is a request for one more piece of information.

How is this different from the 5 Whys?

The 5 Whys assumes the causal link and drills from symptom toward mechanism. Causality assessment decides whether the link exists. They run in that order, and reversing them produces thorough explanations of non-existent problems. Grade first, then drill, and let the grade set how much drilling is justified.

Do we need to grade every incoming complaint?

No, and trying to would be a waste. Grade the complaints that are about to change a decision - anything proposed for a rollback, a hotfix, or a roadmap slot. For the rest, the count is a discovery signal and can stay ungraded. The discipline is about what gets promoted into a decision, not about processing volume.

Who should do the grading?

Whoever can see both the logs and the customer's words, which in most teams means a support engineer and a product manager for ten minutes together. The reason to use a written scale is that it makes the judgement reviewable: someone who disagrees has to say which item they score differently, which is a far more productive argument than one about whether a ticket is real.

Can a single report ever justify shipping a fix?

Yes. A single report with a clean dechallenge and a positive rechallenge is graded certain, and certain is the strongest evidence this framework produces regardless of how many reports back it. One controlled case outranks fifty uncontrolled ones. Volume is a signal about attention; a rechallenge is a signal about causation.

Related Resources

Related Articles

Why Complaint Counts Cannot Become Rates (And What to Compute Instead)

A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.

Feedback Volume Tracks Attention, Not Incidence: How to Read a Complaint Trend

A rise or fall in complaint volume is at least as likely to be a change in how willing people are to report as a change in your product. Here is how to tell them apart.

The Five Whys Technique: How to Find Root Causes in User Research (with AI)

The Five Whys is a root-cause analysis technique that turns surface-level user feedback into actionable insight. Learn how to apply it in interviews and run it with AI-powered probing at scale.

Root Cause Analysis for Customer Research: The Complete Guide

A practical guide to root cause analysis (RCA) for product and customer research — the 5 Whys, fishbone diagrams, and Pareto analysis — and how to find the real driver behind churn, complaints, and product issues.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)

How to rate the severity of usability problems using Nielsen's 0-4 scale, why single-evaluator ratings are unreliable, how to separate severity from priority, and how to replace guessed frequency estimates with measured data.