{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T17:52:31.392Z"},"content":[{"type":"documentation","id":"fad2b2c1-5799-4997-91f7-c1c4e753b23b","slug":"complaint-causality-dechallenge-rechallenge","title":"Did the Feature Cause the Complaint? A Causality Grading Method for Product Feedback","url":"https://www.koji.so/docs/complaint-causality-dechallenge-rechallenge","summary":"A complaint that names a feature asserts a temporal coincidence plus a human guess, and openFDA states that for its own data \"a causal relationship cannot be established between product and reactions listed in a report\". The WHO-UMC system grades a single case on four criteria - time relationship, competing causes, dechallenge (removal) and rechallenge (reintroduction) - returning one of six grades rather than a binary. Collapsing those six to related/not-related discards the two most actionable ones, conditional and unassessable, each of which names a concrete next action. The Naranjo algorithm scores the same reasoning across ten items with bands at 9 or more for definite, 5 to 8 probable, 1 to 4 possible. The key product insight: a positive rechallenge is the strongest single-case causal test and is usually unethical in medicine, yet costs one feature-flag toggle in software, so one controlled case outranks fifty uncontrolled complaints. Koji turns the four tests into structured questions across an affected cohort.","content":"## The short answer\n\nA customer who writes \"since the new editor shipped, my drafts keep vanishing\" has told you two things: drafts vanished, and they believe the editor did it. The first is data. The second is an attribution the customer is not positioned to make, and treating it as data is how teams end up rolling back the wrong release.\n\nDrug safety has spent forty years building the discipline for exactly this: judging whether a single reported event was caused by a specific product, when you cannot run a trial. The method is four questions and a graded verdict, and the strongest of the four tests is one most product teams can run in about ten minutes with a feature flag. A platform like Koji collects the facts those questions need; your flag system supplies the decisive test itself.\n\n## What a report actually asserts\n\n### What a report actually asserts\n\nThe organisations that collect the most reports are the most explicit that a report is not a finding. FDA states plainly that \"FDA does not require that a causal relationship between a product and event be proven\" before a report enters its system. The openFDA documentation for the same data goes further:\n\n> \"a causal relationship cannot be established between product and reactions listed in a report.\"\n\n> \"The information in these reports has not been scientifically or otherwise verified as to a cause and effect relationship\"\n\n> \"Adverse event reports submitted to FDA do not undergo extensive validation or verification.\"\n\nA report, in other words, records a temporal coincidence plus a human guess. That guess carries real information - the customer was there and you were not - but it is a starting hypothesis, and the whole point of causality assessment is to grade it rather than to accept or dismiss it.\n\n## The four questions that grade a single report\n\nThe WHO-UMC system, developed with input from national centres in the WHO Programme for International Drug Monitoring and named for the Uppsala Monitoring Centre, assesses a single case on four criteria: the time relationship between the exposure and the event, whether competing causes can be excluded, what happened when the product was withdrawn, and what happened when it was reintroduced. The last two have names worth learning, because they are the ones with teeth: dechallenge and rechallenge.\n\n### 1. Temporal relationship\n\nDid the event follow the exposure, with a plausible gap? This sounds trivial and is the test that kills the most hypotheses. Pull the actual timestamps rather than trusting the narrative: the deploy time for the affected cohort, and the first occurrence in logs or in the customer's own account history.\n\nThe common failure is a complaint that arrives after a release but describes behaviour that predates it. People notice a problem when something draws their attention to the area, and a release draws attention. If your logs show the same failure at the same rate three weeks before the deploy, the release is exonerated and you have learned something more interesting: the release changed reporting, not behaviour.\n\n### 2. Competing explanations\n\nWhat else changed? Candidates in almost every product incident: a concurrent deploy to another service, a dependency or browser update, a data migration, a permissions change, an expiring credential, a seasonal usage shift, or the customer's own environment. In safety assessment this is called excluding concomitant causes, and it is the criterion that separates \"possible\" from \"probable\".\n\nWrite the competing explanations down explicitly. An undocumented alternative tends to be quietly dropped rather than actually excluded.\n\n### 3. Dechallenge: what happened when it was removed\n\nRemove the suspected cause. Does the problem stop?\n\nIn pharmacovigilance this means stopping the drug. In product terms it means turning the feature off for the affected account, reverting them to the previous version, or moving them out of the experiment arm. A clean dechallenge - problem present, feature removed, problem gone - is substantially stronger evidence than any number of additional reports.\n\nThe weakness of a dechallenge alone is that it is confounded by time. Lots of things resolve on their own. If the problem was going to clear up anyway, removing the feature gets undeserved credit.\n\n### 4. Rechallenge: what happened when it came back\n\nPut it back. Does the problem return?\n\nThis is the strongest single-case test that exists, because it is the one that time cannot fake. A problem that disappears on removal and reappears on reintroduction, in the same account, is very difficult to explain any way other than causally. It converts an anecdote into a controlled experiment with a sample size of one - and an n of 1 with a working control beats an n of 50 without one.\n\nIn medicine, a deliberate rechallenge is often unethical, which is why the WHO-UMC framework treats it as desirable but not required for a \"probable\" verdict. In software it is usually trivial and harmless. That asymmetry is the single biggest opportunity in this article and it is examined below.\n\n## A graded verdict beats a binary one\n\nThe WHO-UMC system does not return yes or no. It returns one of six levels: certain, probable or likely, possible, unlikely, conditional or unclassified, and unassessable or unclassifiable. The last two are not evasions - they are the honest output when the report lacks the information needed to judge it, which is extremely common.\n\nMapped to product work:\n\n| Grade | What it means for a complaint | What to do |\n| --- | --- | --- |\n| Certain | Timing fits, alternatives excluded, dechallenge clean, rechallenge positive | Fix and ship; the investigation is over |\n| Probable | Timing fits, alternatives unlikely, dechallenge clean, no rechallenge | Fix; run the rechallenge if it is cheap |\n| Possible | Timing fits, alternatives not excluded | Investigate; do not roll anything back yet |\n| Unlikely | Timing implausible or a better explanation exists | Close with a note; keep the report for pattern detection |\n| Conditional | Needs more data before it can be graded | Go get the specific missing fact |\n| Unassessable | Report too vague to grade at all | Ask the reporter, or accept it is unusable |\n\n### Why collapsing six grades to two loses what you needed\n\nFor regulatory purposes these six are often collapsed into \"related\" and \"not related\", with certain, probable and possible mapping to related. Product teams do the same thing by instinct, sorting complaints into real and not-real.\n\nThe collapse throws away the two most actionable categories. \"Conditional\" says there is a specific fact that would settle this - a log line, a browser version, one follow-up question - and naming that fact is a concrete next action. \"Unassessable\" says the intake form is failing, which is a fixable process problem rather than a fact about the product. Both disappear the moment you force a binary, and what replaces them is an argument about whether a ticket counts.\n\n## The scored alternative: Naranjo\n\nWhere WHO-UMC relies on expert judgement, the Naranjo algorithm turns the same reasoning into a questionnaire. It is described as \"a questionnaire designed by Naranjo et al. for determining the likelihood of whether an adverse drug reaction (ADR) is actually due to the drug rather than the result of other factors.\"\n\n### The ten-item version, and its score bands\n\nTen items are scored and summed. A total of 9 or more is graded definite; 5 to 8 probable; 1 to 4 possible; and 0 or below doubtful. Two of the ten items are precisely the dechallenge and rechallenge tests. Item three asks: \"Did the adverse reaction improve when the drug was discontinued or a specific antagonist was given?\" Item four asks whether the reaction reappeared when the product was readministered.\n\nThe value of a scored instrument is not that the number is meaningful in itself. It is that two people scoring the same report independently will converge, and that their disagreement localises to a specific item rather than to a general impression. Adapting the ten items to your product once, as a checklist in your issue template, is a half-day of work that pays back permanently.\n\n## Your feature flag is a rechallenge apparatus\n\n### Your feature flag is a rechallenge apparatus\n\nHere is the observation that should change how your team triages.\n\nThe gold standard for establishing single-case causality in medicine is a positive rechallenge, and in medicine it is frequently unavailable because deliberately re-exposing a patient to something that harmed them is unacceptable. Entire methodological literatures exist to work around its absence.\n\nIn software, a rechallenge costs one toggle. If you run feature flags, per-account overrides, or percentage rollouts, you already own the apparatus for the strongest causal test in the field. You can turn the feature off for one complaining account, confirm the problem stops, turn it back on, and confirm it returns - with the customer's permission, inside a support conversation, in an afternoon.\n\nAlmost no product team does this. The default response to an ambiguous complaint is to wait for more complaints, which is the one move that cannot resolve the ambiguity, because every new report shares the same confounding structure as the first. Fifty uncontrolled anecdotes do not add up to one controlled test.\n\n### Why one positive rechallenge beats fifty more complaints\n\nConsider two responses to a single report that the new editor eats drafts.\n\n**Wait for volume.** Four weeks later you have 50 reports. You still do not know whether the editor is at fault, because all 50 arrived through the same channel, from the same self-selected population, after the same release that drew attention to the area. Volume has raised your confidence that *something* is wrong while telling you nothing new about *what*. Worse, the count itself is unreadable as a rate, and it moves with reporting propensity rather than with incidence - two separate problems covered in the companion articles.\n\n**Run the challenge.** Same afternoon: pull the timestamps (test 1), list what else shipped (test 2), disable the editor for that one account and confirm drafts persist (test 3), re-enable and confirm they vanish again (test 4). You now have a \"certain\" grade on a single case, and you can fix it before report 2 arrives. Where the four tests need facts the ticket does not contain, a short Koji study sent to the affected cohort collects them in a day.\n\nThe second path is faster, cheaper, and produces a stronger conclusion. The only reason the first is more common is that waiting feels like gathering evidence.\n\n## Where this sits relative to root cause analysis\n\n### Where this sits relative to root cause analysis\n\nThis method is upstream of techniques like the 5 Whys. Root cause analysis takes a causal link as given and asks why it exists, drilling from symptom toward mechanism. Causality assessment asks whether the link is there at all.\n\nRunning root cause analysis on an ungraded complaint is how teams produce confident, well-documented explanations of things that were never happening. Grade first, then drill. The grade also tells you how much drilling is justified: a \"certain\" earns a full investigation, a \"possible\" earns one clarifying question. A single Koji question put to the affected cohort is often enough to settle whether the link is there at all.\n\n## How Koji handles this\n\nCausality grading needs specific facts from the person who experienced the problem, and a support thread rarely captures them. This is where an interview beats a ticket.\n\n- **The four tests become structured questions.** Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so \"when did you first notice this\" becomes a single_choice on time windows, \"did it stop when we switched you back\" becomes a yes_no, and severity becomes a scale. The grade falls out of the answers instead of out of a debate.\n- **AI follow-ups chase the missing fact.** Most reports arrive \"unassessable\" because one detail is absent. Koji's AI interviewer probes for the file, the step, the browser and the timing automatically, on every response, which is what turns a conditional grade into a gradable one.\n- **Ask the whole affected cohort, not the loudest member.** Rather than waiting for report 50, send a short study to everyone in the affected segment. You get the temporal pattern across the cohort in a day.\n- **Voice or text, unmoderated.** No scheduling, so a causality check does not wait for a researcher's calendar.\n- **Quality scoring flags thin answers** so you know which responses are solid enough to grade and which are not.\n- **The report is live**, so the timing distribution is visible while you still have the flag in your hand.\n\nUsed together with a flag, the sequence is: interview the affected cohort to establish timing and exclude alternatives, then dechallenge and rechallenge one consenting account to settle it. Koji handles the first half; your flag system handles the second.\n\n## Frequently asked questions\n\n### Is it really safe to reproduce a bug on a customer account on purpose?\n\nOnly with explicit, informed consent, on a non-destructive problem, and never where data loss or a financial action is the symptom. For a rendering glitch, a slow load or a confusing flow, most customers are pleased to be asked, because it visibly means someone is working on their issue. Where the symptom is destructive, stop at the dechallenge and accept a probable grade - which is exactly what the WHO-UMC framework does for the same reason.\n\n### What if we cannot reproduce the problem at all?\n\nThen the honest grade is conditional or unassessable, and the useful output is the name of the specific missing fact. Write that fact down and go get it - the exact timestamp, the browser build, the file, the account state. A report you cannot grade is not a report you should either fix or dismiss; it is a request for one more piece of information.\n\n### How is this different from the 5 Whys?\n\nThe 5 Whys assumes the causal link and drills from symptom toward mechanism. Causality assessment decides whether the link exists. They run in that order, and reversing them produces thorough explanations of non-existent problems. Grade first, then drill, and let the grade set how much drilling is justified.\n\n### Do we need to grade every incoming complaint?\n\nNo, and trying to would be a waste. Grade the complaints that are about to change a decision - anything proposed for a rollback, a hotfix, or a roadmap slot. For the rest, the count is a discovery signal and can stay ungraded. The discipline is about what gets promoted into a decision, not about processing volume.\n\n### Who should do the grading?\n\nWhoever can see both the logs and the customer's words, which in most teams means a support engineer and a product manager for ten minutes together. The reason to use a written scale is that it makes the judgement reviewable: someone who disagrees has to say which item they score differently, which is a far more productive argument than one about whether a ticket is real.\n\n### Can a single report ever justify shipping a fix?\n\nYes. A single report with a clean dechallenge and a positive rechallenge is graded certain, and certain is the strongest evidence this framework produces regardless of how many reports back it. One controlled case outranks fifty uncontrolled ones. Volume is a signal about attention; a rechallenge is a signal about causation.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - turning the four causality tests into questions you can total\n- [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) - why waiting for more reports does not resolve an ambiguous one\n- [Feedback Volume Tracks Attention, Not Incidence](/docs/feedback-volume-reporting-propensity) - the trend problem that volume-based triage runs into\n- [The Five Whys Technique](/docs/five-whys-technique-user-research) - the drilling method that belongs after the grade, not before it\n- [Root Cause Analysis for Customer Research](/docs/root-cause-analysis-guide) - the fuller treatment of mechanism once causation is settled\n- [Usability Issue Severity Ratings](/docs/usability-issue-severity-ratings) - scoring how much a graded issue matters","category":"Analysis & Synthesis","lastModified":"2026-09-28T03:48:21.747399+00:00","metaTitle":"Did the Feature Cause the Complaint?","metaDescription":"A complaint naming a feature is a hypothesis, not evidence. Grade it with four tests - and use your feature flag as a rechallenge.","keywords":["did the feature cause the complaint","causality assessment product feedback","dechallenge rechallenge","naranjo algorithm product","complaint triage causality","who-umc causality"],"aiSummary":"A complaint that names a feature asserts a temporal coincidence plus a human guess, and openFDA states that for its own data \"a causal relationship cannot be established between product and reactions listed in a report\". The WHO-UMC system grades a single case on four criteria - time relationship, competing causes, dechallenge (removal) and rechallenge (reintroduction) - returning one of six grades rather than a binary. Collapsing those six to related/not-related discards the two most actionable ones, conditional and unassessable, each of which names a concrete next action. The Naranjo algorithm scores the same reasoning across ten items with bands at 9 or more for definite, 5 to 8 probable, 1 to 4 possible. The key product insight: a positive rechallenge is the strongest single-case causal test and is usually unethical in medicine, yet costs one feature-flag toggle in software, so one controlled case outranks fifty uncontrolled complaints. Koji turns the four tests into structured questions across an affected cohort.","aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}