Back to docs
Research Methods

The Case Definition: The Decision That Fixes or Breaks Every Issue Count (2026)

Before you count affected users you have to decide who counts. Epidemiology has a century of discipline on this, including one rule almost every product team breaks.

Every number you report about an issue - how many users hit it, whether it is getting better, which segment suffers most - is downstream of one decision you probably made in a hurry: what counts as a case. Field epidemiology treats that decision as a formal artifact called a case definition, writes it down, and changes it deliberately at known points in an investigation. Product teams usually leave it implicit, which is why two analysts can report an 8x difference in affected users from the same ticket queue and both be right.

The short answer

Write the definition down before you count. Use a deliberately broad definition while you are still working out what is happening, and a deliberately narrow one when you are testing what caused it. Classify cases into tiers rather than in or out. And never, under any circumstances, define a case by the cause you are trying to evaluate.

What a case definition actually is

The CDC's Principles of Epidemiology course defines it plainly: "A case definition is a set of standard criteria for classifying whether a person has a particular disease, syndrome, or other health condition." In an outbreak investigation it is usually restricted further by time, place and person.

The product translation is direct. A case definition for an issue investigation specifies the observable symptom, the window, and the population - for example, an account that submitted a payment and received a failure response, between the 14th and the 28th, on the web client. What it must not specify is why you think it happened.

Use two definitions, not one

The single most useful thing epidemiology does here is refuse to pick one definition. The CDC describes the staging explicitly. Early on, "investigators may use a 'loose' or sensitive case definition that includes confirmed, probable, and possible cases to characterize the extent of the problem, identify the populations affected, and develop hypotheses about possible causes." The course then describes the second stage just as explicitly: "Later on, when hypotheses have come into sharper focus, the investigator may tighten the case definition by dropping the 'possible' and sometimes the 'probable' category."

The reason for tightening is stated just as plainly: "In analytic epidemiology, inclusion of false-positive cases can produce misleading results."

So the two definitions do different jobs, and using the wrong one for the wrong job is the error:

  • Sensitive, early. You want to know whether there is a problem at all, roughly how big it is, and who is in it. Missing real cases is the expensive mistake here, so cast wide and accept some false positives.
  • Specific, later. You want to know what caused it, or whether a fix worked. Now a false positive is the expensive mistake, because contaminated cases wash out the association you are trying to measure.

The CDC's tiers transfer cleanly. "To be classified as confirmed, a case usually must have laboratory verification. A case classified as probable usually has typical clinical features of the disease without laboratory confirmation. A case classified as possible usually has fewer of the typical clinical features."

TierEpidemiologyProduct research
ConfirmedLaboratory verificationReproduced, or server logs show the failure for that account
ProbableTypical features, no lab confirmationClear symptom description matching the known signature, no log evidence
PossibleFewer typical featuresVague report consistent with the issue but also with several others

The rule almost every product team breaks

Here is the sentence worth printing out. The CDC states: "The case definition must not include the exposure or risk factor you are interested in evaluating. This is a common mistake."

Consider what happens when you ignore it. Your team suspects the new checkout provider. Someone defines the investigation set as tickets mentioning the new checkout. Now 100 percent of your cases involve the new checkout, and the provider looks overwhelmingly implicated. But that result was guaranteed by the definition before any data was examined. You did not measure an association; you asserted one and then counted it.

The correct move is to define the case by the symptom - a failed payment submission in the window - and then measure how many of those cases touched the new provider versus the old one. Only then does the comparison carry information. If 300 of 400 symptom-defined failures ran through the new provider while it handled only 40 percent of traffic, you have a real signal. If they are proportional to traffic, you have just saved yourself a rollback.

This failure is especially common when the investigation set comes from a search query, because a keyword search is an exposure-based definition wearing a convenient disguise. Searching your corpus for the feature name you suspect is the same mistake with a different interface.

A worked example: one queue, eight different answers

Take 1,000 support tickets from a two-week release window. After classification:

  • Confirmed - reproduced in-house or matched to a server-side failure for that account: 48
  • Probable - describes the exact signature, submitted payment then saw a specific error, no log match available: 112
  • Possible - says something like "checkout is broken" with no detail: 260

A sensitive definition counts all three: 420 affected accounts. A specific definition counts confirmed only: 48 affected accounts. From one queue, with no disagreement about any individual ticket, the headline number varies by a factor of 8.75. Both are defensible. Neither is wrong. They answer different questions.

Now the trap. Suppose week one reports the sensitive number, 420, because you were still scoping the problem. Week two, now in causal-analysis mode, you report the specific number, 48. Nobody is lying and nobody mentions the change. The dashboard shows affected accounts falling from 420 to 48 - an 89 percent improvement - produced entirely by a definition change. A fix that did nothing gets credited, and the engineer who actually fixed something next month will be unable to show it.

The discipline that prevents this costs one line: report the definition alongside the count, every time, and when you tighten it, restate the prior period under the new definition as well. Two numbers for the transition period, always.

Changing the definition without corrupting the trend

Changing a case definition mid-investigation is correct practice, not a failure. Doing it silently is the failure. A workable protocol:

  1. Version the definition. A short identifier in the report is enough.
  2. When the definition changes, recount at least one prior period under both the old and new version, and publish both.
  3. Keep the tier assignments, not just the total. If you only store the final count you cannot ever recount, and every historical comparison becomes permanently unavailable.
  4. State the window and population in the definition itself, so a later reader cannot accidentally compare a seven-day count to a fourteen-day one.

How Koji handles this

The reason retrofitting a case definition onto free text is painful is that the defining symptom was never asked of everyone. A support queue only contains people who wrote in, and only tells you what each of them happened to volunteer. Koji lets you move the definition upstream, to collection time.

Because a Koji study can mark a question required, the AI interviewer covers the defining symptom with every participant rather than only the ones who raise it unprompted. That gives you something a ticket queue structurally cannot: a denominator of people who were actually asked. Koji's structured questions make the classification itself unambiguous - a yes_no item for whether the participant hit the failure, a single_choice item for which flow they were in, a scale item for severity - so the case assignment does not depend on parsing a sentence someone typed while annoyed. Those are three of the six types available, alongside open_ended, multiple_choice and ranking.

Koji's analysis layer maps onto the tier structure more neatly than most teams expect. Each extracted answer carries a confidence flag of high, medium or low, which is a close analogue of confirmed, probable and possible, and each one stores the indices of the transcript messages it was drawn from. So when you tighten a definition from "all tiers" to "high confidence only," you can recount the existing corpus instantly instead of re-reading it, and you can show a reviewer exactly which participant words put each case in its tier. That is what makes the honest two-number transition cheap enough to actually do.

One more advantage worth naming: Koji's AI interviewer asks follow-up questions in the moment, so an ambiguous answer can be resolved into a confirmed or ruled-out case during the interview. In a survey or a ticket queue the ambiguity is permanent, and permanent ambiguity is what forces you to choose between a number that overcounts and a number that undercounts.

Common mistakes

  • Defining the case by the suspected cause. Including the exposure guarantees the finding. A keyword search for the feature you suspect is the same error.
  • Reporting a count without its definition. The number is meaningless alone, and the next analyst will unknowingly produce a different one.
  • Tightening the definition silently. This manufactures improvement. It is the single most common way a feedback dashboard lies.
  • Using the strict definition during discovery. You will conclude the problem is small because you excluded most of it, and close the investigation early.
  • Using the loose definition for causal analysis. False positives dilute the association until a real cause looks like noise.
  • Collapsing to in or out. The probable and possible tiers carry the information about how confident you are, and deleting them replaces a measurable uncertainty with an argument.
  • Leaving the window implicit. Half of all apparent trends in issue counts are two different window lengths compared to each other.

Frequently asked questions

What is a case definition in product research?

It is a written set of criteria deciding which users, sessions or reports count as instances of the issue you are investigating. It specifies the observable symptom, the time window and the population, and it deliberately excludes any statement about the suspected cause.

Why would I use two different definitions in one investigation?

Because the two phases have opposite error costs. While scoping, missing real cases is worse, so you want a broad definition. While testing a cause or a fix, false positives are worse, because they dilute the association you are trying to detect. One definition cannot be correct for both.

Why can the case definition not include the cause I suspect?

Because it makes the conclusion circular. If a case is defined as a report mentioning the new checkout, then every case involves the new checkout by construction, and the apparent association is an artifact of your own filter rather than a measurement.

Can I change the case definition partway through?

Yes, and you often should. The requirement is that you version it, and that you recount at least one prior period under both the old and the new definition so the trend stays interpretable. Silent changes create improvements that did not happen.

How do confirmed, probable and possible map onto product issues?

Confirmed means you reproduced it or found server-side evidence for that specific account. Probable means the report matches the known signature but you have no independent confirmation. Possible means the report is consistent with this issue and also with several others.

Does this matter if I only have support tickets to work with?

It matters more. A ticket queue contains only people who wrote in and only the details they volunteered, so you have no denominator of people who were asked. You can still classify into tiers and publish the definition, but you should state plainly that the count is of reports, not of affected users.

Related Resources

Related Articles

Why Complaint Counts Cannot Become Rates (And What to Compute Instead)

A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.

Customer Feedback Categorization: How to Build a Feedback Taxonomy That Scales (2026)

A practical guide to categorizing customer feedback: designing a feedback taxonomy, choosing flat vs. hierarchical tags, avoiding tag sprawl, and using AI auto-tagging to turn thousands of unstructured comments into quantified themes.

Why Your Quarterly Metric Shows a Trend That Is Not There (2026)

Undersampling does not blur a cycle, it counterfeits a different one. How the gap between your measurement waves manufactures smooth trends, flat lines, and reversed directions - and the three-question test that catches it.

Product Feedback Triage: A Framework for Turning Noise Into a Prioritized Backlog

A practical framework for triaging product feedback at scale — capture, dedupe, tag, route, and validate every request before it ever reaches prioritization. Includes a triage workflow, a severity matrix, and an AI-native approach.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Support Ticket Analysis: How to Mine Customer Service Data for Product Insights

A practical guide to systematically extracting product insights from customer support tickets — covering manual coding workflows, AI-powered thematic analysis, and how to tie ticket themes to business impact.