Back to docs
Research Methods

Negative Controls in User Research: Test Your Process on a Signal That Is Not There (2026)

Run your research process where the answer must be nothing. If it still returns a confident finding, the finding is the process. Three controls you can run this quarter.

Answer first: a negative control is a deliberate run of your research process under conditions where the answer must be nothing - a concept you will never build, a segment the finding cannot apply to, an outcome your change could not possibly move. If the process still returns a confident finding, you have learned the most valuable thing available: that your process produces the same output whether or not the signal is there. Laboratory biologists do this as a matter of routine. Epidemiologists borrowed it. Product research almost never does, which is why so many concept tests return encouraging numbers for concepts nobody wants.

Koji makes negative controls affordable, because running one extra arm of a study costs a publish rather than another three weeks of recruiting.

What a negative control actually is

Marc Lipsitch, Eric Tchetgen Tchetgen and Ted Cohen set out the argument for importing this from the lab into observational research in Epidemiology in 2010, in a paper called "Negative Controls: A Tool for Detecting Confounding and Bias in Observational Studies". Their definition of the technique is one sentence and it is the whole idea: negative controls mean "to repeat the experiment under conditions in which it is expected to produce a null result and verify that it does indeed produce a null result".

Their point is that this is not an exotic method. It is "a routine precaution taken in the design of biological laboratory experiments", and it is aimed at exactly the category of problem that no amount of care inside the study can reach: "both suspected and unsuspected sources of spurious causal inference". You cannot adjust for a confounder you have not thought of. You can, however, run the whole machine in a situation where any signal it reports must be spurious.

The paper lists three lab strategies, each of which has an exact product-research counterpart:

  1. Leave out an essential ingredient. If the effect requires neutrophils, run the experiment without neutrophils and confirm nothing happens.
  2. Inactivate the hypothesised active ingredient. Neutralise the cytokine with a specific antibody; the killing should stop.
  3. Check for an effect that would be impossible by the hypothesised mechanism. Run it on a bacterial species the mechanism cannot touch.

They also distinguish two kinds: a negative control outcome, where you keep the exposure and measure something it could not have caused, and a negative control exposure, where you swap the cause for one that could not produce the effect.

The study that shows why this is not academic

The example Lipsitch and colleagues work through is the best argument for the method anyone has produced, because the failure it exposed had survived years of careful adjustment.

Observational studies had repeatedly found that seniors who get a flu vaccine are far less likely to die during flu season. Lisa Jackson and colleagues, writing in the International Journal of Epidemiology in 2006, tested that claim with a negative control built from the calendar: vaccination happens in autumn, influenza circulates in winter, so the protective effect - if it is real - must be concentrated during flu season and absent before it.

They followed a cohort of 72,527 people aged 65 and over across eight years. The relative risk of death for vaccinated compared with unvaccinated people was 0.39 before influenza season, 0.56 during it, and 0.74 after it.

The protection was largest in the window where the vaccine could not have been working. Their conclusion is blunt: "the magnitude of the bias demonstrated by the associations before the influenza season was sufficient to account entirely for the associations observed during influenza season". The people who get vaccinated in autumn are simply healthier than the people who do not.

Two details are worth carrying into product work.

The adjustment made it worse. The authors report that controlling for diagnosis-code covariates "resulted in estimates that were further from the null, in all time periods". The standard remedy - add controls - moved the numbers in the wrong direction, and nothing inside the analysis would have revealed that.

The floor exceeded the signal. A process that returns 0.39 where the true answer must be 1.00 cannot be read at face value when it returns 0.56 somewhere else. Dividing the in-season figure by the out-of-season one gives 1.44 - which is to say that once you deflate by the instrument's own floor, the apparent benefit does not merely shrink, it changes sign. That is an illustrative deflation rather than the authors' own estimate, but it makes the point: an uncalibrated instrument's output is not a measurement.

Jackson and colleagues also ran the second kind of control, the impossible-by-mechanism one: they substituted hospitalisation for injury or trauma as the outcome. Flu vaccination was "protective" against being hospitalised for an injury too.

The three negative controls a product research team can run this quarter

1. The decoy concept (leave out the essential ingredient)

Put a concept into your concept test that you are certain nobody wants and that you will never build. Not a joke - a plausible, competently written concept for a thing your team has independently concluded is dead. Score it with the same instrument, the same participants, the same scale.

Now you have a floor. Here is what that does to a real result set:

ConceptTop-two-box interestRead against the floor
Decoy (will never be built)32 percentThis is zero
Concept A47 percent15 points of real signal
Concept B36 percent4 points - noise
Concept C55 percent23 points - the strongest
Concept D30 percentBelow the floor

Without the decoy row, Concept D reads as nearly a third of the market is interested. With it, D scores below a concept the team knows is worthless, which is the single most useful sentence a concept test can produce. Note also that the ordering does not change but the interpretation does: the gap between B and A is more than three times larger than the raw numbers suggest, because both carry the same 32-point floor.

This is not a fake door test. A fake door measures demand for a feature you are considering building, covered in Fake Door Testing. A decoy is the opposite instrument: a concept you will never build, included solely to find out what your demand test says when the answer is no. The two work well together, and running a fake door without ever having measured its floor is how a 4 percent click-through gets promoted to a roadmap item.

The reason a floor exists at all is well documented. In a meta-analysis in Environmental and Resource Economics in 2005, James Murphy and colleagues gathered 28 stated-preference valuation studies that elicited both hypothetical and real values through the same mechanism, producing 83 observations, and found a "median ratio of hypothetical to actual value of only 1.35, and the distribution has severe positive skewness". Their opening line is a useful corrective in both directions: individuals "are widely believed to overstate their economic valuation of a good by a factor of two or three" - widely believed, and the measured median is 1.35, with a long tail. You cannot guess your instrument's floor from folklore. You have to measure it, and the decoy is how.

2. The placebo segment (inactivate the active ingredient)

Take the finding and re-run the analysis on a segment for which the causal story cannot hold. If your explanation is enterprise admins churn because the SSO setup requires their IT team, run the same analysis on self-serve single-seat accounts, who have no IT team and no SSO. If the same theme comes out with the same strength, the theme is coming from your codebook, your prompt or your analyst, not from the mechanism.

This is the negative control that catches analytic artefacts rather than sampling ones, and it is the only cheap way to find out whether a thematic analysis is reading the data or reading the researcher.

3. The untouchable metric (impossible by mechanism)

Pick an outcome your change could not have affected and check that it did not move. Jackson's team used injury hospitalisation. A product team shipping an onboarding improvement can check whether the improved cohort also shows a lift in billing-page visits, support-article reads for an unrelated product area, or logins to a different surface. If everything improved, you did not improve onboarding; you selected a better cohort.

This one is nearly free, because it requires no new fieldwork - only the discipline of naming the untouchable metric before you look.

Why this belongs at the end of a chain

Most research failure modes are visible somewhere in the data if you know where to look: a bad denominator, a shifting definition, a sample that under-covers a segment. This one is not.

Every individual response in a decoy-inflated concept test is legitimate. Nobody lied. The participants really did rate it a 4. The analysis really did find the theme. The formula was applied correctly. There is no defect to find inside the study, because the process is not malfunctioning - it is functioning exactly as it does when the signal is real. The output is the same whether or not the thing you are studying exists. No sample size fixes that, and a larger sample makes it worse by estimating the artefact more precisely.

That is what distinguishes a negative control from the other quality moves in this corpus. Attention checks catch participants who are not paying attention, covered in Attention Check Questions. Screener precision measures who got into the sample, covered in Screener Accuracy. Both inspect the inputs. A negative control inspects the machine, and it is the only one of the three that can detect a bias nobody anticipated.

How Koji helps

Negative controls fail to happen for an economic reason, not an intellectual one: in traditional research, an extra arm means an extra round of recruiting, scheduling and moderation, which nobody will spend on a concept they have already decided is dead.

  • The decoy arm is nearly free. Koji's AI-moderated interviews run in parallel, so adding a fifth concept or a placebo segment changes the participant count, not the calendar. A control that costs a day gets run; a control that costs three weeks gets cut in the planning meeting.
  • Identical instrument, by construction. A negative control is only valid if the control arm is measured exactly like the real one. Because Koji studies are defined as structured question sets rather than as a moderator's habits, the decoy concept meets the same scale question, the same ranking, the same probes. Human moderators, however disciplined, cannot guarantee that.
  • The six structured question types give you comparable floors. open_ended, scale, single_choice, multiple_choice, ranking, and yes_no each produce a measurable distribution, so the floor is a number rather than an impression. Ranking questions are especially useful here, because a decoy placed in a ranking forces participants to rank it against real options rather than rate everything positively. See Structured Questions in AI Interviews.
  • Automatic thematic analysis makes the placebo-segment test a same-day exercise. Re-running the analysis over a segment the story cannot apply to is a filter and a re-run, not a second synthesis workshop.
  • Voice interviews and real-time reporting mean the floor is measured on the same population, in the same week, under the same conditions as the finding it calibrates - which is the condition Lipsitch and colleagues spend most of their paper trying to approximate in observational data.

A legacy survey tool can technically host a decoy concept. What it cannot do is make the extra arm cheap enough that teams run one every time, which is the difference between a technique that exists and a technique that is used.

Common mistakes

Choosing a decoy that is obviously absurd. The control has to be plausible enough that a participant engages with it seriously. A comedy concept measures the floor of a joke, not the floor of your instrument.

Telling the analysis team which arm is the control. Blind it if you can. A synthesis that knows which concept is the decoy will find reasons for its score.

Running the control once and then trusting it forever. The floor moves with the panel, the season and the wording. Re-measure when any of those change.

Treating a failed negative control as a reason to discard the study. It is a reason to deflate the estimate and to look for the confounder. Jackson and colleagues did not throw away the vaccine data; they used the control to show how much of the effect it could account for.

Only ever running positive checks. A process that has only ever been tested on cases where you expected a finding has never been tested at all. Schedule one negative control per quarter the way you schedule a Koji study refresh, and it stops being the thing that gets cut.

Frequently asked questions

What is a negative control in user research?

It is a deliberate run of the research process where the correct answer is known to be nothing: a concept you will never build, a segment the causal story cannot apply to, or a metric the change could not have moved. The purpose is to see what your process reports when there is nothing to report. Lipsitch and colleagues describe the general form as repeating the experiment under conditions expected to produce a null result and verifying that it does.

How is a decoy concept different from a fake door test?

A fake door test measures demand for something you might actually build. A decoy measures what your demand instrument says about something you definitely will not. Running both gives you a number and a floor to read it against; running only the fake door gives you a number and no scale. The mechanics of fake doors are covered in Fake Door Testing.

Does this replace statistical significance testing?

No, it sits underneath it. Significance testing asks whether an observed difference is larger than chance variation; a negative control asks whether your process produces differences of that size when the true answer is zero. A result can be statistically significant and still sit below the floor your own control established, which is exactly the situation Jackson and colleagues documented.

How many participants does a negative control arm need?

Roughly the same as the arm it calibrates, because you are estimating the floor with comparable precision. In practice teams run the decoy at the same n as each real concept. If that feels expensive, that is the cost signal Koji's parallel AI-moderated interviews are designed to remove.

What if the negative control comes back clean?

Then you have positive evidence that the instrument can return a null, which is worth stating in the readout. It is the same logic as a severe test that a belief survives, discussed in Falsifiable Research Questions: a check you could have failed and did not is informative, and a check you could not have failed is not.

Can I run a negative control on a qualitative study?

Yes, and the placebo-segment version is the most useful thing on this list for qualitative work. Run your codebook over transcripts from a population the finding cannot describe. If the themes reappear at similar strength, the themes are an artefact of the analysis, which is a problem no amount of additional coding will fix. The same logic applied to a whole protocol, using seeded falsehoods, is Research Design Sensitivity.

The bottom line

Every research process has a floor - the finding it returns when nothing is there - and almost no product team has measured theirs. Add one concept you will never build, one segment your story cannot describe, and one metric your change cannot move. If the process lights up on any of them, you have not found a flaw in a study; you have found out what your studies have been reporting all along.

Related Resources

Related Articles

Attention Check Questions: How to Catch Low-Effort Survey Responses Without Annoying Real Participants

Attention check questions catch inattentive, low-effort, and fraudulent survey responses. Learn the main types, how many to use, the pitfalls, and why a conversational AI interview reduces the need for them in the first place.

Fake Door Testing (Painted Door Test): Validate Demand Before You Build

A practical guide to fake door and painted door testing — how to measure real demand for a feature before writing code, what metrics to track, the ethics, and how to learn the why behind every click.

Falsifiable Research Questions: Can Your Study Produce the Answer You Do Not Want? (2026)

A question is falsifiable when you can name the answer that kills the belief and your protocol can produce it. Sharp questions beat big samples by about 78 to 1.

Research Design Sensitivity: Would Your Study Have Caught It If You Were Wrong? (2026)

Coverage is the list of topics your guide touches. Sensitivity is whether any answer could have contradicted you. Measure it by seeding falsehoods into your own plan.

Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted (2026)

A research screener is a diagnostic test. At a 5% target incidence, a screener with 90% sensitivity and 85% specificity delivers a sample that is 76% wrong. How to compute positive predictive value, measure it on your own studies, and raise it.

Smoke Tests and Fake Door Tests: How to Validate Demand Before You Build

Smoke tests and fake door tests measure real user demand for an idea before any code is written. Learn the playbook used by Buffer, Dropbox, and modern product teams — and how to pair it with AI interviews.

Stated vs. Revealed Preferences: Why Customers Say One Thing and Do Another (2026)

Customers routinely say one thing and do another — the say-do gap. This guide explains stated vs. revealed preferences, why the gap exists, what the data shows about its size, and how to design research that gets past what people claim to what they actually do.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.