Back to docs
Participant Recruitment

Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted (2026)

A research screener is a diagnostic test. At a 5% target incidence, a screener with 90% sensitivity and 85% specificity delivers a sample that is 76% wrong. How to compute positive predictive value, measure it on your own studies, and raise it.

A screener is a diagnostic test, and like every diagnostic test its usefulness depends on something most teams never measure: the incidence of your target population in the pool you are screening. A screener with 90% sensitivity and 85% specificity sounds excellent. Run it against a pool where 5% of people are your target and 76% of everyone who passes will be the wrong person. The pass rate tells you nothing about this. The number that does is positive predictive value, and it is the number almost no research team computes.

This guide shows you how to compute it, what it does to your recruiting plan, and the two levers that actually move it.

Your screener is a diagnostic test

Every screener asks the same structural question a medical test asks: given this evidence, does this person belong to the class I care about? That means it has the same four outcomes, and the same four names.

Actually your targetNot your target
Screener says passTrue positiveFalse positive
Screener says failFalse negativeTrue negative

Two properties describe how the test behaves on people whose status you already know:

  • Sensitivity is the share of true targets who pass. If 90 out of every 100 real weekly users correctly report weekly use, sensitivity is 90%.
  • Specificity is the share of non-targets who correctly fail. If 85 out of every 100 monthly users correctly report monthly use, specificity is 85%.

Both are properties of the questions you wrote. Neither one tells you what you actually want to know, which runs in the opposite direction: given that this person passed, what is the probability they are really my target? That is the positive predictive value, and it depends on a third number that has nothing to do with your questions at all.

The third number: incidence

Incidence is the share of your screening pool that genuinely belongs to your target population. Panel vendors call it IR and price on it. Most in-house teams never estimate it, because the pool feels like it is full of the right people. It rarely is.

Here is the arithmetic on a pool of 2,000 people, with a target incidence of 5%, a screener at 90% sensitivity and 85% specificity:

Actually target (100)Not target (1,900)Total
Passed90285375
Failed101,6151,625

375 people pass. Only 90 of them are your target. Positive predictive value is 90 / 375 = 24%. Three out of every four people you are about to interview are not the people you designed the study for, and every one of them passed a screener you would describe as accurate.

Note what happens to the mirror-image number. Negative predictive value is 1,615 / 1,625 = 99.4%. Your screener is superb at telling you who is definitely not a fit and nearly useless at telling you who is. That asymmetry is not a flaw in your questions. It is what low incidence does to every test.

Why smart people get this backwards

This confusion is not a research problem, it is a general problem with how conditional probabilities read. Gerd Gigerenzer and Adrian Edwards documented it in the BMJ with doctors who had an average of 14 years of professional experience. Asked about a colorectal cancer screen with a cancer prevalence of 0.3%, a test sensitivity of 50% and a false positive rate of 3%, they were asked what the probability was that someone who tested positive actually had cancer.

The correct answer is about 5%. As the authors report, the doctors' answers ranged from 1% to 99%, with about half of them estimating the probability as 50% (the sensitivity) or 47% (sensitivity minus false positive rate). In other words, half of a room of experienced clinicians answered with the test's sensitivity when asked for its predictive value.

Gigerenzer and Edwards name the error precisely: "the conditional probability of A given B is confused with that of B given A." They also observe that "many doctors have trouble distinguishing between the sensitivity, the specificity, and the positive predictive value of test — three conditional probabilities" (Gigerenzer G, Edwards A. Simple tools for understanding risks: from innumeracy to insight. BMJ 2003;327:741-744).

When a product manager reads "82% of people who took our screener qualified" and concludes the sample is clean, that is the same substitution. The pass rate is a fact about your questions. It is not a fact about your sample.

Their fix is worth borrowing wholesale: state the numbers as natural frequencies rather than percentages. In their mammography example, out of 1,000 women, "only seven of the 77 women who test positive actually have breast cancer, which is one in 11 (9%)." Seven and 77 refer to the same 1,000 people, so no reference class switch is required to compare them. Write your own screener numbers that way and the problem becomes visible without any statistical training: of the 375 people who passed, 90 are who we wanted.

What incidence does to the plan

Hold the screener constant at 90% sensitivity and 85% specificity, and vary only the incidence of the pool:

Target incidencePass ratePositive predictive value
1%15.8%5.7%
5%18.8%24.0%
10%22.5%40.0%
20%30.0%60.0%
50%52.5%85.7%

Two things in that table are worth staring at.

First, the pass rate barely moves while the predictive value moves by a factor of 15. A pool at 1% incidence and a pool at 10% incidence produce pass rates of 15.8% and 22.5%. Those look like the same study. One of them yields a sample that is 94% wrong and the other yields a sample that is 60% wrong.

Second, the highest-leverage variable is not your screener. Going from 5% to 20% incidence takes PPV from 24% to 60% without changing a single question. Rewriting the screener to reach 95% specificity at the same 5% incidence gets you to 48.6%. Where you recruit outperforms how you screen.

The recruiting plan follows directly. If you need 12 genuine participants at 24% PPV, you need to complete 50 sessions, and to get 50 passers at an 18.8% pass rate you need to screen 267 people. Teams that plan for "12 interviews, so let's screen 40 people" are not under-recruiting by a little. They are under-recruiting by a factor of six, and the gap gets filled with whoever passed.

What low incidence looks like in the wild

A published account from Drexel University makes the funnel concrete. A qualitative team recruiting caregivers of people with dementia who also had chronic leg wounds sent a recruitment email through a national NIH-funded research registry to 1,499 people. Twenty-two indicated interest and were passed to the team for screening. Of those, two were eligible and enrolled, four were not eligible, and 16 never responded (Sefcik JS, Hathaway Z, DiMaria-Ghalili RA. When snowball sampling leads to an avalanche of fraudulent participants in qualitative research. Int J Older People Nurs 2023;18(6):e12572).

Two enrolled participants out of 1,499 contacts is an effective incidence of 0.13%. At that incidence, no screener you can write produces a usable predictive value. The only workable moves are the ones that change the pool, and that study went on to demonstrate what happens when a team reaches for an easier pool instead: the same paper documents an influx of fraudulent participants after the team opened snowball referrals, and concludes that none of the data collected were used for analysis.

The two numbers you can actually measure

You do not know your true incidence, your sensitivity or your specificity, and you never will exactly. You can estimate all three well enough to act.

1. Estimate PPV directly from the sessions you already ran. After each interview, have the moderator or the transcript answer one question: did this person turn out to be the target? Aggregate that across a study and you have a measured positive predictive value, not a modelled one. If 7 of your last 20 sessions were with the wrong person, your PPV is roughly 35% and every plan should assume it.

2. Estimate sensitivity by re-screening people you already know. Take 30 customers whose status you can verify from product data (they demonstrably used the feature weekly last month) and send them the screener cold. The share who pass is your sensitivity, measured on real humans rather than assumed. This is the single most useful hour of research ops most teams have never spent, and it usually returns a number well below the 90% everyone assumes.

With a measured PPV and a measured sensitivity, incidence falls out of the arithmetic, and your recruiting plan stops being a guess.

How to raise predictive value

There are exactly three levers, and they are not equally strong.

Raise the incidence of the pool. This is the strongest lever by a wide margin, and the one that has nothing to do with question wording. Screening your own product's weekly-active list instead of a general panel can take incidence from 3% to 60%, which does more for your sample than any rewrite. See in-product research recruiting for the mechanics, and purposive sampling for choosing the frame deliberately.

Add a second, independent test. Two tests compound only if they are independent, which means the second one must not be another self-report question in the same format. A behavioural confirmation (product telemetry, an uploaded artefact, an account lookup) is independent of a claimed answer in a way that a differently-worded version of the same question is not. This is where an AI-moderated first few minutes earns its keep: an open-ended "walk me through the last time you did this" is a genuinely different test from a multiple-choice "how often do you do this?", because it requires episodic detail rather than a category selection.

Improve specificity, last. Rewriting questions to be harder to fake helps, but as the table above shows, it moves the number least, and as the companion article explains, tightening has a cost that most teams never price. See tightening your screener before you reach for this one.

What this is not

This article is about the base rate of a class of people in a pool — the prevalence that determines whether a passing screener means anything. It is not about the base rate of a product bet succeeding, which is a different question about reference classes for forecasting and is covered in the base rate for a product bet. Same phrase, different denominators.

Nor is it about sensitivity in the psychophysical sense. Signal detection theory uses the same 2x2 and the word sensitivity to describe a respondent deciding whether they noticed something, where the two numbers of interest are discriminability and criterion. Here the thing being tested is not a person's perception but an instrument, and the decisive third number is the prevalence of the target in the pool, which does not appear in that framing at all.

It is also not a guide to writing screener questions. For question types, templates and flow, use screener questions for user research and research screener questions. This article assumes those questions exist and asks what they are actually delivering.

The modern approach: screening as a conversation, not a gate

Traditional survey tools treat screening as a binary gate: the respondent answers five multiple-choice questions, a rule fires, and they are in or out. That design guarantees the arithmetic above, because a fixed set of self-report questions is a single test with a fixed sensitivity and specificity, and nothing downstream ever revisits the decision.

Koji changes the structure of the test rather than the wording of the questions.

  • Structured questions do the deterministic part. Koji supports six question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — so the hard eligibility rules (role, tool, purchase window) are captured as clean, machine-readable fields rather than parsed out of prose. See the structured questions guide for how each type behaves.
  • The AI moderator runs the second, independent test. Because Koji's AI interviewer probes follow-ups in the same session, a claimed answer is immediately tested against an episodic one. Someone who selected "I manage a team of five" and then cannot describe a single thing they did as a manager last week has failed a test that no screener form can administer. That probing is the same mechanism the Drexel team eventually relied on: their nurse scientist detected the fraudulent participants through inconsistencies under follow-up questioning, not through the screening tool.
  • Quality scores make PPV measurable. Every Koji interview carries a 1-5 quality score, so the share of sessions that were genuinely with the target becomes a number you can read off a report instead of a suspicion you carry around.
  • Screening is not a wasted contact. Because screening and interviewing happen in one continuous session rather than as a form followed by a scheduled call days later, the 16 non-responders in the Drexel funnel never form. People who pass go straight into the interview while their intent is live.

The teams that get this right stop treating the screener as a filter to be perfected and start treating it as the first of two tests, with the conversation as the second. That is a structural fix, and structural fixes beat wording fixes at every incidence in the table above.

Frequently asked questions

What is a good positive predictive value for a research screener?

There is no threshold that is good in the abstract, because PPV is not a property of your screener — it is a property of your screener combined with your pool. A PPV of 40% is excellent when you are hunting a 2% population and poor when you are screening your own power users. The useful discipline is to measure it, state it alongside your findings, and plan recruitment volume from it rather than from the number of sessions you want. If you have never measured it, assume it is worse than you think: the first measurement almost always comes in below the team's estimate.

How do I estimate the incidence of my pool without a big study?

Take the last 100 people who entered your screener and, for as many as you can, verify their target status from a source that is not the screener: product analytics, CRM records, support history, an account lookup. The share who genuinely qualify in that verified subset is a serviceable incidence estimate. If you cannot verify any of them from an independent source, that is itself the finding — you are recruiting from a pool whose composition you have no visibility into, and the correct next step is to change pools rather than to sharpen questions.

Does a longer screener produce a better sample?

Not reliably, and often the opposite. Each additional criterion multiplies the chance that a genuine target is wrongly excluded, and the people excluded are not a random subset of your targets. A five-criterion screener at 90% sensitivity per criterion keeps only 59% of the real population. It also does very little to the respondents who are deliberately misrepresenting themselves, because they answer strategically rather than honestly. The full arithmetic is in tightening your screener.

Why does the pass rate not tell me whether my screener is working?

Because the pass rate mixes true positives and false positives into one number, and the mix depends on incidence, which the pass rate cannot see. Two studies with a 19% pass rate can have predictive values of 6% and 40%. A falling pass rate is equally ambiguous: it can mean your criteria got sharper or it can mean your pool got worse. Only a measurement that compares screener verdicts to verified status — a post-session confirmation, a product-data check — distinguishes the two.

Is this the same thing as sampling bias?

Related but distinct. Sampling bias is about who enters your frame at all. Screener predictive value is about what happens to the people already in the frame when you apply a fallible test to them. You can have a perfectly representative pool and still end up with a sample that is 76% wrong, purely from the arithmetic of low incidence. In practice both are usually operating at once, which is why measuring PPV separately is worth the effort.

Can I just interview everyone and sort it out afterwards?

Sometimes, and it is an underrated option when incidence is moderate and session cost is low, which is exactly the situation AI-moderated interviewing creates. If a session costs minutes rather than a scheduled hour, letting borderline candidates through and classifying them from the transcript gives you a much better test than the screener, because it uses far more evidence. The constraint is analysis discipline: you must actually tag which sessions were on-target and exclude them from the findings, rather than letting them quietly widen every theme.

Related Resources

Related Articles

The Base Rate Nobody Measured: Why Every Flag in Your Research Stack Has an Unknown Precision (2026)

Every AI tag, sentiment label and risk score is a diagnostic test whose precision depends on a prevalence nobody measured. Accuracy rises as precision collapses. How to audit the unflagged pile and publish a precision footer.

Purposive Sampling: The Complete Guide to Strategic Participant Selection

A complete guide to purposive (purposeful) sampling in qualitative research — covering all major types, when to use each, how to determine sample size, and how AI tools enable purposive sampling at scale.

In-Product Research Recruiting: Recruit Customer Interview Participants From Inside Your App

Stop paying recruiting panels for participants you already have. Learn how to recruit research participants directly from your product using embedded prompts, in-app banners, email triggers, and personalized AI interview links. Faster, cheaper, and more representative than external panels — with zero scheduling friction.

Screener Questions for User Research: A Complete Guide

Learn how to write effective screener questions that find the right research participants — and how Koji's intake forms and AI interviews make screening faster and more natural.

Tightening Your Screener Makes the Sample Purer and the Findings Worse (2026)

Adding screening criteria rejects 41% of your genuine target population at five criteria, and barely touches respondents who misrepresent themselves. Positive predictive value rises while the share of sessions held with a deceptive respondent nearly triples.

Did Users Actually Notice? Sensitivity vs Criterion in Did-You-Notice Questions (2026)

The percentage of users who say they noticed your change is not a measurement of whether they noticed. Signal detection theory separates detection from willingness to say yes.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.