Tightening Your Screener Makes the Sample Purer and the Findings Worse (2026)
Adding screening criteria rejects 41% of your genuine target population at five criteria, and barely touches respondents who misrepresent themselves. Positive predictive value rises while the share of sessions held with a deceptive respondent nearly triples.
When you discover that most of the people passing your screener are not your target, the reflex is to tighten it: add criteria, add attention checks, add a speeder flag. That reflex improves the metric you were worried about and makes your sample worse. Tightening deletes a large, non-random share of your genuine participants, and it barely touches the respondents who are misrepresenting themselves on purpose — because those two groups fail screening tests for opposite reasons. In a worked model below, five criteria raise positive predictive value from 20% to 45% while the share of your final sessions held with someone actively deceiving you rises from 20% to 55%.
This is the second half of the screening problem. The first half is in screener accuracy; this article is about what happens when you act on it.
Criteria multiply, and they multiply against you
Screening criteria are conjunctive: a candidate has to pass all of them. If each criterion admits a genuine target 90% of the time — a generous assumption, since real people misremember frequencies, use different words for their own role, and round their own usage — then joint sensitivity is the product.
| Criteria | Genuine targets who pass all | Genuine targets lost |
|---|---|---|
| 1 | 90.0% | 10.0% |
| 3 | 72.9% | 27.1% |
| 5 | 59.0% | 41.0% |
| 8 | 43.0% | 57.0% |
A five-criterion screener, each part of which you would describe as accurate, rejects 41% of the population it was designed to find. An eight-criterion screener rejects more of your target population than it admits. Nobody experiences this as a failure, because rejected candidates leave no trace: they see a "thanks, you're not a fit for this study" message and never appear in any report you read.
The false negatives are not a random sample of your targets
If the 41% were a random subset, the cost would be pure recruiting inefficiency and nothing more. They are not random. Screening questions test a specific and unevenly distributed skill: the ability to recognise your own behaviour in someone else's categories, quickly, in a form.
The people who fail that test while genuinely belonging to your population are systematically:
- The ones who use different vocabulary for their own job. A "revenue operations" screener misses the person whose title is "sales systems lead" and who does exactly the job you are studying.
- The ones who are honest about uncertainty. Asked how many times last month they used a feature, a careful person picks the conservative bucket. The frequency thresholds in most screeners punish that.
- The ones whose usage is bursty rather than regular. "Weekly or more" excludes the person who uses your product intensely for three days per release cycle, who may be your most interesting participant.
- The ones for whom the behaviour is unremarkable. Deeply habituated users under-report their own frequency, because habits are hard to count from memory. See recall bias for the underlying mechanism.
Every one of those exclusions is correlated with something you are trying to learn about. Tightening does not shave noise off the edges of your sample; it selects for participants who are fluent at answering screeners about themselves in your terminology. That is a real trait, and it is not the trait you were sampling on.
Meanwhile the strategic respondents sail through
The other half of the inversion is the more uncomfortable one. The standard tightening instruments are close to useless against respondents who are deliberately misrepresenting themselves, and there is a large, well-measured study that says so.
Pew Research Center ran a study across six online sources — three opt-in sample sources, one crowdsourcing platform, and two panels recruited offline by address-based random sampling — comprising more than 60,000 interviews with at least 10,000 from each source. It defined a bogus respondent as someone who took the survey multiple times, reported living outside the target country, gave multiple non sequitur open-ended answers, or always said they approve or favour regardless of what was asked. Bogus respondents ran at 7% in the crowdsourced poll, 5% on average across the three opt-in panels, and 1% in the address-recruited panels (Kennedy C, Hatley N, Lau A, Mercer A, Keeter S, Ferno J, Asare-Marfo D. Assessing the Risks to Online Polls From Bogus Respondents. Pew Research Center, February 2020).
Then it tested the two most common quality checks against those cases. The results are worth reading as sensitivity measurements, because that is what they are:
| Check | Share of all respondents flagged | Share of bogus cases caught |
|---|---|---|
| Attention check ("please choose Silver") | 1.4% | 16% |
| Speeding flag (under 3 min vs 7 min median) | 1.5% | 13% |
| Either check | — | 24% |
The report's own summary of the attention check is blunt: among the bogus cases, "most of them passed the attention check (84%)", and "a standard attention check does not work for detecting the large majority of cases found to be giving the type of low quality, biasing data bogus respondents engage in." Combining both flags still leaves it true that "about three-quarters (76%) of bogus cases pass both the attention check and the fast response check."
The sign inversion, in the source's own data
The speeder flag is the cleanest demonstration that these instruments point the wrong way. The flag encodes an assumption — a fake respondent rushes. Pew measured the opposite: in the crowdsourced sample, the bogus respondents had a longer median completion time than everyone else, 701 seconds versus 489. The people gaming the study took 43% more time than the honest ones, because gaming a study is work.
The report notes that this matches other teams' findings, citing work in which "potential trolls and potentially fraudulent IP addresses take significantly longer on the survey on average," and a study in which crowdsourced workers operating through server farms took nearly twice as long as those not using one.
So the speeder flag removes fast honest respondents and retains slow strategic ones. It is not merely insensitive; on this population it is anti-correlated with the thing it was built to detect. And the purge rate reflects it: those flags removed 29% of the bogus cases in the panels but just 7% in the crowdsourced sample — least effective exactly where the problem was worst.
What tightening actually does to your sample
Put both halves together in one model. Take a pool of 2,000 people: 100 genuine targets, 95 strategic respondents (people who will answer whatever qualifies them, at roughly the 5% rate Pew measured for opt-in sources), and 1,805 honest non-targets. Give each criterion 90% sensitivity for genuine targets, 85% specificity against honest non-targets, and a 95% pass rate for strategic respondents, who answer to qualify rather than to report.
| One criterion | Five criteria | |
|---|---|---|
| Genuine targets who pass | 90.0 | 59.0 |
| Honest non-targets who pass | 270.8 | 0.1 |
| Strategic respondents who pass | 90.2 | 73.5 |
| Total passing | 451.0 | 132.7 |
| Positive predictive value | 20.0% | 44.5% |
| Strategic share of the false positives | 25.0% | 99.8% |
| Strategic share of your final sample | 20.0% | 55.4% |
Read the last two rows together. Tightening more than doubled the metric you were optimising. It also changed the composition of your errors completely: at one criterion, three-quarters of the wrong people are honest respondents who simply misjudged a category, and they will still describe their real experience accurately once the interview starts. At five criteria, essentially every remaining wrong person is someone actively constructing answers, and they will keep constructing them for the next 45 minutes.
An honest false positive dilutes your findings. A strategic false positive fabricates them. Tightening trades the first kind for the second, and the summary statistic gets better while the interview transcripts get worse.
What this looks like when it happens to you
The Drexel case study cited in the companion article is exactly this failure. A qualitative team recruiting caregivers of people with dementia and chronic wounds opened snowball referrals and immediately received a flood of volunteers who completed the screening tool cleanly and were scheduled for interviews. The screening tool caught none of them.
What caught them was a domain expert in conversation. The interviewer — a nurse scientist with extensive qualitative experience — began questioning authenticity by the second interview, on evidence no form could have collected: participants across separate interviews gave near-identical answers about asking a pharmacist what to put on a wound, described treatments only as "medicine" or "some drugs" and could not expand, contradicted themselves under probing (no money for care, then a doctor visiting three times a week), and appeared to be shuffling papers to find answers. Two consecutive participants claimed to be caring for teenage brothers with dementia. None of the data collected were used for analysis.
The lesson is not that the team should have written a stricter screener. It is that the discriminating evidence lived in the follow-up questions, and a screener cannot ask follow-up questions.
What to do instead
1. Cut criteria down to the ones that are genuinely disqualifying. Most screeners carry criteria that would be nice to have. Each one costs you roughly 10% of your real population. Keep the two or three where a wrong answer makes the session worthless; move the rest into the interview as questions you can analyse on, or into quotas you fill loosely.
2. Replace self-report criteria with verification wherever a record exists. A criterion checked against account data, product telemetry or a CRM field is not a test with 90% sensitivity; it is a fact. Every criterion you can move from claimed to verified removes a multiplication from the table above and cannot be gamed.
3. Make the second test different in kind, not stricter in degree. Two self-report questions about the same attribute are not independent tests — someone misreporting on one will misreport on the other. An open-ended request for episodic detail ("walk me through the last time you did this, start to finish") is independent, because it demands the kind of specificity that people who have not lived the experience cannot produce. That is precisely the evidence that exposed the fraudulent participants above.
4. Measure your instruments before you trust them. Run your attention check and speeder flag against sessions you have independently classified as good or bad, and compute their sensitivity on your own data. If your flag catches 15% of the bad sessions, it is a rounding error dressed up as a control, and the cases it does catch may be the honest careless ones you would rather keep.
5. Widen the screen and narrow at analysis. When session cost is low, the strongest design is a permissive screen with rigorous post-hoc classification: let borderline candidates through, then tag each completed session as on-target or off-target from the transcript and exclude the off-target ones from your themes. A transcript is orders of magnitude more evidence than a form, so the classification is far more accurate — and unlike a screener rejection, it is reversible and auditable.
How this differs from fraud detection
Survey fraud and respondent quality covers the tactics — what fraud looks like, which warning signs to watch, how to design a fraud-resistant study. This article is about the measured accuracy of those tactics: how many of the bad cases each one actually catches, and what happens to the good cases you lose along the way. Use that article to build your controls, and this one to decide how much to believe them.
It is also distinct from data annotation quality, which uses gold tasks and honeypots to police a workforce doing repeated labelling work. Gold tasks work there because the annotator returns many times and you can seed known answers across their queue. A research participant appears once, so seeded items are a single-shot test with the same sensitivity problem as everything else on this page.
The modern approach: two tests instead of one
The structural fix is to stop asking one instrument to do two jobs. Koji separates them:
- Structured questions carry the disqualifying rules. The six question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — let you capture hard criteria as clean fields rather than free text, so the deterministic part of screening stays deterministic. The structured questions guide covers when each type is the right instrument.
- The AI moderator administers the independent second test. Because Koji's AI interviewer probes follow-ups in real time, every claimed screener answer is immediately tested against episodic detail in the same session. The failure mode that exposed the Drexel impostors — vague answers that collapse under a specific follow-up — is exactly what conversational probing surfaces, and it costs nothing extra because it happens inside the interview you were running anyway.
- Quality scores turn your controls into measurements. Each interview carries a 1-5 quality score, so "how many of my sessions were with the wrong person" stops being a guess. That is the input you need to compute your own flag sensitivity rather than borrowing Pew's.
- Permissive screening becomes affordable. The widen-and-narrow strategy is only realistic when a session costs minutes rather than a recruiter's hour and a calendar slot. Running interviews around the clock with an AI moderator is what makes a permissive screen the cheap option instead of the reckless one.
Traditional survey platforms offer you a longer list of screening rules and a library of trap questions. Both make the numbers in the table above worse, in the specific direction of leaving you alone in a room with the respondents who are best at passing tests.
Frequently asked questions
Should I remove attention checks from my studies entirely?
No, but price them correctly. Pew measured that a standard attention check catches about 16% of bogus cases while flagging 1.4% of everyone, so it is a cheap filter with low sensitivity, not a guarantee. Keep it if it costs you nothing, never treat a passed attention check as evidence of a good respondent, and do not add a second and third trap question expecting the sensitivity to compound — the cases they miss are missed for a reason, namely that the respondent is reading carefully and answering strategically.
How many screening criteria is too many?
The arithmetic says the cost is roughly 10% of your genuine target population per criterion, so three is usually the practical ceiling for a self-report screener and five is where you are rejecting more real participants than most teams would accept if they could see it. The better question is how many criteria are disqualifying rather than merely desirable. If a wrong answer would not make you throw the session away, it does not belong in the screener; make it an interview question and analyse on it.
Does paying a higher incentive make misrepresentation worse?
It increases the return on qualifying, which is the mechanism at work, but low incentives create their own problems by shifting your sample toward people with unusually low opportunity cost. The more effective lever is to reduce what misrepresentation is worth: verify against records where you can, and use conversational probing so that a fabricated eligibility claim does not survive the first five minutes. Research incentive strategies covers the amount question properly.
If tightening raises positive predictive value, why is it wrong?
It is not wrong, it is incomplete. PPV counts wrong participants but does not distinguish between kinds of wrong. An honest non-target gives you accurate reports about the wrong population, which dilutes findings in a way analysis can often detect and correct. A strategic respondent gives you invented reports about a population they do not belong to, which analysis cannot detect and which contaminates every theme they touch. Tightening improves the count while worsening the mix, so a rising PPV is only good news if the composition of your remaining errors is holding steady — and it usually is not.
How do I estimate my own flag sensitivity without a huge sample?
Take 40 completed sessions, classify each one independently of the flags — a researcher reading the transcript, or a verification against account records — then compare that classification to what your flags said. Forty sessions is enough to tell a 15% sensitivity from a 60% one, which is the decision you actually face. Do it once per collection channel rather than once overall, because Pew's most striking result was that the same flags removed 29% of bogus cases in panels and 7% in the crowdsourced sample.
Are these findings from political polling relevant to B2B product research?
The mechanism transfers even though the population does not. The Pew numbers describe consumer opt-in panels, and your incidence, incentives and fraud rates will differ. What transfers is the structural result: checks built on the assumption that bad respondents are careless will miss respondents who are motivated, because motivation produces careful behaviour. In B2B the strategic respondent is often someone who wants the incentive or the vendor relationship badly enough to claim a role they do not hold, and every word of the analysis above applies unchanged.
Related Resources
- Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted — the arithmetic this article responds to
- Survey Fraud and Respondent Quality — the detection tactics whose accuracy is measured here
- The Base Rate Nobody Measured — why every automated flag has an unknown precision
- Structured Questions Guide — the six question types and which ones belong in a screener
- Recall Bias: How Faulty Memory Distorts Research — why honest people fail frequency criteria
- Research Incentive Strategies: What to Pay and How — setting incentives without buying misrepresentation
Related Articles
The Base Rate Nobody Measured: Why Every Flag in Your Research Stack Has an Unknown Precision (2026)
Every AI tag, sentiment label and risk score is a diagnostic test whose precision depends on a prevalence nobody measured. Accuracy rises as precision collapses. How to audit the unflagged pile and publish a precision footer.
Research Incentive Strategies: What to Pay and How
A practical guide to incentivizing research participants — when to offer compensation, how much to pay, and choosing between cash, gift cards, and product access.
Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)
Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.
Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted (2026)
A research screener is a diagnostic test. At a 5% target incidence, a screener with 90% sensitivity and 85% specificity delivers a sample that is 76% wrong. How to compute positive predictive value, measure it on your own studies, and raise it.
Did Users Actually Notice? Sensitivity vs Criterion in Did-You-Notice Questions (2026)
The percentage of users who say they noticed your change is not a measurement of whether they noticed. Signal detection theory separates detection from willingness to say yes.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)
Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.