{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-25T15:58:15.854Z"},"content":[{"type":"documentation","id":"c9c804cf-2374-45df-99b7-83840a0e7969","slug":"screener-tightening-false-negatives","title":"Tightening Your Screener Makes the Sample Purer and the Findings Worse (2026)","url":"https://www.koji.so/docs/screener-tightening-false-negatives","summary":"Screening criteria are conjunctive, so five criteria at 90% sensitivity each keep only 59% of genuine targets, and the 41% lost are systematically the people who describe their own behaviour in different words. Pew Research Center measured that attention checks catch 16% of bogus respondents and speeder flags 13%, with 76% passing both, and that bogus crowdsourced respondents were slower than honest ones (701 vs 489 seconds). Tightening raises positive predictive value from 20% to 45% while raising the strategic-respondent share of the final sample from 20% to 55%.","content":"When you discover that most of the people passing your screener are not your target, the reflex is to tighten it: add criteria, add attention checks, add a speeder flag. That reflex improves the metric you were worried about and makes your sample worse. Tightening deletes a large, non-random share of your genuine participants, and it barely touches the respondents who are misrepresenting themselves on purpose — because those two groups fail screening tests for opposite reasons. In a worked model below, five criteria raise positive predictive value from 20% to 45% while the share of your final sessions held with someone actively deceiving you rises from 20% to 55%.\n\nThis is the second half of the screening problem. The first half is in [screener accuracy](/docs/screener-accuracy-positive-predictive-value); this article is about what happens when you act on it.\n\n## Criteria multiply, and they multiply against you\n\nScreening criteria are conjunctive: a candidate has to pass all of them. If each criterion admits a genuine target 90% of the time — a generous assumption, since real people misremember frequencies, use different words for their own role, and round their own usage — then joint sensitivity is the product.\n\n| Criteria | Genuine targets who pass all | Genuine targets lost |\n|---|---|---|\n| 1 | 90.0% | 10.0% |\n| 3 | 72.9% | 27.1% |\n| 5 | 59.0% | 41.0% |\n| 8 | 43.0% | 57.0% |\n\nA five-criterion screener, each part of which you would describe as accurate, rejects **41% of the population it was designed to find.** An eight-criterion screener rejects more of your target population than it admits. Nobody experiences this as a failure, because rejected candidates leave no trace: they see a \"thanks, you're not a fit for this study\" message and never appear in any report you read.\n\n## The false negatives are not a random sample of your targets\n\nIf the 41% were a random subset, the cost would be pure recruiting inefficiency and nothing more. They are not random. Screening questions test a specific and unevenly distributed skill: the ability to recognise your own behaviour in someone else's categories, quickly, in a form.\n\nThe people who fail that test while genuinely belonging to your population are systematically:\n\n- **The ones who use different vocabulary for their own job.** A \"revenue operations\" screener misses the person whose title is \"sales systems lead\" and who does exactly the job you are studying.\n- **The ones who are honest about uncertainty.** Asked how many times last month they used a feature, a careful person picks the conservative bucket. The frequency thresholds in most screeners punish that.\n- **The ones whose usage is bursty rather than regular.** \"Weekly or more\" excludes the person who uses your product intensely for three days per release cycle, who may be your most interesting participant.\n- **The ones for whom the behaviour is unremarkable.** Deeply habituated users under-report their own frequency, because habits are hard to count from memory. See [recall bias](/docs/recall-bias) for the underlying mechanism.\n\nEvery one of those exclusions is correlated with something you are trying to learn about. Tightening does not shave noise off the edges of your sample; it selects for participants who are fluent at answering screeners about themselves in your terminology. That is a real trait, and it is not the trait you were sampling on.\n\n## Meanwhile the strategic respondents sail through\n\nThe other half of the inversion is the more uncomfortable one. The standard tightening instruments are close to useless against respondents who are deliberately misrepresenting themselves, and there is a large, well-measured study that says so.\n\nPew Research Center ran a study across six online sources — three opt-in sample sources, one crowdsourcing platform, and two panels recruited offline by address-based random sampling — comprising more than 60,000 interviews with at least 10,000 from each source. It defined a bogus respondent as someone who took the survey multiple times, reported living outside the target country, gave multiple non sequitur open-ended answers, or always said they approve or favour regardless of what was asked. Bogus respondents ran at **7% in the crowdsourced poll, 5% on average across the three opt-in panels, and 1% in the address-recruited panels** (Kennedy C, Hatley N, Lau A, Mercer A, Keeter S, Ferno J, Asare-Marfo D. *Assessing the Risks to Online Polls From Bogus Respondents.* Pew Research Center, February 2020).\n\nThen it tested the two most common quality checks against those cases. The results are worth reading as sensitivity measurements, because that is what they are:\n\n| Check | Share of all respondents flagged | Share of bogus cases caught |\n|---|---|---|\n| Attention check (\"please choose Silver\") | 1.4% | 16% |\n| Speeding flag (under 3 min vs 7 min median) | 1.5% | 13% |\n| Either check | — | 24% |\n\nThe report's own summary of the attention check is blunt: among the bogus cases, \"most of them passed the attention check (84%)\", and \"a standard attention check does not work for detecting the large majority of cases found to be giving the type of low quality, biasing data bogus respondents engage in.\" Combining both flags still leaves it true that \"about three-quarters (76%) of bogus cases pass both the attention check and the fast response check.\"\n\n## The sign inversion, in the source's own data\n\nThe speeder flag is the cleanest demonstration that these instruments point the wrong way. The flag encodes an assumption — a fake respondent rushes. Pew measured the opposite: in the crowdsourced sample, the bogus respondents had a **longer** median completion time than everyone else, 701 seconds versus 489. The people gaming the study took 43% more time than the honest ones, because gaming a study is work.\n\nThe report notes that this matches other teams' findings, citing work in which \"potential trolls and potentially fraudulent IP addresses take significantly longer on the survey on average,\" and a study in which crowdsourced workers operating through server farms took nearly twice as long as those not using one.\n\nSo the speeder flag removes fast honest respondents and retains slow strategic ones. It is not merely insensitive; on this population it is anti-correlated with the thing it was built to detect. And the purge rate reflects it: those flags removed 29% of the bogus cases in the panels but just 7% in the crowdsourced sample — least effective exactly where the problem was worst.\n\n## What tightening actually does to your sample\n\nPut both halves together in one model. Take a pool of 2,000 people: 100 genuine targets, 95 strategic respondents (people who will answer whatever qualifies them, at roughly the 5% rate Pew measured for opt-in sources), and 1,805 honest non-targets. Give each criterion 90% sensitivity for genuine targets, 85% specificity against honest non-targets, and a 95% pass rate for strategic respondents, who answer to qualify rather than to report.\n\n| | One criterion | Five criteria |\n|---|---|---|\n| Genuine targets who pass | 90.0 | 59.0 |\n| Honest non-targets who pass | 270.8 | 0.1 |\n| Strategic respondents who pass | 90.2 | 73.5 |\n| Total passing | 451.0 | 132.7 |\n| **Positive predictive value** | **20.0%** | **44.5%** |\n| Strategic share of the false positives | 25.0% | 99.8% |\n| **Strategic share of your final sample** | **20.0%** | **55.4%** |\n\nRead the last two rows together. Tightening more than doubled the metric you were optimising. It also changed the *composition* of your errors completely: at one criterion, three-quarters of the wrong people are honest respondents who simply misjudged a category, and they will still describe their real experience accurately once the interview starts. At five criteria, essentially every remaining wrong person is someone actively constructing answers, and they will keep constructing them for the next 45 minutes.\n\nAn honest false positive dilutes your findings. A strategic false positive fabricates them. Tightening trades the first kind for the second, and the summary statistic gets better while the interview transcripts get worse.\n\n## What this looks like when it happens to you\n\nThe Drexel case study cited in the companion article is exactly this failure. A qualitative team recruiting caregivers of people with dementia and chronic wounds opened snowball referrals and immediately received a flood of volunteers who completed the screening tool cleanly and were scheduled for interviews. The screening tool caught none of them.\n\nWhat caught them was a domain expert in conversation. The interviewer — a nurse scientist with extensive qualitative experience — began questioning authenticity by the second interview, on evidence no form could have collected: participants across separate interviews gave near-identical answers about asking a pharmacist what to put on a wound, described treatments only as \"medicine\" or \"some drugs\" and could not expand, contradicted themselves under probing (no money for care, then a doctor visiting three times a week), and appeared to be shuffling papers to find answers. Two consecutive participants claimed to be caring for teenage brothers with dementia. None of the data collected were used for analysis.\n\nThe lesson is not that the team should have written a stricter screener. It is that the discriminating evidence lived in the follow-up questions, and a screener cannot ask follow-up questions.\n\n## What to do instead\n\n**1. Cut criteria down to the ones that are genuinely disqualifying.** Most screeners carry criteria that would be nice to have. Each one costs you roughly 10% of your real population. Keep the two or three where a wrong answer makes the session worthless; move the rest into the interview as questions you can analyse on, or into quotas you fill loosely.\n\n**2. Replace self-report criteria with verification wherever a record exists.** A criterion checked against account data, product telemetry or a CRM field is not a test with 90% sensitivity; it is a fact. Every criterion you can move from claimed to verified removes a multiplication from the table above and cannot be gamed.\n\n**3. Make the second test different in kind, not stricter in degree.** Two self-report questions about the same attribute are not independent tests — someone misreporting on one will misreport on the other. An open-ended request for episodic detail (\"walk me through the last time you did this, start to finish\") is independent, because it demands the kind of specificity that people who have not lived the experience cannot produce. That is precisely the evidence that exposed the fraudulent participants above.\n\n**4. Measure your instruments before you trust them.** Run your attention check and speeder flag against sessions you have independently classified as good or bad, and compute their sensitivity on your own data. If your flag catches 15% of the bad sessions, it is a rounding error dressed up as a control, and the cases it does catch may be the honest careless ones you would rather keep.\n\n**5. Widen the screen and narrow at analysis.** When session cost is low, the strongest design is a permissive screen with rigorous post-hoc classification: let borderline candidates through, then tag each completed session as on-target or off-target from the transcript and exclude the off-target ones from your themes. A transcript is orders of magnitude more evidence than a form, so the classification is far more accurate — and unlike a screener rejection, it is reversible and auditable.\n\n## How this differs from fraud detection\n\n[Survey fraud and respondent quality](/docs/survey-fraud-respondent-quality) covers the tactics — what fraud looks like, which warning signs to watch, how to design a fraud-resistant study. This article is about the measured accuracy of those tactics: how many of the bad cases each one actually catches, and what happens to the good cases you lose along the way. Use that article to build your controls, and this one to decide how much to believe them.\n\nIt is also distinct from [data annotation quality](/docs/data-annotation-quality-guide), which uses gold tasks and honeypots to police a workforce doing repeated labelling work. Gold tasks work there because the annotator returns many times and you can seed known answers across their queue. A research participant appears once, so seeded items are a single-shot test with the same sensitivity problem as everything else on this page.\n\n## The modern approach: two tests instead of one\n\nThe structural fix is to stop asking one instrument to do two jobs. Koji separates them:\n\n- **Structured questions carry the disqualifying rules.** The six question types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — let you capture hard criteria as clean fields rather than free text, so the deterministic part of screening stays deterministic. The [structured questions guide](/docs/structured-questions-guide) covers when each type is the right instrument.\n- **The AI moderator administers the independent second test.** Because Koji's AI interviewer probes follow-ups in real time, every claimed screener answer is immediately tested against episodic detail in the same session. The failure mode that exposed the Drexel impostors — vague answers that collapse under a specific follow-up — is exactly what conversational probing surfaces, and it costs nothing extra because it happens inside the interview you were running anyway.\n- **Quality scores turn your controls into measurements.** Each interview carries a 1-5 quality score, so \"how many of my sessions were with the wrong person\" stops being a guess. That is the input you need to compute your own flag sensitivity rather than borrowing Pew's.\n- **Permissive screening becomes affordable.** The widen-and-narrow strategy is only realistic when a session costs minutes rather than a recruiter's hour and a calendar slot. Running interviews around the clock with an AI moderator is what makes a permissive screen the cheap option instead of the reckless one.\n\nTraditional survey platforms offer you a longer list of screening rules and a library of trap questions. Both make the numbers in the table above worse, in the specific direction of leaving you alone in a room with the respondents who are best at passing tests.\n\n## Frequently asked questions\n\n### Should I remove attention checks from my studies entirely?\n\nNo, but price them correctly. Pew measured that a standard attention check catches about 16% of bogus cases while flagging 1.4% of everyone, so it is a cheap filter with low sensitivity, not a guarantee. Keep it if it costs you nothing, never treat a passed attention check as evidence of a good respondent, and do not add a second and third trap question expecting the sensitivity to compound — the cases they miss are missed for a reason, namely that the respondent is reading carefully and answering strategically.\n\n### How many screening criteria is too many?\n\nThe arithmetic says the cost is roughly 10% of your genuine target population per criterion, so three is usually the practical ceiling for a self-report screener and five is where you are rejecting more real participants than most teams would accept if they could see it. The better question is how many criteria are *disqualifying* rather than merely desirable. If a wrong answer would not make you throw the session away, it does not belong in the screener; make it an interview question and analyse on it.\n\n### Does paying a higher incentive make misrepresentation worse?\n\nIt increases the return on qualifying, which is the mechanism at work, but low incentives create their own problems by shifting your sample toward people with unusually low opportunity cost. The more effective lever is to reduce what misrepresentation is worth: verify against records where you can, and use conversational probing so that a fabricated eligibility claim does not survive the first five minutes. [Research incentive strategies](/docs/incentive-strategies) covers the amount question properly.\n\n### If tightening raises positive predictive value, why is it wrong?\n\nIt is not wrong, it is incomplete. PPV counts wrong participants but does not distinguish between kinds of wrong. An honest non-target gives you accurate reports about the wrong population, which dilutes findings in a way analysis can often detect and correct. A strategic respondent gives you invented reports about a population they do not belong to, which analysis cannot detect and which contaminates every theme they touch. Tightening improves the count while worsening the mix, so a rising PPV is only good news if the composition of your remaining errors is holding steady — and it usually is not.\n\n### How do I estimate my own flag sensitivity without a huge sample?\n\nTake 40 completed sessions, classify each one independently of the flags — a researcher reading the transcript, or a verification against account records — then compare that classification to what your flags said. Forty sessions is enough to tell a 15% sensitivity from a 60% one, which is the decision you actually face. Do it once per collection channel rather than once overall, because Pew's most striking result was that the same flags removed 29% of bogus cases in panels and 7% in the crowdsourced sample.\n\n### Are these findings from political polling relevant to B2B product research?\n\nThe mechanism transfers even though the population does not. The Pew numbers describe consumer opt-in panels, and your incidence, incentives and fraud rates will differ. What transfers is the structural result: checks built on the assumption that bad respondents are careless will miss respondents who are motivated, because motivation produces careful behaviour. In B2B the strategic respondent is often someone who wants the incentive or the vendor relationship badly enough to claim a role they do not hold, and every word of the analysis above applies unchanged.\n\n## Related Resources\n\n- [Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted](/docs/screener-accuracy-positive-predictive-value) — the arithmetic this article responds to\n- [Survey Fraud and Respondent Quality](/docs/survey-fraud-respondent-quality) — the detection tactics whose accuracy is measured here\n- [The Base Rate Nobody Measured](/docs/classifier-precision-base-rate-research) — why every automated flag has an unknown precision\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and which ones belong in a screener\n- [Recall Bias: How Faulty Memory Distorts Research](/docs/recall-bias) — why honest people fail frequency criteria\n- [Research Incentive Strategies: What to Pay and How](/docs/incentive-strategies) — setting incentives without buying misrepresentation\n","category":"Research Operations","lastModified":"2026-08-23T03:28:48.605292+00:00","metaTitle":"Screener False Negatives: Why Tightening Backfires (2026)","metaDescription":"Five screening criteria reject 41% of your real target population, while attention checks catch just 16% of bad-faith respondents. What tightening really does to your sample.","keywords":["screener false negatives","tightening screening criteria","attention check sensitivity","professional respondents","conjunctive screening criteria","screener over-qualification","research participant fraud"],"aiSummary":"Screening criteria are conjunctive, so five criteria at 90% sensitivity each keep only 59% of genuine targets, and the 41% lost are systematically the people who describe their own behaviour in different words. Pew Research Center measured that attention checks catch 16% of bogus respondents and speeder flags 13%, with 76% passing both, and that bogus crowdsourced respondents were slower than honest ones (701 vs 489 seconds). Tightening raises positive predictive value from 20% to 45% while raising the strategic-respondent share of the final sample from 20% to 55%.","aiPrerequisites":["Familiarity with screener positive predictive value","A screener with more than one criterion"],"aiLearningOutcomes":["Compute joint sensitivity across conjunctive screening criteria","Recognise which genuine participants a tightened screener systematically excludes","Measure the sensitivity of your own attention and speeder checks","Design a permissive screen with rigorous post-session classification"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}