{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-03T15:27:26.430Z"},"content":[{"type":"documentation","id":"6cb5b73d-fb6f-4574-90ae-8a2cf92fd7be","slug":"implausible-self-report-screening-cutoff","title":"Implausible Answers: Screening Self-Reported Quantities Against an Objective Ceiling (2026)","url":"https://www.koji.so/docs/implausible-self-report-screening-cutoff","summary":"A plausibility screen compares a self-reported quantity against a bound established independently of the report. Validated against doubly labelled water across 22 studies and 429 participants, the standard nutrition cutoff caught about half of under-reporters at near-perfect specificity, improved with stratification rather than threshold loosening, and is reliable for grading population bias but not for individual verdicts.","content":"**Some self-reported numbers cannot be true, and you can show it without calling anyone a liar.** The method is to derive a bound from something you already know independently, compare each reported quantity against that bound, and treat the comparison as a diagnostic test whose operating point you choose deliberately.\n\nNutrition research has been doing exactly this for more than three decades, has validated the screen against an objective physical measurement, and has published how well it performs. That makes it the best available template for product teams holding piles of self-reported usage numbers. The headline finding is not that people misreport. It is that **a plausibility screen grades a dataset well and convicts an individual badly**, and that the way to improve it is to add information about the person rather than to loosen the threshold.\n\n## The logic: a bound you already have\n\nIn 1991, Goldberg and colleagues published a method in the *European Journal of Clinical Nutrition* for identifying diet records that cannot be accurate. The insight is that you never need to know how much someone actually ate. You need a bound implied by something else you know.\n\nIf body weight is stable across the measurement period, energy intake must roughly equal energy expenditure. Expenditure has a floor, basal metabolic rate, which can be measured or predicted from body size. So the ratio of reported intake to basal metabolic rate has a minimum plausible value, and a report below it is not merely surprising, it is physically inconsistent with a person whose weight did not change.\n\nThe paper sets two limits. Goldberg and colleagues put the stricter one at 1.35 times basal metabolic rate when that rate has been measured rather than predicted. The two limits answer different questions. The first \"tests whether reported energy intake measurements can be representative of long-term habitual intake\". The second \"tests whether reported energy intakes are a plausible measure of the food consumed during the actual measurement period\". The second is deliberately more liberal than the first, because it has to absorb the known imprecision of the measurement itself.\n\nThe force of the approach is in how the conclusion is framed: \"Results falling below these limits must be recognized as being incompatible with long-term maintenance of energy balance\". The screen never asks whether the respondent is honest. It asks whether the report is compatible with a constraint established outside the report. That is a far more defensible question, and it is the move worth copying.\n\n## How well does it work? Someone measured that too\n\nA plausibility screen is only as good as its error rates, and in 2000 Black validated this one against doubly labelled water, an objective biochemical measure of total energy expenditure. Twenty-two studies supplied the database, with 429 participants. Classified against the objective measure, under-reporters, acceptable reporters and over-reporters made up \"34, 62 and 4% respectively of all subjects\".\n\nThe performance figures are the valuable part, because they turn a rule of thumb into a diagnostic test with a chosen operating point:\n\n- Using a single cutoff at a physical activity level of 1.55, sensitivity was 0.50 for men and 0.52 for women, with specificity of 1.00 and 0.99.\n- Using a cutoff for an activity level of 1.95, sensitivity rose to 0.76 and 0.85, but specificity fell to 0.87 and 0.78.\n- Assigning participants to low, medium and high activity levels and applying three corresponding cutoffs, \"sensitivity improved to 0.74 and 0.67 without loss of specificity\".\n\nThree lessons sit in those three lines. The conservative setting almost never falsely accuses an accurate reporter, and catches only about half of the misreporters. Loosening the threshold buys sensitivity by spending specificity, which means it starts discarding honest respondents. And **stratifying by a covariate raises sensitivity without paying for it**, which is the only free improvement on offer.\n\n**What the trade-off costs, worked through.** Suppose 500 people answer a question about hours per week spent in your product, and suppose 30 percent of them, 150 people, substantially overstate.\n\nApply the conservative operating point, sensitivity 0.50 and specificity 0.99:\n\n- Correctly flagged: 75 of the 150 overstaters.\n- Wrongly flagged: about 3 or 4 of the 350 accurate reporters.\n- Of roughly 79 flagged responses, about 95 percent are genuinely problematic. But 75 overstaters remain in your data.\n\nNow apply a looser point, sensitivity 0.80 and specificity 0.82:\n\n- Correctly flagged: 120 overstaters.\n- Wrongly flagged: 63 accurate reporters.\n- Of 183 flagged responses, about 66 percent are genuinely problematic.\n\nLoosening caught 45 more real problems and wrongly flagged about 60 more honest respondents - roughly **1.3 accurate respondents sacrificed for each additional misreporter caught**, with precision falling from about 95 percent to about 66 percent. That is the arithmetic behind preferring Black's third option: get information about the person instead.\n\n## Grade the dataset, do not convict the respondent\n\nA companion 2000 paper in the *International Journal of Obesity* is blunt about the limits. On refining the input assumptions, Black reports that \"The effect of these changes is to widen the confidence limits and reduce the sensitivity of the cut-off.\" On what the method is for: \"The Goldberg cut-off can be used to evaluate the mean population bias in reported energy intake\". And on what it is not for: \"Sensitivity for identifying under-reporters at the individual level is limited.\"\n\nThis is the most transferable conclusion in the literature, and it inverts the instinct. Most teams reach for a plausibility screen in order to find the bad rows and delete them. The evidence says the screen is dependable as a measure of how biased the dataset is overall, and undependable as a verdict on any one person.\n\nSo the output of a plausibility screen should be a property of the wave, not a kill list. *This wave shows a 22 percent implausible rate, against 9 percent last quarter* is a usable, honest finding. *These 41 respondents are lying* is not supported by the instrument.\n\nNote also the uncomfortable corollary in the first quotation. Making the uncertainty estimates more honest **widened the confidence limits and reduced sensitivity**. A team that improves its own error estimates will watch its screen catch fewer cases. That is the screen becoming more truthful, not less effective, and it should not be read as a regression.\n\n## The ceilings you already hold\n\nThe method needs an independently established bound. Most product teams already have several and never use them.\n\n| Self-reported quantity | Independent bound you already hold | Implausible when |\n| --- | --- | --- |\n| Hours per week in the product | Session duration from product analytics | Reported hours greatly exceed observed logged-in time |\n| Hours per week on a work task | Contracted working hours | Reported hours across tasks exceed a working week |\n| Times used last month | Days in the period, plan rate limits | Reported count exceeds the maximum the period allows |\n| Team members using it | Seats on the subscription | Reported users exceed provisioned seats |\n| Tenure as a customer | Account creation date | Reported tenure predates the account |\n| Share of work spent in each tool | The other shares the same person reported | The shares sum far above 100 percent |\n\nThe last row is the cheapest and most neglected, because it needs no external data at all. Ask for several shares in one study and check whether they can coexist. The internal-consistency check is free and it catches a different population than any external bound does.\n\nOne caution on the first two rows. Product analytics is a bound, not the truth: a logged-in session is not the same thing as attention, and a user working in a second tool on your output is doing work your telemetry cannot see. Use it to establish what is impossible, not to overwrite what people tell you. The broader version of that argument is in [stated versus revealed preferences](/docs/stated-vs-revealed-preferences).\n\n## Decide the handling rule before you look\n\nA plausibility screen hands you a researcher degree of freedom, because the exclusion rule can be chosen after you have seen which choice moves the headline. Three defensible policies, all of which must be fixed in advance:\n\n- **Report both.** Publish the estimate with and without flagged responses, and let the gap be part of the finding.\n- **Keep and weight.** Retain flagged responses but down-weight them, which avoids throwing away a non-random share of the sample.\n- **Exclude by a stated rule.** Acceptable when the rule, the threshold and the expected flag rate are written down before the data is examined.\n\nChoosing among these after seeing the results is the pattern described in [p-hacking and researcher degrees of freedom](/docs/p-hacking-researcher-degrees-of-freedom), and a plausibility screen is an unusually tempting place to do it.\n\n## How this differs from attention checks and screener fraud\n\nThree different problems get conflated here, and they need different instruments.\n\n- [Attention checks](/docs/attention-check-questions) detect **inattention**: the respondent was not reading. The signal is a wrong answer to an item with an objectively correct response.\n- [Screener tightening and fraud detection](/docs/screener-tightening-false-negatives) address **eligibility**: the respondent should not be in the study, or misrepresented themselves to get in.\n- Plausibility screening addresses **magnitude**: a genuine, attentive, eligible respondent has given a number that cannot be right.\n\nThe third is usually not deception at all. It is estimation error, of the kind produced by asking someone to quantify something they have never counted, which is why exclusion is the wrong default and grading is the right one. Where the misreporting does have a motivated direction, the mechanism is normally [social desirability bias](/docs/social-desirability-bias), and the fix there is question design rather than screening.\n\n## How Koji handles this\n\n- **A quality score on every interview, with a reasoned breakdown.** Koji scores each interview from 1 to 5 and records a rationale alongside component scores for relevance, depth and coverage. That is dataset-level grading in exactly the form the nutrition literature says a plausibility screen should take.\n- **Extraction confidence is stored, not assumed.** Each structured answer Koji extracts carries a confidence level of high, medium or low, so a questionable quantity is marked as questionable rather than silently promoted to a clean number.\n- **Six structured question types make the internal-consistency check easy.** Koji ships six [structured question types](/docs/structured-questions-guide): open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Asking several shares as scale questions in one study gives you the free sum-to-100 check with no external data.\n- **The number and the account of it are stored together.** Koji keeps a structured value and a qualitative answer for each question, so an implausible figure arrives with the participant's own description of how they arrived at it.\n- **Traceability to the source messages.** Every extracted answer records which messages it came from, so a flagged quantity can be checked against the transcript instead of argued about.\n- **The decisive advantage: you can ask in the moment.** A survey can only record an impossible number. Koji's AI interviewer can probe it while the participant is still there, which frequently resolves the implausibility into an ordinary misunderstanding of the question. That converts a discarded row into usable data.\n\nA legacy survey tool like SurveyMonkey gives you the implausible answer and no way to interrogate it. An AI-native platform can ask the follow-up that tells you whether the number was a lie, a guess, or a misread question.\n\n## Frequently asked questions\n\n### What is a plausibility screen for survey data?\n\nIt is a check that compares a self-reported quantity against a bound established independently of the report, and flags answers that are incompatible with that bound. It never asks whether a respondent is honest, only whether the answer can coexist with something else you already know, which makes it far more defensible than a judgement about truthfulness.\n\n### How accurate are these screens?\n\nValidated against an objective measure of energy expenditure across 22 studies and 429 participants, the standard nutrition cutoff caught about half of misreporters at its conservative setting while almost never flagging an accurate reporter. Loosening the threshold raised sensitivity but began discarding honest respondents, and stratifying by activity level improved sensitivity without that cost.\n\n### Should I delete implausible responses?\n\nUsually not. The evidence is that such screens measure population-level bias well but have limited sensitivity at the individual level, so they support a statement about your dataset rather than a verdict on a person. Report the implausible rate, or keep and down-weight flagged responses, and fix the exclusion rule before you look at results.\n\n### What bound should I use for product data?\n\nUse something you already measure independently: logged-in time from product analytics, days available in the reference period, seats on the subscription, or the account creation date. The cheapest check needs no external data at all, which is asking for several shares of work in one study and testing whether they sum to a possible total.\n\n### Is this the same as an attention check?\n\nNo. An attention check detects a respondent who was not reading, using an item with an objectively correct answer. Plausibility screening addresses a genuine, attentive, eligible respondent whose reported magnitude cannot be right, which is usually estimation error rather than inattention or deception.\n\n### Why did my screen get weaker after I improved it?\n\nBecause more honest uncertainty estimates widen the confidence limits and reduce the sensitivity of the cutoff. A screen that assumes its inputs are more precise than they are will appear to catch more. Losing apparent power after a more careful calibration means the screen has become more truthful, and it is not a regression.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and the free internal-consistency check\n- [Digit Heaping](/docs/self-reported-number-heaping-rounding) - why the quantities you are screening are rounded in the first place\n- [Attention Check Questions](/docs/attention-check-questions) - detecting inattention, a different problem with a different instrument\n- [Screener Tightening](/docs/screener-tightening-false-negatives) - eligibility and fraud, the third member of the set\n- [Social Desirability Bias](/docs/social-desirability-bias) - the mechanism behind directional misreporting\n- [Total Survey Error](/docs/total-survey-error-research-budget) - where reporting bias sits in the full error budget\n","category":"Analysis & Synthesis","lastModified":"2026-10-02T03:33:42.434984+00:00","metaTitle":"Implausible Answers: Screening Self-Reported Quantities","metaDescription":"Some self-reported numbers cannot be true. How to build a plausibility cutoff, what its real error rates are, and why it grades data not people.","keywords":["implausible survey answers","plausibility screening","goldberg cutoff","self-report validation","data quality screening","misreporting detection","survey plausibility check"],"aiSummary":"A plausibility screen compares a self-reported quantity against a bound established independently of the report. Validated against doubly labelled water across 22 studies and 429 participants, the standard nutrition cutoff caught about half of under-reporters at near-perfect specificity, improved with stratification rather than threshold loosening, and is reliable for grading population bias but not for individual verdicts.","aiPrerequisites":["Basic familiarity with survey data quality","Access to one independently measured bound such as product analytics"],"aiLearningOutcomes":["Derive a plausibility bound from data you already hold","Read a plausibility screen as a diagnostic test with an operating point","Choose a handling rule before examining results","Distinguish magnitude implausibility from inattention and from fraud"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}