{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-09T13:21:42.457Z"},"content":[{"type":"documentation","id":"ca449eb4-8e66-4e71-b3f2-f15bc78c0052","slug":"publication-bias-product-research","title":"Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)","url":"https://www.koji.so/docs/publication-bias-product-research","summary":"Publication bias is the over-representation of positive findings in a body of evidence, caused by which studies get written up rather than by bad data. Franco, Malhotra and Simonovits (2014) observed a full population of 221 experiments and found 62 percent of strong results published against 21 percent of nulls, with the gap created almost entirely by authors never writing nulls up. Turner et al. (2008) showed the resulting effect-size inflation is about 32 percent. Product research is more exposed than academia because small evidence bases sit inside the fragile range of Rosenthal fail-safe arithmetic. The fix is a study register plus mandatory closed-ended outcome recording.","content":"Publication bias is the systematic over-representation of positive findings in a body of evidence, caused not by bad data but by which studies get written up and circulated. The best direct measurement of it comes from Franco, Malhotra and Simonovits, who tracked a known population of 221 completed social science experiments and found that 62 percent of strong results were published against 21 percent of null results - and, decisively, that the gap was created almost entirely by researchers never writing the null studies up, not by journals rejecting them. Product research is more exposed to this failure than academia is, not less, because a commercial evidence base is small enough that a handful of unreported nulls can reverse its conclusion, and because there is no journal to blame.\n\n**Key takeaways**\n\n- Of 48 null results in a fully observed population of experiments, only 10 were published and 31 were never even written up (Franco, Malhotra and Simonovits, *Science*, 2014).\n- Among studies that were actually written up, null and strong results were published at almost the same rate. The file drawer is filled by authors, not gatekeepers.\n- Selective publication inflates apparent effect sizes by roughly a third: Turner and colleagues found published antidepressant trials showed a mean effect of 0.37 against 0.15 for unpublished ones, a 32 percent overall inflation (*New England Journal of Medicine*, 2008).\n- Rosenthal showed that 15 studies averaging a modest effect can be reversed by just 6 filed-away nulls. Most product research programmes are permanently inside that fragile range.\n- The fix is procedural, not analytical: register the study before it is fielded, and record its outcome whether or not anyone builds a deck.\n\n## What publication bias actually is\n\nThe textbook definition is that the published literature is a biased sample of the research conducted. Robert Rosenthal named the mechanism in 1979 and put it in deliberately extreme terms: journals are filled with the 5 percent of studies that show Type I errors, while the file drawers back at the lab hold the 95 percent that came out non-significant (\"The File Drawer Problem and Tolerance for Null Results\", *Psychological Bulletin*, 1979, 86(3):638-641).\n\nThat framing has kept the problem filed under \"things that happen to journals\". It is the wrong file. Strip out the word \"publication\" and what remains is a general property of any evidence base assembled by people who choose what to escalate: **the record over-represents the outcomes that were rewarding to report.** A product organisation has escalation, rewards, and choice. It therefore has publication bias, with none of the corrective machinery that science eventually built.\n\nThe distinction worth holding onto is that publication bias is not a data-quality problem. Every individual study in the record can be perfectly executed, perfectly analysed, and perfectly honest. The distortion is created entirely by which of them exist in a form anyone can find.\n\n## How big the file drawer is when you can actually see it\n\nMost estimates of publication bias are inferred from the shape of the published record, which is circular. Franco, Malhotra and Simonovits found a way around that. Time-sharing Experiments in the Social Sciences (TESS) runs peer-reviewed experiments on nationally representative samples, which means the funder knows the full population of studies conducted - including the ones that vanished. Their analysis covered a final sample of 221 studies.\n\n| Result strength | Never written up | Written but unpublished | Published | Published share |\n|---|---|---|---|---|\n| Null | 31 | 7 | 10 | 21% |\n| Mixed | 10 | 32 | 40 | 49% |\n| Strong | 4 | 31 | 56 | 62% |\n| Total | 45 | 70 | 106 | 48% |\n\nSource: Franco, Malhotra and Simonovits, \"Publication Bias in the Social Sciences: Unlocking the File Drawer\", *Science*, 2014, 345(6203):1502-1505.\n\nRead the first column before the last one. That is where the finding is.\n\nSixty-five percent of the null studies were never written up at all. Only 4 percent of the strong studies met the same fate. And among the studies that *were* written up, the publication rates converge almost completely: 10 of 17 written-up nulls were published (59 percent) against 56 of 87 written-up strong results (64 percent). A five-point gap at the journal stage; a sixty-one-point gap at the writing stage.\n\n**The gatekeepers were not the problem.** The authors were. When Franco and colleagues followed up with researchers about the abandoned studies, the reasons were mundane: the result was not interesting enough to be worth the effort of writing, or a co-author lost enthusiasm, or the project was overtaken by something more promising. Nobody suppressed anything. Everyone made a locally reasonable decision about where to spend the next week.\n\nThat is exactly the decision a product manager makes about whether to build the readout deck for a study that came back flat.\n\n## What it does to the numbers, not just the catalogue\n\nA missing study is not merely an absence. It changes the value of everything that remains.\n\nTurner and colleagues obtained FDA reviews for all 74 registered trials of 12 antidepressants, covering 12,564 patients, and compared them against the published literature (*New England Journal of Medicine*, 2008, 358:252-260). The FDA judged 38 of the 74 trials (51 percent) positive. In the published literature, 48 of 51 published studies read as positive - 94 percent. The two confidence intervals did not overlap.\n\nThe mechanism was double-sided. Of the 36 trials the FDA did not judge positive, 22 were never published and 11 were published in a form that read as positive. Data from 3,449 patients (27 percent of the total) never appeared at all.\n\nThen the part that should worry anyone who quotes an effect size:\n\n| Evidence base | Mean weighted effect size (Hedges g) | 95% CI |\n|---|---|---|\n| Published studies only | 0.37 | 0.33 to 0.41 |\n| Unpublished studies | 0.15 | 0.08 to 0.22 |\n| All FDA studies pooled | 0.31 | - |\n\nMeta-analysing the published record instead of the complete record inflated the effect by 32 percent overall, and by between 11 and 69 percent depending on the drug. The published literature was not wrong about direction. It was systematically wrong about magnitude, in one direction, by about a third.\n\nTranslate that to a product context. Your internal evidence base says the onboarding redesign pattern works, that the pricing page copy change moves conversion, that the concierge model beats self-serve for a particular segment. Each of those claims is the survivor of a selection process that discarded the flat results. Every effect size in your institutional memory is an overestimate, and the ones you trust most are the ones that survived the most filtering.\n\n## Why product teams are more exposed than researchers are\n\nThe standard reassurance about publication bias is that meta-analysis eventually washes it out. There is real substance to that: when Head and colleagues text-mined p-values across scientific disciplines, they concluded that p-hacking probably does not drastically alter consensus drawn from meta-analyses, because the distorted studies tend to be small and get down-weighted when pooled (*PLoS Biology*, 2015). Our guide to [p-hacking and researcher degrees of freedom](/docs/p-hacking-researcher-degrees-of-freedom) covers that counter-evidence in detail.\n\nThe protection depends entirely on having many independent studies of the same question. Rosenthal quantified how many.\n\nHis fail-safe number asks: how many filed-away null studies would have to exist to reduce a combined significant result to non-significance? For large literatures the answer is reassuring - he calculated that 311 studies averaging a modest effect would require nearly 50,000 hidden nulls to overturn, which is absurd on its face. But he also spelled out the sobering half, and this is the half that applies to you:\n\n- **15 studies** averaging a small effect combine to p = .026. Just **6** filed-away nulls flip that to non-significant.\n- **2 studies** averaging a strong effect combine to p = .002. Just **4** new nulls bring it into the non-significant range.\n\nRosenthal proposed that a body of evidence be considered resistant to the file drawer problem only when the fail-safe number exceeds 5k + 10, where k is the number of known studies.\n\nNow count your own corpus. How many studies has your team run on activation? On willingness to pay for the enterprise tier? On why trial users churn in week two? For most teams the honest answer is between two and fifteen, which is precisely the regime where a small number of quiet nulls can reverse the conclusion. **Commercial research does not have the aggregation that protects academic consensus. It has the fragile end of Rosenthal's arithmetic and a roadmap riding on it.**\n\nThere is a second asymmetry. In science, a false conclusion built on a biased evidence base is eventually corrected by a failed replication, embarrassing but survivable. In product work, nobody runs the replication. The correction arrives eighteen months later disguised as a feature with poor adoption and no obvious cause.\n\n## The five filters between a study and a decision\n\nPublication bias in a company is not one event. It is a funnel, and each stage is a place where a null result is quietly more likely to be lost than a positive one.\n\n1. **The fielding filter.** Studies that seem unlikely to produce a clear answer get deprioritised before they run. This is invisible in every audit, because there is no artifact.\n2. **The write-up filter.** The study ran, the data exists in a tool somewhere, and no one built the deck. This is Franco's 65 percent, and in most organisations it is by far the largest leak.\n3. **The escalation filter.** The readout exists but never leaves the research channel, because \"we found no difference\" is a hard slide to open a leadership review with.\n4. **The narrative filter.** The study is presented, but the flat primary outcome is subordinated to an interesting secondary finding. Turner found exactly this in 11 of the antidepressant publications, where a non-significant prespecified primary outcome was either demoted or omitted and a positive result highlighted in its place.\n5. **The memory filter.** Six months on, the finding that gets cited in the next planning document is the one that made a good story. The null is technically in the repository and functionally gone.\n\nFilters 2 and 5 do the most damage and receive the least attention, because neither produces a decision anyone can point to. Nothing was suppressed. The evidence base just quietly curated itself.\n\n## How to close the file drawer\n\nThe scientific community did not solve publication bias with better statistics. It solved it, to the extent it has, with registration: declare the study before you run it, and the record of its existence outlives your enthusiasm for its result. The same move works internally and is considerably cheaper.\n\n**1. Keep a study register, not just a repository.** A repository holds outputs. A register holds intentions. Every study gets an entry the day it is commissioned, with the question, the decision it feeds, the primary outcome, and the planned sample - and that entry persists whether or not anything is ever written up. Our guide to [building a research repository](/docs/research-repository-guide) covers the storage layer; the register is the row that gets created before there is anything to store.\n\n**2. Make the outcome field mandatory and closed-ended.** Every registered study resolves to one of: supported, not supported, inconclusive, or abandoned before completion. Free-text summaries let a null quietly become \"directionally encouraging\".\n\n**3. Report the denominator in every synthesis.** When someone presents \"our research shows X\", the honest form is \"four of the six studies we ran on this question showed X\". If the denominator is not available, the claim has not been checked for the file drawer. Our guide to [evidence synthesis](/docs/evidence-synthesis-research-findings) sets out how to pool a body of internal studies without inheriting this bias.\n\n**4. Separate \"no effect\" from \"could not tell\".** A large share of internal nulls are not evidence of absence at all; they are studies that were too small to detect anything. That distinction is the entire subject of [equivalence testing](/docs/equivalence-testing-no-difference), and it matters here because a register that files uninformative nulls alongside genuine nulls is not much better than no register at all.\n\n**5. Change who benefits from a null.** This is the only structural fix. As long as a flat result represents six weeks of wasted budget, someone will find something in it. Our guide to [statistical power and minimum detectable effect](/docs/statistical-power-minimum-detectable-effect) explains how to make a null informative in advance, which is what turns it into a publishable internal result.\n\n## How Koji closes the write-up filter automatically\n\nFranco and colleagues located the leak precisely: the studies disappeared at the point where someone had to sit down and write them up. That is a striking finding, because it is not a moral failure or a statistical one. It is a workload problem, and workload problems are the kind that tooling actually fixes.\n\n**The readout is generated, not authored.** With Koji, every study produces a structured report as responses arrive - themes, quotes, per-question breakdowns, quality scores - without anyone deciding it is worth the effort. The artifact whose absence *is* the file drawer exists by default. A flat result is one click from being a permanent, searchable record instead of a folder nobody opened. Traditional tooling makes the opposite trade: SurveyMonkey or a transcript pile gives you raw material that still requires days of human synthesis, and that synthesis cost is exactly the filter that null results fail.\n\n**Consistent instruments make nulls comparable.** Koji studies are built from six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - documented in our [structured questions guide](/docs/structured-questions-guide). Because the same question type produces the same data shape every time, a null in study three is directly comparable with a positive in study seven. Ad hoc instruments make nulls unpoolable, and unpoolable findings are the first to be forgotten.\n\n**Cheap studies make nulls survivable.** A null result is only embarrassing in proportion to what it cost. When an AI-moderated study of 200 people can be fielded and analysed in days rather than six weeks, \"we checked and there was nothing there\" becomes an ordinary, useful sentence rather than a career problem. Teams using AI-assisted research report 60 to 80 percent faster time-to-insight than equivalent manual studies, and the second-order effect of that speed is the one that matters here: it removes the economic pressure that fills the file drawer.\n\n**Registration is the default state.** Because a Koji study exists as a brief before it is fielded, with its questions and target sample declared, the register entry is a by-product of setting the study up rather than an administrative task somebody has to remember. You do not have to build the discipline. You have to not delete the row.\n\n## Honest objections\n\n**\"Our nulls are not hidden, they are in Slack.\"** A finding that exists only in a thread is functionally filed. The test is not whether the information was ever transmitted; it is whether someone assembling the evidence base on that question a year from now will find it. Applying the denominator rule from the previous section usually settles this question quickly and uncomfortably.\n\n**\"Most of our nulls really are uninformative.\"** Often true, and it is a reason to fix the power problem rather than a reason to discard them. An underpowered null is genuinely close to worthless in isolation, but it is not worthless in aggregate: three underpowered nulls on the same question are meaningful evidence, and they can only be pooled if all three were recorded.\n\n**\"This is academic hygiene applied to a commercial setting.\"** The commercial setting is where the stakes for this particular failure are highest. Science tolerates publication bias because it has volume, replication and meta-analysis to absorb it. A product team has one study per question and Rosenthal's fragile arithmetic. The hygiene is not borrowed - it is more necessary here.\n\n## Frequently asked questions\n\n### What is the difference between publication bias and p-hacking?\nThey are different failures at different stages. P-hacking distorts a single study, by choosing analytic options after seeing the data until something crosses the threshold. Publication bias distorts a body of studies, by determining which completed studies enter the record at all. A perfectly pre-registered, perfectly analysed study still contributes to publication bias if its null result is never written up. See [p-hacking and researcher degrees of freedom](/docs/p-hacking-researcher-degrees-of-freedom) for the within-study failure.\n\n### Does publication bias apply to qualitative research?\nYes, and arguably more strongly, because there is no significance threshold to make the selection visible. The qualitative equivalent is the study whose themes were unsurprising, which never gets presented because it confirmed what the team already believed. Nobody records \"we interviewed twelve customers and heard nothing new\", even though that is a real and useful finding about the maturity of your understanding.\n\n### How do I detect publication bias in my own evidence base?\nYou cannot detect it from the record itself, which is the central difficulty. You can only detect it by comparing the record against a list of studies that were commissioned. If you have no such list, the honest position is that you do not know the size of your file drawer. Building the register is therefore both the fix and the only available diagnostic.\n\n### Is a funnel plot useful for internal research?\nRarely. Funnel plots and similar meta-analytic diagnostics need dozens of studies with comparable effect sizes to show the characteristic asymmetry. An internal programme with eight studies on a question does not have the statistical resolution for it. Registration is the practical alternative for corpora of this size.\n\n### Does the file drawer matter if we make decisions from one study anyway?\nIt matters more. A team that decides from single studies is fully exposed to whichever study happened to survive the escalation filter, with no aggregation to average out the selection. The one study that reached the room is not a random draw from the studies you ran - it is the one that had the most quotable result.\n\n### What is a fail-safe number and should I compute one?\nIt is the count of hidden null studies that would be required to overturn a combined result, from Rosenthal (1979), with a suggested threshold of 5k + 10 where k is the number of known studies. You almost certainly should not compute it formally for internal research - the assumptions do not transfer well. Use it as an intuition pump instead: with fewer than about fifteen studies on a question, a small number of unreported nulls can reverse your conclusion, so the register matters more than the arithmetic.\n\n## Related Resources\n\n- [P-Hacking and Researcher Degrees of Freedom](/docs/p-hacking-researcher-degrees-of-freedom) - the within-study counterpart to this failure\n- [Equivalence Testing: How to Prove There Is No Difference](/docs/equivalence-testing-no-difference) - what a null result actually licenses you to claim\n- [Evidence Synthesis](/docs/evidence-synthesis-research-findings) - pooling a body of internal studies without inheriting their selection bias\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and when to use each\n- [Building a UX Research Repository](/docs/research-repository-guide) - the storage layer a study register sits on top of\n- [Statistical Power and Minimum Detectable Effect](/docs/statistical-power-minimum-detectable-effect) - making a null result informative before you field it\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - the same selection logic applied to respondents rather than studies\n- [The Multiple Comparisons Problem](/docs/multiple-comparisons-problem) - false positives manufactured by breadth of testing","category":"Research Methods","lastModified":"2026-08-09T03:18:55.340821+00:00","metaTitle":"Publication Bias in Product Research: The File-Drawer Problem","metaDescription":"Only 21 percent of null results get published, and 65 percent are never written up at all. Learn what selective reporting does to your internal evidence base and how a study register closes the file drawer.","keywords":["publication bias","file drawer problem","file drawer effect","selective reporting","null results","fail-safe number","research evidence base","study register","effect size inflation","reporting bias"],"aiSummary":"Publication bias is the over-representation of positive findings in a body of evidence, caused by which studies get written up rather than by bad data. Franco, Malhotra and Simonovits (2014) observed a full population of 221 experiments and found 62 percent of strong results published against 21 percent of nulls, with the gap created almost entirely by authors never writing nulls up. Turner et al. (2008) showed the resulting effect-size inflation is about 32 percent. Product research is more exposed than academia because small evidence bases sit inside the fragile range of Rosenthal fail-safe arithmetic. The fix is a study register plus mandatory closed-ended outcome recording.","aiPrerequisites":["Familiarity with running or commissioning research studies","Basic understanding of significance testing and effect sizes"],"aiLearningOutcomes":["Explain publication bias as a selection failure rather than a data-quality failure","Quantify the size of the file drawer using the Franco TESS population data","Identify the five filters that remove null findings between a study and a decision","Apply Rosenthal fail-safe reasoning to judge how fragile a small internal evidence base is","Build a study register that records intentions before outcomes","Distinguish a genuine null result from an underpowered one before filing it"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min"}],"pagination":{"total":1,"returned":1,"offset":0}}