{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-09T03:29:39.079Z"},"content":[{"type":"documentation","id":"5caf997c-5d1b-4113-90dc-c5d3651b1dd7","slug":"p-hacking-researcher-degrees-of-freedom","title":"P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)","url":"https://www.koji.so/docs/p-hacking-researcher-degrees-of-freedom","summary":"P-hacking inflates false positives when analytic decisions are made after seeing the data, and it survives every multiple-comparisons correction because the relevant tests were never counted. Simmons, Nelson and Simonsohn (2011) showed that four ordinary choices - a second dependent variable, adding observations, controlling for gender, and dropping one of three conditions - raise the false-positive rate from 5 percent to 60.7 percent, and that optional stopping alone produces significance 22 percent of the time under a true null. John, Loewenstein and Prelec (2012) found 94 percent of 2,155 surveyed psychologists admitted at least one questionable practice, with the most common rated defensible. Gelman and Loken showed the problem applies even to a single reported test, because multiplicity lives in the analyses that would have been run under different data. The remedy is a one-page pre-committed analysis plan naming the decision, primary outcome, comparison family, stopping rule, exclusions, covariates and decision thresholds - a floor rather than a ceiling, which leaves exploration legal but labelled.","content":"**Answer first: p-hacking is what happens when analytic decisions are made after seeing the data. It does not require dishonesty and it does not require running many reported tests. In the landmark simulation by Simmons, Nelson and Simonsohn, four entirely ordinary analytic choices - two outcome measures instead of one, adding ten more respondents, controlling for gender, and dropping one of three conditions - raise the false-positive rate from the nominal 5 percent to 61 percent. The fix is not more statistical sophistication. It is a timestamp: write down the decision, the primary measure, the comparison family, the stopping rule and the exclusion rules before the data arrives, then label everything else exploratory.**\n\nThe multiple comparisons problem is about arithmetic: run enough tests and one will cross the line. P-hacking is about something more uncomfortable. It inflates false positives even when you report exactly one test, because the choice of *which* test to report was itself informed by the data.\n\nThis is the failure mode that survives every correction procedure, because the tests being corrected for were never counted.\n\n## The 61 percent finding\n\nIn \"False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant\" (*Psychological Science*, 2011, 22(11):1359-1366), Joseph Simmons, Leif Nelson and Uri Simonsohn simulated what happens to the false-positive rate when a researcher retains a few common analytic freedoms. Their results:\n\n| Researcher degree of freedom | False-positive rate at p < 0.05 |\n| --- | --- |\n| No flexibility (nominal) | 5.0% |\n| Two dependent variables instead of one (r = 0.50) | 9.5% |\n| Adding 10 more observations if the first test failed | 7.7% |\n| Controlling for gender, or a gender interaction | 11.7% |\n| Running three conditions and reporting any two or all three | 12.6% |\n| **All four combined** | **60.7%** |\n\nTheir summary of the combined case is blunt: it \"would lead to a stunning 61% false-positive rate.\"\n\nLook at what is on that list. Collecting a second outcome measure is good practice. Topping up a sample that came in light is normal operations. Controlling for a demographic is what a careful analyst does. Running three price points instead of two is better design. Every individual item is defensible. Together they turn a 1-in-20 error rate into a coin flip that lands on \"significant\" more often than not.\n\nThe authors also isolated **optional stopping** - checking results as data arrives and stopping when significance appears. A researcher who begins with 10 observations per condition and re-tests after every single additional observation, stopping at significance or at 50 per condition, finds a significant effect **22 percent** of the time when nothing is there. This is precisely how most product teams monitor a running study.\n\nTo prove the point on real humans rather than simulations, they ran an experiment in which participants listened either to \"When I'm Sixty-Four\" by the Beatles or to a control track, then reported their date of birth and their father's age. Controlling for father's age, participants who heard the Beatles song were nearly a year and a half *younger* than the control group - adjusted means of 20.1 versus 21.5 years, F(1, 17) = 4.92, p = 0.040. The result is impossible. Music does not change when you were born. The paper is a demonstration that the standard toolkit, honestly applied, will certify an impossibility.\n\n## These are not fringe practices\n\nLeslie John, George Loewenstein and Drazen Prelec surveyed academic psychologists about their own behaviour in \"Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling\" (*Psychological Science*, 2012, 23(5):524-532). They emailed 5,964 researchers and received 2,155 responses, using an incentive scheme designed to make honest admission rational.\n\n| Practice | Self-admission rate |\n| --- | --- |\n| Failing to report all of a study's dependent measures | 63.4% |\n| Deciding whether to collect more data after checking whether results were significant | 55.9% |\n| Selectively reporting studies that \"worked\" | 45.8% |\n| Deciding whether to exclude data after seeing the impact of doing so | 38.2% |\n| Failing to report all of a study's conditions | 27.7% |\n| Reporting an unexpected finding as having been predicted from the start | 27.0% |\n| Rounding a p-value down (reporting 0.054 as below 0.05) | 22.0% |\n| Stopping data collection early because the desired result appeared | 15.6% |\n| Falsifying data | 0.6% |\n\nTwo things stand out. First, **94 percent of respondents admitted to at least one** of these practices. Second, notice the gap between the top of the table and the bottom. Outright falsification sits at well under 1 percent. The practices that dominate are the ones researchers rated as *defensible* - and in many individual cases they genuinely are.\n\nThat is the real lesson. P-hacking is not a population of bad actors. It is the default behaviour of competent people working without a written plan, in a system that rewards clean results.\n\nThe commercial version is worse in one specific way. In academia, a false positive buys a publication and is eventually caught by a failed replication. In product research, a false positive buys a roadmap commitment, an engineering quarter and a repositioned landing page - and nobody ever runs the replication, so the correction arrives disguised as a disappointing quarter with no obvious cause.\n\n## The garden of forking paths: you do not have to fish to be caught\n\nThe most important refinement of this idea comes from Andrew Gelman and Eric Loken, whose 2013 paper carries its whole argument in the title: \"The Garden of Forking Paths: Why Multiple Comparisons Can Be a Problem, Even When There Is No Fishing Expedition or P-Hacking and the Research Hypothesis Was Posited Ahead of Time.\"\n\nTheir point is subtle and it matters more than the simulation. You do not need to have *run* many analyses for multiplicity to apply. It is enough that you *would have run a different one* had the data come out differently.\n\nImagine a satisfaction study. If the effect had shown up in the overall mean, you would have reported the mean. It did not, but there was a clear gap among new users, so you reported that. Had the gap instead appeared among enterprise accounts, you would have reported that. Had it appeared only in the second wave, you would have framed it as a trend. You ran one test. The garden contained a dozen paths, and the data chose which one you walked.\n\nThis is why \"but I only ran one analysis\" is not a defence, and why no post-hoc correction can save you: the tests you must correct for are counterfactual. They exist in the decision procedure, not in the output log.\n\nGelman and Loken also give the best one-line description of what pre-registration actually does: it is **\"a floor, not a ceiling.\"** A plan is a list of things you intend to do. It does not forbid you from noticing something unexpected - Fleming would still have seen the mould. It only fixes which claims were made before the data spoke.\n\n## Does it matter in practice?\n\nHonesty requires reporting the strongest counter-evidence. Megan Head and colleagues text-mined p-values across scientific disciplines in \"The Extent and Consequences of P-Hacking in Science\" (*PLoS Biology*, 2015, 13(3):e1002106). Using the shape of the p-curve just below the 0.05 threshold, they found evidence of p-hacking in every discipline where they had good statistical power. But their conclusion on consequences was measured: p-hacking \"probably does not drastically alter scientific consensuses drawn from meta-analyses,\" because p-hacked studies tend to be small and receive less weight when results are pooled.\n\nThat is genuine reassurance for a field that accumulates hundreds of studies per question. It is cold comfort for a product team, which typically has *one* study per question and no meta-analysis to wash out the noise. The protective mechanism Head and colleagues identified - aggregation across many independent studies - is exactly the thing commercial research does not have. The selection failure that keeps a commercial evidence base small and skewed is covered in [publication bias in product research](/docs/publication-bias-product-research).\n\nThe scale of the underlying problem is visible in the Open Science Collaboration's \"Estimating the Reproducibility of Psychological Science\" (*Science*, 2015, 349(6251)): of 100 studies from leading journals, 97 percent reported significant original results, but only **36 percent** replicated significantly, and replication effect sizes were roughly **half** the originals. Those originals were peer-reviewed work by trained researchers. A quarterly tracker analysed by one analyst under deadline is not held to a higher standard.\n\n## The full inventory of degrees of freedom in product research\n\nAcademic lists do not map cleanly onto commercial work. Here is the version that matters for a product or insights team:\n\n**Before fielding**\n- Which respondents qualify (screener strictness, quota definitions)\n- How long to run, and whether to extend if numbers look light\n- Whether to include a pilot wave in the final dataset\n\n**Sample and stopping**\n- Topping up a cell that is \"nearly significant\"\n- Stopping a running study when the result looks good\n- Dropping a fielding wave that behaved oddly\n\n**Measure definition**\n- Top-box versus top-two-box versus mean on the same scale item\n- Which of several correlated metrics is \"the\" outcome\n- Recoding a scale midpoint as positive, negative or excluded\n- Index construction: which items go into a composite\n\n**Exclusions**\n- Speeders, straight-liners, failed attention checks - and where exactly the cut-off sits\n- Whether to exclude a segment that \"does not fit the study intent\"\n\n**Comparison choice**\n- Which banner, which contrast, which wave-over-wave window\n- Whether to test overall first or go straight to segments\n\n**Framing**\n- Presenting an exploratory result as the hypothesis\n- Choosing the baseline that makes the change look largest\n\nEach of these is a fork. None of them is misconduct. Collectively they are more than enough to produce a confident, well-presented, false conclusion - and they compound with the pure arithmetic covered in our guide to the [multiple comparisons problem](/docs/multiple-comparisons-problem).\n\n## The fix: a one-page pre-committed analysis plan\n\nThe remedy is not statistical sophistication. It is a short document written before data collection, which converts every decision above from a post-hoc judgement into a pre-registered commitment.\n\nA workable product-research analysis plan fits on one page and has seven fields:\n\n| Field | What it fixes |\n| --- | --- |\n| **The decision** | What action changes based on the result, and who owns it |\n| **Primary outcome** | One metric, defined exactly (including top-box rule and scale handling) |\n| **Comparison family** | The full list of comparisons that can produce a reported finding, and its count |\n| **Sample size and stopping rule** | Target n and the rule for termination, decided in advance |\n| **Exclusion rules** | Speeder threshold, attention-check policy, and what happens if exclusion changes the result |\n| **Covariates** | Which controls will be applied, plus a commitment to report the uncontrolled result too |\n| **Decision thresholds** | What magnitude of effect triggers which action - written before you know the answer |\n\nSimmons, Nelson and Simonsohn later distilled the disclosure half of this into a single sentence, sometimes called the 21-word solution: *We report how we determined our sample size, all data exclusions (if any), all manipulations, and all measures in the study.* Adapted to commercial research, it is a line in the appendix of every readout, and it costs nothing.\n\nThree rules make the plan work rather than becoming theatre:\n\n1. **Exploration stays legal.** The plan does not ban curiosity. It separates confirmatory claims from exploratory ones so that both can appear in the same deck with different labels and different consequences. Exploratory findings generate next quarter's pre-specified test; they do not generate roadmap commitments.\n2. **Deviations are recorded, not hidden.** Plans change for good reasons. Write what changed and why, before you look at the effect the change had.\n3. **Somebody other than the analyst signs it.** A plan reviewed only by its author is a diary. This is the function of a [research peer review gate](/docs/research-peer-review-qa-gate) - a second person confirming that the comparison family and the stopping rule were fixed before fielding.\n\nThe complementary defence is replication rather than correction. If a finding matters enough to build on, the cheapest defensible test is not a cleverer analysis of the existing dataset - it is a fresh, pre-specified study on the surviving hypothesis. That reframing is what makes cheap research a *methodological* asset and not merely a budget one.\n\n## Where the analyst is not the problem\n\nTwo adjacent failures are commonly mislabelled as p-hacking, and the distinction matters because the remedies differ.\n\n**Observer bias** operates during collection, not analysis: an interviewer who expects an answer subtly elicits it. Pre-committing an analysis plan does nothing about it; see [observer bias](/docs/observer-bias) for the interview-side controls.\n\n**Measurement insensitivity** produces the mirror-image error. When a scale is saturated at the top, no analytic freedom will find an effect, and the team concludes there is nothing to find. That is a false negative generated by the instrument, covered in [ceiling and floor effects](/docs/ceiling-floor-effects-research).\n\nBetween them, these three articles cover the full triangle: false positives from breadth, false positives from freedom, false negatives from the instrument.\n\n## How Koji helps\n\nLegacy research tooling is neutral about when decisions get made. It hands you a dataset and a crosstab engine, and it has no idea whether your comparison was chosen in the brief or invented at 11pm before the readout. That neutrality is the vulnerability.\n\n**The brief is the pre-registration.** Koji research briefs capture the decision, the target audience and the questions before a single respondent is recruited. Because the brief exists as a versioned artifact created before fielding, it functions as a timestamp - the thing a post-hoc analysis plan can never credibly reconstruct. This is the single most valuable structural property for defending against forking paths, and it comes from the workflow rather than from analyst discipline.\n\n**Structured questions fix the measure before the data arrives.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Because scale points, choice options and ranking sets are defined at design time, the \"which coding of this variable separates the groups\" fork closes automatically. There is no midpoint to reassign after the fact. Our [structured questions guide](/docs/structured-questions-guide) covers how to choose among the six, and [survey design best practices](/docs/survey-design-best-practices) covers writing them well.\n\n**Consistent AI moderation removes the collection-side forks.** A human moderator running twenty interviews across three weeks drifts - probing harder on the theme that seems promising, which is optional stopping applied to qualitative work. Koji's AI moderator applies the same guide to every participant while still following up adaptively on what each person says, so depth is preserved without the interviewer's evolving hypothesis steering the sample.\n\n**Cheap replication beats clever reanalysis.** This is the decisive advantage. Traditional research economics make the pre-committed plan feel expensive, because if the primary comparison comes back null you have burned six weeks and a large budget with nothing to show. That pressure is what generates p-hacking in the first place. When a fresh, properly specified study takes days rather than weeks, a null primary result is a cheap and informative outcome instead of a career problem, and the incentive to rescue it disappears.\n\n**Thematic analysis reports coverage, not thresholds.** Koji clusters open-ended responses into themes with counts and verbatims. Coverage counts are not threshold crossings, so they are not vulnerable to the same manipulation - though they demand their own honesty, including reporting the participants who contradicted the theme. See [thematic analysis](/docs/thematic-analysis-guide) and [data saturation](/docs/data-saturation-qualitative-research).\n\n## A working checklist\n\nBefore fielding:\n1. Write the one-page plan. Name the decision, the primary outcome and its exact definition.\n2. Count the comparison family and pick your error control, per the [multiple comparisons guide](/docs/multiple-comparisons-problem).\n3. Set the sample size and the stopping rule. Write them down. Do not check results before reaching the target.\n4. Fix exclusion criteria and attention-check policy in advance.\n5. Have a second person sign the plan.\n\nAfter the data arrives:\n6. Run the primary analysis first, and record its result before running anything else.\n7. Report the uncontrolled result alongside any covariate-adjusted one.\n8. Report the result with and without exclusions.\n9. Label every non-pre-specified result \"exploratory\" in the deck itself, not just in the appendix.\n10. Convert the best exploratory findings into next quarter's pre-specified tests.\n\n## The bottom line\n\nP-hacking is not a story about dishonest researchers. Ninety-four percent of surveyed academics admitted at least one questionable practice, and the most common ones were rated defensible by the people who used them. Four ordinary analytic choices are enough to push a 5 percent error rate to 61 percent, and Gelman and Loken showed that you do not even have to make the choices consciously - it is enough that the data would have led you somewhere else.\n\nThe defence is a timestamp, not a technique. Decide what you are testing, how much data you will collect, what counts as the outcome and what you will do about each possible answer - all before the answers exist. Then explore freely, label it honestly, and confirm what matters with a fresh study rather than a cleverer cut of the old one.\n\n**Start free with 10 credits** and run a study where the brief is written first, the six structured question types fix your measures at design time, and confirming a finding costs days instead of a quarter.\n\n## Frequently asked questions\n\n### Is p-hacking the same as fraud?\nNo, and treating it as fraud is why it persists. Data falsification was admitted by 0.6 percent of surveyed researchers; failing to report all dependent measures was admitted by 63.4 percent. The dominant practices are ordinary analytic judgements made without a written plan, and the people making them generally rate them as defensible. The remedy is procedural - a pre-committed plan and honest labelling - not disciplinary.\n\n### Can I still explore my data?\nYes, and you should. A pre-committed analysis plan is a floor, not a ceiling: it fixes which claims were made before the data spoke, and leaves you free to look at anything else. The only requirement is that exploratory findings are labelled as such in the deck and are treated as hypotheses for a future study rather than as conclusions that justify a roadmap decision today.\n\n### What if I need to add respondents because the sample came in light?\nDecide the rule in advance and it is not a problem. \"We will field until we reach 400 completes or 14 days, whichever comes first\" is a legitimate stopping rule. What inflates false positives is *conditional* topping up - checking the result, then adding respondents because the p-value was close. Simulations show that testing repeatedly as data arrives produces a significant result about 22 percent of the time when no effect exists.\n\n### How is p-hacking different from the multiple comparisons problem?\nMultiple comparisons is arithmetic: many tests, one threshold, inevitable false positives. It applies even when every test was pre-specified and fully reported. P-hacking is behavioural: the analytic path was chosen after seeing the data, so the tests you should correct for are the ones you would have run under different data. Correction procedures address the first and are powerless against the second.\n\n### Does p-hacking actually change conclusions, or is it just a theoretical worry?\nBoth are true, in different settings. Head and colleagues found p-hacking widespread across disciplines but concluded it probably does not drastically alter meta-analytic consensus, because p-hacked studies are typically small and get down-weighted when pooled. That protection depends on having many independent studies of the same question. Commercial research usually has one, which is why the risk is more acute for product teams than for the academic fields where the problem was first documented.\n\n### What is the single highest-value change a small team can make?\nWrite down the primary outcome and the stopping rule before fielding, and have one other person confirm it. That single step closes the two most commonly admitted degrees of freedom - undisclosed outcome selection and conditional sample topping up - and it takes about ten minutes. Everything else on the checklist is refinement.\n\n## Related Resources\n\n- [The Multiple Comparisons Problem](/docs/multiple-comparisons-problem) - the arithmetic half of false positives\n- [Ceiling and Floor Effects](/docs/ceiling-floor-effects-research) - the false-negative failure mode\n- [Statistical Significance in Survey Research](/docs/statistical-significance-survey-research) - what a p-value does and does not mean\n- [Research Peer Review: The Pre-Launch QA Gate](/docs/research-peer-review-qa-gate) - who signs the plan before fielding\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that fix your measures at design time\n- [Quasi-Experimental Design](/docs/quasi-experimental-design-guide) - specifying the impact model before the data arrives\n- [Pilot Study Guide](/docs/pilot-study-user-research-guide) - testing the instrument before it counts","category":"Research Methods","lastModified":"2026-08-09T03:24:56.561912+00:00","metaTitle":"P-Hacking and Researcher Degrees of Freedom: The 61 Percent Problem","metaDescription":"Four ordinary analytic choices push the false-positive rate from 5 percent to 61 percent. Learn the full inventory of researcher degrees of freedom in product research and the one-page analysis plan that closes them.","keywords":["p-hacking","researcher degrees of freedom","garden of forking paths","pre-registration","pre-committed analysis plan","questionable research practices","optional stopping","false positive research","HARKing"],"aiSummary":"P-hacking inflates false positives when analytic decisions are made after seeing the data, and it survives every multiple-comparisons correction because the relevant tests were never counted. Simmons, Nelson and Simonsohn (2011) showed that four ordinary choices - a second dependent variable, adding observations, controlling for gender, and dropping one of three conditions - raise the false-positive rate from 5 percent to 60.7 percent, and that optional stopping alone produces significance 22 percent of the time under a true null. John, Loewenstein and Prelec (2012) found 94 percent of 2,155 surveyed psychologists admitted at least one questionable practice, with the most common rated defensible. Gelman and Loken showed the problem applies even to a single reported test, because multiplicity lives in the analyses that would have been run under different data. The remedy is a one-page pre-committed analysis plan naming the decision, primary outcome, comparison family, stopping rule, exclusions, covariates and decision thresholds - a floor rather than a ceiling, which leaves exploration legal but labelled.","aiPrerequisites":["Familiarity with p-values and significance testing","Experience running or reading survey and experiment readouts"],"aiLearningOutcomes":["Recognise researcher degrees of freedom across fielding, sampling, measure definition, exclusions and framing","Explain why a single reported test can still be inflated by the garden of forking paths","Write a one-page pre-committed analysis plan with seven required fields","Set a stopping rule in advance and avoid conditional sample topping up","Separate confirmatory from exploratory findings in a readout and act on them differently","Use cheap replication rather than reanalysis to confirm a finding"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min"}],"pagination":{"total":1,"returned":1,"offset":0}}