Back to docs
Research Methods

P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)

Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.

Answer first: p-hacking is what happens when analytic decisions are made after seeing the data. It does not require dishonesty and it does not require running many reported tests. In the landmark simulation by Simmons, Nelson and Simonsohn, four entirely ordinary analytic choices - two outcome measures instead of one, adding ten more respondents, controlling for gender, and dropping one of three conditions - raise the false-positive rate from the nominal 5 percent to 61 percent. The fix is not more statistical sophistication. It is a timestamp: write down the decision, the primary measure, the comparison family, the stopping rule and the exclusion rules before the data arrives, then label everything else exploratory.

The multiple comparisons problem is about arithmetic: run enough tests and one will cross the line. P-hacking is about something more uncomfortable. It inflates false positives even when you report exactly one test, because the choice of which test to report was itself informed by the data.

This is the failure mode that survives every correction procedure, because the tests being corrected for were never counted.

The 61 percent finding

In "False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant" (Psychological Science, 2011, 22(11):1359-1366), Joseph Simmons, Leif Nelson and Uri Simonsohn simulated what happens to the false-positive rate when a researcher retains a few common analytic freedoms. Their results:

Researcher degree of freedomFalse-positive rate at p < 0.05
No flexibility (nominal)5.0%
Two dependent variables instead of one (r = 0.50)9.5%
Adding 10 more observations if the first test failed7.7%
Controlling for gender, or a gender interaction11.7%
Running three conditions and reporting any two or all three12.6%
All four combined60.7%

Their summary of the combined case is blunt: it "would lead to a stunning 61% false-positive rate."

Look at what is on that list. Collecting a second outcome measure is good practice. Topping up a sample that came in light is normal operations. Controlling for a demographic is what a careful analyst does. Running three price points instead of two is better design. Every individual item is defensible. Together they turn a 1-in-20 error rate into a coin flip that lands on "significant" more often than not.

The authors also isolated optional stopping - checking results as data arrives and stopping when significance appears. A researcher who begins with 10 observations per condition and re-tests after every single additional observation, stopping at significance or at 50 per condition, finds a significant effect 22 percent of the time when nothing is there. This is precisely how most product teams monitor a running study.

To prove the point on real humans rather than simulations, they ran an experiment in which participants listened either to "When I'm Sixty-Four" by the Beatles or to a control track, then reported their date of birth and their father's age. Controlling for father's age, participants who heard the Beatles song were nearly a year and a half younger than the control group - adjusted means of 20.1 versus 21.5 years, F(1, 17) = 4.92, p = 0.040. The result is impossible. Music does not change when you were born. The paper is a demonstration that the standard toolkit, honestly applied, will certify an impossibility.

These are not fringe practices

Leslie John, George Loewenstein and Drazen Prelec surveyed academic psychologists about their own behaviour in "Measuring the Prevalence of Questionable Research Practices With Incentives for Truth Telling" (Psychological Science, 2012, 23(5):524-532). They emailed 5,964 researchers and received 2,155 responses, using an incentive scheme designed to make honest admission rational.

PracticeSelf-admission rate
Failing to report all of a study's dependent measures63.4%
Deciding whether to collect more data after checking whether results were significant55.9%
Selectively reporting studies that "worked"45.8%
Deciding whether to exclude data after seeing the impact of doing so38.2%
Failing to report all of a study's conditions27.7%
Reporting an unexpected finding as having been predicted from the start27.0%
Rounding a p-value down (reporting 0.054 as below 0.05)22.0%
Stopping data collection early because the desired result appeared15.6%
Falsifying data0.6%

Two things stand out. First, 94 percent of respondents admitted to at least one of these practices. Second, notice the gap between the top of the table and the bottom. Outright falsification sits at well under 1 percent. The practices that dominate are the ones researchers rated as defensible - and in many individual cases they genuinely are.

That is the real lesson. P-hacking is not a population of bad actors. It is the default behaviour of competent people working without a written plan, in a system that rewards clean results.

The commercial version is worse in one specific way. In academia, a false positive buys a publication and is eventually caught by a failed replication. In product research, a false positive buys a roadmap commitment, an engineering quarter and a repositioned landing page - and nobody ever runs the replication, so the correction arrives disguised as a disappointing quarter with no obvious cause.

The garden of forking paths: you do not have to fish to be caught

The most important refinement of this idea comes from Andrew Gelman and Eric Loken, whose 2013 paper carries its whole argument in the title: "The Garden of Forking Paths: Why Multiple Comparisons Can Be a Problem, Even When There Is No Fishing Expedition or P-Hacking and the Research Hypothesis Was Posited Ahead of Time."

Their point is subtle and it matters more than the simulation. You do not need to have run many analyses for multiplicity to apply. It is enough that you would have run a different one had the data come out differently.

Imagine a satisfaction study. If the effect had shown up in the overall mean, you would have reported the mean. It did not, but there was a clear gap among new users, so you reported that. Had the gap instead appeared among enterprise accounts, you would have reported that. Had it appeared only in the second wave, you would have framed it as a trend. You ran one test. The garden contained a dozen paths, and the data chose which one you walked.

This is why "but I only ran one analysis" is not a defence, and why no post-hoc correction can save you: the tests you must correct for are counterfactual. They exist in the decision procedure, not in the output log.

Gelman and Loken also give the best one-line description of what pre-registration actually does: it is "a floor, not a ceiling." A plan is a list of things you intend to do. It does not forbid you from noticing something unexpected - Fleming would still have seen the mould. It only fixes which claims were made before the data spoke.

Does it matter in practice?

Honesty requires reporting the strongest counter-evidence. Megan Head and colleagues text-mined p-values across scientific disciplines in "The Extent and Consequences of P-Hacking in Science" (PLoS Biology, 2015, 13(3):e1002106). Using the shape of the p-curve just below the 0.05 threshold, they found evidence of p-hacking in every discipline where they had good statistical power. But their conclusion on consequences was measured: p-hacking "probably does not drastically alter scientific consensuses drawn from meta-analyses," because p-hacked studies tend to be small and receive less weight when results are pooled.

That is genuine reassurance for a field that accumulates hundreds of studies per question. It is cold comfort for a product team, which typically has one study per question and no meta-analysis to wash out the noise. The protective mechanism Head and colleagues identified - aggregation across many independent studies - is exactly the thing commercial research does not have.

The scale of the underlying problem is visible in the Open Science Collaboration's "Estimating the Reproducibility of Psychological Science" (Science, 2015, 349(6251)): of 100 studies from leading journals, 97 percent reported significant original results, but only 36 percent replicated significantly, and replication effect sizes were roughly half the originals. Those originals were peer-reviewed work by trained researchers. A quarterly tracker analysed by one analyst under deadline is not held to a higher standard.

The full inventory of degrees of freedom in product research

Academic lists do not map cleanly onto commercial work. Here is the version that matters for a product or insights team:

Before fielding

  • Which respondents qualify (screener strictness, quota definitions)
  • How long to run, and whether to extend if numbers look light
  • Whether to include a pilot wave in the final dataset

Sample and stopping

  • Topping up a cell that is "nearly significant"
  • Stopping a running study when the result looks good
  • Dropping a fielding wave that behaved oddly

Measure definition

  • Top-box versus top-two-box versus mean on the same scale item
  • Which of several correlated metrics is "the" outcome
  • Recoding a scale midpoint as positive, negative or excluded
  • Index construction: which items go into a composite

Exclusions

  • Speeders, straight-liners, failed attention checks - and where exactly the cut-off sits
  • Whether to exclude a segment that "does not fit the study intent"

Comparison choice

  • Which banner, which contrast, which wave-over-wave window
  • Whether to test overall first or go straight to segments

Framing

  • Presenting an exploratory result as the hypothesis
  • Choosing the baseline that makes the change look largest

Each of these is a fork. None of them is misconduct. Collectively they are more than enough to produce a confident, well-presented, false conclusion - and they compound with the pure arithmetic covered in our guide to the multiple comparisons problem.

The fix: a one-page pre-committed analysis plan

The remedy is not statistical sophistication. It is a short document written before data collection, which converts every decision above from a post-hoc judgement into a pre-registered commitment.

A workable product-research analysis plan fits on one page and has seven fields:

FieldWhat it fixes
The decisionWhat action changes based on the result, and who owns it
Primary outcomeOne metric, defined exactly (including top-box rule and scale handling)
Comparison familyThe full list of comparisons that can produce a reported finding, and its count
Sample size and stopping ruleTarget n and the rule for termination, decided in advance
Exclusion rulesSpeeder threshold, attention-check policy, and what happens if exclusion changes the result
CovariatesWhich controls will be applied, plus a commitment to report the uncontrolled result too
Decision thresholdsWhat magnitude of effect triggers which action - written before you know the answer

Simmons, Nelson and Simonsohn later distilled the disclosure half of this into a single sentence, sometimes called the 21-word solution: We report how we determined our sample size, all data exclusions (if any), all manipulations, and all measures in the study. Adapted to commercial research, it is a line in the appendix of every readout, and it costs nothing.

Three rules make the plan work rather than becoming theatre:

  1. Exploration stays legal. The plan does not ban curiosity. It separates confirmatory claims from exploratory ones so that both can appear in the same deck with different labels and different consequences. Exploratory findings generate next quarter's pre-specified test; they do not generate roadmap commitments.
  2. Deviations are recorded, not hidden. Plans change for good reasons. Write what changed and why, before you look at the effect the change had.
  3. Somebody other than the analyst signs it. A plan reviewed only by its author is a diary. This is the function of a research peer review gate - a second person confirming that the comparison family and the stopping rule were fixed before fielding.

The complementary defence is replication rather than correction. If a finding matters enough to build on, the cheapest defensible test is not a cleverer analysis of the existing dataset - it is a fresh, pre-specified study on the surviving hypothesis. That reframing is what makes cheap research a methodological asset and not merely a budget one.

Where the analyst is not the problem

Two adjacent failures are commonly mislabelled as p-hacking, and the distinction matters because the remedies differ.

Observer bias operates during collection, not analysis: an interviewer who expects an answer subtly elicits it. Pre-committing an analysis plan does nothing about it; see observer bias for the interview-side controls.

Measurement insensitivity produces the mirror-image error. When a scale is saturated at the top, no analytic freedom will find an effect, and the team concludes there is nothing to find. That is a false negative generated by the instrument, covered in ceiling and floor effects.

Between them, these three articles cover the full triangle: false positives from breadth, false positives from freedom, false negatives from the instrument.

How Koji helps

Legacy research tooling is neutral about when decisions get made. It hands you a dataset and a crosstab engine, and it has no idea whether your comparison was chosen in the brief or invented at 11pm before the readout. That neutrality is the vulnerability.

The brief is the pre-registration. Koji research briefs capture the decision, the target audience and the questions before a single respondent is recruited. Because the brief exists as a versioned artifact created before fielding, it functions as a timestamp - the thing a post-hoc analysis plan can never credibly reconstruct. This is the single most valuable structural property for defending against forking paths, and it comes from the workflow rather than from analyst discipline.

Structured questions fix the measure before the data arrives. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Because scale points, choice options and ranking sets are defined at design time, the "which coding of this variable separates the groups" fork closes automatically. There is no midpoint to reassign after the fact. Our structured questions guide covers how to choose among the six, and survey design best practices covers writing them well.

Consistent AI moderation removes the collection-side forks. A human moderator running twenty interviews across three weeks drifts - probing harder on the theme that seems promising, which is optional stopping applied to qualitative work. Koji's AI moderator applies the same guide to every participant while still following up adaptively on what each person says, so depth is preserved without the interviewer's evolving hypothesis steering the sample.

Cheap replication beats clever reanalysis. This is the decisive advantage. Traditional research economics make the pre-committed plan feel expensive, because if the primary comparison comes back null you have burned six weeks and a large budget with nothing to show. That pressure is what generates p-hacking in the first place. When a fresh, properly specified study takes days rather than weeks, a null primary result is a cheap and informative outcome instead of a career problem, and the incentive to rescue it disappears.

Thematic analysis reports coverage, not thresholds. Koji clusters open-ended responses into themes with counts and verbatims. Coverage counts are not threshold crossings, so they are not vulnerable to the same manipulation - though they demand their own honesty, including reporting the participants who contradicted the theme. See thematic analysis and data saturation.

A working checklist

Before fielding:

  1. Write the one-page plan. Name the decision, the primary outcome and its exact definition.
  2. Count the comparison family and pick your error control, per the multiple comparisons guide.
  3. Set the sample size and the stopping rule. Write them down. Do not check results before reaching the target.
  4. Fix exclusion criteria and attention-check policy in advance.
  5. Have a second person sign the plan.

After the data arrives: 6. Run the primary analysis first, and record its result before running anything else. 7. Report the uncontrolled result alongside any covariate-adjusted one. 8. Report the result with and without exclusions. 9. Label every non-pre-specified result "exploratory" in the deck itself, not just in the appendix. 10. Convert the best exploratory findings into next quarter's pre-specified tests.

The bottom line

P-hacking is not a story about dishonest researchers. Ninety-four percent of surveyed academics admitted at least one questionable practice, and the most common ones were rated defensible by the people who used them. Four ordinary analytic choices are enough to push a 5 percent error rate to 61 percent, and Gelman and Loken showed that you do not even have to make the choices consciously - it is enough that the data would have led you somewhere else.

The defence is a timestamp, not a technique. Decide what you are testing, how much data you will collect, what counts as the outcome and what you will do about each possible answer - all before the answers exist. Then explore freely, label it honestly, and confirm what matters with a fresh study rather than a cleverer cut of the old one.

Start free with 10 credits and run a study where the brief is written first, the six structured question types fix your measures at design time, and confirming a finding costs days instead of a quarter.

Frequently asked questions

Is p-hacking the same as fraud?

No, and treating it as fraud is why it persists. Data falsification was admitted by 0.6 percent of surveyed researchers; failing to report all dependent measures was admitted by 63.4 percent. The dominant practices are ordinary analytic judgements made without a written plan, and the people making them generally rate them as defensible. The remedy is procedural - a pre-committed plan and honest labelling - not disciplinary.

Can I still explore my data?

Yes, and you should. A pre-committed analysis plan is a floor, not a ceiling: it fixes which claims were made before the data spoke, and leaves you free to look at anything else. The only requirement is that exploratory findings are labelled as such in the deck and are treated as hypotheses for a future study rather than as conclusions that justify a roadmap decision today.

What if I need to add respondents because the sample came in light?

Decide the rule in advance and it is not a problem. "We will field until we reach 400 completes or 14 days, whichever comes first" is a legitimate stopping rule. What inflates false positives is conditional topping up - checking the result, then adding respondents because the p-value was close. Simulations show that testing repeatedly as data arrives produces a significant result about 22 percent of the time when no effect exists.

How is p-hacking different from the multiple comparisons problem?

Multiple comparisons is arithmetic: many tests, one threshold, inevitable false positives. It applies even when every test was pre-specified and fully reported. P-hacking is behavioural: the analytic path was chosen after seeing the data, so the tests you should correct for are the ones you would have run under different data. Correction procedures address the first and are powerless against the second.

Does p-hacking actually change conclusions, or is it just a theoretical worry?

Both are true, in different settings. Head and colleagues found p-hacking widespread across disciplines but concluded it probably does not drastically alter meta-analytic consensus, because p-hacked studies are typically small and get down-weighted when pooled. That protection depends on having many independent studies of the same question. Commercial research usually has one, which is why the risk is more acute for product teams than for the academic fields where the problem was first documented.

What is the single highest-value change a small team can make?

Write down the primary outcome and the stopping rule before fielding, and have one other person confirm it. That single step closes the two most commonly admitted degrees of freedom - undisclosed outcome selection and conditional sample topping up - and it takes about ten minutes. Everything else on the checklist is refinement.

Related Resources

Related Articles

Quasi-Experimental Design: How to Measure Impact When You Cannot Run an A/B Test (2026)

Most product decisions cannot be randomised. Quasi-experimental designs give you a defensible causal answer anyway. Learn which of the three designs your situation calls for, how to write the impact model before the data arrives, and why interviews are the cheapest confounder detector you have.

Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies

Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.

Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)

Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.

Statistical Significance in Survey Research: A Plain-English Guide (2026)

A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Design Best Practices: From Question Writing to Data Collection

Learn how to design effective surveys with proven best practices for question writing, flow, bias reduction, and data collection — including when to go beyond surveys to AI-powered interviews.