Back to docs
Research Methods

Same Data, Different Answers: The Many-Analysts Problem in Product Research

When 73 teams analyzed identical data to test one hypothesis, over 95 percent of the variance in their results was unexplained. Your analysis is one draw from a distribution you never see.

Give the same dataset and the same question to several competent analysts and they will not return the same answer. Not slightly different confidence intervals - different conclusions, in opposite directions, from honest people with no incentive to disagree. This has now been measured repeatedly, across psychology, neuroimaging, economics and sociology, and the finding is consistent enough to be treated as a property of data analysis rather than a scandal. It has a direct and uncomfortable implication for product research: the readout in your deck is not "what the data says." It is one sample from a distribution of things the data could have said, and nobody in the room can see the rest of it.

The short answer

Analytic variation is not the same problem as p-hacking. P-hacking is one analyst searching a space of choices for the result they want. The many-analysts finding is stranger and harder to fix: independent analysts, pre-registered, with no stake in the outcome and no knowledge of each other's work, land in different places anyway. Nobody cheats and the answers still diverge.

The practical response is not to find a more rigorous analyst. It is to stop reporting a single number as though it were the only one available, and to make the spread part of the finding.

What was actually measured

Three landmark studies define the evidence base. They are worth knowing by their numbers, because the numbers are what make the argument.

StudyFieldDesignHeadline result
Silberzahn et al. (2018), Advances in Methods and Practices in Psychological Science 1(3):337-356Psychology29 teams, 61 analysts, one dataset, one questionEffect estimates ranged from 0.89 to 2.93 in odds-ratio units, median 1.31; 20 of 29 teams found a significant effect, 9 did not
Botvinik-Nezer et al. (2020), Nature 582:84-88Neuroimaging70 teams, one fMRI dataset, nine pre-specified hypothesesNo two teams used identical workflows; for five of the nine hypotheses, between 21.4 and 37.1 percent of teams reported a significant result
Breznau et al. (2022), PNAS 119(44)Sociology161 researchers in 73 teams, 1,253 models, one hypothesisMore than 95 percent of the total variance in numerical results remained unexplained after coding every identifiable decision

Silberzahn: 29 teams, one question about red cards

Twenty-nine research teams comprising 61 analysts were given the same dataset and asked whether soccer referees are more likely to give red cards to players with darker skin tone. The 29 analyses used 21 unique combinations of covariates. Estimated effects ranged from 0.89 to 2.93 in odds-ratio units - that is, from a small negative effect to nearly a tripling of the odds - with a median of 1.31. Twenty teams (69 percent) found a statistically significant positive effect; nine did not.

The part that should worry a research organization is the explanation, or rather its absence. Neither the analysts' prior beliefs about the effect nor their level of expertise explained the variation. Peer ratings of analysis quality did not explain it either. The teams peer-reviewed each other's methods, and the highly-rated approaches did not converge.

Botvinik-Nezer: 70 teams, and no two pipelines alike

Seventy independent teams analyzed a single functional MRI dataset, testing nine hypotheses specified in advance. No two teams chose identical workflows. The result rate varied enormously by hypothesis: one hypothesis produced a significant result for 84.3 percent of teams, three produced significant results for only 5.7 percent, and the remaining five sat between 21.4 and 37.1 percent - the zone where the answer you report is close to a coin flip weighted by your pipeline. Notably, teams whose intermediate statistical maps were highly correlated still reached different conclusions at the hypothesis-testing step. Agreement upstream did not produce agreement downstream.

Breznau: the 95 percent that nobody can explain

The largest and most sobering of the three. 161 researchers in 73 teams tested one hypothesis - that greater immigration reduces public support for social policy - on identical data, having submitted analysis plans in advance. They produced 1,253 converged models, averaging 17.5 per team and ranging from 1 to 124.

Little more than half the reported estimates were not statistically distinguishable from zero, about a quarter were significant and negative, and 16.9 percent were significant and positive. On substantive conclusions, 60.7 percent of team conclusions rejected the hypothesis, 28.5 percent supported it, and 13.5 percent were that the hypothesis was not testable with these data.

The authors then did what nobody had done before: they coded every identifiable decision in every workflow. They found 166 distinct research design decisions, 107 of which were taken by at least three teams, and regressed the results on all of them plus researcher characteristics, expertise and prior beliefs. More than 95 percent of the total variance in numerical results remained unexplained. The divergence is not traceable to a list of choices you could standardize away. The authors call it a hidden universe of uncertainty, visible only when you look at many analyses at once and invisible in any single study.

What this means for a product research readout

Three consequences follow, and each contradicts something most research organizations currently believe.

One. Seniority is not a fix. The standard organizational response to an untrusted number is to have a more experienced person look at it. Silberzahn measured expertise directly and found it did not predict where an analyst landed. Breznau measured expertise, prior beliefs and expectations and found they barely predicted anything. Escalating an analysis to a senior analyst gets you a different draw, not a better one.

Two. Peer review of the method does not converge the answer. In Silberzahn's design, teams reviewed each other's approaches before finalizing. Quality ratings did not explain the spread. A methods review tells you an analysis is defensible; it does not tell you it is the one the other twenty-eight analysts would have produced.

Three. The single readout is the problem, not the analyst who produced it. If a competent analyst could have produced a range of answers, then a deck presenting one answer with a confidence interval is understating uncertainty by an amount nobody has quantified. The confidence interval describes sampling error. It says nothing about analytic variation, which these studies suggest is frequently larger.

What to do about it

The academic remedies are multiverse analysis - Steegen, Tuerlinckx, Gelman and Vanpaemel, Perspectives on Psychological Science 11(5):702-712 (2016) - and specification curve analysis, Simonsohn, Simmons and Nelson, Nature Human Behaviour 4(11):1208-1214 (2020). Both work the same way: instead of choosing one defensible analysis, enumerate all the defensible ones, run them all, and report the distribution. The finding becomes "the effect is positive in 84 percent of reasonable specifications, with a median of X," which is both more honest and, in practice, more persuasive.

Running 1,253 models is not a product-team activity. The ladder below is, and each rung is worth more than the one below it.

RungWhat you doCostWhat it buys
1Write down the analysis decisions you made and the alternatives you rejected15 minutesMakes the forking visible to the reader
2Re-run the primary comparison under 3-5 defensible alternative specifications1-2 hoursAn empirical range instead of a point estimate
3Have a second analyst analyze the same extract independently, without seeing yoursHalf a dayA real second draw from the distribution
4Blind the analysis so specification choices are made before the answer is visible1 hour plus planRemoves preference from the choice of specification
5Full specification curve over the plausible analytic spaceDaysThe distribution itself

Rung 2 is the highest-value rung for almost every team, and it is the one most people have never tried. Pick the three decisions you were genuinely unsure about - the outlier rule, the segment definition, the inclusion window - and run the cross-product. If the conclusion holds in all of them, you have something much stronger than a p-value. If it flips in two of eight, you have learned something that a single analysis would have hidden from you permanently.

Rung 3 has a useful side effect: when two analysts disagree, the disagreement itself is diagnostic. Feed it into an analysis of competing hypotheses rather than negotiating toward the middle. The specification where the two analyses diverge is telling you which assumption is load-bearing.

Where this problem is worse, and where it is smaller

Analytic variation scales with the number of defensible choices, which means it is not uniform across research types.

  • Worst: observational analytics on rich behavioral data. Many covariates, many possible segment definitions, no randomization. This is where the studies above found their spread and where most product analytics lives.
  • Bad: qualitative synthesis without a codebook. The choices are less visible but no fewer. See inter-rater reliability and qualitative research validity.
  • Better: randomized experiments with one pre-specified primary outcome. Randomization removes the covariate question; pre-specification removes most of the rest. See quantitative user research methods.
  • Best: a direct replication. A second study is the only thing that fully separates analytic variation from a real effect. This is why fast studies matter more than clever reanalysis.

The modern approach: making the second draw affordable

Every rung on that ladder is a labor problem. Rung 3 asks for a second analyst - most teams have one analyst. Rung 5 asks for days of modeling. The reason single-analysis readouts dominate is not that anyone believes they are sufficient; it is that the alternative has historically cost more than the decision was worth.

Koji changes the economics on the two rungs that matter most:

  • Consistent automated analysis is a reproducible specification. When the same analysis rubric is applied to every transcript, re-running it under a different definition is a re-run, not a re-read. That makes rung 2 - the alternative-specification check - something you do in an afternoon instead of a sprint.
  • Structured questions shrink the analytic space at the source. All six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - produce values that need no coding decision to become a number. Fewer researcher decisions between the respondent and the chart means fewer forks in the path. This is the cheapest available reduction in analytic variance and it happens at study-design time.
  • Replication becomes a real option. The strongest response to analytic uncertainty is another study, and that has always been the response nobody could afford. AI-moderated interviews put a confirmatory study inside the same decision cycle as the original.
  • Full data export means a second analyst is possible at all. Rung 3 requires handing someone the raw layer. A platform that returns only summaries has made an independent second analysis structurally impossible, which is worth asking any vendor before you build a research programme on them.

Traditional survey tools optimize for the opposite: one dashboard, one default cut, one number per chart. That design makes analytic variation invisible rather than absent.

Frequently asked questions

Is this the same as p-hacking?

No, and the distinction matters. P-hacking involves searching analytic choices for a desired result, usually statistical significance. The many-analysts studies used pre-registered, independent teams with no stake in the outcome and still found wide divergence. P-hacking is a discipline problem with a discipline fix; analytic variation persists after the discipline is applied.

Does more expertise reduce the spread?

The evidence says no. Silberzahn et al. found that neither analyst expertise nor peer ratings of analysis quality explained the variation between teams, and Breznau et al. found that expertise, prior beliefs and expectations barely predicted outcomes. Escalating to a more senior analyst produces a different draw, not a more central one.

If any analysis could be wrong, why analyze at all?

Because the spread is usually informative even when it is wide. In Botvinik-Nezer's data, three hypotheses produced significant results for only 5.7 percent of teams - near-unanimous agreement that there was nothing there. Consistency across specifications is a genuine finding. What you cannot do is treat a single specification's output as if it carried that weight.

What is a multiverse analysis in plain terms?

Enumerate every defensible way to prepare and analyze the data - each outlier rule, each exclusion, each segment definition - run all the combinations, and report the distribution of results rather than one of them. Specification curve analysis is the same idea with a standard way of plotting and testing the resulting distribution.

What is the minimum version I can actually run this week?

Pick the three analysis decisions you were least sure about, run the eight combinations, and put the range in the appendix of your readout. It costs an afternoon and it is the difference between "engagement rose 12 percent" and "engagement rose between 4 and 15 percent across every reasonable way of cutting it."

How does this affect small-sample qualitative research?

It applies, but the fix is different. With ten interviews the analytic choices are about coding and theme construction rather than model specification. The equivalent moves are a shared codebook, a second coder on a subset, and reporting theme prevalence rather than selected quotes.

The bottom line

The many-analysts literature does not show that data analysis is arbitrary. It shows that a single analysis carries more uncertainty than it displays, that the extra uncertainty is not removed by expertise or peer review, and that the only reliable way to see it is to produce more than one analysis. For research teams the practical conclusion is narrow and actionable: report a range, get a second draw when the decision is expensive, and treat the ability to run a fast confirmatory study as the most valuable property your research stack has.

Related Resources

Related Articles

Analysis of Competing Hypotheses: How to Test What Your Research Actually Supports

Most evidence that supports your favorite explanation also supports the ones you never wrote down. ACH is the matrix method that finds the evidence which actually discriminates.

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)

Most teams have dozens of studies and no way to say what they collectively know. Evidence synthesis is the discipline of pooling findings across studies into a single rated conclusion - adapted from GRADE and systematic review practice for product research.

Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement

Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.

P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)

Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.

Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)

Publication bias is not an academic curiosity. In product research it is worse, because nobody rejects your null study - you simply never write it up. Learn how big the file drawer is, what it does to your confidence, and how to build a study register that closes it.

Quantitative User Research: Methods, Examples, and When to Use Them

A complete pillar guide to quantitative user research — the 9 core methods (surveys, A/B testing, analytics, tree testing, SUS, and more), when to use each, sample size rules, and how AI is bridging quant and qual.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.