Why a More Accurate Reviewer Can Make Your Panel Worse (2026)
The diversity prediction theorem, worked through a real panel: collective error equals average individual error minus predictive diversity, and what that means for who you recruit.
Short answer: a panel's error is not the average of its members' errors. It is the average individual error minus the panel's internal disagreement, exactly, as an algebraic identity. That means adding your most accurate reviewer can make the group estimate worse, and adding your least accurate one can make it dramatically better, if the second person is wrong in a direction nobody else was. Hong and Page proved the surprising half of this in the Proceedings of the National Academy of Sciences in 2004: "a team of randomly selected agents outperforms a team comprised of the best-performing agents."
If you have ever picked the three most experienced people for a review panel, you optimised the wrong quantity.
The identity that governs every panel
Scott Page's diversity prediction theorem states it in one line. In his formulation, "The squared error of the collective prediction equals the average squared error minus the predictive diversity."
Written out, with a true value and a set of individual estimates:
- Collective error is how wrong the group's average estimate is, squared.
- Average individual error is how wrong the typical member is, squared, averaged over members.
- Predictive diversity is how much members disagree with each other, measured as the average squared distance from their own group mean.
This is not a tendency or an empirical regularity. It is an identity, true for every set of numbers anyone has ever produced. And it has an immediate consequence that most panel design ignores: there are two ways to make a group more accurate, and recruiting more accurate individuals is only one of them. Increasing disagreement is the other, and it is worth exactly as much, point for point.
A worked panel where the best reviewer makes things worse
Suppose four reviewers are each estimating what percentage of users will adopt a new feature. The truth turns out to be 40 percent.
Start with three reviewers who are all experienced, all sensible, and all anchored on the same internal assumption that adoption will be low. They say 22, 28 and 34.
| Panel | Estimates | Group estimate | Collective error | Average individual error | Predictive diversity |
|---|---|---|---|---|---|
| Base three | 22, 28, 34 | 28 | 144 | 168 | 24 |
| Plus the most accurate candidate (36) | 22, 28, 34, 36 | 30 | 100 | 130 | 30 |
| Plus the least accurate candidate (64) | 22, 28, 34, 64 | 37 | 9 | 270 | 261 |
Check the identity in each row: 168 minus 24 is 144, 130 minus 30 is 100, and 270 minus 261 is 9. The table validates its own construction, which is the useful property of an exact identity.
Now read the two candidates. The first candidate guesses 36, which is off by 4 and makes them the single most accurate person in the room, better than all three incumbents. The second guesses 64, which is off by 24 and makes them the worst forecaster on the panel by a wide margin.
Hiring the accurate one improves the group estimate from 12 points too low to 10 points too low. Hiring the bad one improves it from 12 points too low to 3 points too low, cutting collective error from 144 to 9, a reduction of about 94 percent. It does this while making average individual error substantially worse, from 168 to 270.
The mechanism is not mysterious once the identity is in view. The three incumbents shared a bias. Adding a fourth person who shared it slightly less did almost nothing. Adding a person who erred hard in the opposite direction supplied the disagreement that cancelled the shared error. The panel got worse at forecasting and better at estimating, at the same time.
Why the best performers are the most redundant
The uncomfortable part of Hong and Page's result is the reason it happens. Their explanation is structural: "as the initial pool of problem solvers becomes large, the best-performing agents necessarily become similar in the space of problem solvers." Selecting for ability selects for a particular way of being right, and the more candidates you screen, the more tightly your finalists converge on it.
Their verdict on the best-performing team is blunt: "Their relatively greater ability is more than offset by their lack of problem-solving diversity."
Translate that into research operations. Your most senior researchers went through similar training, read similar sources, and have internalised similar heuristics about what customers do. Each is individually more accurate than a junior colleague. Collectively they are close to being one reviewer consulted three times, and the identity says a panel of near-duplicates has almost no diversity term to subtract, so its collective error is stuck near its average individual error.
This is also why the surprisingly popular literature dismisses plain voting. Prelec, Seung and McCoy note that democratic methods "are biased for shallow, lowest common denominator information, at the expense of novel or specialized knowledge that is not widely shared." A panel of similar experts is a small, expensive democracy with exactly that bias, and scoring answers you cannot verify is the companion move for recovering the informed minority view.
Independence is the load-bearing assumption
Condorcet's jury theorem is the older, binary version of the same story, and it is explicit about the condition. It assumes "each voter has an independent probability p of voting for the correct decision", and then: "If p is greater than 1/2, then adding more voters increases the probability that the majority decision is correct."
The theorem also has a failure branch that deserves more attention than it gets: "if p is less than 1/2, then adding more voters makes things worse: the optimal jury consists of a single voter." A panel whose members are worse than chance on your question is actively harmed by being a panel. Scale amplifies whatever competence you started with, in either direction.
But independence is the assumption that actually breaks in practice, and it breaks in ordinary, well-intentioned ways:
- Reviewers read the same research summary before scoring, so they inherit the same framing.
- The first person to speak in the synthesis meeting anchors everyone else.
- All reviewers sat in the same customer calls, so their private information is the same information.
- One reviewer is known to be the expert, so the others defer.
Every one of these converts independent judgements into correlated ones, which shrinks the diversity term toward zero and quietly removes the benefit you built the panel for. Averaging correlated judgements gives you the look of a panel with the accuracy of one person.
What to do instead of recruiting the best
The identity implies a different recruiting rule: pick for complementary error, not for individual accuracy.
- Collect estimates before any discussion. This is the single highest-value change, because discussion is the main destroyer of independence. Get every number in writing first, then talk.
- Recruit people whose information sources differ. A support lead, a salesperson and a researcher will be wrong about adoption in three different directions. Three researchers will be wrong in one.
- Measure the disagreement and report it. Predictive diversity is computable from the estimates you already collected. If it is near zero, your panel is decoration, and you should say so.
- Do not resolve disagreement too early. A Delphi process deliberately runs multiple rounds, and the rounds are valuable precisely because they preserve independent judgement before converging.
- Keep the outlier's number in the average. The worked table above is what deleting an outlier costs you. The instinct to drop the person who said 64 would have thrown away the entire benefit.
How Koji handles this
Independence is an operational property, not an attitude, and it is mostly destroyed by scheduling and sequencing. That is where Koji helps.
- AI-moderated interviews are structurally independent. Every participant is interviewed separately by Koji's AI consultant, so no participant hears another's answer first. The correlated-judgement problem that wrecks panels does not arise in the raw data.
- Structured questions make diversity measurable. With scale and ranking questions, the spread across respondents is a number you can compute rather than an impression. Koji supports six types, open_ended, scale, single_choice, multiple_choice, ranking and yes_no, and the four non-open types all produce the distributions this identity needs.
- Automatic thematic analysis reports the spread, not just the headline. Koji's real-time reporting shows the distribution of scale answers rather than collapsing them to a mean, so a bimodal panel does not get averaged into a false consensus.
- Customizable AI consultants let you run the same brief past different populations without a moderator drifting between them. That is how you get genuinely different error directions instead of one house view repeated.
- Traceability back to the transcript means you can inspect the dissenting cluster instead of deleting it, which is the practical form of keeping the outlier in the average.
Koji does not make your reviewers smarter. It makes their judgements independent by default and their disagreement visible, which the identity says is worth exactly as much.
Common mistakes
- Screening a panel for accuracy alone. You are maximising one term and ignoring the other, and the ignored term is often larger.
- Discussing before collecting. A pre-meeting Slack thread can zero out your diversity term before anyone writes a number down.
- Dropping outliers as errors. Sometimes they are errors. But an outlier is the only thing that can cancel a shared bias, so removing it needs a reason beyond being far from the others.
- Assuming more reviewers is always better. Condorcet says scale helps only when members are better than chance, and the diversity identity says it helps only when they are not duplicates.
- Reporting the group mean without the spread. A mean of 30 from estimates of 29, 30, 31 and a mean of 30 from estimates of 5, 20, 40, 55 are completely different findings, and only one of them should change your roadmap. See inter-rater reliability for the qualitative analogue.
Frequently asked questions
Does this mean I should recruit less capable reviewers?
No. It means capability and diversity are two separate contributions and you should stop buying only the first. The ideal addition is someone who is both accurate and different. When you have to choose, the identity tells you how to compare them: a candidate who adds more predictive diversity than they add average error will improve the group estimate, even if they are individually the weakest person on the panel.
How do I measure predictive diversity in practice?
Collect every reviewer's numeric estimate independently, compute the group mean, then take each estimate's squared distance from that mean and average those. That number is your predictive diversity, and it is on the same scale as your error terms, so you can compare them directly. With Koji, scale and ranking questions give you the per-respondent values you need without manual collation.
Does the theorem apply to qualitative themes or only to numbers?
The identity itself is arithmetic and needs numeric estimates. The underlying lesson transfers to qualitative work intact: a coding team whose members interpret transcripts the same way has no diversity term, so it will reproduce a shared misreading with high confidence and high agreement. High agreement between similar coders is not evidence of accuracy, which is why inter-rater reliability is a check on consistency rather than on truth.
How many judges do I need before diversity matters?
It matters at three. The worked example in this article uses a base panel of three and a single addition, and the collective error moves by a factor of 16. What changes with larger panels is stability rather than relevance: with more members, both the average error and the diversity term are estimated more precisely, so the identity becomes a more reliable guide to whether an addition will help.
What breaks the theorem?
Nothing breaks the identity, because it is algebra. What breaks the benefit is correlation between judgements. If reviewers share information, framing or deference, their estimates converge, predictive diversity shrinks toward zero, and collective error rises to meet average individual error. Condorcet's version fails in a second way too: if members are individually worse than chance, adding members makes the majority verdict worse rather than better.
How does Koji help me keep judgements independent?
Koji interviews every participant separately with an AI consultant, so nobody is anchored by hearing somebody else's answer, and the AI does not drift the way a human moderator does between sessions. Structured scale and ranking questions then preserve the full distribution of answers in Koji's reporting instead of collapsing it to an average, so you can see whether the independence you designed for actually produced disagreement.
Related Resources
- Structured Questions: The Complete Guide - the scale and ranking question types that make diversity computable
- The Delphi Method - a process built to preserve independent judgement across rounds
- Scoring Answers You Cannot Verify - recovering the informed minority view that voting discards
- Inter-Rater Reliability in Qualitative Research - why coder agreement measures consistency, not accuracy
- Calibration Scoring for Research Teams - scoring individual forecasters once you have collected their estimates
- How Many Interviews Are Enough? - sample sizing for discovery, a different question from panel composition
Related Articles
The Delphi Method: A Complete Guide to Reaching Expert Consensus
A practical guide to the Delphi method — the structured, multi-round technique for building expert consensus through anonymous questionnaires and controlled feedback. Learn the process, panel size, rounds, and modern AI-assisted alternatives.
How Many Interviews Are Enough? A Guide to Sample Size
Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.
Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement
Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.
Calibration Scoring for Research Teams: How to Find Out If Your Insights Were Actually Right (2026)
Research is graded on process and almost never on outcome. Forecasting tournaments solved this with proper scoring rules. Here is how to score a research team on whether its claims came true.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.