{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T17:48:38.950Z"},"content":[{"type":"documentation","id":"3c45e6c3-aa5c-493a-b7a2-8915f500b61d","slug":"judge-diversity-group-accuracy-research","title":"Why a More Accurate Reviewer Can Make Your Panel Worse (2026)","url":"https://www.koji.so/docs/judge-diversity-group-accuracy-research","summary":"The diversity prediction theorem is an exact identity: a group's squared collective error equals average individual squared error minus predictive diversity. So accuracy and disagreement contribute equally to group performance. In a worked panel estimating 40 percent adoption, three reviewers sharing a low bias (22, 28, 34) produce collective error 144; adding the most accurate candidate (36) only reduces it to 100, while adding the least accurate (64) cuts it to 9, about 94 percent, despite raising average individual error from 168 to 270. Hong and Page (PNAS, 2004) showed randomly selected teams can beat best-performer teams because top performers converge in approach. Condorcet's jury theorem requires independent voters with competence above one half; correlated judgements from shared framing, anchoring or deference shrink the diversity term and remove the benefit.","content":"**Short answer:** a panel's error is not the average of its members' errors. It is the average individual error minus the panel's internal disagreement, exactly, as an algebraic identity. That means adding your most accurate reviewer can make the group estimate worse, and adding your least accurate one can make it dramatically better, if the second person is wrong in a direction nobody else was. Hong and Page proved the surprising half of this in the Proceedings of the National Academy of Sciences in 2004: \"a team of randomly selected agents outperforms a team comprised of the best-performing agents.\"\n\nIf you have ever picked the three most experienced people for a review panel, you optimised the wrong quantity.\n\n## The identity that governs every panel\n\nScott Page's diversity prediction theorem states it in one line. In his formulation, \"The squared error of the collective prediction equals the average squared error minus the predictive diversity.\"\n\nWritten out, with a true value and a set of individual estimates:\n\n- **Collective error** is how wrong the group's average estimate is, squared.\n- **Average individual error** is how wrong the typical member is, squared, averaged over members.\n- **Predictive diversity** is how much members disagree with each other, measured as the average squared distance from their own group mean.\n\nThis is not a tendency or an empirical regularity. It is an identity, true for every set of numbers anyone has ever produced. And it has an immediate consequence that most panel design ignores: there are two ways to make a group more accurate, and recruiting more accurate individuals is only one of them. Increasing disagreement is the other, and it is worth exactly as much, point for point.\n\n## A worked panel where the best reviewer makes things worse\n\nSuppose four reviewers are each estimating what percentage of users will adopt a new feature. The truth turns out to be 40 percent.\n\nStart with three reviewers who are all experienced, all sensible, and all anchored on the same internal assumption that adoption will be low. They say 22, 28 and 34.\n\n| Panel | Estimates | Group estimate | Collective error | Average individual error | Predictive diversity |\n| --- | --- | --- | --- | --- | --- |\n| Base three | 22, 28, 34 | 28 | 144 | 168 | 24 |\n| Plus the most accurate candidate (36) | 22, 28, 34, 36 | 30 | 100 | 130 | 30 |\n| Plus the least accurate candidate (64) | 22, 28, 34, 64 | 37 | 9 | 270 | 261 |\n\nCheck the identity in each row: 168 minus 24 is 144, 130 minus 30 is 100, and 270 minus 261 is 9. The table validates its own construction, which is the useful property of an exact identity.\n\nNow read the two candidates. The first candidate guesses 36, which is off by 4 and makes them the single most accurate person in the room, better than all three incumbents. The second guesses 64, which is off by 24 and makes them the worst forecaster on the panel by a wide margin.\n\nHiring the accurate one improves the group estimate from 12 points too low to 10 points too low. Hiring the bad one improves it from 12 points too low to 3 points too low, cutting collective error from 144 to 9, a reduction of about 94 percent. It does this while making average individual error substantially worse, from 168 to 270.\n\nThe mechanism is not mysterious once the identity is in view. The three incumbents shared a bias. Adding a fourth person who shared it slightly less did almost nothing. Adding a person who erred hard in the opposite direction supplied the disagreement that cancelled the shared error. The panel got worse at forecasting and better at estimating, at the same time.\n\n## Why the best performers are the most redundant\n\nThe uncomfortable part of Hong and Page's result is the reason it happens. Their explanation is structural: \"as the initial pool of problem solvers becomes large, the best-performing agents necessarily become similar in the space of problem solvers.\" Selecting for ability selects for a particular way of being right, and the more candidates you screen, the more tightly your finalists converge on it.\n\nTheir verdict on the best-performing team is blunt: \"Their relatively greater ability is more than offset by their lack of problem-solving diversity.\"\n\nTranslate that into research operations. Your most senior researchers went through similar training, read similar sources, and have internalised similar heuristics about what customers do. Each is individually more accurate than a junior colleague. Collectively they are close to being one reviewer consulted three times, and the identity says a panel of near-duplicates has almost no diversity term to subtract, so its collective error is stuck near its average individual error.\n\nThis is also why the surprisingly popular literature dismisses plain voting. Prelec, Seung and McCoy note that democratic methods \"are biased for shallow, lowest common denominator information, at the expense of novel or specialized knowledge that is not widely shared.\" A panel of similar experts is a small, expensive democracy with exactly that bias, and [scoring answers you cannot verify](/docs/incentive-compatible-honest-answers-research) is the companion move for recovering the informed minority view.\n\n## Independence is the load-bearing assumption\n\nCondorcet's jury theorem is the older, binary version of the same story, and it is explicit about the condition. It assumes \"each voter has an independent probability p of voting for the correct decision\", and then: \"If p is greater than 1/2, then adding more voters increases the probability that the majority decision is correct.\"\n\nThe theorem also has a failure branch that deserves more attention than it gets: \"if p is less than 1/2, then adding more voters makes things worse: the optimal jury consists of a single voter.\" A panel whose members are worse than chance on your question is actively harmed by being a panel. Scale amplifies whatever competence you started with, in either direction.\n\nBut independence is the assumption that actually breaks in practice, and it breaks in ordinary, well-intentioned ways:\n\n- Reviewers read the same research summary before scoring, so they inherit the same framing.\n- The first person to speak in the synthesis meeting anchors everyone else.\n- All reviewers sat in the same customer calls, so their private information is the same information.\n- One reviewer is known to be the expert, so the others defer.\n\nEvery one of these converts independent judgements into correlated ones, which shrinks the diversity term toward zero and quietly removes the benefit you built the panel for. Averaging correlated judgements gives you the look of a panel with the accuracy of one person.\n\n## What to do instead of recruiting the best\n\nThe identity implies a different recruiting rule: pick for complementary error, not for individual accuracy.\n\n1. **Collect estimates before any discussion.** This is the single highest-value change, because discussion is the main destroyer of independence. Get every number in writing first, then talk.\n2. **Recruit people whose information sources differ.** A support lead, a salesperson and a researcher will be wrong about adoption in three different directions. Three researchers will be wrong in one.\n3. **Measure the disagreement and report it.** Predictive diversity is computable from the estimates you already collected. If it is near zero, your panel is decoration, and you should say so.\n4. **Do not resolve disagreement too early.** A [Delphi process](/docs/delphi-method-guide) deliberately runs multiple rounds, and the rounds are valuable precisely because they preserve independent judgement before converging.\n5. **Keep the outlier's number in the average.** The worked table above is what deleting an outlier costs you. The instinct to drop the person who said 64 would have thrown away the entire benefit.\n\n## How Koji handles this\n\nIndependence is an operational property, not an attitude, and it is mostly destroyed by scheduling and sequencing. That is where Koji helps.\n\n- **AI-moderated interviews are structurally independent.** Every participant is interviewed separately by Koji's AI consultant, so no participant hears another's answer first. The correlated-judgement problem that wrecks panels does not arise in the raw data.\n- **Structured questions make diversity measurable.** With scale and ranking questions, the spread across respondents is a number you can compute rather than an impression. Koji supports six types, open_ended, scale, single_choice, multiple_choice, ranking and yes_no, and the four non-open types all produce the distributions this identity needs.\n- **Automatic thematic analysis reports the spread, not just the headline.** Koji's real-time reporting shows the distribution of scale answers rather than collapsing them to a mean, so a bimodal panel does not get averaged into a false consensus.\n- **Customizable AI consultants let you run the same brief past different populations** without a moderator drifting between them. That is how you get genuinely different error directions instead of one house view repeated.\n- **Traceability back to the transcript** means you can inspect the dissenting cluster instead of deleting it, which is the practical form of keeping the outlier in the average.\n\nKoji does not make your reviewers smarter. It makes their judgements independent by default and their disagreement visible, which the identity says is worth exactly as much.\n\n## Common mistakes\n\n- **Screening a panel for accuracy alone.** You are maximising one term and ignoring the other, and the ignored term is often larger.\n- **Discussing before collecting.** A pre-meeting Slack thread can zero out your diversity term before anyone writes a number down.\n- **Dropping outliers as errors.** Sometimes they are errors. But an outlier is the only thing that can cancel a shared bias, so removing it needs a reason beyond being far from the others.\n- **Assuming more reviewers is always better.** Condorcet says scale helps only when members are better than chance, and the diversity identity says it helps only when they are not duplicates.\n- **Reporting the group mean without the spread.** A mean of 30 from estimates of 29, 30, 31 and a mean of 30 from estimates of 5, 20, 40, 55 are completely different findings, and only one of them should change your roadmap. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) for the qualitative analogue.\n\n## Frequently asked questions\n\n### Does this mean I should recruit less capable reviewers?\n\nNo. It means capability and diversity are two separate contributions and you should stop buying only the first. The ideal addition is someone who is both accurate and different. When you have to choose, the identity tells you how to compare them: a candidate who adds more predictive diversity than they add average error will improve the group estimate, even if they are individually the weakest person on the panel.\n\n### How do I measure predictive diversity in practice?\n\nCollect every reviewer's numeric estimate independently, compute the group mean, then take each estimate's squared distance from that mean and average those. That number is your predictive diversity, and it is on the same scale as your error terms, so you can compare them directly. With Koji, scale and ranking questions give you the per-respondent values you need without manual collation.\n\n### Does the theorem apply to qualitative themes or only to numbers?\n\nThe identity itself is arithmetic and needs numeric estimates. The underlying lesson transfers to qualitative work intact: a coding team whose members interpret transcripts the same way has no diversity term, so it will reproduce a shared misreading with high confidence and high agreement. High agreement between similar coders is not evidence of accuracy, which is why inter-rater reliability is a check on consistency rather than on truth.\n\n### How many judges do I need before diversity matters?\n\nIt matters at three. The worked example in this article uses a base panel of three and a single addition, and the collective error moves by a factor of 16. What changes with larger panels is stability rather than relevance: with more members, both the average error and the diversity term are estimated more precisely, so the identity becomes a more reliable guide to whether an addition will help.\n\n### What breaks the theorem?\n\nNothing breaks the identity, because it is algebra. What breaks the benefit is correlation between judgements. If reviewers share information, framing or deference, their estimates converge, predictive diversity shrinks toward zero, and collective error rises to meet average individual error. Condorcet's version fails in a second way too: if members are individually worse than chance, adding members makes the majority verdict worse rather than better.\n\n### How does Koji help me keep judgements independent?\n\nKoji interviews every participant separately with an AI consultant, so nobody is anchored by hearing somebody else's answer, and the AI does not drift the way a human moderator does between sessions. Structured scale and ranking questions then preserve the full distribution of answers in Koji's reporting instead of collapsing it to an average, so you can see whether the independence you designed for actually produced disagreement.\n\n## Related Resources\n\n- [Structured Questions: The Complete Guide](/docs/structured-questions-guide) - the scale and ranking question types that make diversity computable\n- [The Delphi Method](/docs/delphi-method-guide) - a process built to preserve independent judgement across rounds\n- [Scoring Answers You Cannot Verify](/docs/incentive-compatible-honest-answers-research) - recovering the informed minority view that voting discards\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - why coder agreement measures consistency, not accuracy\n- [Calibration Scoring for Research Teams](/docs/research-calibration-brier-score) - scoring individual forecasters once you have collected their estimates\n- [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) - sample sizing for discovery, a different question from panel composition\n","category":"Analysis & Synthesis","lastModified":"2026-09-28T04:00:44.319912+00:00","metaTitle":"The Diversity Prediction Theorem: Why Your Best Reviewer Can Hurt a Panel (2026)","metaDescription":"Collective error equals average individual error minus predictive diversity. A worked panel where the worst forecaster cuts error 94 percent.","keywords":["diversity prediction theorem","wisdom of crowds research","condorcet jury theorem","group accuracy","judge diversity","panel composition research"],"aiSummary":"The diversity prediction theorem is an exact identity: a group's squared collective error equals average individual squared error minus predictive diversity. So accuracy and disagreement contribute equally to group performance. In a worked panel estimating 40 percent adoption, three reviewers sharing a low bias (22, 28, 34) produce collective error 144; adding the most accurate candidate (36) only reduces it to 100, while adding the least accurate (64) cuts it to 9, about 94 percent, despite raising average individual error from 168 to 270. Hong and Page (PNAS, 2004) showed randomly selected teams can beat best-performer teams because top performers converge in approach. Condorcet's jury theorem requires independent voters with competence above one half; correlated judgements from shared framing, anchoring or deference shrink the diversity term and remove the benefit.","aiPrerequisites":["Comfort with squared error as a measure of accuracy","Experience running review panels or synthesis sessions"],"aiLearningOutcomes":["State the diversity prediction theorem and verify it arithmetically","Predict whether adding a given reviewer will help or hurt a panel","Identify the practices that destroy judgement independence","Recruit review panels for complementary error rather than individual accuracy"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}