{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-12T14:17:41.039Z"},"content":[{"type":"documentation","id":"64be75dc-fb6a-49d9-b203-d0402ac946e0","slug":"research-calibration-brier-score","title":"Calibration Scoring for Research Teams: How to Find Out If Your Insights Were Actually Right (2026)","url":"https://www.koji.so/docs/research-calibration-brier-score","summary":"Research is graded on process and almost never on whether its conclusions came true. Proper scoring rules from forecasting, principally the Brier score, assign a number to a probabilistic claim once reality settles it, and decompose a track record into calibration and resolution. Evidence from geopolitical forecasting tournaments shows accuracy is trainable; measurement of 1,514 real intelligence forecasts shows the dominant error is underconfidence, worst on the highest-stakes questions. The guide covers making research claims scoreable and running the programme without corrupting it.","content":"\n\n**Answer first: research teams are evaluated on whether their method was sound, never on whether their conclusions turned out to be correct. Forecasting has spent seventy-five years solving exactly this problem, using proper scoring rules - the Brier score chief among them - that assign a number to a probabilistic claim once reality settles it. Adopting them requires one uncomfortable change: your research claims have to be stated precisely enough to be graded, which means abandoning the hedged language that currently protects them.**\n\nThis is the capstone discipline. A pre-specified [stopping rule](/docs/interim-analysis-stopping-rules-research) tells you when to end a study. A [matched control group](/docs/case-control-research-churn-lost-deals) tells you whether a comparison is fair. Neither tells you whether the answer was right. Only scoring does.\n\n## Process quality and outcome accuracy are different things\n\nEvery quality mechanism in research is a process check. Was the sample appropriate? Was the guide well-constructed? Was the analysis defensible? Was it peer reviewed? These are worth doing, and our guides to [blind analysis](/docs/blind-analysis-research) and [the many-analysts problem](/docs/many-analysts-one-dataset) cover why process discipline matters more than most teams assume.\n\nBut notice what process checks cannot do. They cannot distinguish a team that is right 80 percent of the time from a team that is right 50 percent of the time, because both can follow good process. And a research function that has never measured its own hit rate has no evidence for its central claim - that acting on its findings produces better outcomes than not.\n\nThe reason this gap persists is not laziness. It is that outcome grading requires a record of what you predicted, in a form specific enough to be scored, made before the outcome was known. Almost no research team keeps one.\n\n## Enter the proper scoring rule\n\nThe foundational move came from weather forecasting. Glenn Brier, in **\"Verification of forecasts expressed in terms of probability\" (*Monthly Weather Review* 78(1):1-3, 1950)**, proposed scoring a probabilistic forecast by the squared difference between the stated probability and what actually happened, coded as 1 or 0.\n\nFor a single binary claim:\n\n**Brier score = (probability you stated - outcome)^2**\n\nLower is better. A perfect confident call scores 0. A maximally wrong confident call scores 1. Saying 50 percent always scores 0.25, whatever happens.\n\n| Your stated probability | Outcome | Brier score | Reading |\n| --- | --- | --- | --- |\n| 0.90 | Happened (1) | 0.01 | Confident and right |\n| 0.90 | Did not happen (0) | 0.81 | Confident and wrong - very expensive |\n| 0.60 | Happened (1) | 0.16 | Mildly useful |\n| 0.50 | Either | 0.25 | No information, and no risk |\n| 0.20 | Did not happen (0) | 0.04 | Confident against, and right |\n\nThe critical property is that the Brier score is a **proper** scoring rule: your expected score is best when you report your true belief. You cannot game it by hedging toward 50 percent, because 0.25 is a mediocre score you are guaranteed to earn forever. You also cannot game it by overclaiming, because a wrong 0.95 is punished savagely. This is what makes it usable as a team metric rather than a target people learn to manipulate.\n\nAveraged over many claims, the Brier score decomposes into two components worth tracking separately:\n\n- **Calibration** - when you say 70 percent, does it happen about 70 percent of the time? This is about honesty of confidence.\n- **Resolution** - do you say different things about different questions, or is every answer near the base rate? This is about informativeness.\n\nA team can be perfectly calibrated and useless. Predicting the base rate on every question is well calibrated and tells nobody anything. **You need both, and only tracking them separately shows you which one you lack.**\n\n## The evidence that this can be taught\n\nThe most relevant body of evidence comes from a forecasting tournament sponsored by the US intelligence community, in which five university research groups competed to elicit and aggregate accurate probability estimates for geopolitical events.\n\n**Mellers and colleagues (*Psychological Science* 25(5):1106-1115, 2014)** reported three interventions that worked, and none of them was a better algorithm. **Probability training** corrected cognitive biases, encouraged forecasters to use reference classes, and supplied heuristics such as averaging multiple estimates. **Teaming** let forecasters share information and argue about rationales. **Tracking** placed the top 2 percent of performers from year one into elite teams. All three improved both calibration and resolution. The authors framed forecasting as commonly viewed as a statistical problem but improvable through behavioural intervention.\n\n**Mellers and colleagues (*Perspectives on Psychological Science* 10(3):267-281, 2015)** followed the top performers and found that, defying the expectation of regression toward the mean two years running, superforecasters maintained high accuracy across hundreds of questions and a wide range of topics. Their explanation combined cognitive style, task-specific skill, motivation, and an enriched environment - concluding that superforecasters are partly discovered and partly created.\n\nThe transferable claim for a research function is narrow but strong: **accuracy at probabilistic judgement is a trainable skill, the training is cheap and mostly conceptual, and it only works if scores exist.** You cannot train what you do not measure.\n\n## The finding that inverts the usual story\n\nIf you expect the punchline to be \"experts are overconfident,\" the best real-world measurement says otherwise, and this is the most useful single result in the literature for a research audience.\n\n**Mandel and Barnes (*Proceedings of the National Academy of Sciences* 111(30):10984-10989, 2014)** scored **1,514 strategic intelligence forecasts** abstracted from real intelligence reports produced by a working assessment unit - not a laboratory exercise. Both discrimination and calibration were very good. Discrimination was better for senior analysts than junior ones, and better on easier questions.\n\nThe miscalibration that did exist ran in the opposite direction to the stereotype. **It was mainly underconfidence: analysts assigned more uncertainty than was warranted given how well they actually discriminated.** Underconfidence was *more* pronounced on harder forecasts and on forecasts deemed more important for policy decisions. Despite this, there was a shortage of forecasts in the least informative 0.4 to 0.6 band. Simply recalibrating the forecasts substantially reduced the underconfidence.\n\nTranslate that into research practice and it describes a familiar pathology precisely. **The higher the stakes, the more a research team hedges.** \"Directionally, this suggests users may prefer...\" is the underconfident forecast, produced most reliably on exactly the questions where a clear answer is worth the most. Hedging feels like intellectual honesty. Measured against outcomes, it is a systematic error - and unlike overconfidence, nobody ever gets criticised for it, which is why it persists.\n\nThere is a second cost. **A hedged claim cannot be scored, so it cannot be learned from.** A team that never commits to a number never generates the record that would let it improve, which makes hedging self-perpetuating.\n\n## Making a research claim scoreable\n\nThe mechanics are the easy part. The discipline is writing the claim before the outcome and resisting the urge to soften it.\n\n| Typical research claim | Scoreable version |\n| --- | --- |\n| \"Users found the new onboarding confusing\" | \"70 percent confident that day-7 activation for the new onboarding is below the current flow when measured at the end of Q3\" |\n| \"There is strong demand for the integration\" | \"80 percent confident that at least 15 percent of enterprise accounts enable the integration within 90 days of GA\" |\n| \"Price is the main driver of churn\" | \"60 percent confident that the pricing change reduces gross monthly churn by at least 0.5 points within two quarters\" |\n| \"This concept tested well\" | \"75 percent confident this variant beats control on the primary metric in the follow-up experiment\" |\n\nFour requirements make a claim scoreable:\n\n1. **A probability**, not a hedge word. \"Likely\" is not a number, and different readers translate it into wildly different numbers.\n2. **Resolution criteria** fixed in advance - the metric, the threshold, and the measurement window, agreed before the outcome is known.\n3. **A resolution date.** Unresolved forecasts accumulate and quietly become the ones you never grade, which biases your record toward the questions that settled fast.\n4. **A written record with a timestamp**, so the claim cannot drift. Your [research repository](/docs/research-repository-guide) is the natural home; the claim belongs attached to the study that generated it.\n\n### Two habits worth importing with the scoring\n\n**The outside view.** Before estimating from the specifics of your case, ask what happened in the reference class of similar cases. What share of features like this one hit their adoption target? Probability training in the tournaments explicitly taught this move, and it is the cheapest correction available because teams reliably estimate from the vivid particulars of the case in front of them.\n\n**Averaging independent estimates.** Have three people estimate before they discuss. The average of independent estimates is usually better than any individual estimate and better than the number the group converges on after the most senior person speaks first. This is the same logic that makes the [Delphi method](/docs/delphi-method-guide) work.\n\n## Running it: a programme you can start this quarter\n\n1. **Attach a forecast to every study that informs a decision.** One line: probability, metric, threshold, date. Do this at readout, before anyone acts.\n2. **Log it where it cannot be edited.** Timestamped, with the resolution criteria stated. Track predicted-versus-actual as a standing metric, as covered in [activating research insights](/docs/activating-research-insights).\n3. **Resolve on schedule, including the awkward ones.** Grade every forecast whose date has passed, not the ones you remember fondly. Selectively resolving is the [publication bias](/docs/publication-bias-product-research) problem applied to your own track record.\n4. **Review the decomposition quarterly.** Poor calibration means your confidence language is wrong. Poor resolution means you are hedging to the base rate and adding nothing.\n5. **Do not attach the score to individual performance reviews.** The moment a Brier score affects someone's rating, hedging becomes rational and the record stops being informative. Score the function, not the person.\n\nThat last point is the difference between a programme that survives a year and one that is quietly abandoned.\n\n## The modern approach: how Koji helps\n\nOutcome scoring lives or dies on whether the original claim was recorded in a specific, retrievable form. The main reason teams cannot grade their past research is that the findings were prose in a slide deck, and prose is unfalsifiable by default.\n\n**Structured questions produce claims with stable identity.** Koji supports six question types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - and every one carries a stable question ID from the interview plan through moderation and analysis into the report. A finding anchored to an identified item (\"62 percent selected this option in the `multiple_choice` item on switching triggers\") is a fact you can return to in six months and check. A finding that exists only as a sentence in a deck is not. The [structured questions guide](/docs/structured-questions-guide) covers how each type aggregates.\n\n**Quantified findings make thresholds writable.** Koji reports aggregate `scale` distributions and `ranking` positions across respondents rather than leaving you to characterise a mood. That is what lets a readout state a threshold instead of a direction, which is the precondition for a scoreable forecast.\n\n**Speed makes the feedback loop short enough to learn from.** Calibration training works because forecasters get scored repeatedly. A research function that ships four studies a year generates four data points, which is not enough to detect miscalibration in a working lifetime. A function running weekly AI-moderated studies generates enough resolved forecasts within a year for the decomposition into calibration and resolution to mean something. This is the strongest practical argument for research velocity, and it is not the one usually made.\n\n**Re-running the same instrument closes the loop.** Because a study definition and its question IDs persist, you can field the identical instrument after the change ships and compare like with like. The follow-up measurement is what resolves the forecast, and it costs a fraction of the original study.\n\nKoji does not tell you whether you were right. It makes the original claim specific, retrievable, and cheap to re-measure - which is the entire infrastructure requirement for finding out.\n\n## Frequently asked questions\n\n### What is a Brier score in plain terms?\n\nIt is the squared difference between the probability you stated and what actually happened, scored as 1 or 0. Say 90 percent and be right and you score 0.01; say 90 percent and be wrong and you score 0.81. Lower is better, and it is averaged over many forecasts to grade a track record.\n\n### Why not just track how often the team was right?\n\nBecause a hit rate throws away the confidence information. Being right on calls you made at 55 percent is very different from being right on calls you made at 95 percent, and a hit rate cannot tell them apart. A proper scoring rule prices confidence, which is what makes hedging unprofitable.\n\n### Will scoring make my team more conservative?\n\nNot if you use a proper scoring rule and keep it off individual performance reviews. The Brier score is specifically designed so hedging toward 50 percent scores mediocre forever. The real-world risk runs the other way: Mandel and Barnes found professional analysts erred toward underconfidence, most on the highest-stakes questions.\n\n### What is the difference between calibration and resolution?\n\nCalibration asks whether things you call 70 percent happen about 70 percent of the time. Resolution asks whether you say meaningfully different things about different questions or just repeat the base rate. A team can be perfectly calibrated and useless, so track both separately.\n\n### How many forecasts do I need before the score means anything?\n\nEnough that a single lucky call cannot dominate - practically, a few dozen resolved forecasts before reading the decomposition seriously. This is why research cadence matters: four studies a year will not produce a usable record in any reasonable timeframe.\n\n### Can forecasting accuracy actually be improved, or is it a fixed trait?\n\nIt can be improved. In the tournament research, probability training, team collaboration, and tracking top performers all improved both calibration and resolution, and top performers sustained their accuracy across two years rather than regressing to the mean. The training is mostly conceptual - reference classes, averaging independent estimates - and cheap.\n\n## Related Resources\n\n- [Interim Analysis and Stopping Rules](/docs/interim-analysis-stopping-rules-research) - deciding in advance when a study ends, so the claim it produces is not selected on its own result\n- [Case-Control Research for Churn and Lost Deals](/docs/case-control-research-churn-lost-deals) - making a backwards-looking comparison fair enough to be worth scoring\n- [Blind Analysis](/docs/blind-analysis-research) - process discipline that protects the analysis, which scoring alone cannot supply\n- [Publication Bias in Product Research](/docs/publication-bias-product-research) - why selectively resolving your own forecasts recreates the file-drawer problem\n- [Activating Research Insights](/docs/activating-research-insights) - tracking predicted versus actual outcomes as a standing research metric\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types that turn a finding into something you can re-measure\n","category":"Research Methods","lastModified":"2026-08-12T03:24:15.27832+00:00","metaTitle":"Calibration and Brier Scores for Research Teams (2026)","metaDescription":"Score your research team on whether its claims came true, not just whether the method was sound. Brier scores, calibration versus resolution, and the underconfidence finding that inverts the usual story.","keywords":["brier score","calibration research team","forecasting accuracy","proper scoring rule","resolution criteria","superforecasting","reference class forecasting","research track record","probabilistic judgment","outcome accuracy"],"aiSummary":"Research is graded on process and almost never on whether its conclusions came true. Proper scoring rules from forecasting, principally the Brier score, assign a number to a probabilistic claim once reality settles it, and decompose a track record into calibration and resolution. Evidence from geopolitical forecasting tournaments shows accuracy is trainable; measurement of 1,514 real intelligence forecasts shows the dominant error is underconfidence, worst on the highest-stakes questions. The guide covers making research claims scoreable and running the programme without corrupting it.","aiPrerequisites":["Experience presenting research findings to decision makers","Basic comfort with probability"],"aiLearningOutcomes":["Compute and interpret a Brier score for a research claim","Separate calibration from resolution in a track record","Rewrite a hedged research finding as a scoreable forecast","Set resolution criteria and dates that prevent selective grading","Run a team scoring programme without incentivising hedging"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}