Back to docs
Research Methods

Calibration Scoring for Research Teams: How to Find Out If Your Insights Were Actually Right (2026)

Research is graded on process and almost never on outcome. Forecasting tournaments solved this with proper scoring rules. Here is how to score a research team on whether its claims came true.

Answer first: research teams are evaluated on whether their method was sound, never on whether their conclusions turned out to be correct. Forecasting has spent seventy-five years solving exactly this problem, using proper scoring rules - the Brier score chief among them - that assign a number to a probabilistic claim once reality settles it. Adopting them requires one uncomfortable change: your research claims have to be stated precisely enough to be graded, which means abandoning the hedged language that currently protects them.

This is the capstone discipline. A pre-specified stopping rule tells you when to end a study. A matched control group tells you whether a comparison is fair. Neither tells you whether the answer was right. Only scoring does.

Process quality and outcome accuracy are different things

Every quality mechanism in research is a process check. Was the sample appropriate? Was the guide well-constructed? Was the analysis defensible? Was it peer reviewed? These are worth doing, and our guides to blind analysis and the many-analysts problem cover why process discipline matters more than most teams assume.

But notice what process checks cannot do. They cannot distinguish a team that is right 80 percent of the time from a team that is right 50 percent of the time, because both can follow good process. And a research function that has never measured its own hit rate has no evidence for its central claim - that acting on its findings produces better outcomes than not.

The reason this gap persists is not laziness. It is that outcome grading requires a record of what you predicted, in a form specific enough to be scored, made before the outcome was known. Almost no research team keeps one.

Enter the proper scoring rule

The foundational move came from weather forecasting. Glenn Brier, in "Verification of forecasts expressed in terms of probability" (Monthly Weather Review 78(1):1-3, 1950), proposed scoring a probabilistic forecast by the squared difference between the stated probability and what actually happened, coded as 1 or 0.

For a single binary claim:

Brier score = (probability you stated - outcome)^2

Lower is better. A perfect confident call scores 0. A maximally wrong confident call scores 1. Saying 50 percent always scores 0.25, whatever happens.

Your stated probabilityOutcomeBrier scoreReading
0.90Happened (1)0.01Confident and right
0.90Did not happen (0)0.81Confident and wrong - very expensive
0.60Happened (1)0.16Mildly useful
0.50Either0.25No information, and no risk
0.20Did not happen (0)0.04Confident against, and right

The critical property is that the Brier score is a proper scoring rule: your expected score is best when you report your true belief. You cannot game it by hedging toward 50 percent, because 0.25 is a mediocre score you are guaranteed to earn forever. You also cannot game it by overclaiming, because a wrong 0.95 is punished savagely. This is what makes it usable as a team metric rather than a target people learn to manipulate.

Averaged over many claims, the Brier score decomposes into two components worth tracking separately:

  • Calibration - when you say 70 percent, does it happen about 70 percent of the time? This is about honesty of confidence.
  • Resolution - do you say different things about different questions, or is every answer near the base rate? This is about informativeness.

A team can be perfectly calibrated and useless. Predicting the base rate on every question is well calibrated and tells nobody anything. You need both, and only tracking them separately shows you which one you lack.

The evidence that this can be taught

The most relevant body of evidence comes from a forecasting tournament sponsored by the US intelligence community, in which five university research groups competed to elicit and aggregate accurate probability estimates for geopolitical events.

Mellers and colleagues (Psychological Science 25(5):1106-1115, 2014) reported three interventions that worked, and none of them was a better algorithm. Probability training corrected cognitive biases, encouraged forecasters to use reference classes, and supplied heuristics such as averaging multiple estimates. Teaming let forecasters share information and argue about rationales. Tracking placed the top 2 percent of performers from year one into elite teams. All three improved both calibration and resolution. The authors framed forecasting as commonly viewed as a statistical problem but improvable through behavioural intervention.

Mellers and colleagues (Perspectives on Psychological Science 10(3):267-281, 2015) followed the top performers and found that, defying the expectation of regression toward the mean two years running, superforecasters maintained high accuracy across hundreds of questions and a wide range of topics. Their explanation combined cognitive style, task-specific skill, motivation, and an enriched environment - concluding that superforecasters are partly discovered and partly created.

The transferable claim for a research function is narrow but strong: accuracy at probabilistic judgement is a trainable skill, the training is cheap and mostly conceptual, and it only works if scores exist. You cannot train what you do not measure.

The finding that inverts the usual story

If you expect the punchline to be "experts are overconfident," the best real-world measurement says otherwise, and this is the most useful single result in the literature for a research audience.

Mandel and Barnes (Proceedings of the National Academy of Sciences 111(30):10984-10989, 2014) scored 1,514 strategic intelligence forecasts abstracted from real intelligence reports produced by a working assessment unit - not a laboratory exercise. Both discrimination and calibration were very good. Discrimination was better for senior analysts than junior ones, and better on easier questions.

The miscalibration that did exist ran in the opposite direction to the stereotype. It was mainly underconfidence: analysts assigned more uncertainty than was warranted given how well they actually discriminated. Underconfidence was more pronounced on harder forecasts and on forecasts deemed more important for policy decisions. Despite this, there was a shortage of forecasts in the least informative 0.4 to 0.6 band. Simply recalibrating the forecasts substantially reduced the underconfidence.

Translate that into research practice and it describes a familiar pathology precisely. The higher the stakes, the more a research team hedges. "Directionally, this suggests users may prefer..." is the underconfident forecast, produced most reliably on exactly the questions where a clear answer is worth the most. Hedging feels like intellectual honesty. Measured against outcomes, it is a systematic error - and unlike overconfidence, nobody ever gets criticised for it, which is why it persists.

There is a second cost. A hedged claim cannot be scored, so it cannot be learned from. A team that never commits to a number never generates the record that would let it improve, which makes hedging self-perpetuating.

Making a research claim scoreable

The mechanics are the easy part. The discipline is writing the claim before the outcome and resisting the urge to soften it.

Typical research claimScoreable version
"Users found the new onboarding confusing""70 percent confident that day-7 activation for the new onboarding is below the current flow when measured at the end of Q3"
"There is strong demand for the integration""80 percent confident that at least 15 percent of enterprise accounts enable the integration within 90 days of GA"
"Price is the main driver of churn""60 percent confident that the pricing change reduces gross monthly churn by at least 0.5 points within two quarters"
"This concept tested well""75 percent confident this variant beats control on the primary metric in the follow-up experiment"

Four requirements make a claim scoreable:

  1. A probability, not a hedge word. "Likely" is not a number, and different readers translate it into wildly different numbers.
  2. Resolution criteria fixed in advance - the metric, the threshold, and the measurement window, agreed before the outcome is known.
  3. A resolution date. Unresolved forecasts accumulate and quietly become the ones you never grade, which biases your record toward the questions that settled fast.
  4. A written record with a timestamp, so the claim cannot drift. Your research repository is the natural home; the claim belongs attached to the study that generated it.

Two habits worth importing with the scoring

The outside view. Before estimating from the specifics of your case, ask what happened in the reference class of similar cases. What share of features like this one hit their adoption target? Probability training in the tournaments explicitly taught this move, and it is the cheapest correction available because teams reliably estimate from the vivid particulars of the case in front of them.

Averaging independent estimates. Have three people estimate before they discuss. The average of independent estimates is usually better than any individual estimate and better than the number the group converges on after the most senior person speaks first. This is the same logic that makes the Delphi method work.

Running it: a programme you can start this quarter

  1. Attach a forecast to every study that informs a decision. One line: probability, metric, threshold, date. Do this at readout, before anyone acts.
  2. Log it where it cannot be edited. Timestamped, with the resolution criteria stated. Track predicted-versus-actual as a standing metric, as covered in activating research insights.
  3. Resolve on schedule, including the awkward ones. Grade every forecast whose date has passed, not the ones you remember fondly. Selectively resolving is the publication bias problem applied to your own track record.
  4. Review the decomposition quarterly. Poor calibration means your confidence language is wrong. Poor resolution means you are hedging to the base rate and adding nothing.
  5. Do not attach the score to individual performance reviews. The moment a Brier score affects someone's rating, hedging becomes rational and the record stops being informative. Score the function, not the person.

That last point is the difference between a programme that survives a year and one that is quietly abandoned.

The modern approach: how Koji helps

Outcome scoring lives or dies on whether the original claim was recorded in a specific, retrievable form. The main reason teams cannot grade their past research is that the findings were prose in a slide deck, and prose is unfalsifiable by default.

Structured questions produce claims with stable identity. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and every one carries a stable question ID from the interview plan through moderation and analysis into the report. A finding anchored to an identified item ("62 percent selected this option in the multiple_choice item on switching triggers") is a fact you can return to in six months and check. A finding that exists only as a sentence in a deck is not. The structured questions guide covers how each type aggregates.

Quantified findings make thresholds writable. Koji reports aggregate scale distributions and ranking positions across respondents rather than leaving you to characterise a mood. That is what lets a readout state a threshold instead of a direction, which is the precondition for a scoreable forecast.

Speed makes the feedback loop short enough to learn from. Calibration training works because forecasters get scored repeatedly. A research function that ships four studies a year generates four data points, which is not enough to detect miscalibration in a working lifetime. A function running weekly AI-moderated studies generates enough resolved forecasts within a year for the decomposition into calibration and resolution to mean something. This is the strongest practical argument for research velocity, and it is not the one usually made.

Re-running the same instrument closes the loop. Because a study definition and its question IDs persist, you can field the identical instrument after the change ships and compare like with like. The follow-up measurement is what resolves the forecast, and it costs a fraction of the original study.

Koji does not tell you whether you were right. It makes the original claim specific, retrievable, and cheap to re-measure - which is the entire infrastructure requirement for finding out.

Frequently asked questions

What is a Brier score in plain terms?

It is the squared difference between the probability you stated and what actually happened, scored as 1 or 0. Say 90 percent and be right and you score 0.01; say 90 percent and be wrong and you score 0.81. Lower is better, and it is averaged over many forecasts to grade a track record.

Why not just track how often the team was right?

Because a hit rate throws away the confidence information. Being right on calls you made at 55 percent is very different from being right on calls you made at 95 percent, and a hit rate cannot tell them apart. A proper scoring rule prices confidence, which is what makes hedging unprofitable.

Will scoring make my team more conservative?

Not if you use a proper scoring rule and keep it off individual performance reviews. The Brier score is specifically designed so hedging toward 50 percent scores mediocre forever. The real-world risk runs the other way: Mandel and Barnes found professional analysts erred toward underconfidence, most on the highest-stakes questions.

What is the difference between calibration and resolution?

Calibration asks whether things you call 70 percent happen about 70 percent of the time. Resolution asks whether you say meaningfully different things about different questions or just repeat the base rate. A team can be perfectly calibrated and useless, so track both separately.

How many forecasts do I need before the score means anything?

Enough that a single lucky call cannot dominate - practically, a few dozen resolved forecasts before reading the decomposition seriously. This is why research cadence matters: four studies a year will not produce a usable record in any reasonable timeframe.

Can forecasting accuracy actually be improved, or is it a fixed trait?

It can be improved. In the tournament research, probability training, team collaboration, and tracking top performers all improved both calibration and resolution, and top performers sustained their accuracy across two years rather than regressing to the mean. The training is mostly conceptual - reference classes, averaging independent estimates - and cheap.

Related Resources

Related Articles

Activating Research Insights: Turn Findings Into Product Decisions

A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

The Delphi Method: A Complete Guide to Reaching Expert Consensus

A practical guide to the Delphi method — the structured, multi-round technique for building expert consensus through anonymous questionnaires and controlled feedback. Learn the process, panel size, rounds, and modern AI-assisted alternatives.

Same Data, Different Answers: The Many-Analysts Problem in Product Research

When 73 teams analyzed identical data to test one hypothesis, over 95 percent of the variance in their results was unexplained. Your analysis is one draw from a distribution you never see.

Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)

Publication bias is not an academic curiosity. In product research it is worse, because nobody rejects your null study - you simply never write it up. Learn how big the file drawer is, what it does to your confidence, and how to build a study register that closes it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.