Back to docs
Research Operations

Interlaboratory Comparison for Research Teams: How to Find Out If Your Numbers Are Off

You cannot detect your own systematic bias from inside your own process. How to run the research equivalent of a proficiency testing scheme, including assigned values, z and zeta scores, and what to do when a round fails.

Short answer: you cannot detect your own systematic bias from inside your own process. No amount of internal consistency will reveal it, because a constant error is constant in every one of your measurements. The only instrument that finds it is an external comparison: the same question, put through a second independent team, panel or vendor, with the results scored against each other on a published rule. Testing laboratories have run this as a routine obligation for decades under the name proficiency testing, and the statistics are standardised in ISO 13528. Research teams almost never do it, which is why a house style of being wrong can persist for years without anyone noticing.

This guide adapts the laboratory scheme to a product research team, including the scoring rules and what to do when you fail.

The arithmetic that makes internal checks insufficient

The ISO 5725 series decomposes measurement precision into two conditions. Repeatability is what you get when the same operator measures the same thing the same way in the same place over a short interval. Reproducibility is what you get when the measurement is repeated somewhere else, by someone else. The relationship is an identity:

reproducibility variance = between-laboratory variance + repeatability variance
s_R^2 = s_L^2 + s_r^2

Read that as a statement about what your own data can tell you, and the consequence is stark. Every internal quality check you run - re-asking a question, double-coding a slice, having a second analyst review the deck - estimates the repeatability term. None of them can contain the between-laboratory term, because you only have one laboratory. Your internal consistency is a lower bound on your true error, and the gap between that bound and reality is invisible by construction.

The Guide to the Expression of Uncertainty in Measurement (GUM, JCGM 100:2008) puts the same point in one sentence: "an unrecognized systematic effect cannot be taken into account in the evaluation of the uncertainty of the result of a measurement but contributes to its error." The error is there. Your uncertainty statement does not know about it.

The GUM also names the remedy, in a note that reads like it was written for research teams:

"Determining the same measurand by different methods, either in the same laboratory or in different laboratories, or by the same method in different laboratories, can often provide valuable information about the uncertainty attributable to a particular method. In general, the exchange of measurement standards or reference materials between laboratories for independent measurement is a useful way of assessing the reliability of evaluations of uncertainty and of identifying previously unrecognized systematic effects."

What a proficiency testing scheme actually is

The laboratory version works like this. A provider distributes an identical test item to every participating laboratory. Each laboratory measures it using its own normal procedure and reports a number. The provider then establishes an assigned value for the item, scores every participant against it, and publishes the results.

The assigned value is the crux. ISO 13528 recognises five ways of establishing it, and they fall into two families with very different properties:

FamilyHow the assigned value is setStrengthWeakness
Independent referenceA certified reference material, or a measurement by a higher-order reference laboratoryDetects errors shared by every participantOften unavailable, and expensive when it exists
ConsensusDerived from the participants' own reported resultsAlways available; needs no external truthIf everyone shares a bias, the consensus carries it and nobody is flagged

The weakness of the consensus family is the one that transfers most sharply to research. If every team in your industry uses the same panel provider and the same question wording, a consensus value will confirm all of you. Consensus detects outliers, not shared errors. Knowing which of the two you have bought is the difference between a scheme that reassures and a scheme that informs.

The four scores, and which one is honest

Once there is an assigned value, performance is a normalised difference between the participant result and the assigned value. Eurachem, the European network of analytical chemistry organisations, sets out four in its 2024 leaflet on performance assessment:

ScoreDenominatorWhat it tests
D%The assigned value, as a percentageWhether the deviation is inside a permitted relative error
zThe standard deviation for proficiency assessmentWhether the deviation is large relative to what is acceptable for the purpose
zetaThe combined uncertainty of the assigned value and of your own reported uncertaintyWhether your result agrees with the assigned value within the uncertainty you claimed
EnExpanded uncertainties combined, at roughly 95% confidenceThe same test as zeta, in the units metrology comparisons use

The z score is the workhorse. Under ISO 13528:2022 clause 9.4.2 the interpretation is conventional and unambiguous:

  • |z| less than or equal to 2.0 - acceptable, satisfactory performance.
  • 2.0 < |z| < 3.0 - a warning signal, questionable performance.
  • |z| greater than or equal to 3.0 - unacceptable, an action signal, unsatisfactory performance.

But the score research teams should care about most is zeta, and the reason is subtle enough to be worth spelling out. The zeta score does not merely ask whether your number was close to the truth. It asks whether your number was close to the truth given how confident you said you were. As Eurachem notes, an unsatisfactory zeta score "can either be caused by an inappropriate estimation of the measured quantity value, or of its measurement uncertainty, or both."

A team that reports a number 3 points off with honest wide error bars passes. A team that reports a number 1 point off with a confidently narrow error bar fails. That is exactly the right incentive for a research function, and it is the opposite of the incentive most research functions currently operate under, where confident narrow claims are rewarded and hedged ones are treated as weak.

Designing a round for a product research team

You do not need an accreditation body. You need two independent executions of the same question and a rule agreed before the results come back.

  1. Choose the measurand precisely enough to be reproduced. Not "customer satisfaction" but "the proportion of active accounts on the Pro plan who rate onboarding 8 or above on a 0-10 scale, measured in the last 30 days." If two teams cannot read that sentence and measure the same thing, the round will measure your brief, not your process.
  2. Pick the comparison partner. In order of strength: a second research vendor; a second panel provider with your own team running both; two internal teams working blind to each other; the same team using a deliberately different method (interviews against behavioural data). Weaker partners still work - they just constrain what a disagreement can tell you.
  3. Decide the assigned value in advance. If you have an independent reference - a behavioural ground truth, a billing system, a census of the population - use it, and say so. If not, you are running a consensus scheme and you must write down that shared bias will not be detected.
  4. Require an uncertainty statement from every participant. This is the step that makes the round worth running, and the one everyone wants to skip. A result without a claimed uncertainty cannot be scored with zeta, and a scheme without zeta only tests accuracy, never honesty.
  5. Set the standard deviation for proficiency assessment from the decision, not from the data. How wrong could this number be before the decision it feeds changes? That is your denominator. Deriving it from the observed spread of participants just tells you how much you agree with each other.
  6. Score, using the published bands. |z| at or above 3 is an action signal. Warnings are warnings, not failures.
  7. Require documented root-cause analysis and verified corrective action for any action signal. In accredited laboratories a failed round does not automatically cost accreditation, but it does compel a documented investigation and evidence that the fix worked. Import that. A scheme with no consequence attached becomes a scheme nobody prepares for.
  8. Repeat on a fixed cadence. One round is an anecdote. The value is in the series - which is where you find out that you are not randomly wrong but consistently high.

Reading a disagreement without a fight

Rounds fail politically far more often than they fail statistically, so decide in advance what each outcome means.

OutcomeMost likely readingNext action
Both teams agree, both claimed tight uncertaintyEither genuinely good, or a shared method biasIf the assigned value was consensus, you have learned less than it feels like. Seek an independent reference before relaxing.
Teams disagree, both claimed tight uncertaintyAt least one uncertainty statement is wrongDo not argue about who is right first. Argue about who under-declared their uncertainty.
Teams disagree, both claimed wide uncertaintyThe method is genuinely impreciseCorrect behaviour. The finding is about the method, not the teams.
One team is an outlier across several roundsA systematic effect in that team's processRoot-cause investigation. A consistent direction is far more informative than a large one-off.

The third row is the one worth defending in your organisation. Two honest teams disagreeing with wide error bars is a success of the scheme, not an embarrassment. If disagreement is punished, the next round will produce agreement by social means and tell you nothing.

Note also what this is not. The evidence that independent analysts reach different conclusions from identical data is well established and covered in the many-analysts problem. That article establishes that the divergence exists. This one is about running a recurring scheme that tells you where you personally sit in that spread - and, more usefully, whether you sit on the same side of it every time.

How Koji makes a round affordable

The historical objection to any of this is cost: running a study twice, independently, doubles fieldwork, moderation and analysis. That is a genuine objection to the traditional model and a much weaker one now.

  • The brief is reproducible. A Koji study is defined by an explicit research brief and a question set with stable IDs, so handing "the same study" to a second team is a real transfer rather than a game of telephone. Two teams start from the same specification instead of the same vague ambition.
  • Structured questions make the results directly comparable. All six types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - produce structured, extracted answers rather than free text that has to be re-coded before it can be compared. A scale item gives you a distribution on both sides; a ranking item gives you average positions; single_choice and multiple_choice give you frequencies; yes_no gives you a proportion. Comparing two open_ended codebooks is the only part that needs judgement.
  • The second execution costs a fraction of the first. No moderator scheduling, no second agency onboarding. This is what moves the scheme from an annual aspiration to a quarterly routine.
  • Uncertainty statements have something to stand on. Because every theme and every chart point links back to the source interview, a claimed uncertainty can be inspected rather than taken on faith - which is what makes the zeta score meaningful instead of decorative.
  • Blinding is cheap. Two teams can run the same brief without seeing each other's results, which is the condition the whole scheme depends on. See blind analysis for the within-team version of the same idea.

Traditional survey platforms are built on the assumption that you run a study once. Nothing in SurveyMonkey, Typeform or Qualtrics makes a second independent execution cheaper, which is precisely why almost nobody has ever checked their own research against an outside reference.

Frequently asked questions

What is proficiency testing and why would a research team run it?

Proficiency testing is a recurring scheme in which independent laboratories measure the same item and are scored against an assigned value on a published rule. Testing laboratories run it as an accreditation obligation. A research team should run the equivalent because internal quality checks can only estimate repeatability - the variation you get repeating your own process - and are mathematically incapable of revealing a systematic bias shared across everything you do. An external comparison is the only instrument that finds a constant error.

How do I interpret a z score in a proficiency test?

Under ISO 13528:2022, a |z| of 2.0 or below is satisfactory, a |z| between 2.0 and 3.0 is a warning signal indicating questionable performance, and a |z| of 3.0 or above is an action signal indicating unsatisfactory performance requiring documented root-cause analysis. The z score compares your deviation from the assigned value against a standard deviation for proficiency assessment, which should be set from what the decision can tolerate rather than from how much the participants happen to agree.

What is the difference between a z score and a zeta score?

The z score asks whether your result was close enough to the assigned value. The zeta score asks whether your result agrees with the assigned value within the uncertainty you yourself claimed, using the combined uncertainty of the assigned value and your reported uncertainty as its denominator. That makes zeta the honest score: a team that is somewhat off with candid wide error bars passes, while a team that is slightly off with an overconfident narrow error bar fails. For research organisations that reward confident claims, this is the more valuable of the two.

Can I run an interlaboratory comparison with only one research team?

Yes, though with reduced power. In descending order of strength you can compare against a second vendor, a second panel provider with your own team running both, two internal teams working blind to each other, or the same team using a deliberately different method such as interviews versus behavioural data. Each weaker option shrinks the between-laboratory component you can detect, so record which one you used. Even the weakest version beats the alternative, which is having no external check at all.

What if both teams agree - does that prove the number is right?

Not necessarily, and this is the most important limitation to understand before you start. If your assigned value came from the participants themselves rather than an independent reference, agreement only shows that nobody is an outlier. A bias shared by every participant - the same panel provider, the same question wording, the same sampling frame - will pass a consensus scheme unnoticed and will be positively confirmed by it. Whenever an independent reference exists, such as behavioural data or a billing system, use it, and record in the report which family your assigned value came from.

How often should we run a round?

Often enough for a series to accumulate, because the informative pattern is directional rather than dramatic. A single round tells you whether you were an outlier once, which is close to noise. Four rounds tell you whether you are consistently high, and a consistent direction points straight at a systematic effect in your process. Quarterly is a workable cadence for most product research functions, with the strong caveat that a scheme with no documented corrective-action requirement attached will decay into an exercise nobody prepares for.

Related Resources

Related Articles

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

Same Data, Different Answers: The Many-Analysts Problem in Product Research

When 73 teams analyzed identical data to test one hypothesis, over 95 percent of the variance in their results was unexplained. Your analysis is one draw from a distribution you never see.

Measurement System Analysis: How Much of Your Segment Difference Is the Instrument? (2026)

How to separate real variation between customers from variation created by measuring them. The intraclass correlation, the four classes of monitor, probable error, and how to run an honest R&R study on a research metric.

Research Independence: Why the Team That Built the Feature Should Not Grade It

The five threats to independence from professional ethics codes, applied to product research - and why structural independence is a different problem from cognitive bias.

Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies

Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.

Process Controls vs Output Checks: How to Earn the Right to Read Fewer Transcripts

Evidence that your research process worked substitutes for evidence about each individual output. The trade auditors formalized, why existence is not operation, and how reperformance proves a control actually ran.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)

Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.