Back to docs
Analysis & Synthesis

The Barnum Test: How to Tell If a Research Finding Says Anything About Your Users (2026)

A research finding only describes your users if a team with different users would reject it. Use a forced choice against a plausible decoy, not a stakeholder rating.

Answer first: A research finding is only about your users if a team with different users would reject it. The fastest way to check is a forced choice against a plausible decoy, not a rating. Asking stakeholders whether a finding sounds right cannot separate a real insight from a universal one, and that limitation has been measured repeatedly since 1948.

This article gives you the history, the one experiment that proves ratings are the wrong instrument, and a procedure called the swap test that takes about twenty minutes to run.

A 1948 classroom demonstration that should worry every research team

Bertram Forer gave a personality test to 39 of his students, discarded what they had written, and handed every one of them the same sketch: thirteen statements he had assembled from a newsstand astrology book. He asked each student to rate how well it described them on a scale running from 0 (very poor) to 5 (excellent). The class average came out above 4.2. The work appeared as The Fallacy of Personal Validation: A Classroom Demonstration of Gullibility, in the Journal of Abnormal and Social Psychology, volume 44, issue 1, pages 118-123, 1949.

A note on that number, because it is the kind of detail this article is about: reference works print Forer's class mean as either 4.26 or 4.30. Above 4.2 is what they agree on, so that is all this article claims. Paul Meehl gave the phenomenon its lasting name, the Barnum effect, in 1956.

The year before Forer's class, Ross Stagner ran the same manoeuvre on working professionals. He asked a group of personnel managers to take a personality test, then gave each of them generalised feedback that had no relation to their answers and was drawn instead from horoscopes, graphological analyses and similar material. More than half described the assessment as accurate. Almost none described it as wrong.

Neither result is about gullible people. Both are about a measurement problem. The managers and the students were asked to rate accuracy, and a rating is the one instrument that cannot tell a description that fits you from a description that fits everybody.

The experiment that proves a rating is the wrong instrument

In 2008 Alyssa Jayne Wyman and Stuart Vyse ran a double-blind comparison that settles the question for practical purposes. Writing in the Journal of General Psychology, volume 135, issue 3, pages 287-300, they asked 52 college students (38 women, 14 men, mean age 19.3 years, standard deviation 1.3 years) to do two things.

First, pick their own personality summary out of a pair: one true summary derived from the NEO Five-Factor Inventory, one bogus. Second, do the same with a computer-generated astrological natal chart against a bogus alternative.

The results split cleanly along the lines of the instrument used:

  • On the forced choice, participants identified their real NEO Five-Factor Inventory profiles at a greater-than-chance level, and were unable to identify their real astrological summaries.
  • On the accuracy ratings, a Barnum effect appeared for both the psychological and the astrological measures.

Read those two sentences together, because the whole argument of this article sits between them. The rating said the astrology was accurate and said the real instrument was accurate. It could not separate them. The forced choice separated them perfectly: it found signal where there was signal and found none where there was none.

What you askWhat it measuresCan it detect a generic finding?
Does this describe our users? (rating)Plausibility, fluency, agreeableness of the readerNo. Generic text rates as highly as specific text
Which of these two describes our users? (forced choice)Discrimination against a competing specificYes. This is the only one that can fail for the right reason
Does this match your experience? (yes or no)Recognition, which generic statements also triggerNo, for the same reason as a rating

Every informal validation ritual a product team runs is in the top row. The finding goes on a slide, the room nods, somebody says it matches what they hear on calls, and the finding is treated as confirmed. That ritual has the diagnostic power of an astrology rating.

Forer's statements, rewritten as product findings

Two of Forer's thirteen statements, quoted from his sketch, were: You prefer a certain amount of change and variety and become dissatisfied when hemmed in by restrictions and limitations, and You pride yourself as an independent thinker and do not accept others' statements without satisfactory proof.

Now compare the shape of those to findings that appear in real research readouts:

  • Users want flexibility, but they get frustrated when the workflow constrains them.
  • Our buyers are sophisticated and do not trust vendor claims without evidence.
  • Customers value their time and abandon processes that feel unnecessarily long.
  • Teams want visibility into what is happening without being overwhelmed by noise.

Each of those is true. Each would be rated accurate by almost any stakeholder in almost any company. None of them discriminates, which means none of them can be wrong, which means none of them can inform a decision. A statement that no observation could contradict cannot tell you what to build. That connects this problem to falsifiable research questions and severe tests, which handles the same failure at the point the study is designed rather than at the point the finding is reported.

Here is the diagnostic difference in practice:

Generic versionDiscriminating version
Users want faster onboardingUsers who import from a spreadsheet abandon at the column-mapping step, and the ones who abandon have more than 40 columns
Buyers do not trust vendor claimsBuyers forward our pricing page to a finance colleague before the second call, and that colleague asks about overage first
Customers value responsive supportCustomers who file a second ticket on the same issue stop filing a third, and churn two months later

The right-hand column can be false. That is what makes it worth knowing.

The swap test

The swap test is a forced choice between your finding and a decoy. It takes five steps.

  1. State the finding as one sentence. If it will not fit in one sentence, the discriminating detail is probably already missing.
  2. Write a decoy. Replace the discriminating content with a plausible competing specific. Not the negation, and not something absurd: a decoy that a reasonable person could believe about a company like yours. If your finding says abandonment happens at column mapping, the decoy says it happens at the authentication step.
  3. Present both in randomised order to people who know your users, as a forced choice: which one describes our users? Collect confidence separately.
  4. Count. The proportion choosing the real finding is your discrimination rate.
  5. Interpret against chance, which is the step most teams skip.

That last step is arithmetic you can do in advance. With twelve judges, chance alone puts six on the correct answer. To clear a conventional one-sided threshold you need ten of the twelve: the probability of ten or more correct out of twelve by chance is 79 in 4096, about 0.019. Nine of twelve gives 299 in 4096, about 0.073, which does not clear it. With twenty judges the bar is fifteen, about 0.021.

So the honest headline from a swap test is not the room nodded. It is ten of twelve picked the real finding, against a bar of ten. A readout that reports the second sentence is making a claim that can fail.

What a failed swap test actually means

A finding that cannot beat its decoy has one of three problems, and only one of them is a reason to throw it away.

It is genuinely generic. The statement is true of any user of any product in the category. Discard it, or demote it to context.

It is real but stated at the wrong altitude. This is the common case, and it is the reason the swap test earns its place. The discriminating detail existed in the transcripts and was edited out on the way to the slide, usually in the name of making the finding sound broader and more important. The fix is to go back to the evidence and recover the specific, not to delete the insight. This failure mode is the mirror image of the one described in writing research insight statements: a well-formed sentence can be well-formed and still empty.

Your judges do not know your users. If the panel is made up of people who have never spoken to a customer, a failed swap test is a finding about your panel. Run it on the people closest to the accounts.

The swap test measures discrimination, not truth. A finding can beat its decoy and still be wrong, because the judges share the researcher's exposure to the same conversations. Treat it as a floor, not a ceiling, and read it alongside negative controls in user research, which applies the same logic at the data-collection stage by building a condition where the answer cannot exist.

How Koji helps

The swap test is cheap to describe and historically expensive to run, because it needs a second pass over a panel after the analysis is finished. That is a second round of recruiting, scheduling and moderating for a question that feels like administrative overhead. It is the first thing cut.

Koji removes the reason it gets cut. Because the interview is AI-moderated, the control pass costs roughly what the original costs, and it can run while the readout is still being written.

The mechanics map onto Koji's structured question types directly. Koji supports six: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. The choice between them is the whole methodological argument of this article.

StepKoji question typeWhy this one
The swap test itselfsingle_choiceForces discrimination between the finding and the decoy. This is the Wyman and Vyse forced choice
Confidence in the choicescaleUseful alongside the choice, useless instead of it
Four candidate findings at oncerankingReturns an average position per finding, so the generic ones sink
Which specific details the judge recognisesmultiple_choiceSeparates recognition of the detail from acceptance of the headline
Would you change a roadmap decision on this?yes_noConverts agreement into a commitment
Why did you pick that one?open_endedCatches judges who chose correctly for the wrong reason

Two further details matter. Koji's structured interview mode follows the key questions closely rather than following interesting threads, which is what a validation pass needs and what an exploratory pass deliberately does not do. And Koji's analysis attaches a confidence level of high, medium or low to each extracted answer, so a thin or hedged response is visible rather than averaged in.

Finally, this is not a new principle bolted onto the product. Koji's built-in Mom Test framework already ships the anti-pattern Do not accept compliments as validation - dig for facts, and the core principle Dig for specifics - always or never needs examples. Those rules were written to protect you from a participant's polite agreement. The Barnum test applies the identical rule one level up, to your stakeholders' polite agreement, which nobody checks. Koji's six question types make the check a configuration choice rather than a research project. See the structured questions guide for how the six types behave in reports.

Common mistakes

  • Writing an absurd decoy. If the decoy is obviously wrong, everyone picks correctly and the test proves nothing. The decoy has to be a live hypothesis.
  • Reporting the proportion without the chance baseline. Eight of twelve sounds like a pass and is not.
  • Running it on one person. A single judge cannot beat chance in any meaningful sense.
  • Treating a pass as proof the finding is true. It is evidence the finding is specific. Specificity and truth are different properties, and Koji will not collapse them for you.
  • Using a rating because it is easier to collect. This is the exact error Forer demonstrated. A five-point agreement scale on a single statement is the Barnum instrument.
  • Swapping only the noun. Changing users to customers is not a decoy. Change the discriminating claim.

Frequently asked questions

What is the Barnum effect in user research?

The Barnum effect is the tendency to accept a vague, broadly applicable description as a specific and accurate account of oneself or one's own situation. In user research it shows up when a finding that would be true of almost any company is received by stakeholders as a precise description of their users. Forer demonstrated it in 1948 by giving 39 students the same thirteen-statement sketch and collecting accuracy ratings above 4.2 on a 0 to 5 scale.

Why is asking stakeholders whether a finding sounds right not a validation?

Because a rating cannot fail for the right reason. Wyman and Vyse found in 2008 that accuracy ratings produced a Barnum effect for both a real personality inventory and an astrological chart, while a forced choice between a true and a bogus summary correctly separated them. Ratings measure plausibility; only a forced choice against a competing specific measures discrimination.

How do I write a good decoy for the swap test?

Replace the discriminating content of your finding with a plausible competing specific, keeping the sentence structure, length and tone identical. The decoy should be something a reasonable colleague could believe about a company like yours. Avoid negations and avoid anything absurd, because both make the choice easy and the test uninformative.

How many judges does a swap test need?

Enough to beat chance. With twelve judges you need ten correct choices to clear a conventional one-sided threshold, because the probability of ten or more out of twelve by chance is about 0.019. With twenty judges the bar is fifteen. Fewer than about eight judges cannot produce an interpretable result.

What if my finding fails the swap test?

Check which of three things happened before discarding anything. The finding may be genuinely generic, in which case demote it to context. It may be real but stated at too high an altitude, with the discriminating detail lost somewhere between the transcript and the slide, which is the most common case and calls for recovering the specific. Or your judges may not know your users well enough to choose.

Does this mean generic findings are always worthless?

No. A generic statement can be a correct piece of background, a useful framing for an audience new to the problem, or a confirmation that your users are not unusual on some dimension. It is only a problem when it is presented as a research result that should change a decision, because a statement that no observation could contradict cannot favour one decision over another.

The bottom line

Forer showed in 1948 that people rate a universal description as accurate. Wyman and Vyse showed in 2008 that the rating is the problem and a forced choice is the cure. Between those two results sits every readout in which a finding was confirmed because the room agreed with it.

Write the finding. Write a decoy a colleague could believe. Make people choose, count the choices, and compare the count to chance. If your finding cannot beat a plausible alternative, you have learned something more useful than another round of agreement.

Related Resources