Back to docs
Research Methods

Likelihood of Confusion Surveys: The Eveready and Squirt Formats Explained (2026)

The two survey formats courts recognise for trademark confusion, the numbers that have persuaded judges, the attacks each format invites, and how to run the same design on your own sub-brand or packaging change.

Short answer: there are two recognised formats. The Eveready format shows respondents only the junior mark and asks an open question - who puts this out, and what other products do they make - and it suits a senior mark that is already well known. The Squirt format shows both marks and asks whether they come from the same company or different companies, and it suits marks that are not well known and that meet each other on a shelf. Eveready risks measuring nothing when the senior mark is weak. Squirt risks measuring guessing, because a two-option question invites an answer from someone who has no opinion. Choose the format from the market, not from the number you want.

Confusion is an empirical claim about what real people believe, which is why survey evidence has been at the centre of trademark disputes for fifty years. The same two designs answer a question ordinary product teams ask constantly: when we ship this sub-brand, this packaging refresh, or this look-alike feature, do customers think it comes from us or from somebody else?

The Eveready format

The format takes its name from Union Carbide Corp. v. Ever-Ready Inc., 531 F.2d 366 (7th Cir. 1976). Union Carbide sold EVEREADY batteries. A competitor sold EverReady lamps and bulbs. Carbide ran two surveys through a market research expert.

Respondents were shown only the junior product - a picture of an EverReady lamp in one study, a blister pack of EverReady bulbs in the other - and asked:

  1. A screening question designed to eliminate people working in the bulb or lamp industries.
  2. Who do you think puts out the lamp shown here?
  3. What makes you think so?
  4. Please name any other products put out by the same concern.

That is the whole instrument, and its elegance is that it never mentions the senior mark. The respondent supplies it or does not.

The results: 1,009 interviews in the lamp study and 1,014 in the bulb study. Six respondents (0.6 percent) named Union Carbide outright in the lamp study and 13 (1.3 percent) in the bulb study, but 551 (54.6 percent) and 545 (53.7 percent) named Carbide products such as batteries as coming from the same concern. Totals were 55.2 percent and 60.6 percent.

The Seventh Circuit reversed the district court for refusing to credit these surveys, and noted the percentages were substantially higher than those held sufficient in earlier cases, citing figures of 11.4, 25, 18, and 18 and 24 percent. That footnote is the origin of the working rule of thumb practitioners still use: numbers in the mid-teens and above have persuaded courts, and there is no statutory line.

Two details are worth stealing. The court approved of the parties presenting the questionnaire to the judge before fielding, calling it a commendable procedure given the cost of a survey - the trademark equivalent of pre-registering your analysis. And an advertising question in the bulb study biased the results, because the defendant did not advertise mini-bulbs while Carbide advertised heavily, so a question about advertising could only point one way. One badly aimed question cost that study part of its weight.

The Squirt format

SquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980), gives the other design. SquirtCo sold SQUIRT; Seven-Up launched QUIRST. Women 25 and older were intercepted leaving grocery stores, the study was postured as a study of grocery products, the marks were presented by sound alone, and respondents were asked:

  1. Do you think SQUIRT and QUIRST are put out by the same company or by different companies?
  2. What makes you think that?

Of 152 interviews in Chicago, 51 respondents (34 percent) said the same company, 84 (55 percent) said different companies, and 17 (11 percent) said they did not know. Of 476 interviews in Phoenix, 110 (23 percent) said the same company, 161 (34 percent) different, and 205 (43 percent) did not know.

Seven-Up attacked on six grounds, and the list is a free checklist of everything that can go wrong with this design: the first question was dichotomous and therefore encouraged guessing by suggesting it was improper to answer that you did not know; the phrase put out by was ambiguous because the two products shared a bottler in Phoenix; the Phoenix respondents were not even grocery shoppers, let alone soft drink buyers; the drop between Chicago and Phoenix suggested QUIRST was establishing its own identity; soft drinks are chosen visually while the survey used sound only; and survey results require expert interpretation, which SquirtCo did not provide.

The Eighth Circuit still affirmed, on the rule that in evaluating survey evidence technical deficiencies go to the weight to be accorded them, rather than to their admissibility. That rule cuts both ways: your survey will usually get in, and it will usually get discounted for exactly the reasons above.

Choosing between them

Eveready formatSquirt format
What respondents seeJunior mark onlyBoth marks, or both in context
Core questionWho puts this out?Same company or different companies?
Question styleOpen endedClosed, two or three options
Best whenSenior mark is strong and widely knownMarks are less known and compete directly
Main weaknessMeasures nothing if the senior mark is unknownInvites guessing; needs a do not know option and a control
Reported results55.2 and 60.6 percent (Union Carbide)34 percent Chicago, 23 percent Phoenix (SquirtCo)
ControlUsually a control cell with a non-infringing markEssential, because the closed question generates noise

The decision rule is simple. If a respondent who has never heard of the senior brand would be unable to answer the Eveready question, the Eveready format will produce a floor near zero and you will learn nothing. If your respondents could plausibly answer either way, the Squirt format needs a control cell to tell you what the guessing rate is.

Both formats need a do not know option that is offered rather than merely permitted. The 43 percent do not know rate in Phoenix is not a defect. It is the single most informative number in that study, because it tells you how much of the same company response is opinion and how much is politeness.

The universe defeats both formats

Neither design survives a bad sample. In Amstar Corp. v. Domino Pizza, Inc., 615 F.2d 252 (5th Cir. 1980), the plaintiff surveyed women found at home during daylight hours who identified themselves as the household grocery buyer, in ten cities of which eight had no Domino Pizza outlet at all. The court held the proper universe had not been examined and discounted the survey entirely, noting the appropriate universe should include a fair sampling of those purchasers most likely to partake of the goods or services of the alleged infringer. The defendant survey, run on the premises of its own outlets, failed for the mirror-image reason. Two surveys, two opposite sampling errors, both worthless. That is the subject of the companion article on survey universe.

Translating this into product research

Strip out the litigation and these are two general-purpose instruments for measuring source attribution.

  • Eveready, product version. Show a customer your new sub-brand, integration, or packaging with no context and ask who makes it and what else they make. This is the honest test of whether a brand extension carries your equity or launches from zero.
  • Squirt, product version. Show your feature alongside a competitor implementation and ask whether they come from the same company. This is the test of whether a look-alike release is diluting you, and it is the test to run before legal asks for it.

Run either one with a control cell and you have an internal number that will survive a sceptical review, which is more than most brand decks manage.

How Koji fits

Both formats are short instruments with one crucial open question attached, which is exactly the shape that legacy survey tools handle worst. A form captures who do you think puts this out and then stops. The follow-up - what makes you think so - is the answer that tells you whether the respondent is reasoning from the logo, the name, the colour, or a guess.

Koji runs these as AI-moderated conversations in voice or text. You set the closed question as single_choice or yes_no, add scale items for confidence, use ranking where you need a preference order among several marks, use multiple_choice for the aided recall list, and set the probe as open_ended - and the AI asks the follow-up itself, on every respondent, adapting to what they said rather than firing a fixed second question. All six structured question types are covered in the structured questions guide.

Because there is no moderator to schedule, a control cell costs the same as the test cell, which is the practical reason most commercial confusion studies never run one. Analysis is automatic: Koji clusters the reason-why answers into themes, so the report tells you not just that 34 percent said the same company, but that most of them said it because of a shape, not a name. Reports are live as interviews land, so you can stop a study that is clearly reading zero rather than paying for all 400 completes.

For a contested matter, retain a qualified survey expert. Use platforms like Koji for the pilot, the wording test, and the ongoing internal record.

Common mistakes

  • Picking the format that flatters the case. The format follows the market. An Eveready study on an unknown senior mark is a designed null result.
  • Omitting the do not know option. A forced binary manufactures the confusion it measures.
  • Skipping the control cell. Without it, your 20 percent has no denominator of noise.
  • Asking about advertising when only one party advertises. The Union Carbide bulb survey lost weight on exactly this.
  • Using the wrong sensory channel. Soft drinks are chosen by sight; a sound-only study invites the objection Seven-Up made.
  • Fielding without a pilot. Ambiguous phrasing such as put out by is only visible once real people answer it.

Frequently asked questions

What percentage of confusion is enough?

There is no fixed threshold. The Seventh Circuit in Union Carbide pointed to earlier cases crediting 11.4, 18, 24, and 25 percent, and treated its own 55 percent results as far above them. Practitioners commonly regard the mid-teens as the zone where a survey starts to persuade, with a control cell subtracted first.

Should I use the Eveready or the Squirt format?

Use Eveready when the senior mark is strong enough that an uninformed respondent could name it unprompted. Use Squirt when both marks are relatively unknown and consumers encounter them side by side. If neither condition holds cleanly, run both cells.

Do technical flaws make a survey inadmissible?

Usually not. Courts in both Union Carbide and SquirtCo applied the rule that technical deficiencies go to weight rather than admissibility. The practical effect is that a flawed survey gets in and then gets discounted, which is often worse than not running it.

Can I run a confusion survey without a lawyer?

For internal product decisions, yes, and you should. For a dispute, no. Design and interpretation by unqualified people is a known basis for exclusion, and in Elliott v. Google two surveys were excluded because counsel designed them.

How large should the sample be?

The classic cases ran roughly 150 to 1,000 per cell. Base size per cell drives precision, so a 400-person study split across a test and a control cell is generally more useful than an 800-person single-cell study.

How does this differ from ordinary name testing?

Name testing asks whether a name is appealing, pronounceable, and free of bad associations. A confusion survey asks whether people attribute it to somebody else. They use different questions and are often run together.

Related Resources

Related Articles

Brand Research Interviews: How to Understand Brand Perception Through Conversation

A complete guide to running qualitative brand research interviews — covering brand perception, positioning validation, competitive differentiation, and brand equity — using AI-moderated conversations at scale.

Genericness Surveys: The Teflon and Thermos Formats, and How a Brand Loses Its Name (2026)

How genericness is measured: the Teflon classification format, the Thermos imaginary-situation format, the exact results from DuPont, American Thermos, Elliott v. Google and Booking.com, and why the question format decides the answer.

Name Testing: How to Validate a Product or Brand Name With Real Customers

Name testing is the research method for choosing a product, brand, or feature name by measuring how real customers react to it. This guide covers what to measure, how to avoid the classic "pick the favorite" trap, and how to run name testing at scale with AI interviews.

Open-Ended vs. Closed-Ended Questions: Examples and When to Use Each

Open-ended questions reveal the "why" in respondents'' own words; closed-ended questions deliver clean, countable data. Learn the difference, see examples of both, and discover why the best research pairs them — and how AI captures both at once.

Secondary Meaning Surveys: How to Prove a Descriptive Name Points to One Source (2026)

A practical guide to secondary meaning and acquired distinctiveness surveys: the legal target, the numbers courts have accepted, why the control term decides everything, and how to run the same study on your own brand in days instead of months.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Evidence in Court: Daubert, FRE 702, and Research That Survives Cross-Examination

The standard courts apply to survey evidence is a free quality bar for ordinary product research. Here is the checklist, the attacks it defeats, and why the control group is a legal instrument as much as a statistical one.

Survey Universe: How to Define Who Counts Before You Collect a Single Answer (2026)

The universe is the population whose opinion is actually relevant to your claim. Get it wrong and no sample size, weighting or analysis can rescue the study. A protocol, four documented failures, and how to enforce it at the door.