{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-11T12:30:27.902Z"},"content":[{"type":"documentation","id":"8736345a-3a77-4378-a249-e87514551b14","slug":"likelihood-of-confusion-survey","title":"Likelihood of Confusion Surveys: The Eveready and Squirt Formats Explained (2026)","url":"https://www.koji.so/docs/likelihood-of-confusion-survey","summary":"Two survey formats are recognised for trademark confusion. The Eveready format (Union Carbide Corp. v. Ever-Ready Inc., 531 F.2d 366) shows only the junior mark and asks who puts this out plus what makes you think so; it produced 55.2 percent and 60.6 percent across 1,009 and 1,014 interviews, and the Seventh Circuit noted earlier cases had credited 11.4 to 25 percent. The Squirt format (SquirtCo v. Seven-Up, 628 F.2d 1086) shows both marks and asks same company or different; it produced 34 percent in Chicago (n=152) and 23 percent in Phoenix (n=476) with a 43 percent do-not-know rate, and drew six documented attacks including that the dichotomous question encourages guessing. Technical deficiencies go to weight rather than admissibility. A bad universe voids either format, as in Amstar v. Domino Pizza.","content":"**Short answer: there are two recognised formats. The Eveready format shows respondents only the junior mark and asks an open question - who puts this out, and what other products do they make - and it suits a senior mark that is already well known. The Squirt format shows both marks and asks whether they come from the same company or different companies, and it suits marks that are not well known and that meet each other on a shelf. Eveready risks measuring nothing when the senior mark is weak. Squirt risks measuring guessing, because a two-option question invites an answer from someone who has no opinion. Choose the format from the market, not from the number you want.**\n\nConfusion is an empirical claim about what real people believe, which is why survey evidence has been at the centre of trademark disputes for fifty years. The same two designs answer a question ordinary product teams ask constantly: when we ship this sub-brand, this packaging refresh, or this look-alike feature, do customers think it comes from us or from somebody else?\n\n## The Eveready format\n\nThe format takes its name from Union Carbide Corp. v. Ever-Ready Inc., 531 F.2d 366 (7th Cir. 1976). Union Carbide sold EVEREADY batteries. A competitor sold EverReady lamps and bulbs. Carbide ran two surveys through a market research expert.\n\nRespondents were shown only the junior product - a picture of an EverReady lamp in one study, a blister pack of EverReady bulbs in the other - and asked:\n\n1. A screening question designed to eliminate people working in the bulb or lamp industries.\n2. Who do you think puts out the lamp shown here?\n3. What makes you think so?\n4. Please name any other products put out by the same concern.\n\nThat is the whole instrument, and its elegance is that it never mentions the senior mark. The respondent supplies it or does not.\n\nThe results: 1,009 interviews in the lamp study and 1,014 in the bulb study. Six respondents (0.6 percent) named Union Carbide outright in the lamp study and 13 (1.3 percent) in the bulb study, but 551 (54.6 percent) and 545 (53.7 percent) named Carbide products such as batteries as coming from the same concern. Totals were 55.2 percent and 60.6 percent.\n\nThe Seventh Circuit reversed the district court for refusing to credit these surveys, and noted the percentages were substantially higher than those held sufficient in earlier cases, citing figures of 11.4, 25, 18, and 18 and 24 percent. That footnote is the origin of the working rule of thumb practitioners still use: numbers in the mid-teens and above have persuaded courts, and there is no statutory line.\n\nTwo details are worth stealing. The court approved of the parties presenting the questionnaire to the judge before fielding, calling it a commendable procedure given the cost of a survey - the trademark equivalent of pre-registering your analysis. And an advertising question in the bulb study biased the results, because the defendant did not advertise mini-bulbs while Carbide advertised heavily, so a question about advertising could only point one way. One badly aimed question cost that study part of its weight.\n\n## The Squirt format\n\nSquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980), gives the other design. SquirtCo sold SQUIRT; Seven-Up launched QUIRST. Women 25 and older were intercepted leaving grocery stores, the study was postured as a study of grocery products, the marks were presented by sound alone, and respondents were asked:\n\n1. Do you think SQUIRT and QUIRST are put out by the same company or by different companies?\n2. What makes you think that?\n\nOf 152 interviews in Chicago, 51 respondents (34 percent) said the same company, 84 (55 percent) said different companies, and 17 (11 percent) said they did not know. Of 476 interviews in Phoenix, 110 (23 percent) said the same company, 161 (34 percent) different, and 205 (43 percent) did not know.\n\nSeven-Up attacked on six grounds, and the list is a free checklist of everything that can go wrong with this design: the first question was dichotomous and therefore encouraged guessing by suggesting it was improper to answer that you did not know; the phrase put out by was ambiguous because the two products shared a bottler in Phoenix; the Phoenix respondents were not even grocery shoppers, let alone soft drink buyers; the drop between Chicago and Phoenix suggested QUIRST was establishing its own identity; soft drinks are chosen visually while the survey used sound only; and survey results require expert interpretation, which SquirtCo did not provide.\n\nThe Eighth Circuit still affirmed, on the rule that in evaluating survey evidence technical deficiencies go to the weight to be accorded them, rather than to their admissibility. That rule cuts both ways: your survey will usually get in, and it will usually get discounted for exactly the reasons above.\n\n## Choosing between them\n\n| | Eveready format | Squirt format |\n| --- | --- | --- |\n| What respondents see | Junior mark only | Both marks, or both in context |\n| Core question | Who puts this out? | Same company or different companies? |\n| Question style | Open ended | Closed, two or three options |\n| Best when | Senior mark is strong and widely known | Marks are less known and compete directly |\n| Main weakness | Measures nothing if the senior mark is unknown | Invites guessing; needs a do not know option and a control |\n| Reported results | 55.2 and 60.6 percent (Union Carbide) | 34 percent Chicago, 23 percent Phoenix (SquirtCo) |\n| Control | Usually a control cell with a non-infringing mark | Essential, because the closed question generates noise |\n\nThe decision rule is simple. If a respondent who has never heard of the senior brand would be unable to answer the Eveready question, the Eveready format will produce a floor near zero and you will learn nothing. If your respondents could plausibly answer either way, the Squirt format needs a control cell to tell you what the guessing rate is.\n\nBoth formats need a do not know option that is offered rather than merely permitted. The 43 percent do not know rate in Phoenix is not a defect. It is the single most informative number in that study, because it tells you how much of the same company response is opinion and how much is politeness.\n\n## The universe defeats both formats\n\nNeither design survives a bad sample. In Amstar Corp. v. Domino Pizza, Inc., 615 F.2d 252 (5th Cir. 1980), the plaintiff surveyed women found at home during daylight hours who identified themselves as the household grocery buyer, in ten cities of which eight had no Domino Pizza outlet at all. The court held the proper universe had not been examined and discounted the survey entirely, noting the appropriate universe should include a fair sampling of those purchasers most likely to partake of the goods or services of the alleged infringer. The defendant survey, run on the premises of its own outlets, failed for the mirror-image reason. Two surveys, two opposite sampling errors, both worthless. That is the subject of the companion article on [survey universe](/docs/survey-universe-definition).\n\n## Translating this into product research\n\nStrip out the litigation and these are two general-purpose instruments for measuring source attribution.\n\n- **Eveready, product version.** Show a customer your new sub-brand, integration, or packaging with no context and ask who makes it and what else they make. This is the honest test of whether a brand extension carries your equity or launches from zero.\n- **Squirt, product version.** Show your feature alongside a competitor implementation and ask whether they come from the same company. This is the test of whether a look-alike release is diluting you, and it is the test to run before legal asks for it.\n\nRun either one with a control cell and you have an internal number that will survive a sceptical review, which is more than most brand decks manage.\n\n## How Koji fits\n\nBoth formats are short instruments with one crucial open question attached, which is exactly the shape that legacy survey tools handle worst. A form captures who do you think puts this out and then stops. The follow-up - what makes you think so - is the answer that tells you whether the respondent is reasoning from the logo, the name, the colour, or a guess.\n\nKoji runs these as AI-moderated conversations in voice or text. You set the closed question as single_choice or yes_no, add scale items for confidence, use ranking where you need a preference order among several marks, use multiple_choice for the aided recall list, and set the probe as open_ended - and the AI asks the follow-up itself, on every respondent, adapting to what they said rather than firing a fixed second question. All six structured question types are covered in the [structured questions guide](/docs/structured-questions-guide).\n\nBecause there is no moderator to schedule, a control cell costs the same as the test cell, which is the practical reason most commercial confusion studies never run one. Analysis is automatic: Koji clusters the reason-why answers into themes, so the report tells you not just that 34 percent said the same company, but that most of them said it because of a shape, not a name. Reports are live as interviews land, so you can stop a study that is clearly reading zero rather than paying for all 400 completes.\n\nFor a contested matter, retain a qualified survey expert. Use platforms like Koji for the pilot, the wording test, and the ongoing internal record.\n\n## Common mistakes\n\n- **Picking the format that flatters the case.** The format follows the market. An Eveready study on an unknown senior mark is a designed null result.\n- **Omitting the do not know option.** A forced binary manufactures the confusion it measures.\n- **Skipping the control cell.** Without it, your 20 percent has no denominator of noise.\n- **Asking about advertising when only one party advertises.** The Union Carbide bulb survey lost weight on exactly this.\n- **Using the wrong sensory channel.** Soft drinks are chosen by sight; a sound-only study invites the objection Seven-Up made.\n- **Fielding without a pilot.** Ambiguous phrasing such as put out by is only visible once real people answer it.\n\n## Frequently asked questions\n\n### What percentage of confusion is enough?\n\nThere is no fixed threshold. The Seventh Circuit in Union Carbide pointed to earlier cases crediting 11.4, 18, 24, and 25 percent, and treated its own 55 percent results as far above them. Practitioners commonly regard the mid-teens as the zone where a survey starts to persuade, with a control cell subtracted first.\n\n### Should I use the Eveready or the Squirt format?\n\nUse Eveready when the senior mark is strong enough that an uninformed respondent could name it unprompted. Use Squirt when both marks are relatively unknown and consumers encounter them side by side. If neither condition holds cleanly, run both cells.\n\n### Do technical flaws make a survey inadmissible?\n\nUsually not. Courts in both Union Carbide and SquirtCo applied the rule that technical deficiencies go to weight rather than admissibility. The practical effect is that a flawed survey gets in and then gets discounted, which is often worse than not running it.\n\n### Can I run a confusion survey without a lawyer?\n\nFor internal product decisions, yes, and you should. For a dispute, no. Design and interpretation by unqualified people is a known basis for exclusion, and in Elliott v. Google two surveys were excluded because counsel designed them.\n\n### How large should the sample be?\n\nThe classic cases ran roughly 150 to 1,000 per cell. Base size per cell drives precision, so a 400-person study split across a test and a control cell is generally more useful than an 800-person single-cell study.\n\n### How does this differ from ordinary name testing?\n\nName testing asks whether a name is appealing, pronounceable, and free of bad associations. A confusion survey asks whether people attribute it to somebody else. They use different questions and are often run together.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types behind both formats\n- [Survey Universe: Defining Who Counts](/docs/survey-universe-definition) - the sampling error that voided both surveys in Amstar\n- [Secondary Meaning Surveys](/docs/secondary-meaning-survey) - proving a descriptive name identifies one source\n- [Genericness Surveys](/docs/genericness-survey-teflon-thermos) - the Teflon and Thermos formats\n- [Name Testing Research](/docs/name-testing-research) - validating a name with real customers\n- [Open-Ended vs Closed-Ended Questions](/docs/open-ended-vs-closed-ended-questions) - why the reason-why probe carries the finding\n- [Survey Evidence in Court](/docs/survey-evidence-court-daubert) - control groups and the admissibility record","category":"Research Methods","lastModified":"2026-08-11T03:22:37.854648+00:00","metaTitle":"Likelihood of Confusion Surveys: Eveready vs Squirt Formats","metaDescription":"The two trademark confusion survey formats courts recognise, the exact questions and percentages from Union Carbide and SquirtCo, the attacks each design invites, and how to run them on your own brand.","keywords":["likelihood of confusion survey","Eveready survey format","Squirt survey format","trademark confusion research","Union Carbide Ever-Ready","SquirtCo Seven-Up","brand confusion testing","source attribution research"],"aiSummary":"Two survey formats are recognised for trademark confusion. The Eveready format (Union Carbide Corp. v. Ever-Ready Inc., 531 F.2d 366) shows only the junior mark and asks who puts this out plus what makes you think so; it produced 55.2 percent and 60.6 percent across 1,009 and 1,014 interviews, and the Seventh Circuit noted earlier cases had credited 11.4 to 25 percent. The Squirt format (SquirtCo v. Seven-Up, 628 F.2d 1086) shows both marks and asks same company or different; it produced 34 percent in Chicago (n=152) and 23 percent in Phoenix (n=476) with a 43 percent do-not-know rate, and drew six documented attacks including that the dichotomous question encourages guessing. Technical deficiencies go to weight rather than admissibility. A bad universe voids either format, as in Amstar v. Domino Pizza.","aiPrerequisites":["Familiarity with survey question design","Basic understanding of brand and trademark concepts"],"aiLearningOutcomes":["Choose between the Eveready and Squirt formats based on market conditions","Write the exact question sequence each format requires","Anticipate the six standard attacks on a confusion survey","Adapt both formats to internal sub-brand and packaging decisions"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}