Genericness Surveys: The Teflon and Thermos Formats, and How a Brand Loses Its Name (2026)
How genericness is measured: the Teflon classification format, the Thermos imaginary-situation format, the exact results from DuPont, American Thermos, Elliott v. Google and Booking.com, and why the question format decides the answer.
Short answer: genericness is decided by the primary significance of the mark to the relevant public, and it is measured with one of two formats. The Teflon format teaches respondents the difference between a brand name and a common name, then asks them to classify a list of names including yours. The Thermos format puts respondents in an imaginary situation and asks what they would call the thing or ask for. The two formats routinely give opposite-looking answers about the same word, and the case law shows courts consistently preferring the Teflon format because it asks the legally correct question. Which format you choose is not a methodological preference. It is the finding.
This is the failure mode nobody plans for. Aspirin, cellophane, escalator, thermos: each was once a trademark, and each was appropriated by the public until the word named the product rather than the producer. Trademark lawyers call it genericide. For a product team, the same measurement answers a friendlier question: has our name become the category, and is that an asset or the beginning of a loss?
The legal target
Under 15 U.S.C. 1064(3), the primary significance of the registered mark to the relevant public rather than purchaser motivation shall be the test for determining whether the registered mark has become the generic name of goods or services. The Second Circuit put the same idea plainly in King-Seeley Thermos Co. v. Aladdin Industries, 321 F.2d 577 (2d Cir. 1963): a mark is not generic merely because it has some significance to the public as an indication of the nature or class of an article; to become generic the principal significance of the word must be its indication of the nature or class of an article, rather than an indication of its origin.
The Ninth Circuit reduced it to a memorable pair of questions in Elliott v. Google, Inc., 860 F.3d 1151 (9th Cir. 2017): the who-are-you / what-are-you test. If the public primarily understands the mark as answering who are you, it is a brand. If it answers what are you, it is generic. The same opinion added a constraint that survey designers get wrong constantly: a claim of genericness must always relate to a particular type of good or service. IVORY is arbitrary for soap and generic for elephant tusks; the question is meaningless without a genus.
The Teflon format
The format is named for E. I. DuPont de Nemours and Co. v. Yoshida International, Inc., 393 F. Supp. 502 (E.D.N.Y. 1975). DuPont commissioned a survey in which the interviewer first explained the difference between a brand name and a common name using the example Chevrolet and automobile, then asked whether each of eight names was a brand name or a common name.
| Name | Brand | Common | Do not know |
|---|---|---|---|
| STP | 90 | 5 | 5 |
| JELLO | 75 | 25 | 1 |
| COKE | 76 | 24 | - |
| TEFLON | 68 | 31 | 2 |
| THERMOS | 51 | 46 | 3 |
| ASPIRIN | 13 | 86 | - |
| MARGARINE | 9 | 91 | 1 |
| REFRIGERATOR | 6 | 94 | - |
Percentages as reported in the opinion. The list is doing something clever. MARGARINE and REFRIGERATOR are known generics; STP and JELLO are known brands. They calibrate the instrument, prove the respondents can perform the task, and give the court a scale on which 68 means something. In effect the practice items are the control group, and they are the reason the Teflon format survived fifty years while looser designs did not.
The inversion that should change how you read every survey
The same opinion contains the sharpest lesson in trademark research, and it generalises far beyond trademarks.
The defendant ran two nationwide studies of adult women. Among those aware of non-stick cookware, 86.1 percent named only TEFLON when asked what the name of these pots and pans is, and 71.7 percent gave only TEFLON as the name they would use to describe them to a store clerk or a friend, rising to 79.3 percent counting those who added other answers. DuPont ran a study of its own, Survey A, asking whether respondents knew a brand name or trademark for such coatings: 48 percent of the entire sample answered TEFLON, and 68 percent of those respondents knew no other word or term for the coatings.
Read together, that looks like a generic word. Then DuPont ran Survey B, the Teflon classification format, and 68 percent called TEFLON a brand name against 31 percent common name.
The court chose Survey B, holding that in the other studies respondents were, by the design of the questions, more often than not focusing on supplying a name without regard to whether the principal significance of the name supplied indicated the nature of the article or its origin. The word did not change. The question did.
The generalisable framework: when a question does not distinguish naming from source-identifying, respondents default to naming, and you will measure genericness that is not there. The same defect shows up in product research every week. Ask which tool people use for a job and you get the category leader; ask which company makes the tool they use and you get a different and much smaller number. The gap between those two numbers is not noise. It is the measurement.
The Thermos format
The other design puts respondents in a situation and asks what they would say. American Thermos Products Co. v. Aladdin Industries, 207 F. Supp. 9 (D. Conn. 1962) supplies both a good version and a bad one.
The defendant commissioned a study of 3,300 interviewees, conducted according to the standards in the Judicial Conference Handbook of Recommended Procedures for the Trial of Protracted Cases. The results: about 75 percent of American adults familiar with containers that keep contents hot or cold call such a container a thermos; about 12 percent know that thermos has trademark significance; about 11 percent use the term vacuum bottle. The court accepted it and found the mark generic; the Second Circuit affirmed.
The plaintiff ran its own survey of 3,650 people and asked them to name any trademark or brand names they were familiar with for vacuum bottles or insulated containers. Roughly a third answered Thermos. The court gave it little weight, observing that the question focused the mind of the interviewee upon trademarks or brand names and left little or no opportunity for the revelation of a generic use, and that the sample skewed toward higher education and income and therefore did not constitute a representative cross-section.
Same case, same year, two surveys, one useless. The difference was that one question let the answer come out and the other told the respondent what kind of answer to give.
When the two formats disagree: Elliott v. Google
Elliott argued the GOOGLE mark had become generic because most people use google as a verb. His admissible survey was a Thermos-style study: 251 respondents were asked what word or phrase they would use to tell a friend to search for something on the internet, and over half used google as a verb.
Google answered with a Teflon survey: a little over 93 percent of respondents classified Google as a brand name.
The Ninth Circuit affirmed summary judgment for Google, holding that verb use does not automatically constitute generic use, and adopting the district court distinction between a discriminate verb (google it, meaning use Google) and an indiscriminate one (google it, meaning search anywhere). It also noted the Teflon survey offers comparative evidence as to how consumers primarily understand the word irrespective of its grammatical function - that is, the classification format tests the legal question and the imaginary-situation format tests linguistic habit.
Two of Elliott three surveys never reached the merits at all: they were excluded because they had been designed and conducted by counsel, who is not qualified to design or interpret surveys.
The modern boundary: Booking.com
USPTO v. Booking.com B.V., 591 U.S. 549 (2020) held that a term styled generic.com is a generic name for a class of goods or services only if the term has that meaning to consumers. Consumer perception, not a per se rule, decides it. The Court cautioned in the same breath that surveys can be helpful evidence of consumer perception but require care in design and interpretation.
The survey found 74.8 percent thought Booking.com was a brand name and 23.8 percent thought it was generic. Justice Breyer, dissenting, pointed out that the same survey found 33 percent thought Washingmachine.com - which corresponds to no company at all - was a brand, and 60.8 percent thought it generic, arguing that the instrument tested familiarity rather than genericness. Whichever side you find convincing, the methodological point is unavoidable: a classification survey without an implausible control item cannot separate brand recognition from brand meaning.
A survey is one input, not the verdict
In Princeton Vanguard, LLC v. Frito-Lay North America, Inc., 786 F.3d 960 (Fed. Cir. 2015), both sides ran Teflon surveys on PRETZEL CRISPS. One found 41 percent brand, 41 percent category, 18 percent unsure; the other found 55 percent brand, 36 percent generic. The Board criticised the universe of one as underinclusive and ultimately gave controlling weight to dictionary definitions and evidence of use by the media, third parties and the applicant itself. The Federal Circuit vacated on a different ground - the Board had analysed the constituent words rather than the mark as a whole - but the lesson stands. Dictionaries, media usage, competitor usage and your own marketing copy are all evidence, and your own copy is the one you control.
What to do about it as a product team
| Situation | Format to run | What a bad result means |
|---|---|---|
| Your brand is becoming the verb for the category | Teflon classification, with calibration items | Brand equity is converting into category vocabulary |
| You are naming a genuinely new category | Thermos imaginary-situation, before launch | You are about to name the category after yourself |
| A competitor uses your name descriptively | Teflon, plus a media and usage audit | Enforcement is overdue |
| You want an early-warning metric | Teflon classification inside the annual tracker | A multi-year slope, not a single reading |
The early-warning version is the one most teams should actually run. King-Seeley lost Thermos partly because the company acted too late; the court found the generic use had become so firmly impressed in everyday language that extraordinary efforts starting in the mid-1950s came too late. A classification item costs almost nothing to carry in an existing tracker and turns a catastrophic discrete event into a slope you can watch.
How Koji fits
A Teflon survey is a classification task plus an explanation, and a Thermos survey is an open elicitation. Both are trivially easy to field badly and surprisingly hard to field well, because the wording carries the entire result.
With Koji you build the classification as single_choice items across a calibration list, add multiple_choice for aided category vocabulary, use scale for confidence, use ranking to see which name respondents reach for first, add a yes_no gate for category awareness, and set the follow-up as open_ended so the AI asks why on every classification. That is all six structured question types in one instrument, and the structured questions guide documents how each is asked, rendered and analysed.
The advantage over a form is specific to this topic. In a form, a respondent who classifies your name as a common name gives you one datapoint. In a Koji interview, the AI immediately asks what makes you say that, and the answer distinguishes a respondent who has never heard of you from one who thinks the word simply means the product. Those two people count the same in a survey and mean entirely different things for your brand. Analysis is automatic, reports update as interviews complete, and voice or text is the respondent choice - so the study runs in days without a moderator.
For contested proceedings, retain a qualified expert. Use platforms like Koji to carry the classification item in your tracker year after year, which is the evidence that actually prevents genericide rather than litigating it.
Common mistakes
- No calibration items. Without known brands and known generics on the list, 68 percent is uninterpretable.
- No implausible control. Washingmachine.com is the whole of the Booking.com dissent.
- Asking for a name instead of a classification. The DuPont inversion: 86 percent said Teflon and 68 percent still called it a brand.
- Testing the word with no genus. Elliott failed partly because he claimed google was generic for the act of searching rather than for search engines.
- Treating verb use as proof. It is not, and a discriminate verb is evidence of strength.
- Skewed samples. The American Thermos plaintiff study reached a higher-educated, higher-income slice and lost weight for it.
- Running it once. Genericide is a slope, not an event.
Frequently asked questions
What is the difference between a Teflon survey and a Thermos survey?
A Teflon survey teaches the brand name versus common name distinction and asks respondents to classify a list of names. A Thermos survey puts respondents in an imaginary situation and asks what they would call the product or how they would ask for it. Teflon tests the legal question directly; Thermos tests linguistic habit.
Which format do courts prefer?
Teflon, where the two conflict. In DuPont the court relied on the classification survey over three name-elicitation studies, and in Elliott v. Google the Ninth Circuit noted the Teflon survey provided comparative evidence about primary significance while the Thermos survey only showed verb use.
What percentage makes a mark generic?
There is no fixed cut-off, but the reported cases give a usable scale. Thermos was found generic when about 75 percent called the product a thermos and about 12 percent knew it had trademark significance. Teflon survived at 68 percent brand and Google at 93 percent brand.
Does it hurt my brand if people use it as a verb?
Not by itself. Elliott v. Google holds that verb use does not automatically constitute generic use, and distinguishes discriminate verb use, where the speaker has your product in mind, from indiscriminate use, where they do not. Measure which one you have.
Can I stop my brand from becoming generic?
Partly, and early action matters more than intensity. King-Seeley lost Thermos because generic use was already embedded in everyday language by the time the company mounted a serious education campaign. Providing and promoting a generic term for the category alongside your brand is the standard defence.
Should an ordinary product team run a genericness survey?
Yes, as a cheap recurring item rather than a project. One classification question with calibration names inside your annual brand tracker converts a slow structural risk into a trend line, and it costs a fraction of a standalone study.
Related Resources
- Structured Questions Guide - the six question types used in a classification instrument
- Secondary Meaning Surveys - the same measurement pointed the other way
- Likelihood of Confusion Surveys - the Eveready and Squirt formats
- Survey Universe: Defining Who Counts - why an underinclusive universe sank a Teflon survey
- Brand Tracking Studies - where the early-warning item belongs
- Name Testing Research - choosing a name that will not become the category
- Survey Evidence in Court - admissibility and the control group
Related Articles
Brand Tracking Studies: How to Measure Brand Health Over Time (2026)
A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.
Likelihood of Confusion Surveys: The Eveready and Squirt Formats Explained (2026)
The two survey formats courts recognise for trademark confusion, the numbers that have persuaded judges, the attacks each format invites, and how to run the same design on your own sub-brand or packaging change.
Name Testing: How to Validate a Product or Brand Name With Real Customers
Name testing is the research method for choosing a product, brand, or feature name by measuring how real customers react to it. This guide covers what to measure, how to avoid the classic "pick the favorite" trap, and how to run name testing at scale with AI interviews.
Perceptual Mapping: How to Visualize Brand Positioning
A complete guide to perceptual mapping — what it is, the attribute-based vs MDS approaches, how to build one step by step, real examples, and how AI-moderated research collects the perception data fast.
Secondary Meaning Surveys: How to Prove a Descriptive Name Points to One Source (2026)
A practical guide to secondary meaning and acquired distinctiveness surveys: the legal target, the numbers courts have accepted, why the control term decides everything, and how to run the same study on your own brand in days instead of months.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Evidence in Court: Daubert, FRE 702, and Research That Survives Cross-Examination
The standard courts apply to survey evidence is a free quality bar for ordinary product research. Here is the checklist, the attacks it defeats, and why the control group is a legal instrument as much as a statistical one.
Survey Universe: How to Define Who Counts Before You Collect a Single Answer (2026)
The universe is the population whose opinion is actually relevant to your claim. Get it wrong and no sample size, weighting or analysis can rescue the study. A protocol, four documented failures, and how to enforce it at the door.