Back to docs
Research Methods

Survey Universe: How to Define Who Counts Before You Collect a Single Answer (2026)

The universe is the population whose opinion is actually relevant to your claim. Get it wrong and no sample size, weighting or analysis can rescue the study. A protocol, four documented failures, and how to enforce it at the door.

Short answer: the universe is the population whose opinion your claim is about; the sample is who you actually reached. Every other quality control in research - sample size, weighting, significance testing, careful analysis - operates inside the universe and is powerless outside it. A study that surveys the wrong people is not a weak study, it is a study about a different question, and courts throw those out entirely rather than discounting them. The rule that generalises: the universe is a property of the claim, not of the study, so a single screener can be perfectly valid for one question and worthless for the next one you ask with it.

This is the capstone of the trademark survey cluster, and it is the piece that decides whether the other three studies survive contact with an adversary. It is also the least glamorous topic in research methodology and the one that voids the most work.

The rule, stated by a court

Amstar Corp. v. Domino Pizza, Inc., 615 F.2d 252 (5th Cir. 1980) contains the cleanest statement of the standard anywhere:

One of the most important factors in assessing the validity of an opinion poll is the adequacy of the survey universe, that is, the persons interviewed must adequately represent the opinions which are relevant to the litigation. The appropriate universe should include a fair sampling of those purchasers most likely to partake of the goods or services of the alleged infringer.

Substitute your own word for the last clause - the users most likely to encounter the feature, the buyers most likely to renew, the applicants most likely to be rejected - and you have the test for any research programme.

Four documented ways to get it wrong

1. Sampling where the phenomenon does not exist. Amstar surveyed in ten cities. Eight had no Domino Pizza outlet at all, and the outlets in the remaining two had been open for less than three months. Respondents were women found at home during six daylight hours who identified themselves as the household member primarily responsible for grocery buying - while the survey evidence in the same case showed Domino Pizza customers were 85.6 percent under 35 years old, 61 percent single, and 63.3 percent male, with 75 percent of stores near college campuses or military bases. The Fifth Circuit discounted the plaintiff survey entirely.

2. Sampling only inside your own funnel. The defendant in the same case ran its survey on the premises of Domino Pizza outlets, and the court rejected it for the mirror-image reason. Everyone in that sample had already chosen the product. This is the single most common error in commercial product research, because the easiest list to survey is the customer list, and the customer list is a record of the people your product did not fail.

3. Reusing a screener across a different claim. In Zatarains, Inc. v. Oak Grove Smokehouse, Inc., 698 F.2d 786 (5th Cir. 1983), the screener asked how many times a month respondents fry fish or other seafood, and only those answering three or more times continued. That was fine for the Fish-Fri question. The same instrument then asked about Chick-Fri, and the Fifth Circuit observed that the qualifier seemed highly unlikely to provide an adequate sample of potential consumers of the chicken product, and that the survey provided nothing more than data regarding perceptions of fish friers about products used for frying chicken. Same respondents, same fieldwork, same expert: valid evidence for one mark, worthless for the other. This is the deepest lesson in the cluster. The universe belongs to the claim, not to the study, so every new question you bolt onto an existing panel silently inherits a screener that was never designed for it.

4. Reaching people who are not in the market at all. In SquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980), one of the six attacks on the Phoenix cell was that respondents were not even grocery shoppers, much less purchasers of soft drinks. The related tell in that study is the 43 percent do-not-know rate in Phoenix against 11 percent in Chicago. A do-not-know rate that swings by a factor of four between cells is usually a universe problem wearing a question-wording costume.

Two further cases round out the pattern. In Princeton Vanguard v. Frito-Lay, 786 F.3d 960 (Fed. Cir. 2015), one side survey was criticised specifically because the universe of survey participants was underinclusive. And in Elliott v. Google, 860 F.3d 1151 (9th Cir. 2017), two of three surveys never reached the substance because they were designed and conducted by counsel rather than by someone qualified to design or interpret surveys - a reminder that who defines the universe is itself part of the record.

The protocol

  1. Write the claim as a sentence before you write a question. Among X, Y percent believe Z. If you cannot fill in X precisely, you do not have a universe yet.
  2. Derive the universe from X, not from availability. Availability sampling is how you end up surveying grocery buyers about a pizza brand.
  3. Ask whether the claim is about buyers, users, prospects, or leavers. These are four different universes and they routinely give opposite answers.
  4. Decide the geography and the time window explicitly. Zatarains won New Orleans, not America. A universe with no time bound produces a number with no expiry date, which is how five-year-old surveys end up in evidence.
  5. Build the screener from the universe definition, and write both down together. A screener stored separately from its rationale will be reused for a claim it does not fit.
  6. Include the people who did not convert. If the only way into your sample is a completed purchase, your universe excludes the failure you are trying to study by construction.
  7. Report the funnel. Invited, screened out and why, started, completed. The screen-out reasons are often the most informative table in the report.
  8. Re-derive the universe for every new claim, even on the same panel. This is step one again, and it is the step the Chick-Fri survey skipped.

Claim to universe, in practice

The claim you want to makeCorrect universeThe wrong universe people actually use
Our onboarding is confusingEveryone who started onboarding in the window, including those who abandonedUsers who completed onboarding and are still active
This price is acceptableProspects who evaluated and did not buy, plus recent buyersCurrent subscribers
Our age gate is accurateEveryone who hit the gate, including those it rejectedAccounts that were successfully created
Customers are satisfiedA random sample of the active basePeople who answered the in-app prompt
The new sub-brand reads as oursCategory buyers who could plausibly meet both brandsYour newsletter list
Our name is still a brand, not a categoryThe general consuming public for that genusYour own customers, who are guaranteed to know you

The last row is worth dwelling on. A genericness or distinctiveness study run on your own customers is guaranteed to succeed and guaranteed to be worthless, because the whole question is what people who are not already yours think the word means.

Why this is a survivorship problem in a suit

Every one of the wrong universes above shares a structure: the sampling frame is defined by an outcome that is downstream of the thing being measured. Ask completers about onboarding and you have excluded the abandoners; ask subscribers about price and you have excluded everyone who found it too expensive; ask successful accounts about an age gate and you have excluded, by construction, every adult it wrongly rejected. The logs cannot distinguish a wrongly rejected adult from a correctly blocked minor, which is why only research reaches that group at all - a point developed in the age assurance guide.

This is survivorship bias with a sampling frame attached, and it is invisible in the output. The report looks complete. Every percentage has a denominator. Nothing in the deck signals that the relevant population was never contacted. That is precisely why courts treat a universe defect as fatal rather than as a discount: there is no statistical repair, because the missing data was never missing at random. It was never eligible.

How Koji fits

Universe definition is enforcement, not intention, and enforcement happens at the door.

Koji applies the screen at intake rather than in the crosstabs. Study intake forms and screener logic decide who proceeds, so a respondent outside the universe never becomes a completed interview you later have to argue about. Because the screen is part of the study configuration rather than a spreadsheet step, the definition travels with the study and shows up in the report.

The screener itself uses the same six structured question types as the main instrument - yes_no for eligibility gates, single_choice for the category-purchase question, multiple_choice for brands considered, scale for frequency or recency, ranking where relative preference decides eligibility, and open_ended when you need to know why someone falls outside - all documented in the structured questions guide. The AI interviewer can also probe an ambiguous screening answer instead of forcing it into a bucket, which is something a form physically cannot do.

Three further properties matter for this specific problem. Recruitment is not limited to your logged-in base, so studies can reach the abandoners and the never-converted rather than only the survivors. Reports are generated automatically as interviews complete, with the screening funnel visible, so an underfilled cell is obvious on day one rather than at readout. And because a study costs days rather than weeks, running a second cell - the control, the leavers, the general population - stops being a budget negotiation. That is the practical reason most commercial research has no control cell and no leaver sample: not ignorance, cost. Platforms like Koji remove the cost, which removes the excuse.

Common mistakes

  • Choosing the universe after seeing the results. If the base changes to make a number look better, the number is no longer evidence. In Amstar the parties disputed whether the confusion level was 6.8 percent or 0.4 percent purely on the basis of which respondents were excluded from the denominator.
  • Confusing sample size with sample validity. A thousand wrong people is worse than a hundred right ones, because it looks authoritative.
  • Weighting your way out. Weighting adjusts for known imbalance within a universe. It cannot conjure a group you never sampled.
  • Letting the panel provider define it. The provider optimises for fill rate.
  • Silent reuse. The Chick-Fri failure, and the most likely one to happen inside a healthy research programme.
  • No do-not-know option. Without it you cannot tell an out-of-universe respondent from an opinion.

Frequently asked questions

What is a survey universe?

The universe is the entire population whose opinion is relevant to the claim being made - for example, prospective purchasers of a product category in a defined geography and time window. The sample is the subset you actually interview. The universe is defined before fielding and determines who is eligible.

What is the difference between a universe and a sampling frame?

The universe is the population you care about in principle. The sampling frame is the concrete list or mechanism you use to reach it, such as a customer database or a panel. Most bias enters through the gap between the two, because frames are built from records of past success.

Can a large sample compensate for the wrong universe?

No. Sample size reduces random error within a universe; it does nothing about being outside it. Amstar discounted surveys entirely on universe grounds without reaching the sample size question, and a larger sample would have made the error more precise, not less wrong.

How do I define the universe for internal product research?

Write the claim as a sentence first: among whom, what percentage, believe what. The whom is the universe. Then check whether your recruitment route can actually reach those people, and if it can only reach converted users, say so in the report.

Can I reuse one screener across several studies?

Only after re-deriving the universe for each claim. Zatarains is the warning: the same fish-frying screener supported a finding for one mark and produced worthless data for the other, in the same survey, on the same day.

What should a report say about the universe?

The definition itself, the screener questions verbatim, the number invited, the number screened out with reasons, the number completed, and any group the recruitment route could not reach. That last line is the one that protects the reader, and it is the one that is almost always missing.

Related Resources

Related Articles

Genericness Surveys: The Teflon and Thermos Formats, and How a Brand Loses Its Name (2026)

How genericness is measured: the Teflon classification format, the Thermos imaginary-situation format, the exact results from DuPont, American Thermos, Elliott v. Google and Booking.com, and why the question format decides the answer.

Likelihood of Confusion Surveys: The Eveready and Squirt Formats Explained (2026)

The two survey formats courts recognise for trademark confusion, the numbers that have persuaded judges, the attacks each format invites, and how to run the same design on your own sub-brand or packaging change.

Sampling Bias: Types, Examples, and How to Avoid It

Sampling bias is when some people in your population are systematically more likely to end up in your sample than others — quietly invalidating your findings. Learn the six main types, classic examples, and how to build a representative sample at scale.

Secondary Meaning Surveys: How to Prove a Descriptive Name Points to One Source (2026)

A practical guide to secondary meaning and acquired distinctiveness surveys: the legal target, the numbers courts have accepted, why the control term decides everything, and how to run the same study on your own brand in days instead of months.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Evidence in Court: Daubert, FRE 702, and Research That Survives Cross-Examination

The standard courts apply to survey evidence is a free quality bar for ordinary product research. Here is the checklist, the attacks it defeats, and why the control group is a legal instrument as much as a statistical one.

Survey Response Bias: The 7 Types That Distort Your Data (and How to Reduce Them)

Response bias is the systematic distortion in how people answer research questions — from telling you what they think you want to hear, to agreeing with everything, to misremembering. This guide breaks down the seven most common response biases and how to reduce each one.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.