Advertising Claim Substantiation: How to Design Survey Research That Backs a Marketing Claim
A claim like "9 out of 10 customers recommend us" is a regulated assertion, and the evidence has to exist before the ad runs. This is how to design the study so the number survives a challenge from a regulator, a competitor, or a self-regulatory body.
Answer first: if a marketing claim is objective and capable of being proved, you must possess and rely on a reasonable basis for it before the ad is disseminated — post-hoc evidence is not a defence. For a survey to serve as that basis, four things must hold: the population was properly chosen and defined, the sample was representative of that population, the data were accurately reported, and the data were analysed in accordance with accepted statistical principles. Most marketing surveys fail on the first one, because the population that was surveyed (existing happy customers) is not the population the claim is about (customers generally, or buyers in the category).
This guide is the quantitative counterpart to Using Research Quotes in Marketing. That article covers individual testimonials. This one covers the aggregate number — the "87% of teams say", the "rated #1 for ease of use", the "twice as fast as the leading alternative" — and how to design a study that can carry it.
The legal standard, in one paragraph
The FTC Policy Statement Regarding Advertising Substantiation sets the rule: advertisers and agencies must have a reasonable basis for advertising claims before they are disseminated. Objective claims represent, explicitly or by implication, that the advertiser has a reasonable basis supporting them, and that representation is itself material to consumers. Failing to possess and rely on a reasonable basis for an objective claim is an unfair and deceptive act under Section 5 of the FTC Act.
The UK position is stated even more directly. CAP Code rule 3.7 requires that before distributing or submitting a marketing communication for publication, marketers must hold documentary evidence to prove claims that consumers are likely to regard as objective and that are capable of objective substantiation. Rule 3.6 adds that communications must not mislead by implying that expressions of opinion are objective claims.
Two words carry most of the weight: before and objective.
What level of evidence counts as reasonable
The Commission determines what constitutes a reasonable basis by weighing six factors:
- The type of claim
- The product
- The consequences of a false claim
- The benefits of a truthful claim
- The cost of developing substantiation for the claim
- The amount of substantiation experts in the field believe is reasonable
The sixth is the one that decides most disputes about survey evidence. If competent researchers in your field would expect a representative sample and a non-leading instrument before believing a percentage, then a poll of your newsletter list does not qualify — regardless of how many people answered.
There is also a trap in the phrasing of the ad itself. When the substantiation claim is express — "tests prove", "studies show", "doctors recommend" — the firm is expected to have at least the advertised level of substantiation. Saying "research shows" when you ran an informal internal poll raises your own evidentiary bar. If you have one survey, say "in a survey of 412 customers", not "studies show".
And one more, from the Endorsement Guides: consumer endorsements themselves are not competent and reliable scientific evidence (16 CFR 255.2(a)). Fifty enthusiastic interviews do not add up to a statistic. They are the reason to run the study, not the study.
The four factors a survey has to satisfy
Courts assessing survey evidence apply a compact standard. The Manual for Complex Litigation (Fourth) section 11.493 states that sampling methods must conform to generally recognised statistical standards, and that relevant factors include whether:
- the population was properly chosen and defined;
- the sample chosen was representative of that population;
- the data gathered were accurately reported; and
- the data were analysed in accordance with accepted statistical principles.
Self-regulatory review — the sort a competitor triggers by challenging your claim — applies substantively the same test, with added attention to whether questions were clear and non-leading and whether interviewing was conducted objectively.
Design against these four factors from the start. Retro-fitting them is impossible, because the first two are decided the moment you choose who to invite.
Factor 1: define the universe from the claim, not the list
Write the claim first, then read the noun. The noun is your universe.
| Claim | Universe the claim asserts | Common wrong universe |
|---|---|---|
| "Most teams switch within a week" | All customers who migrated | Customers who completed migration successfully |
| "Marketers prefer X to Y" | Marketers in the category, including non-customers | Your own customers, who already chose X |
| "9 out of 10 users would recommend us" | All active users | Users who opened the NPS email |
| "Rated #1 for ease of use" | A defined comparison set, all rated by the same people | A single-brand study with no comparison |
That second row is where most claims die. A preference claim about a market cannot be substantiated from a survey of people who have already bought your product; the sample is selected on the outcome. If the claim mentions a category, the survey has to reach the category. If you only have access to your own customers, narrow the claim to your own customers and say so in the ad: "In a survey of 412 Koji customers…" is defensible; "marketers prefer" is not.
Factor 2: make the sample representative, then prove it
Representativeness is a property of the recruitment process, not of the sample size. A million self-selected responses are still self-selected. Three controls do most of the work:
- Sample from a defined frame. Draw at random from all eligible customers or a properly constructed panel, not from whoever is easiest to reach. See Quota Sampling for the practical approach when random sampling is not available.
- Document non-response. Report the invite count, the completion count, and the completion rate. If responders differ from non-responders on anything you can observe, say so and consider weighting. Undocumented nonresponse bias is the most common reason a claim gets withdrawn.
- Do not screen on enthusiasm. Excluding people who are unlikely to answer favourably is fatal, whether it happens by design or by an eager sales team choosing the invite list. Related traps are covered in Sampling Bias and Survivorship Bias in Customer Research.
Base size then determines how precisely you can state the number. At 95% confidence, with a proportion near 50%:
| Completed responses | Margin of error |
|---|---|
| 100 | ±9.8 points |
| 200 | ±6.9 points |
| 400 | ±4.9 points |
| 1,000 | ±3.1 points |
A "9 out of 10" claim from n=100 has a margin of error wide enough that the true value could be 80%. Round down when you publish, and never state a percentage to one decimal place from a base under 400 — the precision is fictional and it signals inexperience to anyone reviewing the file.
Factor 3: report the data accurately
Accurate reporting means the published number is reproducible from the raw file by someone who did not run the study. Three habits achieve it:
- Publish the denominator. "87% of respondents" is incomplete. "87% of 412 respondents who had used the feature in the last 30 days" is a claim someone can check.
- State the base every time it changes. If a question was only asked of a subgroup, the base is the subgroup, not the sample.
- Do not promote a top-two box to a stronger word. "Satisfied or very satisfied" is not "love". Combining scale points is standard practice; renaming them is not.
Question wording sits underneath all of this. A leading, loaded, or double-barrelled item invalidates the number no matter how good the sample is — see How to Write Unbiased Survey Questions and Question Order Bias. For claim work specifically, two extra rules apply: ask the preference or rating question before any question that reveals the sponsor, and never place a benefit statement in the stem of the question that measures agreement with that benefit.
Factor 4: analyse it properly
For a comparative claim, "higher" is not enough — the difference must be statistically meaningful and stated as such. Run the test, record it, and keep the output in the evidence file. Statistical Significance in Survey Research covers the mechanics; Cross-Tabulation Analysis covers reading subgroup differences without inventing them.
The most common analysis failure in claim work is subgroup shopping: running twenty cuts, finding the one where your product wins, and advertising that cut without saying it was one of twenty. If the claim rests on a subgroup, the ad must describe the subgroup, and the study should have specified it in advance.
Comparative claims raise the bar again
Comparisons with an identifiable competitor are the most challenged category of claim, and the CAP Code addresses them specifically. Rule 3.32 requires that comparisons with an identifiable competitor must not mislead about either the advertised product or the competing product. Rule 3.33 requires that they compare products meeting the same needs or intended for the same purpose. Rule 3.34 requires that they objectively compare one or more material, relevant, verifiable and representative features, which may include price.
Read as design requirements, those three rules mean:
- Same needs, same purpose. Comparing your platform against a tool built for a different job fails at the threshold, however true the numbers are.
- Material and representative features. Winning on a feature nobody weights heavily does not support a general superiority claim. Establish importance first — Key Driver Analysis and MaxDiff are the standard tools.
- Verifiable. Someone else must be able to check it. Publish the method, the date, the base, and the exact question.
For head-to-head preference tests, the design rules are strict and unglamorous: the same respondents evaluate both products, exposure order is rotated to neutralise order effects, brand identity is masked where the claim is about the product rather than the brand, and the sponsor is not revealed until the end. A monadic test where each group sees only one product cannot support a preference claim — nobody expressed a preference.
The evidence file
Assemble this before the ad runs, not when the letter arrives. Regulators and self-regulatory bodies ask for substantially this list, and the file has to be dated before publication.
| Item | Why it is asked for |
|---|---|
| Written claim, in final ad wording | The claim under review is the one consumers saw, including implied claims |
| Study objective and pre-specified analysis plan | Shows the cut was not chosen after seeing the data |
| Universe definition and eligibility criteria | Factor 1 |
| Sampling frame, invite counts, completion rate | Factor 2 |
| Full questionnaire, in field order, with routing | Proves the question was not leading and the sponsor was masked |
| Fieldwork dates and mode | Claims age; a 2023 study rarely supports a 2026 claim |
| Raw data export and the analysis that produces the number | Factor 3 |
| Significance tests for any comparison | Factor 4 |
| Name and qualifications of whoever ran it | Speaks to factor 6 of the reasonable-basis test |
A useful internal rule: no claim ships until a second person has reproduced the number from the raw export. That is a specific application of the pre-launch review gate — and it catches transcription and base-size errors at a rate that surprises teams the first year they run it.
How Koji makes claim-grade evidence practical
The awkward truth about conversational research is that a transcript cannot support a percentage. A quote is not a denominator. This is exactly why Koji studies combine both in a single session rather than forcing a choice between depth and countability.
Every Koji study can carry all six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. For claim work, that division of labour is the whole game:
single_choiceandyes_noproduce clean, unambiguous denominators. The preference claim comes from here.scalegives you distributions you can report honestly, including the top-two-box definition you publish alongside the number.rankingestablishes which attributes are material — the CAP 3.34 requirement — before you claim superiority on one.open_endedwith AI follow-up probing captures why, in the participant own words, and produces the quotes that sit beside the number. Because the AI probes rather than accepting a one-word answer, the qualitative layer is genuinely explanatory rather than decorative.
Because the structured answers are typed fields rather than free text, the number is computed from data, not from someone reading transcripts and counting — which is the difference between an auditable claim and an assertion. Export the raw responses as CSV or JSON straight into the evidence file, and the reproducibility test is a five-minute job.
Three further advantages matter for this specific use case. The AI interviewer asks every participant the same questions in the same way, which removes the moderator variance that interviewer bias introduces into human-run claim studies. Fieldwork that takes a panel agency three weeks completes in a day or two, so the study is fresh when the campaign launches rather than stale. And no moderator is present at all, which removes the social desirability pull that makes people overstate preference to a human being from the sponsoring company.
Start with 10 free credits and run the study before you write the headline. That order is not just good practice — it is the legal requirement.
Frequently asked questions
Can I use qualitative interviews to support a percentage claim?
No. Qualitative research identifies what people think and why; it does not estimate how many. The FTC Endorsement Guides state directly that consumer endorsements are not competent and reliable scientific evidence (16 CFR 255.2(a)). Use interviews to find the claim worth testing, then test it with a structured, representative study.
How large does the sample need to be?
Large enough that the margin of error is smaller than the precision your claim implies. At 95% confidence and a proportion near 50%, n=400 gives roughly ±4.9 points and n=1,000 about ±3.1 points. There is no universal minimum, but publishing a percentage from a base under 100 invites a challenge you will lose.
Can I survey only my own customers?
Yes, if the claim is limited to your own customers and the ad says so. "In a 2026 survey of 412 Koji customers, 87% said…" is defensible. The same data cannot support "marketers prefer" or any claim about the wider category, because the sample was selected on the outcome.
Do I have to disclose the sample size in the ad?
Practice varies by market and medium, but disclosing base size, fieldwork date, and universe is the norm for survey-based claims and is what self-regulatory review expects to see. The safer habit is to publish a short methodology line next to the claim and hold the full evidence file internally.
What happens if a competitor challenges the claim?
You are asked to produce the substantiation you held before the ad ran. Evidence developed afterwards is not a substitute — the reasonable basis doctrine requires prior substantiation, though a regulator may consider later evidence when deciding whether pursuing a case serves the public interest. In practice, the outcome turns on whether your file is dated, complete, and reproducible.
How long does a claim stay substantiated?
Only while the underlying facts hold. Product changes, competitor changes, and market drift all erode a claim. Set an expiry date on every substantiated claim — annually is common, sooner for comparative claims in fast-moving categories — and re-field before the date rather than after a challenge.
Related resources
- Structured Questions Guide — the six question types that produce countable, auditable denominators
- Using Research Quotes in Marketing — the endorsement-law companion to this article
- Statistical Significance in Survey Research — testing the difference behind a comparative claim
- How to Write Unbiased Survey Questions — the wording rules that keep an instrument defensible
- Survey Weighting — correcting a skewed sample before you publish a number from it
- Key Driver Analysis — establishing that the attribute you win on is material
Related Articles
Key Driver Analysis: How to Find What Actually Drives Customer Satisfaction
A complete guide to key driver analysis (KDA) — how to use correlation and regression to identify which factors most influence satisfaction, loyalty, and NPS, how to read an importance-performance matrix, and how AI shortens the path from data to decision.
Sampling Bias: Types, Examples, and How to Avoid It
Sampling bias is when some people in your population are systematically more likely to end up in your sample than others — quietly invalidating your findings. Learn the six main types, classic examples, and how to build a representative sample at scale.
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions
A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.
Survey Weighting: How to Correct a Skewed Sample
A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.