Health and Nutrition Claims: What Consumer Research Can and Cannot Substantiate
The FTC says consumer surveys are never sufficient to substantiate a health benefit claim - and in the same document says consumer surveys are valuable for determining what claim you made. Both are true. Learn where the line sits and how to build the perception evidence file.
Start with the sentence that should govern every conversation about research and health marketing. In its Health Products Compliance Guidance, the FTC states that anecdotal evidence about individual consumer experiences, including surveys of consumer experiences, are never sufficient to substantiate claims about the effects of a health product.
Now hold that beside a second statement from the same document: extrinsic evidence such as consumer surveys and copy tests can be valuable in determining how consumers interpret implied claims.
Both are correct, and together they define the job precisely. Consumer research can never tell you that your product works. It is one of the only tools that tells you what you claimed - and under FTC law you are responsible for substantiating every reasonable interpretation of your marketing, not just the sentence you intended.
This guide covers that boundary, the evidence the efficacy side actually requires, and how to build the perception file that almost no health brand has.
Two files, two owners
| Efficacy file | Perception file | |
|---|---|---|
| Question | Does the product produce the benefit? | What benefit do consumers understand us to be promising? |
| Standard | Competent and reliable scientific evidence | Net impression on a reasonable member of the audience |
| Evidence | Randomised controlled human clinical trials | Copy tests, comprehension studies, consumer interviews |
| Owner | Clinical, scientific affairs | Insights, marketing |
| Typically funded | Heavily | Almost never |
Enforcement usually finds the gap in the second column. A brand holds a genuine study supporting a modest, carefully worded finding, and its advertising conveys something considerably larger through imagery, placement, and emphasis. The science was real; the claim that was actually communicated was never substantiated because nobody ever measured what it was.
What the efficacy file requires
Know this side even though research is not how you build it, because it determines which claims are worth testing at all.
The FTC's standard is competent and reliable scientific evidence, defined as tests, analyses, research, or studies that (1) have been conducted and evaluated in an objective manner by experts in the relevant disease, condition, or function, and (2) are generally accepted in the profession to yield accurate and reliable results. The evidence must also be sufficient in quality and quantity, judged in light of the entire body of relevant and reliable evidence.
As a general matter, health benefit substantiation needs to be randomised, controlled human clinical testing. The guidance is explicit about what does not clear the bar:
- Epidemiological or observational studies can show association but do not prove causation. High-quality epidemiologic evidence is accepted only where experts consider it an acceptable substitute and RCTs are not feasible.
- Animal and in vitro studies may provide supporting background but are not sufficient without human RCT confirmation.
- Consumer experience, including surveys of it, is never sufficient - genuine experiences may reflect placebo effects or unrelated change.
- A practitioner's observation about effects on patients is likewise anecdotal.
- Public health recommendations from medical organisations are not substantiation; they reflect judgement on available evidence, not a finding of causation.
The quality criteria are equally specific: a control group, appropriate randomisation, blinding, statistically significant results, and results that are clinically meaningful rather than merely detectable. Studies using multiple outcome measures should report all outcomes with statistical adjustment, and post hoc analysis departing from the original protocol is flagged as a p-hacking indicator that generally does not provide reliable substantiation.
The amount and type of substantiation required scales with the type of product, the type of claim, the consequences of a false claim, the benefits of a truthful one weighed against the cost of substantiation, and the amount of substantiation experts in the field consider reasonable. Health claims that consumers cannot verify themselves - a benefit subject to placebo effect, or one relating to a naturally varying condition - are held to a more exacting standard.
If your claim cannot be supported at that level, no amount of consumer research rescues it. The correct response is to change the claim.
What the perception file requires
Here is where research is not merely useful but close to irreplaceable.
Claims are assessed on the net impression conveyed by all elements of the advertisement together - text, product name, charts, graphs, and images. When an advertisement supports more than one reasonable interpretation, the advertiser must substantiate each of them. And claims are evaluated from the standpoint of the intended audience, with the guidance noting that some audiences may be particularly susceptible to exaggerated claims.
The guidance illustrates implied claims with an example worth memorising: a weight-loss brochure showing doctors in white lab coats looking through microscopes, molecular structures, and a stack of medical journals likely conveys an implied claim that the product has been clinically proven effective - with no such words appearing. Visual vocabulary makes claims.
The disclosure threshold you can actually measure
The single most useful sentence in the guidance for a research team is this: the ultimate test of a disclosure is the net impression consumers take away, and if a significant minority of consumers take a misleading claim from an ad despite a disclosure, the disclosure is not sufficient.
That is an empirical threshold. Not a drafting standard, not a font-size rule - a statement about a measurable proportion of an audience. It means:
- The pass mark is not "most people understood." A minority misreading is a failure.
- The only way to know is to test, because there is no armchair method of estimating the size of a misreading minority.
- If an effective disclosure is not achievable, the guidance is direct: modify the claim so a disclosure is unnecessary, or do not make it.
Qualifiers that do not qualify
The guidance is unusually specific about language that fails, and each item is a testable hypothesis:
- Vague hedges are inadequate. Saying a product "may" have a benefit or "helps" achieve it is not enough.
- Science-stage words read as praise. Consumers are likely to interpret "promising," "preliminary," "initial," or "pilot" as positive product attributes rather than substantial disclaimers - particularly when the study is positively touted in the same advertisement.
- Qualifying emerging science is very difficult. The guidance acknowledges directly that it is hard to adequately convey the uncertain and limited nature of support for a claim based on still-emerging science.
- Placement matters. A fine-print disclosure at the bottom of an advertisement is not clear and conspicuous; a disclosure immediately next to the claim, in the same font size and high contrast, is far more likely to work.
- Disclosures buried in terms and conditions do not cure a claim made in the advertisement.
Every one of those is a hypothesis you can test in an afternoon, and the results routinely surprise the people who wrote the copy.
Designing the study
Stage 1 - unaided takeaway. Show the real asset. Ask what this product does and who it is for, open-ended, before any health vocabulary enters the conversation. Record verbatim. This is the stage that cannot be skipped or reordered.
Stage 2 - implied claim inventory. Test the specific interpretations you suspect. Does this tell you the product treats a condition, or supports a normal function? Does it tell you the product was clinically tested? Does it suggest results without diet or exercise change? Binary questions here give countable rates per claim.
Stage 3 - disclosure effectiveness. Present the asset with your qualifier in place and measure how many still take away the unsupported claim. This directly operationalises the significant-minority threshold and is the stage that produces the number worth reporting.
Stage 4 - audience check. Because claims are judged from the standpoint of the intended audience, run the study on the population you are actually targeting. A claim that reads modestly to a general sample can read very differently to people managing the condition it touches.
Koji's six structured question types cover the whole design:
| Stage | Question type | Output |
|---|---|---|
| Unaided takeaway | open_ended | Verbatim comprehension with AI follow-up on the reasoning |
| Implied claim testing | yes_no | Countable rate per implied claim |
| Claim strength | scale | How strong a promise the copy reads as |
| Benefit attribution | multiple_choice | Which benefits consumers assign to the product |
| Condition inference | single_choice | Whether they read it as treating a condition |
| Benefit weighting | ranking | Which claimed benefit drives purchase |
The unaided stage is the expensive one to do properly, because a useful answer needs a follow-up question tailored to whatever the participant just said - what specifically gave you that impression? Koji's AI moderator asks that of every participant simultaneously, which is what makes a well-sequenced comprehension study something you can field before a launch rather than after a warning letter.
The trap to avoid
Do not let a positive perception study drift into the efficacy file. A finding that customers believe the product helps them is a fact about belief. Presented as support for the benefit itself, it is precisely the anecdotal consumer-experience evidence the guidance rules out.
Label the files separately, store them separately, and write the limitation into the report. A research report that states plainly what it does not establish is more credible internally and considerably more defensible externally. The same discipline applies to customer testimonial interviews - see using research quotes in marketing for the endorsement rules that attach the moment a quote enters an advertisement.
Note too that FDA labeling categories - including structure/function claims - do not govern the FTC's assessment of those claims in advertising. Clearing a claim under one framework does not clear it under the other.
A workable sequence
- Inventory every health-adjacent claim across pack, site, ads, and influencer briefs.
- For each, confirm the efficacy file meets the competent and reliable standard. Downgrade or delete anything that does not.
- Run the perception study on the surviving claims, on the target audience, with assets as they will appear.
- Where a significant minority takes away an unsupported claim, rewrite - then re-test the rewrite rather than assuming it worked.
- Retain both files with dates and asset versions in a research repository.
For the shared structure across claim types, see green claims research and advertising claim substantiation; for reporting qualitative counts honestly, see qualitative research validity.
Frequently asked questions
Can consumer research substantiate a health benefit claim?
No. The FTC guidance states that anecdotal evidence about individual consumer experiences, including surveys of consumer experiences, is never sufficient to substantiate claims about the effects of a health product, because genuine experiences may reflect placebo effects or factors unrelated to the product. Health benefit substantiation generally requires randomised, controlled human clinical testing.
Then why run consumer research at all in this category?
Because you must substantiate every reasonable interpretation your marketing conveys, and you cannot substantiate a claim you have not identified. The same guidance notes that extrinsic evidence such as consumer surveys and copy tests can be valuable in determining how consumers interpret implied claims. Research builds the perception file, not the efficacy file.
What counts as an adequate disclosure?
The test is the net impression consumers take from the advertisement with the disclosure present. The guidance states that if a significant minority of consumers still take a misleading claim despite the disclosure, the disclosure is not sufficient. Vague hedges like "may" or "helps" are inadequate, and words such as promising, preliminary, initial, or pilot are likely to be read as positive attributes rather than as disclaimers.
Do our images and product name matter, or just the claim text?
They matter. Net impression is assessed across all elements including text, product name, charts, graphs, and images. The guidance gives the example of a brochure showing doctors in lab coats, microscopes, and medical journals conveying an implied clinically-proven claim without saying so in words.
Does FDA structure/function compliance cover us with the FTC?
No. The guidance is explicit that FDA labeling rules for structure/function claims do not govern the FTC's assessment of those claims in advertising. The two frameworks operate independently and a claim can be compliant under one and deceptive under the other.
Who should we recruit for a health claim perception study?
The audience you actually target, because claims are evaluated from the standpoint of the intended audience. A claim that reads modestly to a general population sample can convey a much stronger promise to people managing the condition it relates to, and that is the group whose interpretation governs.
Related Resources
- Structured Questions Guide - the six question types and when to use each
- Advertising Claim Substantiation - designing survey research that backs a marketing claim
- Green Claims Research - the same two-file structure for sustainability claims
- AI Claims Substantiation - the engineering burden and the perception burden for AI copy
- Using Research Quotes in Marketing - endorsement rules for testimonials
- HCP Research Guide - recruiting and interviewing healthcare professionals
- AI Research for Pharma and Life Sciences - the wider life sciences programme
Build your perception file. New Koji accounts include 10 credits - enough to run an unaided takeaway study with AI follow-up probing and find out which claim your packaging is actually making.
Related Articles
AI Claims Substantiation: How to Prove an AI-Powered Claim Before You Advertise It
Every "AI-powered" claim carries two burdens of proof: an engineering burden (does the system do it?) and a perception burden (what do customers hear?). Learn how to build both evidence files with consumer research before you ship the copy.
AI Customer Research for Pharma & Life Sciences
How pharma, biotech, and medical device teams use AI interviews to gather HCP and patient insights at scale — faster fielding, deeper qualitative signal, and compliance-aware workflows.
Green Claims Research: How to Substantiate a Sustainability Claim With Consumer Perception Evidence
Every environmental claim carries two substantiation burdens: the science burden and the perception burden. Lab data answers the first. Only research answers the second, and it is the one companies fail. Here is how to design a green claim perception study.
HCP Research: How to Recruit and Interview Physicians, Nurses and Other Healthcare Professionals
Healthcare professionals are the hardest audience in research to reach and the most heavily regulated to pay. Here is how to recruit them, set defensible honoraria, handle Sunshine Act reporting and adverse events, and run interviews that respect a twelve-hour shift.
Qualitative Research Validity and Reliability: How to Build Studies You Can Trust
A practical guide to Lincoln and Guba's trustworthiness framework — credibility, transferability, dependability, and confirmability — and how to build each into your qualitative research studies.
Using Research Quotes in Marketing: FTC Endorsement Guides, Material Connections, and the Reviews Rule
The moment a customer quote leaves your research repository and appears in an advertisement, it stops being data and becomes an endorsement. Three obligations attach immediately: the words must be faithful, the experience must be typical or disclosed, and any material connection must be visible.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Advertising Claim Substantiation: How to Design Survey Research That Backs a Marketing Claim
A claim like "9 out of 10 customers recommend us" is a regulated assertion, and the evidence has to exist before the ad runs. This is how to design the study so the number survives a challenge from a regulator, a competitor, or a self-regulatory body.