AI Claims Substantiation: How to Prove an AI-Powered Claim Before You Advertise It
Every "AI-powered" claim carries two burdens of proof: an engineering burden (does the system do it?) and a perception burden (what do customers hear?). Learn how to build both evidence files with consumer research before you ship the copy.
Before you publish the words "AI-powered," you need two separate evidence files: an engineering file proving the system does what you say, and a perception file proving your customers do not hear a bigger promise than you made. Most teams build the first and skip the second. Every enforcement action in the FTC's Operation AI Comply sweep turned on a gap in one of them, and the most instructive case turned on the absence of a study nobody thought to run.
This guide covers how to identify the claims your AI marketing actually makes, how to design the consumer research that documents them, and where the line sits between a claim you can support and one you cannot.
Why AI claims are a distinct problem
Marketers have substantiated performance claims for decades. AI claims are harder for three structural reasons.
The capability is invisible. A customer can inspect a jacket and judge whether it is warm. Nobody can inspect a model. When a claim cannot be verified by the person hearing it, the burden of proof sits entirely with the advertiser, and regulators hold those claims to a stricter standard.
The words are elastic. "AI-powered" has no agreed technical meaning. A regex and a frontier model can both be described that way in a press release. That elasticity is exactly what makes the phrase legally dangerous, because the meaning that counts is not yours - it is the one your audience takes away.
The category is saturated with hype. When every competitor claims the same thing, buyers calibrate upward. The same sentence conveys a larger promise in 2026 than it did in 2022, which means claim copy that was accurate when written can drift into deception without a single word changing.
On 25 September 2024 the FTC announced Operation AI Comply, a law enforcement sweep of five actions against companies that, in the Commission's framing, seized on AI hype. Then-Chair Lina Khan summarised the theory of the sweep in one sentence: there is no AI exemption from the laws on the books. The sweep has continued across administrations, and the practical lesson for product marketing teams is narrower and more useful than the headlines suggested.
The case that should change how you plan research
The most instructive action was against DoNotPay, which marketed an AI service as "the world's first robot lawyer." The FTC's complaint alleged the company promised consumers could "sue for assault without a lawyer" and "generate perfectly valid legal documents in no time."
The allegation worth pinning to your wall is this one: the complaint alleged the company did not conduct testing to determine whether its AI chatbot's output was equal to the level of a human lawyer, and had not hired or retained any attorneys. The proposed settlement required a payment of $193,000, notice to subscribers who bought the service between 2021 and 2023, and a prohibition on claiming the ability to substitute for any professional service without evidence.
Read that as a research brief. The company made a comparative capability claim - as good as a professional - and the enforcement theory was not that the comparison came out badly. It was that the comparison was never run. The missing artifact was a study.
That gives you a clean rule that applies far beyond legal tech:
A claim that your AI performs at the level of a human, a competitor, or a prior process is a comparative claim, and a comparative claim requires a comparison you actually conducted.
The AI claim ladder
Not all AI claims carry the same burden. Sorting your copy onto a ladder tells you how much evidence each line needs before it ships.
| Rung | Example claim | What you must be able to show |
|---|---|---|
| 1. Composition | "Built with machine learning" | The technology is genuinely used in the product, not just in a prototype or roadmap |
| 2. Capability | "Our AI drafts your summary" | The feature performs the described function reliably in normal use |
| 3. Quantified outcome | "Cuts analysis time by 80%" | A measurement, with a defined baseline, sample, and method |
| 4. Comparative | "As accurate as a senior analyst" | A head-to-head study against the named comparator |
| 5. Substitution | "Replaces your research team" | Everything above, plus evidence the substitution holds across the full job, not one task |
Most teams write rung-4 and rung-5 copy while holding rung-1 evidence. The ladder is useful precisely because it makes that mismatch visible in a review meeting, before legal sees it and after it is still cheap to fix.
Note that rungs 3 through 5 are engineering questions. You answer them with benchmarks, evaluations, and instrumented usage data - not with customer interviews. Consumer research cannot tell you whether your model is accurate. It answers the other question entirely.
The perception burden: what did they actually hear?
Advertising law does not assess the sentence you wrote. It assesses the net impression a reasonable member of your audience takes from the whole communication - text, product name, imagery, and layout together. When a piece of marketing supports more than one reasonable interpretation, the advertiser is responsible for substantiating each of them.
This is where consumer research becomes the load-bearing tool, because implied claims are not knowable from the armchair. A few patterns that reliably produce a gap between intended and received meaning in AI marketing:
- Autonomy inflation. "AI-assisted review" is routinely heard as "no human involved." The word "assisted" carries almost no weight against the word "AI."
- Scope creep from an example. A demo showing one workflow is heard as a claim about every workflow in that category.
- Accuracy transfer. A benchmark figure attached to one narrow task is heard as a general reliability rate for the product.
- Credentialed imagery. Visual cues - lab aesthetics, chart motifs, a clinical or professional setting - can convey a rigor claim on their own, independent of the copy.
- Named-comparator drift. "Faster than manual analysis" is often heard as "faster and at least as good," importing a quality claim you never made.
Each of these is measurable. None of them is guessable.
Designing the claim perception study
The goal is not to ask people whether they like the copy. It is to reconstruct, in their own words and before you lead them, what they believe you promised.
Stage 1 - unaided takeaway (the only stage that cannot be skipped). Show the real asset at real size for a natural amount of time. Then ask an open-ended question: what is this product telling you it can do? Record verbatim. Any prompt naming a capability contaminates everything after it, which is why the sequencing matters more than the wording.
Stage 2 - aided probing. Now test specific interpretations you suspect exist. Did the ad tell you a human reviews the output? Did it tell you how accurate the tool is? Yes/no questions here give you countable rates.
Stage 3 - the claim-strength calibration. Present the claim alongside a scale to capture how strong a promise it reads as. This converts "some people over-read it" into a distribution you can compare across copy variants.
Stage 4 - disclosure check. If your copy relies on a qualifier or footnote, test whether it survives contact with the audience. A qualifier that nobody recalls is not a qualifier.
Koji's structured questions map onto these stages directly, and mixing them in one study is what makes the output usable as evidence rather than as anecdote. The platform supports six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and a claim study uses most of them:
| Stage | Question type | What it produces |
|---|---|---|
| Unaided takeaway | open_ended | Verbatim comprehension, with AI follow-up probing the reasoning behind it |
| Interpretation testing | yes_no | Countable rate per implied claim |
| Claim strength | scale | Distribution showing how big a promise the copy reads as |
| Attribute attribution | multiple_choice | Which capabilities customers assign to the product |
| Comparator selection | single_choice | What they think you are comparing yourself against |
| Feature priority | ranking | Which claimed capability drives the decision |
The reason the open-ended stage matters most is also the reason it has historically been the expensive one. A written survey box gets you eight words and no follow-up. A human moderator gets you depth but costs weeks of scheduling for a sample large enough to report rates. Koji's AI moderator probes each unaided answer conversationally - asking what specifically gave them that impression - across every participant simultaneously, which is what makes a properly sequenced comprehension study something you can run in an afternoon rather than a quarter.
Reading the results without fooling yourself
Report the over-reading rate, not the average. The relevant number is the proportion of participants who took away a claim you cannot substantiate. A mean score hides exactly the tail that creates liability.
Treat a minority takeaway as a finding. If a meaningful share of your audience hears a promise you did not make, that is a problem with the copy, not with those participants. The standard is not whether most people got it right.
Report coverage honestly. With a qualitative sample you are documenting that an interpretation exists and why, not estimating its national prevalence. Say "11 of 40 participants" rather than "28% of consumers." See qualitative research validity for how to frame this properly, and the survey sample size guide if you need to attach a margin of error to the aided stages.
Keep the losing rounds. A copy variant that tested badly and was changed is evidence of a diligent process. Teams routinely delete these, which is precisely backwards.
Date and version everything. Because category expectations drift, a comprehension study has a shelf life. Re-run it when the claim changes, when the product changes, or when the category language shifts around you.
What research cannot do here
Be honest with yourself about the boundary, because getting it wrong is worse than not running the study.
Consumer research cannot establish that your AI is accurate, fast, or better than a human. No number of interviews substitutes for an evaluation. If your claim is on rung 3, 4, or 5 of the ladder, the evidence has to come from measurement - see AI evaluation datasets and golden sets and LLM-as-a-judge vs. human evaluation for how that side is built.
What consumer research does, and what nothing else does, is establish which claims you made. Those are two different files, and you need both.
A workable process
- Inventory every AI claim across site, ads, deck, onboarding, and app copy. Sort each onto the ladder.
- For rungs 3 to 5, confirm the engineering evidence exists and is current. Kill or downgrade any claim where it does not.
- Run the perception study on the assets as they will actually appear.
- Rewrite anything with a material over-reading rate, then re-test the rewrite.
- Store both files together with dates, versions, and the assets tested, in a research repository your legal team can find without asking you.
Teams that do this discover the same thing repeatedly: the fix is usually one adjective, and they would never have found it by arguing about the copy in a document.
Frequently asked questions
Does saying "AI-powered" require substantiation even if we do use AI?
Yes. Using AI somewhere in the product supports a composition claim, but "AI-powered" in context often conveys more - that AI performs the specific task being advertised, at a level worth paying for. Because advertisers are responsible for each reasonable interpretation of their marketing, the question is what the phrase conveys in your specific layout, not whether the underlying technology exists somewhere in your stack.
What made the DoNotPay case different from an ordinary false advertising case?
The alleged failure was evidentiary rather than performative. According to the FTC's complaint the company had not tested whether its output matched the level of a human lawyer before making that comparison in its marketing. The lesson for research planning is that a comparative claim creates an obligation to run the comparison, and the absence of the study is itself the exposure.
Can customer interviews substantiate that our AI is accurate?
No, and you should resist any temptation to present them that way. Customer experience is subject to placebo effects, selection, and recall error. Accuracy is established through evaluation against a labelled dataset. Interviews establish what your marketing communicated, which is a different and equally necessary piece of evidence.
How large a sample does a claim perception study need?
It depends on which stage you are reporting. The unaided takeaway stage is qualitative - 20 to 40 well-probed conversations will surface the interpretations that exist and explain why they form. If you intend to report a rate for a specific implied claim, size the aided stage like any other proportion estimate and report a margin of error alongside it.
Should we test the claim or the whole page?
Test the asset as the customer will encounter it. Net impression is assessed across the entire communication, including imagery, product name, and adjacent copy. A claim that is defensible in isolation can become misleading beside a chart or a testimonial, and testing the sentence alone will miss exactly that interaction.
How often should we re-run this?
Whenever the claim changes, the product changes, or roughly annually for standing claims. AI marketing language inflates quickly across a category, so a phrase that read modestly at launch can read as a much larger promise a year later without anyone editing it.
Related Resources
- Structured Questions Guide - the six question types and when to use each
- Advertising Claim Substantiation - designing survey research that backs a marketing claim
- Green Claims Research - the same two-burden structure applied to sustainability claims
- Using Research Quotes in Marketing - endorsement rules for customer quotes
- AI Evaluation Datasets and Golden Sets - building the engineering side of the evidence
- Message Testing with AI Interviews - testing microcopy and labels with real users
- Research Peer Review - the pre-launch QA gate for study design
Run your first claim perception study free. New Koji accounts include 10 credits - enough to field a full unaided-takeaway study with AI follow-up probing on every response, and see exactly what your customers think you promised.
Related Articles
Evaluation Datasets for AI Products: How to Build a Golden Set from Real User Research (2026)
How to construct and maintain the golden dataset your AI evals run against — sizing and confidence intervals, the four-bucket structure, label-error rates in published benchmarks, sourcing acceptance criteria from real users, and versioning against overfitting.
Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)
Six methods for testing whether your words actually work — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale instead of one participant at a time.
Green Claims Research: How to Substantiate a Sustainability Claim With Consumer Perception Evidence
Every environmental claim carries two substantiation burdens: the science burden and the perception burden. Lab data answers the first. Only research answers the second, and it is the one companies fail. Here is how to design a green claim perception study.
Qualitative Research Validity and Reliability: How to Build Studies You Can Trust
A practical guide to Lincoln and Guba's trustworthiness framework — credibility, transferability, dependability, and confirmability — and how to build each into your qualitative research studies.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
Using Research Quotes in Marketing: FTC Endorsement Guides, Material Connections, and the Reviews Rule
The moment a customer quote leaves your research repository and appears in an advertisement, it stops being data and becomes an endorsement. Three obligations attach immediately: the words must be faithful, the experience must be typical or disclosed, and any material connection must be visible.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Advertising Claim Substantiation: How to Design Survey Research That Backs a Marketing Claim
A claim like "9 out of 10 customers recommend us" is a regulated assertion, and the evidence has to exist before the ad runs. This is how to design the study so the number survives a challenge from a regulator, a competitor, or a self-regulatory body.