{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-07T14:23:54.833Z"},"content":[{"type":"documentation","id":"c931b4de-01d9-4222-8b02-c62884570309","slug":"ai-claims-substantiation-research","title":"AI Claims Substantiation: How to Prove an AI-Powered Claim Before You Advertise It","url":"https://www.koji.so/docs/ai-claims-substantiation-research","summary":"AI marketing claims carry two separate burdens of proof: an engineering burden (the system performs as described, established through evaluation) and a perception burden (customers do not take away a larger promise, established through consumer research). The FTC Operation AI Comply sweep announced 25 September 2024 targeted deceptive AI claims; in the DoNotPay action the alleged failure was that the company never tested whether its output matched a human lawyer before making that comparison. Claims sort onto a five-rung ladder from composition to substitution, with evidence requirements rising at each rung. A claim perception study runs in four stages: unaided takeaway, aided probing, claim-strength calibration, and disclosure check.","content":"Before you publish the words \"AI-powered,\" you need two separate evidence files: an **engineering file** proving the system does what you say, and a **perception file** proving your customers do not hear a bigger promise than you made. Most teams build the first and skip the second. Every enforcement action in the FTC's Operation AI Comply sweep turned on a gap in one of them, and the most instructive case turned on the absence of a study nobody thought to run.\n\nThis guide covers how to identify the claims your AI marketing actually makes, how to design the consumer research that documents them, and where the line sits between a claim you can support and one you cannot.\n\n## Why AI claims are a distinct problem\n\nMarketers have substantiated performance claims for decades. AI claims are harder for three structural reasons.\n\n**The capability is invisible.** A customer can inspect a jacket and judge whether it is warm. Nobody can inspect a model. When a claim cannot be verified by the person hearing it, the burden of proof sits entirely with the advertiser, and regulators hold those claims to a stricter standard.\n\n**The words are elastic.** \"AI-powered\" has no agreed technical meaning. A regex and a frontier model can both be described that way in a press release. That elasticity is exactly what makes the phrase legally dangerous, because the meaning that counts is not yours - it is the one your audience takes away.\n\n**The category is saturated with hype.** When every competitor claims the same thing, buyers calibrate upward. The same sentence conveys a larger promise in 2026 than it did in 2022, which means claim copy that was accurate when written can drift into deception without a single word changing.\n\nOn 25 September 2024 the FTC announced **Operation AI Comply**, a law enforcement sweep of five actions against companies that, in the Commission's framing, seized on AI hype. Then-Chair Lina Khan summarised the theory of the sweep in one sentence: there is no AI exemption from the laws on the books. The sweep has continued across administrations, and the practical lesson for product marketing teams is narrower and more useful than the headlines suggested.\n\n## The case that should change how you plan research\n\nThe most instructive action was against **DoNotPay**, which marketed an AI service as \"the world's first robot lawyer.\" The FTC's complaint alleged the company promised consumers could \"sue for assault without a lawyer\" and \"generate perfectly valid legal documents in no time.\"\n\nThe allegation worth pinning to your wall is this one: the complaint alleged the company **did not conduct testing to determine whether its AI chatbot's output was equal to the level of a human lawyer**, and had not hired or retained any attorneys. The proposed settlement required a payment of $193,000, notice to subscribers who bought the service between 2021 and 2023, and a prohibition on claiming the ability to substitute for any professional service without evidence.\n\nRead that as a research brief. The company made a **comparative capability claim** - as good as a professional - and the enforcement theory was not that the comparison came out badly. It was that the comparison was never run. The missing artifact was a study.\n\nThat gives you a clean rule that applies far beyond legal tech:\n\n> A claim that your AI performs at the level of a human, a competitor, or a prior process is a comparative claim, and a comparative claim requires a comparison you actually conducted.\n\n## The AI claim ladder\n\nNot all AI claims carry the same burden. Sorting your copy onto a ladder tells you how much evidence each line needs before it ships.\n\n| Rung | Example claim | What you must be able to show |\n| --- | --- | --- |\n| 1. Composition | \"Built with machine learning\" | The technology is genuinely used in the product, not just in a prototype or roadmap |\n| 2. Capability | \"Our AI drafts your summary\" | The feature performs the described function reliably in normal use |\n| 3. Quantified outcome | \"Cuts analysis time by 80%\" | A measurement, with a defined baseline, sample, and method |\n| 4. Comparative | \"As accurate as a senior analyst\" | A head-to-head study against the named comparator |\n| 5. Substitution | \"Replaces your research team\" | Everything above, plus evidence the substitution holds across the full job, not one task |\n\nMost teams write rung-4 and rung-5 copy while holding rung-1 evidence. The ladder is useful precisely because it makes that mismatch visible in a review meeting, before legal sees it and after it is still cheap to fix.\n\nNote that rungs 3 through 5 are engineering questions. You answer them with benchmarks, evaluations, and instrumented usage data - not with customer interviews. Consumer research cannot tell you whether your model is accurate. It answers the other question entirely.\n\n## The perception burden: what did they actually hear?\n\nAdvertising law does not assess the sentence you wrote. It assesses the **net impression** a reasonable member of your audience takes from the whole communication - text, product name, imagery, and layout together. When a piece of marketing supports more than one reasonable interpretation, the advertiser is responsible for substantiating each of them.\n\nThis is where consumer research becomes the load-bearing tool, because implied claims are not knowable from the armchair. A few patterns that reliably produce a gap between intended and received meaning in AI marketing:\n\n- **Autonomy inflation.** \"AI-assisted review\" is routinely heard as \"no human involved.\" The word \"assisted\" carries almost no weight against the word \"AI.\"\n- **Scope creep from an example.** A demo showing one workflow is heard as a claim about every workflow in that category.\n- **Accuracy transfer.** A benchmark figure attached to one narrow task is heard as a general reliability rate for the product.\n- **Credentialed imagery.** Visual cues - lab aesthetics, chart motifs, a clinical or professional setting - can convey a rigor claim on their own, independent of the copy.\n- **Named-comparator drift.** \"Faster than manual analysis\" is often heard as \"faster and at least as good,\" importing a quality claim you never made.\n\nEach of these is measurable. None of them is guessable.\n\n## Designing the claim perception study\n\nThe goal is not to ask people whether they like the copy. It is to reconstruct, in their own words and before you lead them, what they believe you promised.\n\n**Stage 1 - unaided takeaway (the only stage that cannot be skipped).** Show the real asset at real size for a natural amount of time. Then ask an open-ended question: what is this product telling you it can do? Record verbatim. Any prompt naming a capability contaminates everything after it, which is why the sequencing matters more than the wording.\n\n**Stage 2 - aided probing.** Now test specific interpretations you suspect exist. Did the ad tell you a human reviews the output? Did it tell you how accurate the tool is? Yes/no questions here give you countable rates.\n\n**Stage 3 - the claim-strength calibration.** Present the claim alongside a scale to capture how strong a promise it reads as. This converts \"some people over-read it\" into a distribution you can compare across copy variants.\n\n**Stage 4 - disclosure check.** If your copy relies on a qualifier or footnote, test whether it survives contact with the audience. A qualifier that nobody recalls is not a qualifier.\n\nKoji's [structured questions](/docs/structured-questions-guide) map onto these stages directly, and mixing them in one study is what makes the output usable as evidence rather than as anecdote. The platform supports six types - **open_ended, scale, single_choice, multiple_choice, ranking, and yes_no** - and a claim study uses most of them:\n\n| Stage | Question type | What it produces |\n| --- | --- | --- |\n| Unaided takeaway | open_ended | Verbatim comprehension, with AI follow-up probing the reasoning behind it |\n| Interpretation testing | yes_no | Countable rate per implied claim |\n| Claim strength | scale | Distribution showing how big a promise the copy reads as |\n| Attribute attribution | multiple_choice | Which capabilities customers assign to the product |\n| Comparator selection | single_choice | What they think you are comparing yourself against |\n| Feature priority | ranking | Which claimed capability drives the decision |\n\nThe reason the open-ended stage matters most is also the reason it has historically been the expensive one. A written survey box gets you eight words and no follow-up. A human moderator gets you depth but costs weeks of scheduling for a sample large enough to report rates. Koji's AI moderator probes each unaided answer conversationally - asking what specifically gave them that impression - across every participant simultaneously, which is what makes a properly sequenced comprehension study something you can run in an afternoon rather than a quarter.\n\n## Reading the results without fooling yourself\n\n**Report the over-reading rate, not the average.** The relevant number is the proportion of participants who took away a claim you cannot substantiate. A mean score hides exactly the tail that creates liability.\n\n**Treat a minority takeaway as a finding.** If a meaningful share of your audience hears a promise you did not make, that is a problem with the copy, not with those participants. The standard is not whether most people got it right.\n\n**Report coverage honestly.** With a qualitative sample you are documenting that an interpretation exists and why, not estimating its national prevalence. Say \"11 of 40 participants\" rather than \"28% of consumers.\" See [qualitative research validity](/docs/qualitative-research-validity) for how to frame this properly, and the [survey sample size guide](/docs/survey-sample-size-guide) if you need to attach a margin of error to the aided stages.\n\n**Keep the losing rounds.** A copy variant that tested badly and was changed is evidence of a diligent process. Teams routinely delete these, which is precisely backwards.\n\n**Date and version everything.** Because category expectations drift, a comprehension study has a shelf life. Re-run it when the claim changes, when the product changes, or when the category language shifts around you.\n\n## What research cannot do here\n\nBe honest with yourself about the boundary, because getting it wrong is worse than not running the study.\n\nConsumer research **cannot** establish that your AI is accurate, fast, or better than a human. No number of interviews substitutes for an evaluation. If your claim is on rung 3, 4, or 5 of the ladder, the evidence has to come from measurement - see [AI evaluation datasets and golden sets](/docs/ai-evaluation-dataset-golden-set) and [LLM-as-a-judge vs. human evaluation](/docs/llm-as-a-judge-vs-human-evaluation) for how that side is built.\n\nWhat consumer research does, and what nothing else does, is establish **which claims you made**. Those are two different files, and you need both.\n\n## A workable process\n\n1. Inventory every AI claim across site, ads, deck, onboarding, and app copy. Sort each onto the ladder.\n2. For rungs 3 to 5, confirm the engineering evidence exists and is current. Kill or downgrade any claim where it does not.\n3. Run the perception study on the assets as they will actually appear.\n4. Rewrite anything with a material over-reading rate, then re-test the rewrite.\n5. Store both files together with dates, versions, and the assets tested, in a [research repository](/docs/research-repository-guide) your legal team can find without asking you.\n\nTeams that do this discover the same thing repeatedly: the fix is usually one adjective, and they would never have found it by arguing about the copy in a document.\n\n## Frequently asked questions\n\n### Does saying \"AI-powered\" require substantiation even if we do use AI?\n\nYes. Using AI somewhere in the product supports a composition claim, but \"AI-powered\" in context often conveys more - that AI performs the specific task being advertised, at a level worth paying for. Because advertisers are responsible for each reasonable interpretation of their marketing, the question is what the phrase conveys in your specific layout, not whether the underlying technology exists somewhere in your stack.\n\n### What made the DoNotPay case different from an ordinary false advertising case?\n\nThe alleged failure was evidentiary rather than performative. According to the FTC's complaint the company had not tested whether its output matched the level of a human lawyer before making that comparison in its marketing. The lesson for research planning is that a comparative claim creates an obligation to run the comparison, and the absence of the study is itself the exposure.\n\n### Can customer interviews substantiate that our AI is accurate?\n\nNo, and you should resist any temptation to present them that way. Customer experience is subject to placebo effects, selection, and recall error. Accuracy is established through evaluation against a labelled dataset. Interviews establish what your marketing communicated, which is a different and equally necessary piece of evidence.\n\n### How large a sample does a claim perception study need?\n\nIt depends on which stage you are reporting. The unaided takeaway stage is qualitative - 20 to 40 well-probed conversations will surface the interpretations that exist and explain why they form. If you intend to report a rate for a specific implied claim, size the aided stage like any other proportion estimate and report a margin of error alongside it.\n\n### Should we test the claim or the whole page?\n\nTest the asset as the customer will encounter it. Net impression is assessed across the entire communication, including imagery, product name, and adjacent copy. A claim that is defensible in isolation can become misleading beside a chart or a testimonial, and testing the sentence alone will miss exactly that interaction.\n\n### How often should we re-run this?\n\nWhenever the claim changes, the product changes, or roughly annually for standing claims. AI marketing language inflates quickly across a category, so a phrase that read modestly at launch can read as a much larger promise a year later without anyone editing it.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and when to use each\n- [Advertising Claim Substantiation](/docs/survey-claim-substantiation-advertising) - designing survey research that backs a marketing claim\n- [Green Claims Research](/docs/green-claims-consumer-perception-research) - the same two-burden structure applied to sustainability claims\n- [Using Research Quotes in Marketing](/docs/research-quotes-marketing-ftc-endorsement-guides) - endorsement rules for customer quotes\n- [AI Evaluation Datasets and Golden Sets](/docs/ai-evaluation-dataset-golden-set) - building the engineering side of the evidence\n- [Message Testing with AI Interviews](/docs/content-testing-guide) - testing microcopy and labels with real users\n- [Research Peer Review](/docs/research-peer-review-qa-gate) - the pre-launch QA gate for study design\n\n---\n\n**Run your first claim perception study free.** New Koji accounts include 10 credits - enough to field a full unaided-takeaway study with AI follow-up probing on every response, and see exactly what your customers think you promised.","category":"Research Methods","lastModified":"2026-08-07T03:19:15.488877+00:00","metaTitle":"AI Claims Substantiation: Proving an AI-Powered Claim | Koji","metaDescription":"How to substantiate an AI-powered marketing claim: the engineering burden vs the perception burden, lessons from FTC Operation AI Comply, and how to design a claim comprehension study.","keywords":["ai claims substantiation","ai washing","ai-powered claim","operation ai comply","claim perception research","ai marketing claims","implied claims"],"aiSummary":"AI marketing claims carry two separate burdens of proof: an engineering burden (the system performs as described, established through evaluation) and a perception burden (customers do not take away a larger promise, established through consumer research). The FTC Operation AI Comply sweep announced 25 September 2024 targeted deceptive AI claims; in the DoNotPay action the alleged failure was that the company never tested whether its output matched a human lawyer before making that comparison. Claims sort onto a five-rung ladder from composition to substitution, with evidence requirements rising at each rung. A claim perception study runs in four stages: unaided takeaway, aided probing, claim-strength calibration, and disclosure check.","aiPrerequisites":["Basic familiarity with your product marketing assets","Understanding of qualitative vs quantitative evidence"],"aiLearningOutcomes":["Sort every AI claim onto a five-rung evidence ladder","Identify implied claims your marketing makes unintentionally","Design a four-stage claim perception study","Distinguish what consumer research can and cannot substantiate"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}