Back to docs
Research Methods

Money-Back and Satisfaction Guarantees: Testing What Your Promise Actually Promises

The FTC Guides say you should use Satisfaction Guarantee or Money Back Guarantee only if you refund the full purchase price on request. Here is how to test what buyers believe your guarantee covers, and whether claimants got it.

Answer first: the FTC's Guides for the Advertising of Warranties and Guarantees (16 CFR Part 239) say a seller should use the terms "Satisfaction Guarantee," "Money Back Guarantee," or "Free Trial Offer" only if it refunds the full purchase price at the purchaser's request, and should disclose any material limitations clearly and prominently. That creates two research questions no legal review can answer: what do buyers believe your guarantee covers after they have seen it with its disclosures, and did the people who actually claimed it get what they expected. Platforms like Koji answer both in one study, because AI-moderated interviews scale to the prospect population and go deep enough with the smaller claimant population to explain what went wrong.

Why a guarantee is a research problem

A guarantee is unusual among marketing claims. Most claims describe the product. A guarantee describes what happens if the product disappoints, which means it is a promise about a future interaction with your company that the buyer has never had. They have to infer its terms entirely from a short phrase and whatever fine print sits near it.

Inference is where things go wrong. "Money Back Guarantee" is four words. The operational reality behind it might involve a 30-day window, a restocking fee, return shipping at the buyer's expense, original packaging, proof of purchase, and a refund that excludes the shipping they originally paid. Each of those may be lawful and each may be disclosed. The question the Guides put front and centre is whether disclosure produced accurate belief, and that has never been answerable from inside the marketing department.

What Part 239 actually says

The Guides were issued at 50 FR 18470 on May 1, 1985, under the FTC's Section 5 authority, and they remain the clearest official statement of what these phrases are supposed to mean. Four sections do the work.

239.3(a), the full-price rule. A seller or manufacturer "should use the terms Satisfaction Guarantee, Money Back Guarantee, Free Trial Offer, or similar representations in advertising only if the seller or manufacturer, as the case may be, refunds the full purchase price of the advertised product at the purchaser's request."

Two phrases carry the weight. "Full purchase price" sets the amount. "At the purchaser's request" sets the trigger, and it is a low one. It is not at the company's discretion, not subject to a reason being accepted, and not conditioned on the company agreeing the product failed.

239.3(b), the limitations rule. An advertisement mentioning a satisfaction guarantee "should disclose, with such clarity and prominence as will be noticed and understood by prospective purchasers, any material limitations or conditions that apply." The FTC's own examples are instructive because they show disclosure integrated into the claim rather than appended to it: "If not completely satisfied with Acme Spot Remover, return the unused portion within 30 days for a full refund," and "Money Back Guarantee! Just return the ABC watch in its original package and ABC will fully refund your money."

239.4, the lifetime rule. If an advertisement uses "lifetime," "life," or similar words to describe duration, it "should disclose, with such clarity and prominence as will be noticed and understood by prospective purchasers, the life to which the representation refers." The Commission's examples make the ambiguity explicit: a muffler guarantee measured by the life of the car it is installed in, transferable even if the owner sells or gives the car away, versus a battery guarantee good only for as long as the original purchaser owns the car.

239.5, the performance rule. A seller "should advertise that a product is warranted or guaranteed only if the seller or manufacturer, as the case may be, promptly and fully performs its obligations under the warranty or guarantee."

That last section is the one most teams skip, and it is the most consequential. It converts your refund operation into part of your advertising claim. A guarantee that reads perfectly and is honoured slowly is not a communication problem. It is a claim your operation is not supporting.

Three things to measure

1. The full-price test

Ask buyers what "full refund" includes. Offer a multiple_choice list: the item price, sales tax, the shipping I paid to receive it, the shipping I pay to send it back, any restocking fee waived, the full amount charged to my card.

Buyers routinely include outbound shipping and often include return shipping. The Guides speak of the "full purchase price of the advertised product," which is narrower. The gap between what your guarantee returns and what buyers believe it returns is a number you can produce in an afternoon, and it predicts a specific downstream event: the refund that was processed correctly and still produced an angry customer, because the amount was lower than expected.

This is the same mechanism as the remedy expectation gap in warranty comprehension research. In both cases the company performs exactly as documented and the customer experiences a broken promise, because the promise they heard was never the promise that was written.

2. The lifetime attribution question

239.4 exists because "lifetime" is one of the most reliably misread words in commerce. Ask it directly as a single_choice: this product carries a lifetime guarantee. Whose or what lifetime? My lifetime, the expected working life of the product, as long as I own it, as long as the company sells it, or I do not know.

Most buyers pick their own lifetime. Almost no guarantee means that. If your answer is "the reasonable service life of the product as determined by us," the distance between that and the modal buyer answer is the disclosure you are missing, and 239.4 says clarity and prominence sufficient to be noticed and understood is the standard. Understood is an empirical word.

3. The claimant experience

239.5 makes performance part of the claim, so a guarantee study that only talks to prospects is half a study. The other half is the population that invoked the guarantee.

PopulationWhat it tells youTypical size
Prospects who have seen the ad but not boughtWhether the claim as advertised creates accurate beliefLarge; scale for reliable rates
Buyers who never claimedWhether the guarantee affected the purchase, and whether they knew it appliedMedium
Claimants whose refund was approvedWhether "promptly and fully" describes what happenedSmall, high signal
Claimants who abandoned mid-processThe most valuable and least studied group of allVery small, highest signal

That last row deserves emphasis. Someone who started a refund and gave up does not appear in your refund rate, your approval rate, or your resolution time. Every operational metric you have is computed over people who finished. Abandonment is invisible in exactly the way a belief-based failure always is, and it is measurable only by asking. The parallel to subscribers who believe they cancelled but did not, covered in auto-renewal and cancellation research, is exact.

The specificity paradox

Here is a genuine tension in this area of law that is worth understanding before you write a guarantee.

15 U.S.C. 2303(b) provides that the Magnuson-Moss designation requirements do not apply to "statements or representations which are similar to expressions of general policy concerning customer satisfaction and which are not subject to any specific limitations." A broad, unconditional statement of goodwill sits outside the full-or-limited labelling regime entirely.

Attach specific limitations and you may lose that shelter, while simultaneously picking up the 239.3(b) duty to disclose those limitations clearly and prominently.

So the legal gradient and the commercial gradient run in opposite directions. A guarantee is legally lightest when it is vaguest, and commercially strongest when it is most specific. "We stand behind everything we sell" carries little regulatory freight and little persuasive force. "Return it within 60 days for a full refund including the shipping you paid, no reason required" is a real reason to buy and a real set of obligations.

Research is what lets you move toward specificity deliberately instead of accidentally. Test two or three candidate wordings, measure what each one causes buyers to believe, then choose the one whose induced belief matches what you will actually do. Choosing wording by legal caution alone reliably produces a guarantee nobody finds persuasive; choosing it by marketing appeal alone reliably produces one you cannot honour.

Designing the study

Recruit people who plausibly buy in your category and have not seen your guarantee before. Show the guarantee exactly as it appears in the wild, including the disclosure, at the size and placement a real buyer would encounter. A guarantee tested as isolated body copy is not the guarantee your customers see.

StageQuestion typeWhat it produces
Unaided read-back of what the guarantee promisesopen_ended with AI probingThe belief in the buyer's own words
What "full refund" includesmultiple_choiceThe full-price gap
Whose lifetime, if you use the wordsingle_choiceThe 239.4 attribution gap
Whether a reason is required to claimyes_noWhether "at the purchaser's request" landed
Days you believe you have to claimopen_endedWindow accuracy against the real term
Confidence in each beliefscaleConfident-and-wrong, the escalation predictor
Which guarantee terms most affect purchaserankingWhat deserves prominence
Claimants only: what actually happenedopen_ended with AI probingThe 239.5 performance record

Two design rules matter more than the rest.

Test the ad, not the policy page. 239.2 addresses advertisements that mention a warranty, and its television footnote is a useful calibration of what the Commission considers adequate: a disclosure made simultaneously with or immediately following the warranty claim in the audio, or, if in the video portion, remaining on screen for at least five seconds. That is the Commission telling you, in a rare quantitative aside, roughly how much exposure it takes for a disclosure to register. If your limitation appears for less time or with less prominence than the claim it qualifies, you are testing the wrong artefact by testing the policy page.

Interview the abandoners first. They are the smallest group and carry the most information per conversation. Because Koji runs interviews without a moderator and without scheduling, a population of eleven people scattered across time zones is a study you can actually field rather than one you give up on.

Why AI interviews suit guarantee research

Guarantee research has a shape that breaks traditional tooling. The prospect side needs scale, because a comprehension rate on five separate beliefs is a quantitative estimate. The claimant side needs depth, because "the refund took a while" is not a finding and "I sent it back, heard nothing for eleven days, emailed twice, and only got a response after I posted about it" is. Historically those were two different studies with two different budgets, and most teams ran neither.

Koji collapses them. The same study can carry structured items that aggregate into rates and open questions where the AI interviewer probes automatically until the answer is specific. All six structured question types described in structured questions in AI interviews (open_ended, scale, single_choice, multiple_choice, ranking, and yes_no) run inside one conversation, so you get the comprehension rate and the explanation from the same respondent rather than stitching a survey to an interview programme afterwards.

The candor point is sharper here than almost anywhere. You are asking people to describe a bad experience with your company, sometimes one they escalated angrily. A moderated call with someone who works for you invites politeness, and politeness is precisely the failure mode that makes claimant research useless. A substantial literature on self-disclosure finds people report more freely, and more critically, when they believe an automated system rather than a person is receiving the answer. For a study whose value lies entirely in the complaints, that is not a nicety.

Interviews run by voice or text at the participant's convenience, and results land as conversations complete. A guarantee rewrite can be tested and read inside a week rather than a quarter, which is the difference between research that informs the decision and research that arrives to validate it.

What good looks like

  • Full-refund composition accuracy above 75 percent. Below that, your refund amount will surprise people who read your ad correctly.
  • Reason-required belief matching reality. If your guarantee is genuinely no-questions-asked, buyers who think they must justify the return are under-claiming, which sounds like savings and behaves like churn.
  • Lifetime attribution: modal answer matches your actual term. If it does not, 239.4 is telling you to fix the disclosure.
  • Claim window accuracy within a few days of the real term, in both directions. People who think the window is shorter than it is abandon early.
  • Zero claimants describing an undisclosed condition. A condition first encountered during the claim is the specific failure 239.3(b) is written to prevent.

Re-run after any change to the wording, the placement, or the refund workflow. Under 239.5 the workflow is part of the claim, so an operations change is a claim change even when no copy moved.

The honest limit

Consumer research cannot tell you whether your guarantee advertising complies with Part 239 or Section 5 of the FTC Act. That is a legal determination made by counsel, and the Guides themselves note they do not anticipate every unfair or deceptive practice and do not limit the Commission's authority to act under Section 5.

What research does is answer the questions the Guides pose in explicitly empirical language. "Noticed and understood by prospective purchasers" is a measurement. "Material limitations" is a judgment that gets easier when you know which limitations buyers failed to anticipate. And "promptly and fully performs" is a claim about your operation that only the people who tested it can confirm. Keep the perception file and the compliance file separate, and neither will be asked to do the other's job.

Frequently asked questions

What does the FTC say a money-back guarantee has to include?

Under 16 CFR 239.3(a), a seller or manufacturer should use the terms "Satisfaction Guarantee," "Money Back Guarantee," "Free Trial Offer," or similar representations only if it refunds the full purchase price of the advertised product at the purchaser's request. Two phrases carry the meaning: "full purchase price" sets the amount, and "at the purchaser's request" sets a low trigger that does not depend on the company accepting a reason. Section 239.3(b) then requires disclosing any material limitations or conditions with such clarity and prominence as will be noticed and understood by prospective purchasers.

Can I put conditions on a satisfaction guarantee?

Yes. The Guides do not prohibit conditions; they require that material limitations be disclosed clearly and prominently, and the FTC's own examples show conditions integrated into the claim rather than buried, such as returning the unused portion within 30 days or returning the product in its original package. The research question is different from the legal one: a conditional guarantee that is fully disclosed is still a different product in the buyer's mind from an unconditional one, so you should test what buyers believe after seeing the claim and the disclosure together.

What does "lifetime" mean in a lifetime guarantee?

Legally it means whatever you disclose it to mean. 16 CFR 239.4 requires that an advertisement using "lifetime" or "life" disclose the life to which the representation refers, clearly and prominently. The Commission's own examples show a muffler guarantee running for the life of the car, transferable if the owner sells or gives it away, and a battery guarantee running only while the original purchaser owns the car. Neither is the buyer's own lifetime, which is what most buyers assume, and that gap is exactly what a single_choice attribution question measures.

Why should I interview people who claimed the guarantee, not just prospects?

Because 16 CFR 239.5 makes performance part of the advertising claim: a seller should advertise that a product is guaranteed only if it promptly and fully performs its obligations. That means your refund operation is inside the claim, and only the people who invoked it can tell you whether "promptly and fully" describes what happened. The highest-value group is the one that started a refund and abandoned it, because they appear in no operational metric you have. Refund rate, approval rate, and resolution time are all computed over people who finished.

How large a sample do I need for guarantee research?

It depends which population. For prospect comprehension, where you are producing rates on several distinct beliefs, plan for 150 to 300 people in your real target market. For claimant and abandoner interviews the useful number is much smaller, often 10 to 25, because you are after mechanism rather than prevalence. Report that side as coverage rather than proportion: "nine of fourteen claimants described the same undisclosed condition" is a finding, whereas "64 percent" implies a precision the design does not support.

Does a satisfaction guarantee fall under the Magnuson-Moss Warranty Act?

Not necessarily. 15 U.S.C. 2303(b) provides that the designation requirements do not apply to statements similar to expressions of general policy concerning customer satisfaction that are not subject to any specific limitations. Adding specific limitations can move you out of that shelter while adding disclosure duties under 16 CFR 239.3(b). This produces the specificity paradox: a guarantee is legally lightest when vaguest and commercially strongest when most specific. Testing candidate wordings lets you choose a point on that gradient deliberately rather than by accident, but the legal classification itself is a question for counsel.


Ready to test your guarantee? Sign up for Koji and get 10 free credits to run your first study. Show buyers the real ad, score what they believe it promises, and interview the claimants who found out the hard way.

Related Resources

Related Articles

AI Claims Substantiation: How to Prove an AI-Powered Claim Before You Advertise It

Every "AI-powered" claim carries two burdens of proof: an engineering burden (does the system do it?) and a perception burden (what do customers hear?). Learn how to build both evidence files with consumer research before you ship the copy.

Auto-Renewal and Cancellation Research: Testing Whether Subscribers Actually Understood

ROSCA requires clear disclosure, express informed consent, and a simple cancellation mechanism - and defines none of those terms. Learn how to test subscription sign-up and cancel flows for comprehension, and why the vacated Click-to-Cancel rule is still the best research brief available.

Health and Nutrition Claims: What Consumer Research Can and Cannot Substantiate

The FTC says consumer surveys are never sufficient to substantiate a health benefit claim - and in the same document says consumer surveys are valuable for determining what claim you made. Both are true. Learn where the line sits and how to build the perception evidence file.

How to Design Post-Purchase Surveys That Increase Repeat Buying

Learn how to design post-purchase surveys that measure satisfaction, improve the buying experience, identify cross-sell opportunities, and turn one-time buyers into loyal repeat customers using AI-powered conversational follow-up.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.