Lemon Law Research: Measuring When a Buyer Decides a Product Is Beyond Repair
Federal law promises a refund or replacement after a reasonable number of repair attempts, then never defines the number. Here is how to measure the attempt at which your customers actually give up, and why it arrives before the statute does.
Answer first: the Magnuson-Moss Warranty Act entitles a buyer to elect a refund or a replacement once a product still fails after "a reasonable number of attempts" to fix it, and Congress handed the job of defining that number to the Federal Trade Commission, which has never issued the rule. State lemon laws filled the vacuum with numeric presumptions, typically four repair attempts, two for a safety defect, or thirty cumulative days out of service. Every one of those numbers describes when your legal exposure opens. None of them describes when your customer stops believing the product can be fixed, which is almost always earlier. That second number is measurable, it belongs to you, and the only way to get it is to interview owners between repair attempts. Tools like Koji run that as a short AI-moderated interview triggered after each service event, so you see the confidence curve fall attempt by attempt instead of discovering it in a demand letter.
The number Congress never defined
15 U.S.C. 2304(a) sets the federal minimum standards a warrantor must meet to call a warranty "full." Subsection (a)(4) is the buyback provision:
if the product (or a component part thereof) contains a defect or malfunction after a reasonable number of attempts by the warrantor to remedy defects or malfunctions in such product, such warrantor must permit the consumer to elect either a refund for, or replacement without charge of, such product or part.
The same sentence continues with an unusual admission of incompleteness: "The Commission may by rule specify for purposes of this paragraph, what constitutes a reasonable number of attempts to remedy particular kinds of defects or malfunctions under different circumstances." That rule was never written. Fifty years after the Act passed in January 1975, the operative threshold in the federal statute remains an undefined adjective.
This is the third federal consumer-protection standard in the post-purchase family that turns on a fact nobody is required to measure. The disclosure mandate in 2302(a) requires "simple and readily understood language" and defines no test, which is the subject of warranty comprehension research. The guarantee rules in 16 CFR Part 239 require disclosures noticed and understood by prospective purchasers, covered in money-back and satisfaction guarantee research. Here the undefined term is not a communication standard but a patience standard, and it is the one with the largest financial consequence attached.
Two adjacent provisions matter when you design research around it. 2304(b)(1) says a warrantor may not impose any duty other than notification on a consumer as a condition of securing a remedy, unless the warrantor can demonstrate that the duty is reasonable. And 2304(c) carves out failures caused by consumer damage or "unreasonable use (including failure to provide reasonable and necessary maintenance)." Both of those turn into research questions the moment a dispute starts: did the owner know notification was all that was required, and does the owner believe they used the product normally?
What the states put in the vacuum
Because the federal number stayed blank, state legislatures wrote their own, mostly for motor vehicles. California is the clearest example. Civil Code 1793.22, the Tanner Consumer Protection Act, creates a rebuttable presumption that a reasonable number of attempts has been made if, within 18 months of delivery or 18,000 miles, whichever comes first, any of the following happens:
| Tanner Act trigger (Cal. Civ. Code 1793.22(b)) | Threshold | What it implies about the buyer |
|---|---|---|
| Nonconformity likely to cause death or serious bodily injury | 2 or more repair attempts | Safety collapses trust roughly twice as fast |
| Same nonconformity, any severity | 4 or more repair attempts | Four is the legislative guess at exhausted patience |
| Vehicle out of service for repair | More than 30 cumulative days | Downtime counts even without a failed fix |
Three details in that statute are worth borrowing even if you sell something other than cars. First, the presumption attaches to the same nonconformity, so a product that fails four different ways does not trigger it, although a customer experiencing four failures certainly feels like it should. Second, the buyer must have directly notified the manufacturer at least once, but only if the manufacturer "clearly and conspicuously disclosed" that requirement in the warranty or owner manual, which makes the enforceability of your own condition depend on a comprehension fact. Third, the presumption is rebuttable, not conclusive.
The federal Act contains a related gap on the remedy side. Under 16 CFR 701.1(e), where a warrantor offers repair, replacement, or refund, the choice among them belongs to the warrantor, not the buyer, until the statutory buyback provision flips it. 2304(a)(4) is the one moment in the scheme where the election passes to the consumer. Buyers rarely know either half of that rule, which is why the transition feels arbitrary to them and mechanical to you.
The give-up point arrives before the statutory threshold
Here is the finding that changes how a product team should think about repeat repairs. The statutory threshold and the trust threshold are two different numbers, and they are not close together.
The statute counts attempts. Your customer counts disappointments. By the second failed fix, most owners have already stopped forecasting a working product and started forecasting a fight. They keep bringing the product in not because their confidence recovered, but because the alternative is losing the money entirely. Attempts three and four are frequently performed on a customer who has already decided, privately, that the product is a lemon.
That gap has an operational consequence most warranty programs never see. A repair attempt that succeeds technically can still fail commercially. The part is replaced, the fault code clears, the ticket closes at 100 percent first-time-fix, and the owner never buys the brand again. Your service metrics record a win. Your renewal rate records the loss eighteen months later, with no field to join the two records on.
Call the measurement the patience curve: the share of owners who still believe the product will be permanently fixed, plotted against the number of repair attempts they have experienced. It is a single scale question asked immediately after each service event, and it produces a decay curve with a knee in it. The knee is your real threshold. Everything to the right of it is repair work performed on a relationship that has already ended.
Two companies with identical four-attempt legal exposure can have knees at attempt two and attempt three respectively, and the one with the earlier knee is losing customers it will never see in a lemon-law claim, because most disappointed owners do not file. They sell the product, write a review, and leave.
Substantial impairment is a subjective standard by design
The parallel remedy under sales law makes the research case even more directly. UCC 2-608(1) lets a buyer revoke acceptance of goods "whose non-conformity substantially impairs its value to him," where the acceptance happened either on the reasonable assumption that the nonconformity would be cured and it was not, or without discovery of the defect because discovery was difficult or the seller gave assurances.
Read those last two words again. The drafters did not write "substantially impairs its value," which would be an objective market test. They wrote value to him. The standard is deliberately indexed to the individual buyer, which means the legal question and the research question are the same question. What counts as substantial impairment for a delivery driver whose van is the business is not what counts for a weekend owner, and the statute is explicitly fine with that.
That gives you a defensible reason to measure impairment per segment rather than assume it. A defect that a lab classifies as cosmetic can be substantially impairing to a buyer who purchased primarily on appearance, and a defect a lab classifies as functional can be tolerable to a buyer who never uses the affected mode. The severity ranking in your engineering triage system was written by engineers. The severity ranking that governs revocation was written, unknowingly, by your customers.
Two more sales-law provisions shape the timeline. Under 2-608(2), revocation must occur within a reasonable time after the buyer discovers or should have discovered the ground for it, and is not effective until the buyer notifies the seller. And under 2-607(3)(a), a buyer who has accepted goods must notify the seller of a breach within a reasonable time after discovering it "or be barred from any remedy." Silence forfeits the claim. This is the mechanism explored in more depth in implied warranty research, and it means the population of owners who formally complain is a floor, not a count.
What to ask, and when
The design that works is short, repeated, and event-triggered rather than periodic. You are building a per-attempt panel, not a satisfaction survey. Keep it to five or six questions so it survives being sent after every service event, and use the structured types so the trend is scoreable rather than anecdotal. See structured questions in AI interviews for the full set of six.
| What you need to know | Question type | Why this type |
|---|---|---|
| Confidence the product is now permanently fixed | scale | Produces the patience curve across attempts |
| Whether this failure is the same one as last time | single_choice | The Tanner presumption attaches to the same nonconformity |
| What the owner could not do while the product was down | open_ended | Impairment in the buyer own terms, per 2-608 |
| Days of use lost since purchase | open_ended | Owner-side counterpart to the 30-day trigger |
| Which outcome the owner now wants | single_choice | Repair, replacement, refund, or already decided to leave |
| Rank what would restore trust | ranking | Forces a trade-off between speed, compensation, and explanation |
The scale question is the spine. Ask it identically every time so the values are comparable across attempts, and resist the temptation to reword it for tone after a bad quarter. The single_choice on desired outcome is the early-warning field: the first time an owner selects refund over repair, the relationship has changed even if the ticket has not.
One caution on the confidence scale. If most of your owners answer at the top of it after attempt one, you have a ceiling and you will not be able to see improvement or decline among your most loyal customers. Check the top-box share before you trust any trend from it, and pair the scale with an open_ended follow-up so the reasoning survives even when the number saturates.
Running it as an AI-moderated interview
The reason this study is rare is not that teams do not want the data. It is that the traditional way to get it does not fit the trigger. A moderated interview cannot be scheduled within hours of a service visit at the volume repair events occur, and by the time a researcher books a session two weeks later the memory has smoothed over. A static survey fits the trigger but collects the useless version of the answer, because "it still is not right" is where a form stops and where an interview should start.
An AI-moderated interview solves the shape of the problem specifically. It can be triggered the moment a work order closes, it runs at any hour in the owner own time zone, it takes voice or text, and when an owner says the fix did not hold, it asks what exactly came back and how soon, without a moderator being awake. Koji generates those follow-ups automatically from the response, so the depth does not depend on which owners happened to get a human.
A workable setup looks like this:
- Build one study and reuse it for every attempt, adding a screening question that captures the attempt number, so all responses land in a single comparable dataset.
- Trigger the invitation from your service system when a repair order closes, not on a weekly batch.
- Keep it under six minutes. Owners who have just had a second failure will not give you twenty.
- Let owners answer by voice. Frustration carries information that a text box flattens, and voice responses run longer.
- Read the report by attempt number rather than in aggregate. The aggregate hides the knee, which is the entire point of the study.
Because analysis is automatic, the practical cost of asking after every attempt instead of every quarter is close to zero, which is what makes the per-attempt panel feasible at all. That is the difference between this and the survey program it replaces: the same budget buys a curve instead of a single number.
Reading the results without fooling yourself
Three failure modes account for most bad conclusions in this study.
Surviving-owner bias. Owners who reach attempt four are, by construction, the ones who did not abandon the product at attempt two. Confidence measured at attempt four is measured on a filtered population and will look better than the truth. Always report the denominator: how many owners entered at attempt one, and how many are still answering.
Blaming the customer too early. 2304(c) does let a warrantor decline where the failure came from consumer damage or unreasonable use, and it is genuinely tempting to route repeat failures into that bucket. Ask the usage questions before you know the diagnosis, not after, or you will find the maintenance answer you went looking for.
Counting tickets instead of episodes. A customer who calls three times about one fault is one episode with three contacts, and a customer who never calls back after the second failure is one episode with zero further contacts and the worst outcome in your dataset. Reconstructing episodes from ticket logs is the subject of support ticket analysis, and it is worth doing before you conclude that repeat failures are rare.
The output you want at the end is a single sentence with two numbers in it: the attempt at which the presumption attaches in your jurisdictions, and the attempt at which your owners stop believing. When the second number is lower than the first, and it usually is, you have quantified the cost of treating a legal threshold as a service target.
Frequently asked questions
Does the federal lemon-law provision apply to products other than cars?
Yes. 15 U.S.C. 2304 applies to consumer products covered by a full written warranty, not only vehicles. What is vehicle-specific is the state layer: most state lemon statutes with numeric presumptions, including the California Tanner Act, are written for motor vehicles. If you sell appliances, electronics, or equipment, the federal reasonable-number-of-attempts standard applies to your full warranties with no number attached to it at all, which is a stronger reason to establish your own defensible threshold from data.
How many repair attempts is a reasonable number?
Federal law does not say. 15 U.S.C. 2304(a)(4) expressly authorized the FTC to define it by rule and the rule was never issued, so there is no federal number. State presumptions for vehicles commonly sit at four attempts for an ordinary defect, two where the defect is likely to cause death or serious bodily injury, or more than thirty cumulative days out of service. Treat those as the point where legal exposure begins, and measure your own customers separately to find the point where confidence collapses, which typically arrives sooner.
What is the difference between a lemon-law claim and revocation of acceptance?
They are separate routes to the same practical outcome. A lemon-law claim runs through a warranty statute and usually depends on a repair-attempt count. Revocation of acceptance runs through UCC 2-608 and depends on whether the nonconformity substantially impairs the value of the goods to that particular buyer, plus timely notice and no substantial change in the condition of the goods. The second standard is subjective by design, which is why it is measurable with research and the first is mostly measurable with service records.
Can we survey owners about this instead of interviewing them?
You can, and you will get the number without the reason. The valuable content in this study is what the owner says immediately after "I do not think it will ever be right," and a form has no way to ask the follow-up. An AI-moderated interview asks it automatically on every response, which is why platforms like Koji produce usable per-attempt evidence at survey-scale volume. Use structured questions for the scoreable spine and let the conversation carry the explanation.
When should the interview be triggered?
When the repair order closes, not on a calendar. The specific memory you need, what failed, how long it was down, what the service interaction was like, decays within days and is contaminated by the next event. Event-triggered invitations also give you a clean attempt number to attach to each response, which is what makes the patience curve possible.
Does asking owners about giving up create the very risk we are trying to avoid?
The evidence points the other way. Owners who have experienced repeat failures have already formed the judgment; the question does not plant it. What asking does is surface the decision while the outcome is still reversible, and give you the chance to offer a replacement before the owner reaches a lawyer or a review site. The silent second-failure owner is the expensive one precisely because nobody asked.
Ready to measure your own threshold? Sign up for Koji and get 10 free credits to run a per-attempt study on your last hundred repair events. Trigger it when the work order closes, let the AI ask the follow-ups, and read the confidence curve by attempt number in hours rather than quarters.
Related Resources
- Structured Questions in AI Interviews - the six question types that make a repeated per-attempt study scoreable
- Warranty Comprehension Research - the disclosure standard that governs what owners were told before the failure
- Implied Warranty Research - the notice rule that quietly bars most claims before they are made
- Extended Warranty and Service Contract Research - what happens when the coverage was bought separately
- Right to Repair Research - measuring whether an owner could have fixed it themselves
- Product Recall Notice Research - the funnel model this study borrows its shape from
- Support Ticket Analysis - rebuilding episodes from contacts before you count failures
Related Articles
Money-Back and Satisfaction Guarantees: Testing What Your Promise Actually Promises
The FTC Guides say you should use Satisfaction Guarantee or Money Back Guarantee only if you refund the full purchase price on request. Here is how to test what buyers believe your guarantee covers, and whether claimants got it.
Product Recall Notice Research: Testing Whether Safety Notices Reach, Convince, and Move Owners
Recall correction rates sit in the single digits while awareness can be above 80 percent. Those are different variables. Here is how to find which stage of the recall funnel is leaking, using the CPSC evidence base.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Support Ticket Analysis: How to Mine Customer Service Data for Product Insights
A practical guide to systematically extracting product insights from customer support tickets — covering manual coding workflows, AI-powered thematic analysis, and how to tie ticket themes to business impact.
Warranty Comprehension Research: Testing Whether Buyers Understand What Your Warranty Covers
The Magnuson-Moss Warranty Act requires warranty terms in simple and readily understood language, but never defines the test. Here is how to measure what buyers actually understand about coverage, remedy, and process.
5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)
A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.