Back to docs
Research Methods

Product Recall Notice Research: Testing Whether Safety Notices Reach, Convince, and Move Owners

Recall correction rates sit in the single digits while awareness can be above 80 percent. Those are different variables. Here is how to find which stage of the recall funnel is leaking, using the CPSC evidence base.

Answer first: a recall correction rate measures one thing, the proportion of recalled units refunded, replaced, or repaired. It does not tell you whether owners heard the notice, believed it applied to them, checked their product, or decided the remedy was worth the trouble. Those are four separate failures with four different fixes, and a single percentage cannot distinguish them. The only way to locate the leak is to interview owners at each stage. Platforms like Koji make that practical, because recall populations are large, dispersed, and time-critical in a way that moderated research has never been able to serve.

The number everyone quotes, and what it hides

The U.S. Consumer Product Safety Commission measures recalls using the correction rate: the proportion of product units recalled that have been refunded, replaced, or repaired. In November 2020, the Government Accountability Office reported (GAO-21-56) that CPSC relies on this single performance metric, and warned that "using a single measure may not allow CPSC to accurately gauge the effectiveness of all its recalls," noting that for cheap products consumers may simply throw the item away rather than seek a refund or replacement. The same report found that only 61 percent of firms had submitted required progress reports more than 75 percent of the time for recalls closed between February 2016 and May 2020.

Reported participation rates are correspondingly grim. Figures presented at CPSC's 2017 recall effectiveness workshop put average consumer participation at roughly 6 percent across product types, around 4 percent for products retailing under $20, and roughly 32 percent for products priced at $10,000 or above.

Read alongside the underlying research, those numbers say something more specific than "recalls do not work."

The recall funnel, from CPSC's own evidence

CPSC's Recall Effectiveness Research: A Review and Summary of the Literature on Consumer Motivation and Behavior, prepared for the Commission and dated July 2003, reconstructs a national survey conducted for the agency's 1980 Recall Effectiveness Task Force. The subject was a series of recalls of hair dryers lined with asbestos, a low-priced product with a short useful life and therefore an expected low measured effectiveness rate. The survey produced a complete funnel:

StageResultWhat a failure here means
Aware of the general hazard85 percent of hair dryer ownersNotification reached people
Checked their own dryer44 percent of those awareAwareness did not convert to self-relevance
Found an affected unitabout one in five of those who checkedIdentification worked
Stopped using it85 percent of affected ownersPersuasion worked
Used the recall remedy5 of 27 affected respondentsThe remedy itself was not worth the cost

Of the remaining affected owners, nine discarded the unit and thirteen stopped using it without throwing it away.

That funnel is the whole argument. Awareness was 85 percent. Remedy uptake among confirmed affected owners was under a fifth. A correction rate computed on this recall would have looked like a communications catastrophe, and communications was the one stage that worked. The failures were at self-relevance, where more than half of aware owners never checked, and at the remedy, where a mail-in repair was not worth the effort for a cheap appliance.

Awareness and correction are different variables, and the industry reports only the one it can count. Any recall programme that responds to a low correction rate by buying more notification is treating the stage that is least likely to be broken.

There is also a measurement point hiding in the last row. Twenty-two of twenty-seven affected owners eliminated their exposure to the hazard, by binning the dryer or by shelving it. Five appear in the correction rate. The 1980 Task Force saw this clearly and recommended that recall effectiveness measurement acknowledge "the existence of uncountable, but appropriate consumer responses to recall messages," and that successful communication of a hazard be treated as an indicator in its own right where response is likely to be understated. GAO made essentially the same criticism forty years later. A consumer who throws away a hazardous $12 product has behaved perfectly and counts as a failure.

What actually predicts participation

The literature is unusually consistent about which variables move recall effectiveness, and message quality is not at the top of the list.

A 1978 CPSC quantitative study of 97 recalls accepted for closeout by July 1976 identified seven variables strongly related to effectiveness: product sale price, average useful life, number of affected units, time in distribution, the percentage of units in consumers' hands, the type of recall action, and the level of direct consumer notification. It also described the conditions under which a recall was likely to be "very ineffective": products priced under two dollars, products with average useful lives under two years, cases where units in distribution exceeded 100,000, and recalls of products that had been in distribution for over five years. Notably, the study did not find a strong relationship between the nature or severity of the hazard and effectiveness, though its authors cautioned that severity classification was highly subjective.

Economists Dennis Murphy and Paul Rubin, working in 1988 with roughly 100 CPSC recalls from the early 1980s, built a predictive model with strong explanatory power from a handful of variables including the proportion of units in consumers' hands, the proportion of consumers directly notified, the months between the end of distribution and the start of the recall, and whether the remedy involved a home repair.

That last variable is the actionable one. According to the CPSC review, providing an in-home repair rather than a remedy requiring return of the product to the manufacturer or retailer was associated with an increase of 14 percentage points in the average recall effectiveness rate.

The cost-of-compliance cliff

The warnings literature the CPSC review surveys makes an even sharper point, and it generalises well beyond safety.

Dingus, Hathaway and Hunn (1991) unobtrusively observed 920 racquetball players. Signs on the court door and the front wall, in ANSI format with a pictograph and a list of eye-injury statistics, instructed players to wear eye protection. The only manipulation was where the goggles were. With goggles in a box just outside the door, 60 percent complied. With goggles at a checkout booth 60 feet away, or not provided at all, compliance was zero percent.

The same team ran a second study in which participants took home a "new formulation" cleaning product to use for a week. Where gloves were provided with the product, roughly 87 percent wore them. Where they were not provided, compliance was 25 percent. The authors concluded that "cost must be very low to achieve the highest possible compliance with a warning's intent. Increasing the cost even a seemingly minor amount can have devastating effects on compliance."

Sixty feet took compliance from 60 percent to zero. The message did not change.

Persuasion has a much smaller dynamic range than friction. A recall study that tests only the notice is measuring the variable with the least leverage. The variables with leverage are the ones in the remedy: whether a box and a prepaid label arrive, whether a technician comes to the house, whether the replacement ships before the return, whether the owner loses use of something they need.

That last factor has its own evidence. Warner (1980) studied the 1976 Corning recall of electric percolators and found that among consumers who did not return the product as instructed, nearly half cited the lack of an alternative unit as a reason, with a further 4 percent citing the inconvenience of returning it. People do not surrender a working coffee pot, or a car seat, or a space heater in January, on the strength of a well-written notice. They surrender it when they have something to use instead.

The 1980 warning that research only recently learned to answer

The single most useful methodological point in the entire CPSC review is a caveat rather than a finding.

Heisler and Bernstein (1980) surveyed consumers about automotive recalls. Around 11 percent of respondents reported difficulty obtaining repairs, typically parts unavailability or excessive completion times, and about 30 percent offered suggestions for making the process easier, including faster service, more cooperative dealers, emphasising that the repair is free, and compensating owners for inconvenience.

When non-participants were asked their major reasons for not responding, 23 percent gave answers reflecting no perceived hazard or no motivation to comply: "didn't have time," "didn't believe anything was wrong," inconvenience, "didn't think it was important," "laziness," and "don't drive the vehicle much." The authors labelled this owner apathy and concluded that negative owner attitudes were a significant cause of non-compliance.

Then they added the caveat that matters:

Respondents may have felt pressure to blame manufacturers and dealers for their behavior rather than admitting their own apathy.

That is a 1980 statement of social desirability bias contaminating the exact measurement a recall programme most needs. Ask people why they ignored a safety warning and a portion will produce a socially acceptable answer, blaming the dealer, the parts, or the letter, rather than admitting they could not be bothered. If your remediation plan is built on those answers, you will fix the dealer network and the letter, and the apathy will be waiting for you on the next recall.

There is now a substantial body of research on self-disclosure showing that people report more openly, with measurably lower impression management, when they believe an automated system rather than a person is receiving their answers. The recall literature identified the problem in 1980 and had no practical way to reduce it at scale. That is precisely the gap an AI-moderated interview closes: nobody is sitting there to be disappointed in you, so "I saw the email and ignored it" becomes a sayable sentence.

Designing the study

Recall research has a hard constraint that shapes everything: it is time-critical. A study that reports in six weeks is a post-mortem. The design below is built to be fielded in days.

StageQuestion typeWhat it produces
Unaided awareness of the recallyes_noWhether notification reached them at all
Where they first heard, if awaresingle_choiceWhich channel is carrying the notice
Whether they believed it applied to their unityes_noThe self-relevance leak
Whether they checked, and howopen_ended with AI probingWhether identification instructions are executable
Perceived seriousness of the hazardscaleWhether the risk framing landed
What they did, and why they stopped thereopen_ended with AI probingThe real reason for the drop-off
Barriers to using the remedymultiple_choiceLoss of use, effort, cost, packaging, distrust
Which remedy they would actually userankingWhether to fund in-home repair or advance replacement

Three rules make it work.

Measure every stage, not just the outcome. The funnel is the deliverable. Reporting "we improved awareness" when the leak is at self-relevance wastes the next budget cycle.

Test identification, not just comprehension. Most notices ask owners to find a model or date code on the product. That is a task, and tasks fail. Have participants actually attempt it and describe what they see. This is where notices with technically perfect language quietly fail, because the code is under the unit, printed in grey on grey, or absent from the packaging the owner still has.

Interview the non-compliers, and design for candour. They are the population the correction rate is made of, and they are the population most likely to give you a face-saving answer. Ask what they did rather than why they failed, keep the framing non-judgmental, and expect the honest reasons to be mundane.

Why AI interviews fit recall research specifically

Recall research fails on traditional tooling for three reasons, and they are all structural rather than budgetary.

Speed. Recruiting, scheduling, and moderating a study takes longer than the window in which findings can change a live recall. Koji studies field the day they are written, and results accumulate as conversations complete.

Scale plus depth together. You need a reliable funnel, which is a quantitative estimate across several stages, and you need the reason behind each drop-off, which is qualitative. Historically that meant a survey that could not ask why and an interview programme that could not reach enough people. Koji runs both in one conversation: structured items aggregate into stage-by-stage rates while the AI interviewer probes automatically whenever an answer is too vague to act on. All six question types described in structured questions in AI interviews (open_ended, scale, single_choice, multiple_choice, ranking, and yes_no) live in the same interview.

Population reach. Recall populations are not a customer list. They include second-hand owners, gift recipients, and people who never registered the product, which is exactly the group 16 CFR 701.4 anticipates in the warranty context and which we cover in warranty comprehension research. Voice and text interviews that run at the participant's convenience reach people a scheduled call never will.

A note on interpretation. Self-reported awareness is a fragile measure, prone to the memory problems described in recall bias, and people will sometimes affirm awareness of a notice they never saw. Anchor the awareness question to a specific artefact, show the actual notice, and ask what they remember about it rather than whether they remember it.

What good looks like

  • A funnel with a named leak, not a single participation percentage.
  • Identification task success above 90 percent. If owners cannot find the model code, nothing downstream can work.
  • Self-relevance conversion above 70 percent among the aware. The 1980 hair dryer figure of 44 percent is the benchmark to beat, and it is beatable, because it reflected a notice that described a hazard class rather than a product an owner could recognise.
  • A remedy ranked first by owners that you are actually funding. If owners rank in-home repair first and you are running mail-in, the Murphy-Rubin 14-point differential is the size of the prize.
  • Barrier reasons collected without an evaluator present, so that apathy shows up as apathy rather than as a complaint about the dealer.

The honest limit

This research does not establish whether your recall is legally adequate, and it does not substitute for the CPSC processes that govern reporting, corrective action plans, and notice content. Those are regulatory obligations determined with counsel and the Commission.

What it does is answer the question the correction rate structurally cannot: which stage of the recall failed. That question has a different answer for every recall, the answer determines where the money should go, and the only people who hold it are the owners who did not respond.

Frequently asked questions

What is a typical product recall response rate?

Low, and highly dependent on product price. Figures presented at CPSC's 2017 recall effectiveness workshop put average consumer participation at roughly 6 percent across product types, around 4 percent for products retailing under $20, and roughly 32 percent for products priced at $10,000 or above. But the headline number is misleading on its own, because CPSC measures the correction rate, meaning units refunded, replaced, or repaired. In the 1980 Task Force hair dryer survey, 85 percent of owners were aware of the hazard and 85 percent of affected owners stopped using the product, while only 5 of 27 used the official remedy. High awareness and a low correction rate are entirely compatible.

Why does the correction rate understate recall effectiveness?

Because it counts only one of several appropriate consumer responses. A consumer who throws away a hazardous inexpensive product has eliminated the hazard perfectly and appears in the statistics as a failure. GAO-21-56 made this point in November 2020, noting that for cheap products consumers may simply throw the item away rather than seek a refund or replacement, and recommending CPSC explore measures beyond the correction rate. CPSC's own 1980 Recall Effectiveness Task Force had already recommended acknowledging "uncountable, but appropriate consumer responses" and treating successful hazard communication as an indicator in its own right.

Does improving the wording of a recall notice increase participation?

Less than most teams expect, because friction has far more leverage than persuasion. Dingus, Hathaway and Hunn (1991) observed 920 racquetball players under identical warning signage and varied only where the eye protection sat: 60 percent complied when goggles were in a box outside the door, and zero percent complied when they were 60 feet away or not supplied. In the recall context, the CPSC literature review reports that offering an in-home repair rather than requiring return of the product was associated with a 14 percentage point increase in average effectiveness. Message testing is worth doing, but the remedy design is where the large gains are.

How do I reach owners who never registered the product?

You largely cannot reach them through your own records, which is why recall research should not be run only on your customer list. Second-hand owners, gift recipients, and non-registrants are a meaningful share of the affected population and behave differently from registered buyers. Recruit through panels and category screeners rather than your CRM, and expect awareness and self-relevance to be markedly lower in that group. This is the same population problem that 16 CFR 701.4 anticipates for warranty registration cards.

Why interview non-compliers rather than survey them?

Because the answers you need are the ones people are least willing to volunteer, and a survey cannot follow up. Heisler and Bernstein (1980) found 23 percent of automotive recall non-participants gave reasons reflecting no perceived hazard or motivation, including "didn't have time," "didn't believe anything was wrong," and "laziness," and the authors explicitly warned the figure was probably understated because "respondents may have felt pressure to blame manufacturers and dealers for their behavior rather than admitting their own apathy." Research on self-disclosure consistently finds people report more openly when they believe an automated system rather than a person is receiving the answer, which is why an AI-moderated interview is a better instrument for this specific measurement.

How quickly can a recall notice study be fielded?

That is the binding constraint, and it is why traditional methods rarely get used for live recalls. Recruiting, scheduling, moderating, transcribing, and coding a conventional study takes longer than the window in which findings can change the recall. An AI-moderated study can be written and fielded the same day, with participants responding by voice or text whenever suits them and results accumulating as conversations complete, so a funnel read is available in days. For a live safety recall, that difference determines whether the research informs the corrective action or documents it afterwards.


Ready to test your recall notice? Sign up for Koji and get 10 free credits to run your first study. Show owners the real notice, have them attempt the identification task, and find out which stage of the funnel is losing people, in days rather than weeks.

Related Resources

Related Articles

Accessibility Compliance Research: What WCAG, the ADA, and the European Accessibility Act Require You to Test

WCAG conformance is an audit standard, not proof your product works for disabled users. Here is what the EAA, ADA Title II, and Section 504 actually demand in 2026 — and how to run the user research that closes the gap.

Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)

Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Support Ticket Analysis: How to Mine Customer Service Data for Product Insights

A practical guide to systematically extracting product insights from customer support tickets — covering manual coding workflows, AI-powered thematic analysis, and how to tie ticket themes to business impact.

Trust and Safety Research: How to Study Harm, Reporting Flows, and Moderation With Real Users

Trust and safety teams run on tickets and telemetry, which only show the harm that got reported. Here is how to research the harm that did not — the four research objects, the DSA and Online Safety Act obligations that require it, and a safeguarding protocol that protects participants.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.