Back to docs
Research Operations

FCA Consumer Duty Customer Research: How to Evidence Consumer Understanding and Fair Value

The FCA's 2026 reviews were blunt: sales data and an absence of complaints prove nothing about consumer understanding. Here is how to test communications with real customers, hit a comprehension target, and build a board-report evidence pack that survives scrutiny.

Answer first: The FCA's Consumer Duty does not ask whether your communications were sent. It asks whether customers understood them, and it expects you to have tested that with real customers and documented the result. In its 13 March 2026 review of the consumer understanding outcome, the FCA identified weak evidence of communication testing as a leading area for improvement and stated plainly that firms relying on sales data or the absence of complaints have no reliable assurance of understanding. One month later, in its 16 April 2026 observations on second-year board reports, it told firms to move beyond management-information dashboards to analysis that draws conclusions — and specifically to deepen the evidence on consumer understanding and support. In practice this makes structured comprehension testing a recurring research obligation for every UK regulated firm, not a project.

What the Duty demands, evidentially

The Consumer Duty (PRIN 2A) took effect on 31 July 2023 for new and existing open products and 31 July 2024 for closed products. It sets a consumer principle, three cross-cutting rules — act in good faith, avoid causing foreseeable harm, and enable and support customers to pursue their financial objectives — and four outcomes. Each outcome implies a different research question:

OutcomeThe regulatory questionWhat research has to produce
Products and servicesDoes the product meet the needs of the identified target market?Evidence that real customers in that segment have the need, and that distribution is reaching them
Price and valueIs the price reasonable relative to the benefits?Evidence of what customers believe they are buying and which benefits they actually value or never use
Consumer understandingCan customers make informed decisions?Tested comprehension of the actual communications, with pass criteria and a record of redrafting
Consumer supportCan customers realise the benefits and act without unreasonable friction?Evidence on where support journeys stall, for whom, and what the customer did next

Notice what all four have in common: the evidence has to come from customers, not from process documentation. That is the shift most firms have still not fully made.

The 2026 reviews: what the FCA praised and criticised

13 March 2026 — consumer understanding review. The FCA published findings on how firms approach the consumer understanding outcome, split into good practice and areas for improvement.

Good practice observedAreas for improvement
Analysing insight from multiple sources — call listening, complaints, chat transcripts, website analytics, drop-off data and surveysWeak evidence of communication testing
Testing communications with real customers both before and after changes, using proportionate methods including short surveys, comprehension checks, callbacks, A/B testing and feedback during digital trialsInaccessible or overly complex communications
Documenting what changed, why, and what effect it hadInsufficient consideration of diverse customer needs
Setting explicit comprehension targets — in one case at least 80% correct recall of key points — and redrafting until the target was metWeak monitoring, and unclear accountability for who decides what and how

16 April 2026 — board report observations. Reviewing second-year reports, the FCA asked firms to move past MI dashboards to analysis that draws conclusions, to document board challenge, to monitor outcomes delivered through distribution chains, and to deepen the evidence on consumer understanding and support.

The pattern across both is consistent: the regulator is no longer satisfied by activity. It wants a testable claim, the test, the result, and the decision that followed.

How to actually run a comprehension test

A comprehension test is not a satisfaction survey. You are not asking whether customers liked the letter; you are measuring whether they can correctly answer questions about what it means for them.

1. Choose the communication and the key points. Take the real artefact — the annual statement, the renewal notice, the arrears letter, the fee disclosure — and list the three to six key points a customer must take away. If you cannot list them, the communication has no defined purpose and that is the first finding.

2. Define pass criteria before fielding. The FCA highlighted a firm using at least 80% correct recall of key points. Pick your threshold, write it down, and treat it as a gate rather than a metric to report.

3. Show the artefact, then test. Present the actual text or screen, then ask closed questions with one correct answer per key point, plus open questions on interpretation.

4. Probe the wrong answers. The score tells you the communication failed; only the reasoning tells you which clause caused it. This is where a static survey stops and an interview keeps going.

5. Redraft and retest. The evidence the FCA wants is the pair: the before result, the change, the after result. One-shot testing produces a number, not a demonstration of improvement.

6. Record the decision trail. Who approved the redraft, on what evidence, and when. Unclear accountability was named as an area for improvement.

Mapping the test to structured question types

Koji's six structured question types let one instrument produce the comprehension score and the diagnosis together:

ElementQuestion typePurpose
Key-point recallsingle_choiceOne correct option per key point — this is what generates the pass rate
Action comprehensionyes_no"Based on this letter, do you need to do anything before 30 September?"
ConfidencescaleSelf-rated clarity, which usefully diverges from actual score and exposes false confidence
Perceived relevance of termsmultiple_choiceWhich listed features the customer believes apply to them
Priority of informationrankingWhich points customers think matter most, versus which you intended to foreground
Interpretationopen_endedTheir reading in their own words, with AI follow-up probing on any incorrect answer

The last row is the differentiator. Platforms like Koji ask the follow-up automatically: when a customer selects the wrong answer about a fee, the AI interviewer asks what they thought the fee covered and where in the document they looked. Ten minutes of conversation produces both the 72%-versus-80% pass rate and the sentence that tells the drafting team which clause to rewrite. A survey tool gives you the first and leaves you guessing at the second; a moderated session gives you both but at twenty times the cost per participant and only during business hours.

The mechanics of comprehension and wording tests generalise beyond financial services — see content testing for the underlying method.

Diverse customer needs and vulnerability

"Insufficient consideration of diverse customer needs" was one of the FCA's named improvement areas, and it is the hardest to evidence with conventional research, because the customers least likely to complete a twelve-minute online survey are precisely the ones the Duty is most concerned with.

Practical steps:

  • Sample deliberately for characteristics of vulnerability — health conditions, low resilience, low capability, negative life events — rather than hoping a general panel covers them. See researching hard-to-reach audiences.
  • Offer voice as well as text. Voice interviews remove the reading and typing burden, which matters for customers with low literacy, dyslexia, visual impairment, or motor difficulty. Offering both modes on the same study is itself evidence of accommodating diverse needs.
  • Make participation async. Fixed appointment slots systematically exclude shift workers, carers and people managing health conditions.
  • Report outcomes split by vulnerability characteristic. The Duty asks whether outcomes differ between groups. An aggregate pass rate hides exactly the disparity the regulator is looking for.
  • Follow accessible-research practice throughout. See accessibility research and, for the wider compliance picture, accessibility compliance research.

Evidence for price and value, and for consumer support

Price and value. Fair value assessments are usually built from cost, margin and benchmark data. What they typically lack is the customer's side: which benefits customers know they have, which they use, and which they would not miss. A short annual study using multiple_choice for feature awareness, ranking for value ordering, and open_ended for the reasoning gives your assessment a customer-evidence limb it probably does not have — and directly addresses whether customers understand what they are paying for.

Consumer support. The support outcome fails at friction points, not in aggregate satisfaction. Interview customers immediately after a support interaction, a claim, a cancellation attempt, or a failed self-service journey, and ask what they were trying to achieve, where they stopped, and what they did next. Triggering an interview from the event is what makes this feasible: Koji studies can be launched from webhooks or a CRM record so the interview arrives while the experience is fresh, rather than in the next quarterly wave.

The board-report evidence pack

Given the April 2026 observations, a defensible pack for the consumer understanding and support outcomes contains:

  1. The inventory — which communications and journeys were tested this period, and which were not, with a reason.
  2. The pass criteria and results — before-and-after comprehension rates against a stated threshold, split by segment and by vulnerability characteristic.
  3. The verbatim evidence — customer quotes showing how a communication was misread. This is what turns a dashboard into analysis.
  4. The decision trail — what was redrafted, by whose decision, on what evidence, and what the retest showed.
  5. The residual risk — communications that still fail the threshold, with owner and remediation date.
  6. Distribution-chain evidence — outcomes for customers acquired through intermediaries, not just direct.
  7. Method and data lineage — instrument, sample, dates, and where raw responses are retained. See research data retention and deletion.

Koji supports the lineage requirement directly: every interview retains its full transcript alongside the structured answers, and data export in CSV or JSON gives you an auditable record to attach to the assessment. Composite quality scores from 1 to 5 let you show that low-effort responses were excluded from the pass-rate calculation rather than quietly diluting it.

On data handling: interviews are async and link-based, so no third-party meeting recorder sits in the chain, and a data processing agreement is available to business customers. For firm-specific requirements on hosting, retention or sub-processors — which any FCA-regulated firm should put to every vendor in writing — contact the Koji team, and run a DPIA covering the processing before you field a study on vulnerable customers.

Cadence and anti-patterns

Test before a change and after it, then monitor continuously. Annual-only testing produces evidence that is eleven months stale for most of the year, and the FCA specifically criticised weak monitoring. Set a refresh trigger on any communication that is materially redrafted, any product change, and any spike in complaints or drop-off — see insight decay and when to re-run a study.

Avoid these:

  • Treating sales volume or complaint absence as evidence of understanding. Named explicitly by the FCA as providing no reliable assurance.
  • Testing satisfaction instead of comprehension. "Was this letter clear?" is not a comprehension test; customers routinely rate unclear documents as clear.
  • Testing with staff or a general panel. Comprehension is population-specific. Test with holders of that product.
  • Reporting an aggregate pass rate only. Split by segment and vulnerability characteristic or you have hidden the finding.
  • Stopping at the score. Without the reasoning behind wrong answers, you cannot fix the wording — and the retest will fail too.
  • No documented owner. Unclear accountability was one of the FCA's named improvement areas.

Frequently asked questions

Does the Consumer Duty require customer research? It does not name research as an activity, but it requires firms to evidence that customers can make informed decisions and are achieving good outcomes. In its 13 March 2026 review the FCA identified weak evidence of communication testing as an area for improvement and said reliance on sales data or an absence of complaints gives no reliable assurance — so in practice, testing with real customers is how the consumer understanding outcome gets evidenced.

What comprehension target should we set? There is no prescribed figure. The FCA's review highlighted a firm using at least 80% correct recall of key points and redrafting until it was met. What matters is that you set a threshold in advance, apply it as a gate, and record the before-and-after result rather than reporting a single unbenchmarked score.

How is a comprehension test different from a customer satisfaction survey? A satisfaction survey asks whether customers found a communication clear; a comprehension test measures whether they can correctly answer questions about what it means for them. The two frequently disagree, which is why self-rated clarity should be captured alongside actual recall rather than instead of it.

How do we evidence consideration of diverse customer needs? Sample deliberately for characteristics of vulnerability rather than relying on a general panel, offer both voice and text participation, keep the study async so it does not exclude shift workers and carers, and report results split by vulnerability characteristic. An aggregate pass rate conceals exactly the disparity the Duty asks about.

How often should communications be tested? Before and after any material change, plus continuous monitoring on live communications, with a refresh triggered by product changes or spikes in complaints and drop-off. The FCA criticised weak monitoring in 2026, and annual-only testing leaves your evidence stale for most of the year.

What should go in the board report on consumer understanding? A tested-communications inventory with reasons for any omissions, stated pass criteria and before-and-after results split by segment and vulnerability, verbatim evidence showing how communications were misread, the decision trail for each redraft, residual risks with owners and dates, distribution-chain outcomes, and the method and data lineage. The April 2026 observations asked firms to move from dashboards to analysis that draws conclusions.

Related Resources

Related Articles

Accessibility Research: How to Include Users with Disabilities in Your Studies

A practical guide to designing and conducting accessible user research — how to recruit participants with disabilities, adapt your methods, and use async AI interviews to remove barriers to participation.

AI Customer Research for Banking & Financial Services

How retail banks, credit unions, and wealth firms use AI interviews to understand customers — onboarding friction, trust, channel preferences, and product fit — at scale and in days.

AI-Powered Customer Research for Insurance Companies (2026)

How insurers run policyholder, claims, retention, and product research at scale with AI interviews - voice or text, automatically analyzed, and compliance-aware.

Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)

Six methods for testing whether your words actually work — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale instead of one participant at a time.

DPIA for User Research: When You Need One and How to Write It (2026)

A practical guide to Data Protection Impact Assessments for customer and user research: the Article 35 triggers, the WP29 nine criteria, what belongs in each section, and a worked example for AI-moderated interviews.

Research Data Retention and Deletion: How Long Should You Keep Interview Data?

There is no universal legal number - which is exactly why having no retention schedule is itself the compliance failure. A tiered, per-artifact schedule for recordings, transcripts, quotes, and reports, plus how to handle deletion requests without losing your insights.

How Long Is User Research Valid? Insight Decay and When to Re-Run a Study

Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.