Vulnerable Customer Research: How to Evidence Good Outcomes Under FCA Rules
The FCA told firms in 2025 that customers in vulnerable circumstances still report worse outcomes — and that outcomes monitoring and customers who never disclose are the two hardest gaps. Both are research problems. Here is how to design a programme that closes them.
If you rely on your vulnerability flags to evidence outcomes for customers in vulnerable circumstances, your evidence is drawn from a biased sample and your monitoring is measuring the wrong population. Most customers in vulnerable circumstances never disclose. The FCA has explicitly named both problems — how to monitor outcomes, and how to reach customers who do not tell you — and neither is solved by better reporting on the flags you already hold. They are solved by primary research designed to find vulnerability rather than waiting for it to be declared.
This guide is about that research programme. It sits alongside the Consumer Duty guide, which covers comprehension testing for the consumer understanding outcome; here the subject is vulnerability as a sampling and monitoring problem across all four Duty outcomes.
What the FCA actually expects
FG21/1, Guidance for firms on the fair treatment of vulnerable customers, is issued under the Principles rather than as Handbook rules — which is why firms so often under-invest in it and then find it applied firmly at supervision. It sets expectations in four areas:
- Understanding the needs of vulnerable customers in your target market and customer base
- Making sure staff have the skills and capability to recognise and respond to those needs
- Responding to those needs through product and service design, flexible customer service and communications
- Monitoring and assessing whether you are meeting and responding to those needs
Areas 1 and 4 are research obligations in everything but name. Area 3 cannot be evidenced without them.
The guidance frames vulnerability through four drivers, and the practical value of the model is that it is situational — most of these are states people move through, not permanent traits.
| Driver | What it covers | Typical research implication |
|---|---|---|
| Health | Physical or mental health conditions, cognitive impairment | Journey length, memory load, ability to complete in one sitting |
| Life events | Bereavement, job loss, relationship breakdown, caring responsibilities | Timing-sensitive research; consent and duty of care |
| Resilience | Irregular income, over-indebtedness, low savings | Financial stress changes decision-making under any communication |
| Capability | Low literacy or numeracy, limited digital skills, low financial knowledge | Method itself can exclude the population you need |
Because vulnerability is situational, a single annual study is structurally incapable of monitoring it. Someone bereaved in March is a different research subject in September.
The 2025 review verdict, and what it means for your research
The FCA published its review of firms' treatment of customers in vulnerable circumstances on 7 March 2025, updated in December 2025. It was substantial work: interviews with 29 firms across 12 markets, a survey of 725 firms, commissioned quantitative and qualitative consumer research with around 1,500 participants, and analysis of the 2020, 2022 and 2024 Financial Lives waves.
The verdict was uncomfortable. Consumers in vulnerable circumstances continue to report poorer outcomes than other consumers, and the gap is widest for people with multiple vulnerability characteristics. The Consumer Duty has visibly renewed firms' focus, but the progress seen in firm-level assessments has not yet shown up in the UK-wide data.
Firms themselves asked the FCA for four things: more sector-specific case studies, better guidance on outcome monitoring methods, approaches for customers who do not disclose vulnerability, and recognition of how vulnerability intersects with Equality Act protected characteristics. The FCA chose not to rewrite FG21/1, publishing good-practice case studies instead — which means the expectation is unchanged and the burden of designing the monitoring sits with you.
Three things follow directly for a research programme:
- Aggregate outcome metrics are not enough. A firm-level satisfaction or complaints number cannot show a gap it is not split by.
- Disclosed flags are the wrong denominator. They measure who told you, which is a different question from who is affected.
- Multiple characteristics need to be visible. If your analysis treats vulnerability as one binary, you cannot see the group with the worst outcomes.
Problem one: the customers who never tell you
Non-disclosure is the central methodological problem. People do not disclose because they do not recognise the label, because they fear consequences for credit or access, because the channel gave them no natural opening, or because they were in a hurry and the disclosure question sat behind a menu.
The fix is not to ask harder. It is to measure the drivers rather than the label:
- Ask about circumstances, not categories. "In the last twelve months, have you experienced any of the following?" with a
multiple_choicelist of concrete life events outperforms any question containing the word vulnerable. - Capture capability behaviourally. A
scalequestion on how confident someone felt completing the last step of a journey tells you more than a self-declared digital skills rating. - Probe the ones that matter. In Koji, an
open_endedfollow-up fires automatically when a confidence score is low — "you said you were not sure what would happen next; what did you do at that point?" — which is where the actual failure story lives. - Let people answer by voice. Typing is itself a capability filter, and a spoken answer from someone with low literacy carries detail no form field would have captured.
Run this in the research instrument, and you get a vulnerability signal for every participant, not only the ones already flagged. That is the denominator the FCA is asking about, and it lets you do the analysis that matters: comparing outcomes between customers you had flagged, customers you had not flagged but who show characteristics, and customers who show none.
Almost every firm that runs this comparison for the first time finds the same thing — the undisclosed group looks materially worse than the flagged group, because the flagged group has been receiving support.
Problem two: monitoring outcomes rather than reporting activity
Outcomes monitoring fails in a predictable way: firms report what they did (calls handled, flags recorded, training completed) instead of what happened to customers. Activity is easy to count and proves nothing.
A workable monitoring design measures each Duty outcome, split by vulnerability characteristic:
| Outcome | What to measure with customers | Method |
|---|---|---|
| Products and services | Did the product still fit after circumstances changed? | Triggered study after a life-event signal |
| Price and value | Did the customer understand what they were paying and why? | Comprehension test plus scale value perception |
| Consumer understanding | Can the customer correctly state what happens next? | Recall test with pre-set pass criteria |
| Consumer support | Could the customer get help, first time, through their channel of choice? | Post-interaction study across channels |
Two design rules carry most of the weight. First, trigger studies on events rather than a calendar — a bereavement notification, a missed payment, a power of attorney registration, a complaint about a communication. Vulnerability is situational and your evidence should be too. Second, always report split, never only aggregate. An 84% satisfaction figure that hides 61% among customers with three or more characteristics is worse than no figure, because it creates false assurance — exactly the pattern the FCA criticised when it said reliance on sales data or an absence of complaints provides no reliable assurance.
Problem three: your research method is probably excluding them
This is the failure that quietly invalidates everything above. Standard research methods select against precisely the customers you need:
- Scheduled video interviews exclude shift workers, carers and anyone without reliable connectivity or a quiet room.
- Panel recruitment over-represents confident, digitally fluent, repeat participants.
- Long written surveys select for literacy and stamina.
- App downloads and account creation exclude low-capability users at the first step.
- Incentives paid only by digital transfer exclude the unbanked and underbanked.
An inclusive design does the opposite. Make participation asynchronous so it fits around caring responsibilities and irregular shifts. Offer voice as a first-class option, not a fallback — for someone with low literacy or a visual impairment, speaking is not an accommodation, it is the usable path. Keep sessions short and allow them to be resumed. Write at a genuinely plain reading level. Offer the study in the languages your customer base actually speaks. And recruit deliberately for the characteristics rather than hoping a general panel contains them.
This is the strongest structural argument for AI-moderated research in this domain: a study that runs by voice or text, asynchronously, in any language, with no moderator to schedule and no software to install, removes most of the access barriers that make vulnerable customers hard to research at all. Koji runs exactly that shape of study, and because the AI probes automatically, a short session still reaches the depth a moderator would have needed a booked hour for.
Building the evidence pack
What a board or supervisor needs to see, in order:
- Who you researched, and how you found them — including how you identified characteristics beyond disclosed flags.
- Outcome results split by driver and by number of characteristics, with the comparison against the non-vulnerable group stated explicitly.
- Verbatim evidence of where the journey failed, quoted, with the transcript retained.
- What changed as a result, with dates and owners — the decision trail, not just the finding.
- Re-test results after the change, which is what converts a finding into evidence of improvement.
- Residual gaps, named, with dates. Boards get more credit for a known gap with an owner than for an unblemished dashboard.
- Method and data lineage — how participants were sampled, what was asked, where transcripts and exports live.
Add the intersection the FCA called out: report where vulnerability characteristics overlap with Equality Act protected characteristics, because that is where both regulatory and reputational risk concentrate.
Common mistakes
- Treating vulnerability as a permanent customer attribute rather than a situational state.
- Monitoring only customers who disclosed, then reporting the result as coverage of vulnerable customers.
- Collapsing all vulnerability into one binary flag, which hides the multiple-characteristic group with the worst outcomes.
- Reporting activity metrics as outcome evidence.
- Running the research with methods that structurally exclude low-capability customers, and never noticing.
- Testing once a year, when the drivers are events that occur continuously.
- Collecting health and financial hardship data without a lawful basis, a retention limit and a deletion route — vulnerability data is often special category data and deserves a data protection impact assessment.
Frequently asked questions
Does the FCA require firms to do research with vulnerable customers? FG21/1 does not name research as an activity, but it requires firms to understand the needs of vulnerable customers in their customer base and to monitor whether those needs are being met. Neither can be evidenced from internal activity metrics alone, and the FCA's March 2025 review found that firms themselves asked for better guidance on outcome monitoring methods. Primary research is how those two expectations get satisfied in practice.
How do we research customers who never disclose their vulnerability? Stop asking for the label and measure the drivers instead. Ask about concrete circumstances in the last twelve months with a multiple_choice list, capture confidence and comprehension behaviourally with scale questions, and probe low scores with open_ended follow-ups. That produces a vulnerability signal for every participant, letting you compare outcomes across flagged customers, unflagged customers with characteristics, and everyone else.
What are the four drivers of vulnerability in FG21/1? Health, life events, resilience and capability. Health covers physical and mental health conditions; life events covers bereavement, job loss, relationship breakdown and caring responsibilities; resilience covers irregular income, over-indebtedness and low savings; capability covers low literacy, numeracy, digital skills or financial knowledge. Most are situational states rather than permanent traits, which is why point-in-time annual research cannot monitor them.
How often should we research vulnerable customer outcomes? Continuously, triggered by events rather than by the calendar — bereavement notifications, missed payments, power of attorney registrations, complaints about communications — plus before and after any material change to a product, journey or communication. An annual study leaves the evidence stale for most of the year and cannot capture a situational driver.
How do we stop our research method from excluding the customers we need? Make it asynchronous so it fits around shifts and caring responsibilities, offer voice as a first-class option rather than a fallback, keep sessions short and resumable, write at a genuinely plain reading level, offer the languages your customer base speaks, avoid app installs and account creation, and recruit deliberately for the characteristics instead of relying on a general panel. Then check the achieved sample against your customer base and report the gap honestly.
What should we report to the board on vulnerable customer outcomes? Outcome results split by driver and by number of characteristics with an explicit comparison to non-vulnerable customers, how the sample was identified beyond disclosed flags, verbatim evidence of journey failures, the decision trail for what changed, re-test results after those changes, named residual gaps with owners and dates, and the intersection with Equality Act protected characteristics. Aggregate satisfaction figures without splits create false assurance rather than evidence.
Related resources
- Structured Questions Guide — the six question types behind driver measurement and comprehension tests
- FCA Consumer Duty Customer Research — comprehension testing for the consumer understanding outcome
- Accessibility Research Guide — including users with disabilities in your studies
- AI Customer Research for Banking & Financial Services — the sector view
- AI-Powered Customer Research for Insurance Companies — insurance-specific outcome monitoring
- DPIA for User Research — assessing risk before you collect health and hardship data
- Research Data Retention and Deletion — retention limits for special category data
Related Articles
Accessibility Research: How to Include Users with Disabilities in Your Studies
A practical guide to designing and conducting accessible user research — how to recruit participants with disabilities, adapt your methods, and use async AI interviews to remove barriers to participation.
AI Customer Research for Banking & Financial Services
How retail banks, credit unions, and wealth firms use AI interviews to understand customers — onboarding friction, trust, channel preferences, and product fit — at scale and in days.
AI-Powered Customer Research for Insurance Companies (2026)
How insurers run policyholder, claims, retention, and product research at scale with AI interviews - voice or text, automatically analyzed, and compliance-aware.
Anonymizing Customer Interview Data: A Practical Guide for Privacy-Safe Research
Five operational techniques for handling PII in AI customer interviews — from intake-time anonymization to stakeholder-safe quote sharing — without sacrificing research signal.
FCA Consumer Duty Customer Research: How to Evidence Consumer Understanding and Fair Value
The FCA's 2026 reviews were blunt: sales data and an absence of complaints prove nothing about consumer understanding. Here is how to test communications with real customers, hit a comprehension target, and build a board-report evidence pack that survives scrutiny.
DPIA for User Research: When You Need One and How to Write It (2026)
A practical guide to Data Protection Impact Assessments for customer and user research: the Article 35 triggers, the WP29 nine criteria, what belongs in each section, and a worked example for AI-moderated interviews.
Research Data Retention and Deletion: How Long Should You Keep Interview Data?
There is no universal legal number - which is exactly why having no retention schedule is itself the compliance failure. A tiered, per-artifact schedule for recordings, transcripts, quotes, and reports, plus how to handle deletion requests without losing your insights.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.