Crisis and Service-Disruption Customer Research
How to run customer research during and after an outage, recall, security incident, or public controversy — the three research windows, what to ask in each, the ethics of interviewing people mid-harm, and how to measure trust recovery.
Crisis research has three windows, and they are not interchangeable: triage during the incident (hours), trust damage and message testing in the first week to month, and behaviour at around 90 days. Most teams do only the last one, by which point they are measuring the aftermath of decisions they already made blind. The constraint has always been speed — conventional recruiting and moderated fieldwork takes three to six weeks, and a crisis window closes in days. That constraint is the part that changed.
This guide covers customer-facing disruptions of any kind: outages and degradations, product recalls, security and privacy incidents, pricing shocks, sudden feature removals, and public controversies. For structured internal postmortems of AI system failures specifically, see AI Incident Postmortem User Research.
Window 1: triage, during the incident
The job here is operational, not strategic. You need to know within hours:
- Who is actually affected, as opposed to who is on the affected infrastructure. These sets are never the same size.
- How severe the impact is per segment — blocked entirely, degraded, or inconvenienced.
- Whether the workaround works. Support will tell you they gave one out. Only customers can tell you whether it functions in their environment.
- What people want said. Frequency, channel, and level of technical detail. Teams consistently over-index on technical explanation and under-index on expected time to resolution.
What not to do in this window: send a satisfaction survey. An NPS request that lands while someone's service is down measures the outage and nothing else, and it reads as oblivious. If you field anything during an incident, label it plainly as incident research, keep it under five minutes, and lead with acknowledgement.
| Window | Timing | Primary question | Instrument |
|---|---|---|---|
| Triage | Hours to 72h | Who is hurt, how badly, what do they need said? | Short AI-moderated interview, affected customers only |
| Trust damage | Day 3 to day 30 | What did this change about how they see us? | Depth interviews + scaled trust items, all three segments |
| Message testing | Overlaps window 2 | Which version of the message reduces perceived risk? | Comparative interviews on 2-3 message variants |
| Behaviour | 60 to 120 days | Did it change what they actually do? | Interviews joined to renewal and usage data |
Window 2: trust damage and message testing
This is the window that gets skipped, and it is the one with the most leverage. Between roughly day three and day thirty, memory is intact, people still use their own words rather than the phrasing your comms team introduced, and the recovery message has not yet been sent to everyone.
Segment into three groups and never blend them:
- Directly affected — experienced the failure themselves.
- Aware but unaffected — heard about it from status pages, press, or peers.
- Unaware — a genuine control group, and the group that tells you whether broad communication helps or merely informs people who had not noticed.
The gap between groups 1 and 2 is your reputational spillover. If aware-but-unaffected customers report almost as much trust damage as affected ones, your problem is narrative rather than technical, and engineering fixes alone will not close it.
What to measure
- Perceived recurrence risk — "how likely is this to happen again" predicts churn far better than satisfaction with the incident response does.
- Attribution — do they blame your engineering, your priorities, or bad luck? Attribution to priorities is the most damaging and the hardest to reverse.
- Threshold — what would have to happen, or happen again, for them to start evaluating alternatives. Ask it directly; people answer it directly.
- Communication assessment — separately for timeliness, honesty, and usefulness. These come apart. Fast and evasive scores worse than slow and candid.
- Repair expectations — what would meaningfully make it right, asked open-endedly before you offer options. The answers are frequently cheaper than the credit you were about to issue.
Testing the message before you send it
Show two or three versions of the recovery communication to a small sample of affected customers and ask, for each, what it makes them believe about whether the problem will recur. This is the single highest-return study in the playbook, and it routinely produces an uncomfortable finding: the version legal preferred reads as evasive and increases perceived risk.
There is broader evidence for the underlying instinct. The 2026 Edelman Trust Barometer, based on nearly 34,000 respondents across 28 countries, found that 75% of people endorse companies consulting people with different views when making decisions, and 74% endorse constructively engaging with those who criticise the company. The reputational upside is in visibly asking, not only in what you conclude.
Window 3: behaviour, at around 90 days
Stated forgiveness and renewal behaviour diverge, sharply and predictably. People are generous in interviews and unsentimental at renewal. So the third wave joins research to actual data:
- Compare renewal, expansion, and usage between affected and unaffected cohorts.
- Interview the affected customers who did renew — quiet retention is under-researched and tells you what actually repaired the relationship.
- Interview those who churned, and specifically ask whether the incident was cause or catalyst. It is usually the catalyst for a decision that was already forming, which is a very different remediation problem.
- Re-run the same scaled trust items from window 2 so you have a trend rather than a snapshot. Identical wording matters more than perfect wording.
If you run only one wave, you get a number with nothing to compare it against — the classic failure of crisis measurement. Two waves with the same instrument beat one thorough wave every time.
Question design in a crisis
Crisis interviews fail in characteristic ways. The fixes are specific:
- Acknowledge in the introduction. State plainly that you are researching the incident. Pretending otherwise makes the first two minutes awkward and contaminates everything after them.
- Let people vent first. Open with something like "tell me what the last week has been like from your side." The complaint is going to come out regardless; getting it out early means later answers are not carrying it.
- Never pre-load the defence. "Given that this was caused by an upstream provider, how do you feel about our response?" is not a question, it is a press release with a question mark.
- Separate the incident from the response. Many customers are fine about the failure and furious about the communication. One combined question hides that entirely, and it is the most actionable split in crisis research.
- Ask for the counterfactual. "What would we have had to do for this to be a non-event for you?" produces more usable answers than any satisfaction scale.
- Cap scaled items. Three or four, and only after the open questions. Scales in a crisis mostly measure current mood; the reasoning is in the prose.
Ethics: when not to research
- Do not research people mid-harm. If customers are actively losing money, revenue, or patients, support them; do not interview them. Wait.
- Do not incentivise into resentment. A gift card offered to someone who just lost a day of work often reads as an attempt to buy silence. Offer something proportionate, or offer nothing and be honest about why you are asking.
- Do not disguise research as support. A support conversation that turns into a study without consent damages both.
- Do not research where the incident is legally live. If the matter is under investigation or litigation is anticipated, coordinate with counsel first — and note that research records created in that period are likely discoverable. See Legal Hold and E-Discovery for Research Data.
Why the speed constraint is the whole problem
Traditional crisis research does not fail because researchers ask bad questions. It fails because by the time a panel is recruited, sessions are scheduled, moderators run them, and transcripts are analysed, the apology has been sent, the credits have been issued, and the decisions are all made. Three to six weeks is the normal cycle, and a crisis window is measured in days.
This is the use case where AI-moderated research is not merely more efficient but categorically different:
- A study goes live in hours. You configure the brief and publish; there is no moderator roster and no scheduling round.
- Interviews run around the clock. Affected customers are dealing with the incident during their working day and will talk at 10pm. An always-available AI interviewer meets them there — see Always-On User Interviews.
- The AI probes the actual grievance. This matters more in a crisis than anywhere else, because each customer's version of the harm is different. A fixed SurveyMonkey or Typeform form asks everyone the same six questions; Koji's interviewer follows up on what the specific person just said, which is how you find out that the real damage was to their credibility with their own customers rather than to their uptime.
- Voice captures what text flattens. Frustration, resignation, and relief are audible. In crisis research, affect is data.
- Analysis is continuous. Reports build as interviews complete rather than after fieldwork closes, so leadership can watch the picture form during the incident instead of receiving it afterwards.
- Structured items make waves comparable. Koji's structured questions — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — let you hold identical scaled trust items across waves 2 and 3 while the open-ended conversation stays adaptive. That combination, fixed measures plus adaptive probing, is exactly what trust-recovery tracking needs and exactly what a survey tool cannot give you.
Common mistakes
- Only measuring after it is over. The most valuable window is during.
- Blending affected and unaffected customers. Hides both the damage and the spillover.
- Sending NPS mid-incident. Measures the incident, damages the relationship.
- Testing the message on internal staff. They know the mitigations and cannot un-know them.
- Measuring once. No baseline, no trend, no way to tell recovery from decay.
- Ignoring the customers who stayed. Retention research is where the repair mechanism is visible.
- Letting comms write the questions. Every question becomes a defence.
Related Resources
- AI Incident Postmortem User Research — structured postmortems for AI system failures
- Customer Retention Research — the 90-day behavioural wave in depth
- Customer Signals — detecting problems before they become incidents
- Always-On User Interviews — why 24/7 fielding matters when the window is days
- Brand Perception Survey Guide — measuring reputational spillover
- Customer Experience Benchmarking — the pre-crisis baseline you will wish you had
- Structured Questions Guide — holding measures identical across waves
Frequently asked questions
Should we survey customers during an active outage? Not with a satisfaction survey. Sending an NPS or CSAT request to someone whose service is currently down measures the outage rather than the relationship, and it reads as tone-deaf at the worst possible moment. What is appropriate during an incident is short, clearly-labelled triage research that acknowledges the problem and asks what the person needs — impact severity, workaround viability, and what they want communicated.
How soon after an incident should we run research? Run triage research during the incident if you need operational answers within hours. Run the substantive trust and communications research in the first seven to thirty days, while memory is intact and the language people use is still their own. Then run a behavioural wave at around 90 days, because stated forgiveness and actual renewal behaviour diverge, and the second number is the one that matters commercially.
How do you ask about a crisis without leading the witness? Acknowledge the incident in the introduction so nobody has to pretend it did not happen, then ask open questions before any scaled ones. Let people vent first — an early open-ended question that invites the complaint gets it out of the way so later answers are not carrying it. Avoid framings that pre-load your defence, like asking whether people understand that the outage was caused by a third party. Ask what they experienced, then what they concluded.
Who should be in the sample during a crisis? Three groups, tracked separately: customers directly affected, customers aware of the incident but not affected, and customers unaware of it. These groups produce very different answers, and blending them hides both the severity of the damage among the affected and the reputational spillover among the merely aware. Sizing the unaware group also tells you whether broad communication would help or would simply inform people who had not noticed.
Can you test crisis communications before sending them? Yes, and it is the highest-return research in the whole playbook. Show two or three versions of the message to a small sample of affected customers and ask what each one makes them believe about whether the problem will recur. Teams routinely discover that the version legal preferred reads as evasive and increases the perception of risk. This can be turned around in a few hours with AI-moderated interviews.
How fast can Koji field a crisis study? A study can be created, configured, and published in well under a day, and because Koji's AI interviewer runs voice and text conversations around the clock with no moderator to schedule, responses arrive as soon as the link reaches people. Analysis runs continuously rather than after fieldwork closes, so you can watch the picture form during the incident instead of receiving it after the decision has been made.
Related Articles
AI Incident Postmortems: How to Investigate Model Failures with User Evidence (2026)
Logs tell you what your model output. They cannot tell you what it cost the person on the other end. A practical guide to running blameless AI incident postmortems with real user evidence - and meeting the reporting clocks that now apply.
Always-On User Interviews: Run 24/7 With an AI Moderator
Run user interviews around the clock without a researcher on every call. An AI moderator interviews participants whenever they show up — across timezones, in voice or text, with results scored and themed automatically.
How to Run Brand Perception Surveys That Reveal What Customers Really Think
The complete guide to brand perception and brand tracking surveys. Learn how to measure awareness, sentiment, associations, and positioning using Koji's conversational approach to uncover authentic brand perceptions.
Customer Experience Benchmarking: How to Measure Against Industry Standards
A complete guide to CX benchmarking — how to measure your customer experience performance against competitors and industry standards using both quantitative metrics and qualitative interviews.
Customer Retention Research: The Complete 2026 Playbook for Reducing Churn Before It Happens
A practitioner's guide to customer retention research — how to combine churn interviews, stay interviews, NPS follow-ups, and continuous voice-of-customer programs to reduce churn 25% or more. Includes question templates, sampling frameworks, and how AI-moderated research scales retention listening across your entire customer base.
Customer Signals: Building an Always-On Insight Layer for Product Decisions
What customer signals are, where they come from, and how to build an always-on signal layer that continuously feeds product decisions — instead of relying on quarterly research projects.
NPS Benchmarks 2026: Net Promoter Score by Industry (Complete Reference)
Compare your NPS to 2026 industry benchmarks for SaaS, ecommerce, financial services, healthcare, and more. Includes what counts as "good", scoring math, and how to dig into the "why" behind your score with AI follow-up interviews.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.