Back to docs
Use Cases

Trust and Safety Research: How to Study Harm, Reporting Flows, and Moderation With Real Users

Trust and safety teams run on tickets and telemetry, which only show the harm that got reported. Here is how to research the harm that did not — the four research objects, the DSA and Online Safety Act obligations that require it, and a safeguarding protocol that protects participants.

Answer first: Trust and safety research is the study of how users actually experience harm, reporting, and enforcement on your platform — as distinct from the moderation telemetry that only ever shows you the harm someone bothered to flag. It has four distinct objects: the harmed user's experience, the usability of your reporting and appeals flows, the comprehension of your policies, and the working experience of your reviewers. Under the EU Digital Services Act and the UK Online Safety Act it is no longer purely voluntary: risk assessments are expected to be informed by evidence about real user experience, and "we had no complaints" is explicitly not evidence. Async, anonymous AI-moderated interviews are unusually well suited to this work, because the people with the most to tell you are the ones least willing to say it to a stranger on a video call.

The blind spot in moderation telemetry

Every trust and safety function has a dashboard: reports received, actions taken, appeal rates, median time to decision, prevalence estimates from sampled review. It is genuinely useful, and it is systematically biased in one direction — it can only measure harm that entered the pipeline.

What it cannot tell you:

  • How many users experienced something harmful and did not report it, and why not
  • Whether people could find the reporting flow, or gave up halfway through it
  • Whether the categories in your report form match the harms users actually experience, or force them into the nearest wrong box
  • What users believed would happen after reporting, versus what did
  • Whether your community guidelines are comprehensible to the people expected to follow them
  • Why some harmed users leave silently instead of escalating

Those are all research questions, and none of them is answerable from log data. This is the same structural gap that makes non-user research necessary in a commercial context: the people who churned in silence are invisible to your own instrumentation.

Consumer sentiment data suggests the gap is wide. Surveys of consumers have found roughly 64% believing companies do not remove enough harmful content while about 32% feel moderation has gone too far — a split that no single prevalence metric can reconcile, and one that only qualitative work can explain. The market has responded: content moderation spending is estimated at around USD 11.6 billion in 2025, forecast to reach roughly USD 13.3 billion in 2026 and USD 26 billion by 2031, with TELUS Digital's Safety in Numbers research reporting 58% of companies treating moderation as a higher priority than a year earlier. Spending is rising faster than understanding.

The four research objects

Most teams conflate these, then wonder why one study cannot answer everything. Separate them.

ObjectCore questionTypical methodWho you recruit
Harmed usersWhat happened, what did you do, what did you need?Retrospective interviews, critical incident techniqueUsers who filed reports; users who did not but experienced harm
Reporting and appeals flowsCan people find it, complete it, and understand the outcome?Usability testing, walk-through interviewsReporters, appellants, and never-reporters
Policy comprehensionDo users understand the rules they are held to?Comprehension testing on real policy textGeneral population of your user base
Reviewers and moderatorsWhat makes a decision hard, slow, or distressing?Anonymous internal interviewsYour own T&S staff and vendor reviewers

The critical incident technique is the single best-fitting method for the first object: it anchors the participant to one specific, recent, memorable event rather than asking them to generalise about safety in the abstract — which is where these interviews usually go wrong.

The regulatory driver: this is now expected evidence

EU Digital Services Act. Under Article 34, providers of very large online platforms and search engines must diligently identify, analyse and assess systemic risks arising from the design, functioning, and use of their service — at least annually, and in any event before deploying functionality likely to have a critical impact on identified risks. The enumerated categories cover illegal content, fundamental rights, civic discourse and electoral processes, gender-based violence, public health, and minors including behavioural addiction. These assessments are expected to be informed by research and by the views of affected individuals and civil society, not solely by internal telemetry. Article 35 then requires mitigation measures that are reasonable, proportionate and effective, and tailored to the specific risks identified — which is impossible to demonstrate if you never established how the risk manifests for real users.

UK Online Safety Act. In-scope services were required to complete illegal-content risk assessments by 24 July 2025, with safety duties applying from 25 July 2025, and services likely to be accessed by children to complete children's risk assessments by 31 July 2025 following Ofcom's Protection of Children Codes published on 24 April 2025. Ofcom has signalled it expects providers of categorised (Category 1 and 2A) services to supply copies of their latest risk assessment records by October 2026, with the register of categorised services expected in July 2026. Services adopting measures other than those in the Codes must keep records demonstrating an equivalent standard — and user evidence is the most direct way to do that.

The through-line for both regimes: an assessment grounded in what users actually experience is defensible; one grounded in the absence of complaints is not. For the parallel AI-specific obligations, see the EU AI Act and user research and AI governance frameworks for research.

Designing the study: risk category to question type

Koji's six structured question types let one instrument carry both the countable figure your risk assessment needs and the mechanism your product team needs.

Research questionQuestion typeWhy
Did you experience X in the last 90 days?yes_noClean incidence denominator for the assessment
What did you do next?single_choiceReported / blocked / left / did nothing — the drop-off shape
Why did you not report it?multiple_choiceFriction inventory: could not find it, did not think it would help, feared retaliation, felt embarrassed
How confident are you that the platform would act?scaleTrust-in-enforcement trend line, quarter over quarter
Rank these harms by how much they affect your experiencerankingForced prioritisation — far more decision-useful than five separate severity scales
Walk me through what happened and what you neededopen_endedThe mechanism, with AI follow-up probing

Set follow-up depth deliberately here. For sensitive items, one gentle follow-up is often the right ceiling; for flow-usability items, two or three probes get you to the actual point of abandonment.

The safeguarding protocol

Research on harm can harm. Treat this as non-negotiable, and write it into the study before you field it.

  1. Never require a participant to recount an incident in detail. Ask what they needed and what the platform did, and let them choose how much of the event to narrate. Make every sensitive question optional.
  2. Signpost support. Include relevant helplines and reporting routes on the interview landing page and in the completion screen, not only in the consent text.
  3. Do not collect the harmful content itself. You are researching the experience, not gathering evidence. Collecting illegal material creates a handling obligation you do not want and cannot discharge inside a research tool.
  4. Minimise identifiability. Run without identifying fields wherever the research question allows it, and never join safety-research responses to account records without explicit, specific consent.
  5. Consent must be specific about use. State plainly that responses inform safety policy and may be summarised in regulatory documentation. See research consent form templates and interview recording consent laws.
  6. Run a DPIA. Data about experiences of harm is, in practice, special-category-adjacent and high risk under GDPR. See DPIA for user research.
  7. Apply stricter rules with minors. Age-appropriate language, guardian consent where required, and no open-ended prompts that could elicit disclosure of ongoing abuse without a route to support. See user research with children and teens.

Why async AI interviews fit this work unusually well

This is one of the few research domains where an AI moderator is not merely faster but methodologically preferable:

  • Disclosure without a witness. Participants disclose stigmatised experiences more readily when no human is watching. A moderated video call maximises social pressure at exactly the wrong moment; an async text interview minimises it.
  • The participant controls the pace. They can pause, leave, come back. Nobody is waiting on the other end of a scheduled hour.
  • It runs when harm is fresh. Reports arrive at 3am on a Sunday. A link-based interview triggered by a report event catches the experience within hours, not after the next fieldwork cycle.
  • Depth at population scale. T&S questions need thousands of responses for a defensible incidence figure and deep narrative for the mechanism. AI moderation is the only approach that delivers both from one instrument.
  • Consistency for the record. Every participant gets the same core questions asked the same way, which matters when the output feeds a regulatory document — while probing still adapts to what each person says.
  • No third-party recording tool. Because interviews are async and link-based, there is no meeting recorder in the loop — one fewer processor handling accounts of harm.

Koji's quality scoring also matters more here than elsewhere: safety studies attract both genuine distress and bad-faith responses, and a composite 1–5 score across relevance, depth and question coverage lets you separate substantive testimony from noise before anything reaches a policy decision.

Researching your reviewers

The most neglected object of the four. Moderators and reviewers hold the tacit knowledge of where your policy is ambiguous — the edge cases they escalate, the categories they distrust, the queues where the guidance contradicts itself. TELUS Digital's research indicates around 44% of organisations now run hybrid human-plus-automation models with about 21% remaining human-led, and reviewer wellbeing is a recognised structural pressure on the workforce.

Treat this as internal research with strong anonymity guarantees: it must never read as performance evaluation. Ask which decisions were hardest and why, which policy lines they cannot apply consistently, what the automation gets wrong, and what support they actually need. Run it anonymously — see anonymous employee research with AI interviews — and aggregate to queue level rather than individual level.

Testing whether anyone understands your rules

Community guidelines are usually written by policy and legal teams, then never tested on the population expected to follow them. Enforcement against a rule nobody comprehends produces appeal volume, not compliance. Test the actual text: show a real policy excerpt, then ask participants to judge specific scenarios against it, using single_choice for their verdict and open_ended for their reasoning. Disagreement rates by scenario tell you exactly which clause is failing. The mechanics are the same as any comprehension study — see content testing.

Anti-patterns

  • Treating report volume as prevalence. Report volume is a function of reporting-flow usability and trust in enforcement at least as much as of actual harm.
  • Recruiting only reporters. The never-reporters are the finding. Recruit them deliberately — see researching hard-to-reach audiences.
  • Asking about "safety" in the abstract. Anchor to specific recent incidents or you will collect opinions about the internet.
  • Running it once, for the audit. Both the DSA and the OSA contemplate ongoing assessment; annual-only research means your evidence is stale for eleven months of every year.
  • Letting a single harrowing account set policy. Pair narrative depth with incidence figures from the same instrument, and label severity separately from confidence.

Frequently asked questions

What is trust and safety research? It is user research aimed at how people experience harm, reporting, enforcement and policy on a platform — covering harmed users, the usability of reporting and appeals flows, comprehension of the rules, and the working experience of reviewers. It complements moderation telemetry, which can only measure harm that was reported.

Is trust and safety research legally required? Not as a named activity, but it is effectively expected evidence. The DSA requires very large platforms to assess systemic risks at least annually informed by research and the views of affected individuals, and the UK Online Safety Act requires risk assessments that Ofcom can inspect, with non-Code measures needing records demonstrating an equivalent standard. Relying on an absence of complaints is not a defensible basis for either.

How do you research harm without re-traumatising participants? Make sensitive questions optional, never require detailed recounting of the incident, focus on what the participant needed and what the platform did, signpost support resources on the landing and completion screens, and avoid collecting the harmful content itself. Async formats help because the participant controls the pace and can stop at any point.

Why use AI-moderated interviews for safety topics rather than human moderators? Because disclosure of stigmatised experience is higher when no human is present, the participant sets the pace, and the interview can run within hours of the incident at any time of day. AI moderation also delivers consistent core questioning across thousands of participants — which matters when the output feeds a regulatory document — while still probing each person's specific account.

Should we interview our own moderators? Yes. Reviewers hold the tacit knowledge of where policy is ambiguous and where automation fails, and they are the least-researched group in most T&S organisations. Run it anonymously, aggregate to queue level rather than individual level, and separate it explicitly from performance management.

How do we turn safety research into a risk assessment? Use one instrument that produces both a countable incidence figure (via yes_no and single_choice items) and narrative mechanism (via open_ended items with AI probing), then map each finding to the risk categories your regime enumerates, label severity and confidence separately, and record the mitigation you chose and why. That combination is what makes an assessment auditable rather than assertive.

Related Resources

Related Articles

Anonymous Employee Research with AI Interviews: Get the Honest Feedback Surveys Miss

Run truly anonymous employee research at scale with AI voice and text interviews. Capture honest feedback on culture, leadership, retention risk, and engagement — without HR ever knowing who said what. Koji removes intake forms, strips identifiers, and still produces aggregated themes and quotes you can act on.

Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)

Six methods for testing whether your words actually work — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale instead of one participant at a time.

Critical Incident Technique: The Interview Method That Captures What Really Matters

Learn how to use the Critical Incident Technique (CIT) to uncover the specific moments that shape user experience. Developed by Flanagan (1954), CIT interviews collect real incidents — not generalizations — to reveal actionable patterns in user behaviour.

DPIA for User Research: When You Need One and How to Write It (2026)

A practical guide to Data Protection Impact Assessments for customer and user research: the Article 35 triggers, the WP29 nine criteria, what belongs in each section, and a worked example for AI-moderated interviews.

The EU AI Act and User Research: What AI-Moderated Interviews Actually Require (2026)

AI-moderated customer interviews sit in the EU AI Act's limited-risk transparency tier, not the high-risk tier. Here is exactly what Article 50 requires from 2 August 2026, the two things that escalate a study to high-risk, and a compliance checklist you can run this week.

Non-User Research: How to Interview the People Who Never Chose You

Your roadmap is built entirely from the opinions of people who said yes. A guide to researching the four types of non-user — rejecters, the unaware, the DIY crowd, and the constrained — including how to recruit them and what to ask.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

User Research With Children and Teens: COPPA, Parental Consent, and Assent

Researching under-13s triggers COPPA verifiable parental consent - including a separate consent before any child data trains AI. Here is the compliance path, the parent-mediated pattern most teams should use instead, and how to design sessions that actually work with young participants.