{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-07-31T02:07:50.536Z"},"content":[{"type":"documentation","id":"25540909-0086-4363-9729-3707fd7033ce","slug":"trust-and-safety-research","title":"Trust and Safety Research: How to Study Harm, Reporting Flows, and Moderation With Real Users","url":"https://www.koji.so/docs/trust-and-safety-research","summary":"Trust and safety research studies how users experience harm, reporting, enforcement and policy — as distinct from moderation telemetry, which can only measure harm that was reported. It has four objects: harmed users, the reporting and appeals flows, policy comprehension, and reviewers themselves. Regulatory drivers: DSA Article 34 requires very large online platforms and search engines to assess systemic risks at least annually and before deploying functionality with critical impact, covering illegal content, fundamental rights, civic discourse and elections, gender-based violence, public health and minors including behavioural addiction, informed by research and the views of affected individuals; Article 35 requires mitigation tailored to identified risks. Under the UK Online Safety Act, illegal-content risk assessments were due 24 July 2025 with duties from 25 July 2025, children's risk assessments by 31 July 2025 after Ofcom's Protection of Children Codes of 24 April 2025; Ofcom expects Category 1 and 2A record copies by October 2026 and the categorised register in July 2026. The safeguarding protocol: never require detailed recounting, signpost support, do not collect the harmful content, minimise identifiability, obtain specific consent, run a DPIA, and apply stricter rules for minors. Async AI-moderated interviews suit this domain because disclosure is higher without a human witness, the participant controls pace, interviews can run within hours of an incident, and one instrument yields both incidence figures and narrative mechanism. Anti-patterns: treating report volume as prevalence, recruiting only reporters, asking about safety in the abstract, running research once for an audit, and letting a single account set policy.","content":"**Answer first:** Trust and safety research is the study of how users actually experience harm, reporting, and enforcement on your platform — as distinct from the moderation telemetry that only ever shows you the harm someone bothered to flag. It has four distinct objects: the harmed user's experience, the usability of your reporting and appeals flows, the comprehension of your policies, and the working experience of your reviewers. Under the EU Digital Services Act and the UK Online Safety Act it is no longer purely voluntary: risk assessments are expected to be informed by evidence about real user experience, and \"we had no complaints\" is explicitly not evidence. Async, anonymous AI-moderated interviews are unusually well suited to this work, because the people with the most to tell you are the ones least willing to say it to a stranger on a video call.\n\n## The blind spot in moderation telemetry\n\nEvery trust and safety function has a dashboard: reports received, actions taken, appeal rates, median time to decision, prevalence estimates from sampled review. It is genuinely useful, and it is systematically biased in one direction — it can only measure harm that entered the pipeline.\n\nWhat it cannot tell you:\n\n- How many users experienced something harmful and did not report it, and why not\n- Whether people could *find* the reporting flow, or gave up halfway through it\n- Whether the categories in your report form match the harms users actually experience, or force them into the nearest wrong box\n- What users believed would happen after reporting, versus what did\n- Whether your community guidelines are comprehensible to the people expected to follow them\n- Why some harmed users leave silently instead of escalating\n\nThose are all research questions, and none of them is answerable from log data. This is the same structural gap that makes [non-user research](/docs/non-user-research) necessary in a commercial context: the people who churned in silence are invisible to your own instrumentation.\n\nConsumer sentiment data suggests the gap is wide. Surveys of consumers have found roughly 64% believing companies do not remove enough harmful content while about 32% feel moderation has gone too far — a split that no single prevalence metric can reconcile, and one that only qualitative work can explain. The market has responded: content moderation spending is estimated at around USD 11.6 billion in 2025, forecast to reach roughly USD 13.3 billion in 2026 and USD 26 billion by 2031, with TELUS Digital's *Safety in Numbers* research reporting 58% of companies treating moderation as a higher priority than a year earlier. Spending is rising faster than understanding.\n\n## The four research objects\n\nMost teams conflate these, then wonder why one study cannot answer everything. Separate them.\n\n| Object | Core question | Typical method | Who you recruit |\n|---|---|---|---|\n| **Harmed users** | What happened, what did you do, what did you need? | Retrospective interviews, critical incident technique | Users who filed reports; users who did not but experienced harm |\n| **Reporting and appeals flows** | Can people find it, complete it, and understand the outcome? | Usability testing, walk-through interviews | Reporters, appellants, and never-reporters |\n| **Policy comprehension** | Do users understand the rules they are held to? | Comprehension testing on real policy text | General population of your user base |\n| **Reviewers and moderators** | What makes a decision hard, slow, or distressing? | Anonymous internal interviews | Your own T&S staff and vendor reviewers |\n\nThe [critical incident technique](/docs/critical-incident-technique) is the single best-fitting method for the first object: it anchors the participant to one specific, recent, memorable event rather than asking them to generalise about safety in the abstract — which is where these interviews usually go wrong.\n\n## The regulatory driver: this is now expected evidence\n\n**EU Digital Services Act.** Under Article 34, providers of very large online platforms and search engines must diligently identify, analyse and assess systemic risks arising from the design, functioning, and use of their service — at least annually, and in any event before deploying functionality likely to have a critical impact on identified risks. The enumerated categories cover illegal content, fundamental rights, civic discourse and electoral processes, gender-based violence, public health, and minors including behavioural addiction. These assessments are expected to be informed by research and by the views of affected individuals and civil society, not solely by internal telemetry. Article 35 then requires mitigation measures that are reasonable, proportionate and effective, and *tailored to the specific risks identified* — which is impossible to demonstrate if you never established how the risk manifests for real users.\n\n**UK Online Safety Act.** In-scope services were required to complete illegal-content risk assessments by 24 July 2025, with safety duties applying from 25 July 2025, and services likely to be accessed by children to complete children's risk assessments by 31 July 2025 following Ofcom's Protection of Children Codes published on 24 April 2025. Ofcom has signalled it expects providers of categorised (Category 1 and 2A) services to supply copies of their latest risk assessment records by October 2026, with the register of categorised services expected in July 2026. Services adopting measures other than those in the Codes must keep records demonstrating an equivalent standard — and user evidence is the most direct way to do that.\n\nThe through-line for both regimes: an assessment grounded in what users actually experience is defensible; one grounded in the absence of complaints is not. For the parallel AI-specific obligations, see [the EU AI Act and user research](/docs/eu-ai-act-user-research-compliance) and [AI governance frameworks for research](/docs/ai-governance-frameworks-research).\n\n## Designing the study: risk category to question type\n\nKoji's six [structured question types](/docs/structured-questions-guide) let one instrument carry both the countable figure your risk assessment needs and the mechanism your product team needs.\n\n| Research question | Question type | Why |\n|---|---|---|\n| Did you experience X in the last 90 days? | `yes_no` | Clean incidence denominator for the assessment |\n| What did you do next? | `single_choice` | Reported / blocked / left / did nothing — the drop-off shape |\n| Why did you not report it? | `multiple_choice` | Friction inventory: could not find it, did not think it would help, feared retaliation, felt embarrassed |\n| How confident are you that the platform would act? | `scale` | Trust-in-enforcement trend line, quarter over quarter |\n| Rank these harms by how much they affect your experience | `ranking` | Forced prioritisation — far more decision-useful than five separate severity scales |\n| Walk me through what happened and what you needed | `open_ended` | The mechanism, with AI follow-up probing |\n\nSet follow-up depth deliberately here. For sensitive items, one gentle follow-up is often the right ceiling; for flow-usability items, two or three probes get you to the actual point of abandonment.\n\n## The safeguarding protocol\n\nResearch on harm can harm. Treat this as non-negotiable, and write it into the study before you field it.\n\n1. **Never require a participant to recount an incident in detail.** Ask what they needed and what the platform did, and let them choose how much of the event to narrate. Make every sensitive question optional.\n2. **Signpost support.** Include relevant helplines and reporting routes on the interview landing page and in the completion screen, not only in the consent text.\n3. **Do not collect the harmful content itself.** You are researching the experience, not gathering evidence. Collecting illegal material creates a handling obligation you do not want and cannot discharge inside a research tool.\n4. **Minimise identifiability.** Run without identifying fields wherever the research question allows it, and never join safety-research responses to account records without explicit, specific consent.\n5. **Consent must be specific about use.** State plainly that responses inform safety policy and may be summarised in regulatory documentation. See [research consent form templates](/docs/research-consent-form-templates) and [interview recording consent laws](/docs/interview-recording-consent-laws).\n6. **Run a DPIA.** Data about experiences of harm is, in practice, special-category-adjacent and high risk under GDPR. See [DPIA for user research](/docs/dpia-user-research).\n7. **Apply stricter rules with minors.** Age-appropriate language, guardian consent where required, and no open-ended prompts that could elicit disclosure of ongoing abuse without a route to support. See [user research with children and teens](/docs/user-research-with-children-teens).\n\n## Why async AI interviews fit this work unusually well\n\nThis is one of the few research domains where an AI moderator is not merely faster but *methodologically preferable*:\n\n- **Disclosure without a witness.** Participants disclose stigmatised experiences more readily when no human is watching. A moderated video call maximises social pressure at exactly the wrong moment; an async text interview minimises it.\n- **The participant controls the pace.** They can pause, leave, come back. Nobody is waiting on the other end of a scheduled hour.\n- **It runs when harm is fresh.** Reports arrive at 3am on a Sunday. A link-based interview triggered by a report event catches the experience within hours, not after the next fieldwork cycle.\n- **Depth at population scale.** T&S questions need thousands of responses for a defensible incidence figure and deep narrative for the mechanism. AI moderation is the only approach that delivers both from one instrument.\n- **Consistency for the record.** Every participant gets the same core questions asked the same way, which matters when the output feeds a regulatory document — while probing still adapts to what each person says.\n- **No third-party recording tool.** Because interviews are async and link-based, there is no meeting recorder in the loop — one fewer processor handling accounts of harm.\n\nKoji's [quality scoring](/docs/understanding-quality-scores) also matters more here than elsewhere: safety studies attract both genuine distress and bad-faith responses, and a composite 1–5 score across relevance, depth and question coverage lets you separate substantive testimony from noise before anything reaches a policy decision.\n\n## Researching your reviewers\n\nThe most neglected object of the four. Moderators and reviewers hold the tacit knowledge of where your policy is ambiguous — the edge cases they escalate, the categories they distrust, the queues where the guidance contradicts itself. TELUS Digital's research indicates around 44% of organisations now run hybrid human-plus-automation models with about 21% remaining human-led, and reviewer wellbeing is a recognised structural pressure on the workforce.\n\nTreat this as internal research with strong anonymity guarantees: it must never read as performance evaluation. Ask which decisions were hardest and why, which policy lines they cannot apply consistently, what the automation gets wrong, and what support they actually need. Run it anonymously — see [anonymous employee research with AI interviews](/docs/anonymous-employee-research-ai-interviews) — and aggregate to queue level rather than individual level.\n\n## Testing whether anyone understands your rules\n\nCommunity guidelines are usually written by policy and legal teams, then never tested on the population expected to follow them. Enforcement against a rule nobody comprehends produces appeal volume, not compliance. Test the actual text: show a real policy excerpt, then ask participants to judge specific scenarios against it, using `single_choice` for their verdict and `open_ended` for their reasoning. Disagreement rates by scenario tell you exactly which clause is failing. The mechanics are the same as any comprehension study — see [content testing](/docs/content-testing-guide).\n\n## Anti-patterns\n\n- **Treating report volume as prevalence.** Report volume is a function of reporting-flow usability and trust in enforcement at least as much as of actual harm.\n- **Recruiting only reporters.** The never-reporters are the finding. Recruit them deliberately — see [researching hard-to-reach audiences](/docs/hard-to-reach-participants-research).\n- **Asking about \"safety\" in the abstract.** Anchor to specific recent incidents or you will collect opinions about the internet.\n- **Running it once, for the audit.** Both the DSA and the OSA contemplate ongoing assessment; annual-only research means your evidence is stale for eleven months of every year.\n- **Letting a single harrowing account set policy.** Pair narrative depth with incidence figures from the same instrument, and label severity separately from confidence.\n\n## Frequently asked questions\n\n**What is trust and safety research?**\nIt is user research aimed at how people experience harm, reporting, enforcement and policy on a platform — covering harmed users, the usability of reporting and appeals flows, comprehension of the rules, and the working experience of reviewers. It complements moderation telemetry, which can only measure harm that was reported.\n\n**Is trust and safety research legally required?**\nNot as a named activity, but it is effectively expected evidence. The DSA requires very large platforms to assess systemic risks at least annually informed by research and the views of affected individuals, and the UK Online Safety Act requires risk assessments that Ofcom can inspect, with non-Code measures needing records demonstrating an equivalent standard. Relying on an absence of complaints is not a defensible basis for either.\n\n**How do you research harm without re-traumatising participants?**\nMake sensitive questions optional, never require detailed recounting of the incident, focus on what the participant needed and what the platform did, signpost support resources on the landing and completion screens, and avoid collecting the harmful content itself. Async formats help because the participant controls the pace and can stop at any point.\n\n**Why use AI-moderated interviews for safety topics rather than human moderators?**\nBecause disclosure of stigmatised experience is higher when no human is present, the participant sets the pace, and the interview can run within hours of the incident at any time of day. AI moderation also delivers consistent core questioning across thousands of participants — which matters when the output feeds a regulatory document — while still probing each person's specific account.\n\n**Should we interview our own moderators?**\nYes. Reviewers hold the tacit knowledge of where policy is ambiguous and where automation fails, and they are the least-researched group in most T&S organisations. Run it anonymously, aggregate to queue level rather than individual level, and separate it explicitly from performance management.\n\n**How do we turn safety research into a risk assessment?**\nUse one instrument that produces both a countable incidence figure (via yes_no and single_choice items) and narrative mechanism (via open_ended items with AI probing), then map each finding to the risk categories your regime enumerates, label severity and confidence separately, and record the mitigation you chose and why. That combination is what makes an assessment auditable rather than assertive.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how each aggregates into a report\n- [Critical Incident Technique](/docs/critical-incident-technique) — the best-fitting method for interviewing about specific harmful events\n- [Anonymous Employee Research with AI Interviews](/docs/anonymous-employee-research-ai-interviews) — how to interview your own reviewers safely\n- [Content Testing](/docs/content-testing-guide) — testing whether users comprehend your community guidelines\n- [User Research with Children and Teens](/docs/user-research-with-children-teens) — additional safeguards for minors\n- [DPIA for User Research](/docs/dpia-user-research) — the assessment you need before fielding a harm study\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) — the parallel obligations for AI-moderated research","category":"Use Cases","lastModified":"2026-07-30T03:23:11.886253+00:00","metaTitle":"Trust and Safety Research: Studying Harm and Reporting Flows (2026)","metaDescription":"Moderation telemetry only shows harm that got reported. How to research the harm that did not — four research objects, DSA and Online Safety Act obligations, and a safeguarding protocol.","keywords":["trust and safety research","content moderation user research","online safety user research","DSA systemic risk assessment research","researching online harms","reporting flow usability","moderator research","community guidelines comprehension testing"],"aiSummary":"Trust and safety research studies how users experience harm, reporting, enforcement and policy — as distinct from moderation telemetry, which can only measure harm that was reported. It has four objects: harmed users, the reporting and appeals flows, policy comprehension, and reviewers themselves. Regulatory drivers: DSA Article 34 requires very large online platforms and search engines to assess systemic risks at least annually and before deploying functionality with critical impact, covering illegal content, fundamental rights, civic discourse and elections, gender-based violence, public health and minors including behavioural addiction, informed by research and the views of affected individuals; Article 35 requires mitigation tailored to identified risks. Under the UK Online Safety Act, illegal-content risk assessments were due 24 July 2025 with duties from 25 July 2025, children's risk assessments by 31 July 2025 after Ofcom's Protection of Children Codes of 24 April 2025; Ofcom expects Category 1 and 2A record copies by October 2026 and the categorised register in July 2026. The safeguarding protocol: never require detailed recounting, signpost support, do not collect the harmful content, minimise identifiability, obtain specific consent, run a DPIA, and apply stricter rules for minors. Async AI-moderated interviews suit this domain because disclosure is higher without a human witness, the participant controls pace, interviews can run within hours of an incident, and one instrument yields both incidence figures and narrative mechanism. Anti-patterns: treating report volume as prevalence, recruiting only reporters, asking about safety in the abstract, running research once for an audit, and letting a single account set policy.","aiPrerequisites":["A platform with user-generated content, messaging, or user-to-user interaction","A completed or in-progress risk assessment under the DSA, the UK Online Safety Act, or an internal safety framework"],"aiLearningOutcomes":["Separate the four objects of trust and safety research and scope studies accordingly","Explain what moderation telemetry structurally cannot measure","Map DSA Article 34 risk categories and UK Online Safety Act duties to research questions","Design an instrument that yields both incidence figures and narrative mechanism","Apply a seven-point safeguarding protocol when researching harm","Interview reviewers and test policy comprehension without creating performance risk"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}