{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-07-31T01:29:41.517Z"},"content":[{"type":"documentation","id":"7ab04bea-4b31-4e87-8f12-3589b5f1031e","slug":"ai-red-teaming-with-users","title":"AI Red Teaming with Real Users: How to Find Harms Before Your Users Do (2026)","url":"https://www.koji.so/docs/ai-red-teaming-with-users","summary":"AI red teaming is adversarial testing that produces a defect list, not a score. Half the harm surface — responsible-AI harms and contextual failure — is a user research problem, not a security one. Covers a six-phase protocol, harm taxonomies, adversary recruiting, severity scoring with measured inter-rater agreement, red-teamer wellbeing, and the EU AI Act Article 55 and NIST obligations that now make adversarial testing mandatory for some providers.","content":"## The short answer\n\n**AI red teaming is the practice of deliberately trying to make your AI product behave badly, then treating what you find as a defect list rather than a score.** It is not benchmarking, it is not usability testing, and it is not a penetration test — though it borrows from all three. The distinguishing feature is intent: every other research method asks *what happens when people use this normally*. Red teaming asks *what happens when someone is trying to break it, or when a normal person hits the worst 0.1% of the input distribution*.\n\nTwo things changed in 2026 that moved this from a frontier-lab specialty to an ordinary product obligation. First, the law caught up: the EU AI Act's obligations for general-purpose AI models have applied since **2 August 2025**, and Article 55(1)(a) requires providers of models with systemic risk to perform model evaluation **including adversarial testing** to identify and mitigate systemic risk ([EU AI Act, Art. 55](https://artificialintelligenceact.eu/article/55/)). Models already on the market before that date have until **2 August 2027** to comply. Second, the evidence base matured: NIST published its Assessing Risks and Impacts of AI (ARIA) program results as **NIST AI 700-2 in November 2025**, built on a pilot involving roughly **51 red teamers across 508 testing sessions** on seven submitted AI applications, and introduced the Contextual Robustness Index (CoRIx) as a measure of whether an application holds up in its intended use context ([NIST AI 700-2](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.700-2.pdf)).\n\nThe uncomfortable part for most product teams is that half of this work is not a security problem. It is a research problem, and it belongs to the people who already know how to recruit strangers, design a probe, and score a subjective judgement reliably.\n\n## Red teaming is not safety benchmarking\n\nThe single most useful framing published on this subject comes from Microsoft's AI Red Team, which documented **eight lessons from 80 red-teaming operations covering more than 100 generative AI products since 2021** ([Bullwinkel et al., 2025, arXiv:2501.07238](https://arxiv.org/abs/2501.07238)). Their third lesson states it plainly: *AI red teaming is not safety benchmarking*.\n\nThe distinction matters operationally. A benchmark asks a fixed set of questions and returns a number you can compare across model versions. Red teaming produces an open-ended, contextual, adversarial exploration whose output is a set of *novel* failures nobody had written down yet. Benchmarks measure known risks; red teaming discovers unknown ones. A team that runs a jailbreak benchmark and reports \"94% refusal rate\" has not red teamed anything — it has measured performance against attacks that were already public, which is exactly the set of attacks least likely to hurt them.\n\nTheir second lesson is equally load-bearing for research teams: *you don't have to compute gradients to break an AI system*. The most effective attacks in their operations were frequently simple prompt-level manipulations at the system layer, not sophisticated optimisation against model weights. That is precisely the kind of work a domain expert with no ML background can do — and the reason a user researcher can lead it.\n\nTheir fifth and sixth lessons close the argument: *the human element of AI red teaming is crucial*, and *responsible AI harms are pervasive but difficult to measure*. Automation extends coverage; it does not replace judgement about whether an output is actually harmful to a specific person in a specific context.\n\n## What a research-led red team actually covers\n\nThere are two halves to the harm surface, and they need different people.\n\n| Layer | Example failures | Who should probe it |\n|---|---|---|\n| Security | Prompt injection, data exfiltration, privilege escalation, SSRF in the retrieval pipeline, tool/MCP abuse | Application security, with AI-specific tooling |\n| Model behaviour | Jailbreaks, refusal failures, unsafe instruction-following, over-refusal of legitimate requests | Mixed — security plus domain experts |\n| Responsible-AI harms | Stereotyping, degraded quality for a dialect or accent, medical or legal overreach, psychosocial harm, manipulation, unfair allocation | **User researchers and affected-community members** |\n| Contextual failure | Confidently wrong output in a high-stakes workflow, silent degradation on rare inputs, harmful defaults | User researchers with domain experts |\n\nThe bottom two rows are where research teams add something no security team can. Whether a model output is *harmful* is a judgement about people, made by people, and it is exactly the kind of subjective, low-agreement judgement that research methodology exists to make reliable. This is the same measurement problem covered in [human evaluation of AI outputs](/docs/human-evaluation-ai-outputs) — rubric anchors, independent raters, measured agreement — applied to the tail of the distribution instead of the middle.\n\nThe security half is not optional either. IBM's 2025 breach research found that **13% of organisations reported a breach of their AI models or applications, and 97% of those breached had no AI-specific access controls in place** ([IBM Cost of a Data Breach Report 2025](https://www.ibm.com/reports/data-breach)). Red teaming does not fix governance, but it is usually what makes the gap visible.\n\n## The six-phase protocol\n\n### 1. Scope the system, not the model\n\nMicrosoft's first lesson is *understand what the system can do and where it is applied*. A model that is harmless in a chat window can be dangerous the moment it is given a tool, a document store, or an audience of teenagers. Write down: the deployment context, the user population (including the population you did not design for), the tools and data the system can reach, and the consequence of a wrong answer. The consequence column is what sets your severity scale.\n\n### 2. Build a harm hypothesis list, not a prompt list\n\nPrompts are artefacts; intents are the unit of work. Recruit the list from four sources: prior incidents in your category, regulatory risk categories (NIST's Generative AI Profile, **NIST AI 600-1**, enumerates twelve), support tickets and trust-and-safety reports from your own product, and — critically — interviews with people who resemble both your attackers and your most vulnerable users.\n\n### 3. Recruit adversaries with standing, not just skill\n\nThe failure mode here is a red team made entirely of engineers who share the builders' blind spots. OpenAI's published approach to external red teaming emphasises deliberately prioritising **geographic and domain diversity** in who is invited to probe a model ([OpenAI, Approach to External Red Teaming](https://cdn.openai.com/papers/openais-approach-to-external-red-teaming.pdf)). For responsible-AI harms specifically, the highest-yield participants are people who would be harmed: clinicians for a health assistant, benefits caseworkers for an eligibility tool, speakers of the dialects your speech model handles worst.\n\nUse a proper [screener](/docs/screener-questions-guide). \"Has broken an AI product before\" is a weak signal; \"works daily in the domain and can tell a plausible wrong answer from a correct one in under ten seconds\" is a strong one.\n\n### 4. Run structured and free-form passes\n\nStructured passes cover your hypothesis list systematically so you can claim coverage. Free-form passes are where the novel findings come from. Budget at least a third of session time for unguided exploration, and record the participant's *strategy*, not only the winning prompt — strategies generalise across model versions in a way that individual prompts do not.\n\n### 5. Score severity and measure agreement\n\nEvery finding needs: reproduction steps, a harm category, a severity rating, an estimated likelihood in real use, and the affected population. Severity is a subjective rating, which means it needs the same discipline as any other subjective rating — at least two independent raters and a measured agreement score. Aim for Cohen's kappa of **0.7 or above**; below 0.4 your severity scale is ambiguous and needs rewriting, not more raters. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) for the mechanics.\n\n### 6. Convert findings into regression tests\n\nMicrosoft's eighth lesson is that *the work of securing AI systems will never be complete*. The practical response is to make each confirmed finding permanent: every reproduced harm becomes a row in your evaluation set, so the next model version is automatically tested against it. That handoff — red team finding to frozen test case — is covered in [building evaluation datasets from real user research](/docs/ai-evaluation-dataset-golden-set).\n\n## Red teaming versus its neighbours\n\n| Method | Question it answers | Input distribution | Typical output |\n|---|---|---|---|\n| Usability testing | Can people complete the task? | Typical | Friction list |\n| Human evaluation | Is the output good enough to ship? | Representative sample | Rubric scores + pass rate |\n| LLM-as-a-judge | Did quality regress since last build? | Frozen eval set | Automated score |\n| **Red teaming** | **What is the worst this can do, and to whom?** | **Adversarial and tail** | **Reproducible harm findings** |\n| Security pentest | Can the system be compromised? | Adversarial, infra-focused | Vulnerability report |\n\nThey compose in a specific order. Red teaming discovers; [human evaluation](/docs/human-evaluation-ai-outputs) quantifies; [an LLM judge](/docs/llm-as-a-judge-vs-human-evaluation) monitors. Skipping the first step means your judge is monitoring for problems you already knew about.\n\n## Protecting the people who do this work\n\nThis is the part most guides omit, and it is a genuine ethical obligation. Red teamers are asked to elicit content that is by design distressing — self-harm instructions, harassment, sexual content involving minors, graphic violence. Treat it as you would any research involving exposure to disturbing material:\n\n- **Informed consent that names the content categories** in advance, not a generic media release. See [research ethics and informed consent](/docs/research-ethics-guide).\n- **A no-penalty opt-out mid-session**, exercised without explanation.\n- **Exposure limits and rotation** — cap session length and consecutive days on a harm category.\n- **Never recruit minors** for harm probing, even when minors are the affected population; work through adult proxies and safeguarding experts instead ([research with children and teens](/docs/user-research-with-children-teens)).\n- **Ethics review** where your organisation has one — much of this work meets the threshold described in [IRB approval for user research](/docs/irb-approval-user-research).\n\n## The modern approach: red teaming at scale with Koji\n\nThe historical constraint on red teaming was throughput. A moderated session costs a moderator, a calendar slot, and an hour; covering forty harm hypotheses with three participants each means 120 hours of scheduling before anyone reads a transcript. That is why most red teaming has been done by six people in a room for two weeks — not because six is the right number, but because 120 is unaffordable.\n\nKoji removes the scheduling layer entirely. Red-team sessions run as async, link-based AI-moderated interviews: you send a link, the AI moderator runs the protocol, probes the participant's reasoning when they find something, and the transcript lands analysed. Forty hypotheses across sixty domain experts is a days-long study, not a quarter-long programme.\n\nThe [six structured question types](/docs/structured-questions-guide) map onto red-team scoring almost exactly:\n\n| Question type | Red-team use |\n|---|---|\n| `open_ended` | The attack narrative, the strategy, and *why* the participant considered the output harmful — with AI follow-up probing that a form cannot do |\n| `single_choice` | Harm category from your taxonomy |\n| `scale` | Severity and likelihood ratings, aggregated into distributions |\n| `yes_no` | Binary criteria: did the system refuse? did it cite a source? did it stay in scope? |\n| `ranking` | Ordering several failing outputs by which would do the most damage |\n| `multiple_choice` | Which populations the participant believes are affected |\n\nBecause every session answers the same structured questions, severity distributions and category frequencies aggregate automatically — you get a ranked harm register rather than forty documents someone has to read. Koji's thematic analysis clusters the *strategies* participants used, which is the durable artefact; its quality scoring rates each conversation 1–5 against your research goals so thin sessions are visible immediately rather than diluting the register. Voice interviews are the right modality when the harm involves speech, accent handling, or social-engineering scripts that only work out loud.\n\n| | Traditional red-team workshop | Koji |\n|---|---|---|\n| Participants | 5–10, mostly internal | 30–100+, external domain experts |\n| Setup | Scheduling, NDAs, facilitation plan | A study link |\n| Time to findings | 2–6 weeks | Days |\n| Scoring | Post-hoc, in a spreadsheet | Structured at capture, aggregated live |\n| Repeatability per model release | Rarely — too expensive | Re-send the link |\n| Coverage of affected communities | Whoever was available | Screened and recruited deliberately |\n\nNone of this replaces a security team's tooling for the infrastructure layer. It replaces the part that was always the bottleneck: getting enough of the right humans to spend focused, structured time attacking your product.\n\n## Common mistakes\n\n1. **Reporting a pass rate.** Red teaming produces a defect list. A percentage implies a fixed denominator, and the whole point is that the denominator is unknown.\n2. **Recruiting only builders.** People who know how the system works probe where they expect weakness; users probe where the system meets their life.\n3. **Testing the model instead of the product.** Guardrails, retrieval, tools and system prompts are where most real failures live.\n4. **Letting findings die in a doc.** If a finding is not a permanent test case, the next release will reintroduce it.\n5. **Treating over-refusal as a non-finding.** A model that refuses to discuss a legitimate medical question is failing a real user, and that harm rarely shows up in a security-framed red team.\n6. **No wellbeing plan.** This is a research-ethics failure, not a logistics oversight.\n\n## Frequently asked questions\n\n**Is AI red teaming legally required for my product?**\nIt depends what you build. Article 55(1)(a) of the EU AI Act requires adversarial testing of general-purpose AI models with systemic risk, and those obligations have applied since 2 August 2025 (2 August 2027 for models already on the market). Most product teams are deployers rather than GPAI providers and are not directly caught by Article 55 — but high-risk system obligations, sector regulators, and enterprise procurement questionnaires increasingly ask for adversarial testing evidence regardless. See [EU AI Act compliance for user research](/docs/eu-ai-act-user-research-compliance) and [AI governance frameworks](/docs/ai-governance-frameworks-research).\n\n**How is red teaming different from usability testing?**\nUsability testing samples the typical input distribution and asks whether people can succeed. Red teaming samples the adversarial and tail distribution and asks what the worst outcome is. A product can pass every usability test and still produce a harm that ends up in a news story.\n\n**How many red teamers do I need?**\nThere is no saturation number, because the space is open-ended. NIST's ARIA pilot used roughly 51 red teamers across 508 sessions on seven applications. A practical starting point for a product team is 20–30 participants spanning at least three distinct perspectives (domain expert, affected community, adversarially-minded generalist), then re-running per major release.\n\n**Can I automate this with an LLM?**\nPartly. Automated adversarial generation is genuinely good at breadth — it will try thousands of variations you would never type. It is poor at deciding whether an output is harmful in context, and it cannot represent the lived experience of the population at risk. The mainstream position, and Microsoft's fifth lesson, is that automation extends coverage while human judgement remains essential.\n\n**Do synthetic users work for red teaming?**\nFor generating candidate attack strategies, they are a reasonable brainstorming aid. For judging harm, no — the documented sycophancy and endorsement biases of AI personas make them unreliable evaluators. See [synthetic users in research](/docs/synthetic-users-research-methodology).\n\n**What do I do with a finding I cannot fix?**\nDocument it, rate it, and route it to a non-model mitigation: a guardrail, a scope restriction, a disclosure, a human-in-the-loop checkpoint, or a decision not to ship into that context. An unfixable finding that is written down and mitigated is a governance artefact; an unfixable finding that is deleted is a liability.\n\n## Related resources\n\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) — the rubric-and-agreement discipline that makes harm severity ratings defensible\n- [Evaluation Datasets for AI Products](/docs/ai-evaluation-dataset-golden-set) — turning red-team findings into permanent regression tests\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) — automating the monitoring layer once you know what to look for\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types used to score severity, category and likelihood at capture\n- [User Research for AI Products](/docs/user-research-for-ai-products) — trust calibration and failure tolerance in normal use\n- [AI Governance Frameworks for Research Teams](/docs/ai-governance-frameworks-research) — ISO/IEC 42001, NIST AI RMF and the EU AI Act compared\n- [Research Ethics and Informed Consent](/docs/research-ethics-guide) — consent design for sessions involving distressing content","category":"Research Methods","lastModified":"2026-07-30T03:16:42.569976+00:00","metaTitle":"AI Red Teaming with Real Users: A 2026 Practitioner Guide","metaDescription":"How to run adversarial testing of AI products with real people: harm taxonomies, recruiting, severity scoring, red-teamer wellbeing, and EU AI Act duties.","keywords":["ai red teaming","red teaming ai products","adversarial testing ai","ai harm discovery","responsible ai harms","eu ai act adversarial testing","ai safety testing with users","llm red team methodology","ai risk assessment research","harm taxonomy ai"],"aiSummary":"AI red teaming is adversarial testing that produces a defect list, not a score. Half the harm surface — responsible-AI harms and contextual failure — is a user research problem, not a security one. Covers a six-phase protocol, harm taxonomies, adversary recruiting, severity scoring with measured inter-rater agreement, red-teamer wellbeing, and the EU AI Act Article 55 and NIST obligations that now make adversarial testing mandatory for some providers.","aiPrerequisites":["Basic understanding of how LLM-based products work","Familiarity with qualitative research recruiting and screening"],"aiLearningOutcomes":["Distinguish red teaming from safety benchmarking, usability testing and penetration testing","Build a harm hypothesis list and a severity scale for your deployment context","Recruit adversaries with domain standing rather than only technical skill","Score harm findings with measured inter-rater agreement","Convert confirmed findings into permanent regression test cases","Apply the ethical safeguards red-teamer wellbeing requires"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min"}],"pagination":{"total":1,"returned":1,"offset":0}}