{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-01T17:00:54.760Z"},"content":[{"type":"documentation","id":"95912610-7d0d-4d8f-89d9-341d4ebb4fa9","slug":"multilingual-research-guide","title":"Multi-Language User Research: How to Interview Participants in Any Language","url":"https://www.koji.so/docs/multilingual-research-guide","summary":"Koji supports AI voice and text interviews in 15+ languages including Spanish, French, German, Japanese, Hindi, Portuguese, and more. This guide covers setting up single-language studies, running multi-market research with separate links, using URL language parameters for embedded research, localizing research briefs, and synthesizing findings across language groups. Platforms like Koji make multilingual research as easy and affordable as single-language research.","content":"# Multi-Language User Research: How to Interview Participants in Any Language\n\nMost research teams run their user interviews in English — even when their customers are not. The result is systematic blind spots: non-English-speaking markets receive less research, their needs are underrepresented in product decisions, and entire customer segments become invisible in qualitative data.\n\nAI-powered interview platforms like Koji change the economics of multilingual research. Instead of hiring bilingual moderators, coordinating cross-timezone sessions, or outsourcing translation, teams can run voice and text interviews in 15+ languages with the same quality, depth, and automated analysis they get from English-language studies.\n\nThis guide covers how to set up multilingual studies in Koji, best practices for cross-language research, and how to synthesize findings across language groups.\n\n## Why Multilingual Research Matters\n\nThe business case for multilingual research is clear:\n\n- **51% of internet users prefer content in their native language** (W3Techs), and this preference extends strongly to research participation\n- Non-English speakers are consistently underrepresented in qualitative research, even at global companies with international customer bases\n- Critical cultural nuances — how users describe problems, what emotional language they use for pain points, how they frame success — get lost when participants are interviewed in a second language\n- GDPR and data localization requirements make European market research increasingly important for any team selling into the EU\n\nTraditional barriers to multilingual research are real: hiring bilingual moderators costs 2–4x standard rates, translation workflows add 1–2 weeks to research timelines, and most analysis tools struggle with non-English transcripts. Platforms like Koji eliminate all three barriers by running the entire interview workflow natively in the target language.\n\n## Languages Supported\n\nKoji's AI interviewer supports the following languages in voice mode:\n\n- English (US, UK, Australian)\n- Spanish (Latin American and Castilian)\n- French\n- German\n- Dutch\n- Japanese\n- Hindi\n- Portuguese (Brazilian and European)\n- Italian\n- Polish\n- Turkish\n- Korean\n- Mandarin Chinese\n- Arabic\n- Swedish\n\n**Text mode** supports all of the above plus dozens of additional languages through Google Gemini's multilingual capabilities — covering virtually all major world languages.\n\nWhen you set a language on your study, Koji automatically:\n1. Generates the AI interviewer's system prompt in that language\n2. Selects an appropriate native-speaker voice profile (voice mode)\n3. Configures speech-to-text transcription for that language\n4. Translates the interview UI, buttons, and participant-facing instructions\n\n## Setting Up a Multilingual Study\n\n### Option A: Single-Language Study\n\nIf all your participants speak the same language, create one study and set the language in **Customize → Interaction Mode → Default Language**.\n\nSteps:\n1. Create a new study in Koji\n2. Describe your research goal in the target language — Koji's AI will generate the brief in that language\n3. Go to **Customize → Interaction Mode**\n4. Set **Default Language** to your target language (e.g., Spanish, French, German)\n5. Update your landing page headline and description in the same language\n6. Run a test interview to verify the AI speaks naturally\n7. Launch and share your interview link with your target participants\n\n**Pro tip:** Write your research brief in the same language as the interview. The AI uses your brief as its source of truth. A brief written natively in Spanish produces more natural Spanish interview flow than an English brief that the system has to bridge across languages.\n\n### Option B: Multi-Market Study with Separate Links\n\nFor research spanning multiple language markets, the recommended approach is to create one study per language. This gives you:\n\n- Clean, separate data per language for market-by-market analysis\n- Language-specific landing pages with culturally appropriate copy\n- Separate transcript analysis per language group\n- Cleaner report generation per market with no cross-language noise\n\nSetup:\n1. Create your base study in your primary language\n2. Create a duplicate study for each additional language\n3. Set the language for each duplicate in the Customize tab\n4. Customize each landing page in the appropriate language\n5. Distribute language-specific interview links to the right audiences\n\nYou can import participants for each language study via CSV, and personalized links will take each participant directly to the correct language study. See [Importing Participants via CSV](/docs/importing-participants-csv) for the import workflow.\n\n### Option C: Language Parameter in the URL\n\nFor embed or in-product use cases where you know a user's locale, you can pass a language parameter in the interview URL to override the default:\n\n- `koji.so/i/your-study?lang=es` → Spanish interview\n- `koji.so/i/your-study?lang=fr` → French interview\n- `koji.so/i/your-study?lang=de` → German interview\n\nThis is particularly useful when embedding Koji in your product — you can detect the user's locale and pass the appropriate `?lang=` parameter so they always receive the interview in their language without needing separate studies.\n\n## Writing Research Briefs for Multilingual Studies\n\nThe research brief is the most important document in any Koji study. For multilingual work, a few additional considerations apply:\n\n**Localize your brief, do not just translate it.**\nDirect translation of English research questions often sounds unnatural in other languages. Idioms, phrasing, and conceptual framing vary significantly across languages and cultures. Where possible, use Koji's AI assistant to generate the brief natively in the target language, or have a native speaker review the translated brief before launch.\n\n**Be explicit about cultural context.**\nIf you are researching a behavior that differs across cultures — how different markets approach financial decisions, privacy expectations, or relationship-driven buying behavior — add that context to the Problem Context section of your brief.\n\n**Consider formality register.**\nSeveral languages have formal and informal registers: Japanese (keigo vs. casual), German (Sie vs. du), French (vous vs. tu), Korean (formal vs. informal speech levels). Specify the appropriate register in your brief. B2B research in Japanese typically requires formal keigo; consumer research for a younger demographic may be more casual.\n\n**Localize screening questions.**\nIntake form fields and screening questions should be in the local language. A French participant filling out an English screening form has a worse experience and may provide less accurate responses.\n\n## Analyzing Multilingual Results\n\n### Transcripts in Native Language\n\nAll transcripts in Koji are captured in the original interview language. If you run a Spanish study, you will see Spanish transcripts. AI analysis — quality scores, themes, sentiment, insights — is generated in the same language as the interview.\n\n### Cross-Language Synthesis\n\nFor research spanning multiple language markets, Koji's report generation can synthesize findings across studies. When you generate a report that draws on multiple studies, the AI identifies common themes across language groups and surfaces market-specific differences.\n\nFor more targeted cross-language comparison, use the AI Consultant in your Insights Dashboard:\n- \"What themes appeared in the German study but not the French study?\"\n- \"Are the same pain points appearing consistently across all markets?\"\n- \"How do Japanese participants describe this problem differently from US participants?\"\n\n### English Summaries from Non-English Research\n\nIf your primary working language is English but your participants are not, you can request that Koji generate report summaries and executive findings in English even when the underlying transcripts are in other languages. Configure this preference at the report generation stage.\n\nThis workflow is common for global research teams: run interviews natively in local languages for authentic responses, then synthesize and report in English for internal stakeholders.\n\n## Use Cases for Multilingual Research\n\n**International market expansion:**\nBefore entering a new market, run discovery interviews with potential customers in their native language. Understand their current alternatives, pain points, and buying criteria without the filter of a second language distorting your data.\n\n**Localization validation:**\nAfter translating your product or marketing materials, interview users in the local language to verify whether the translation resonates naturally or sounds awkward and foreign.\n\n**Global employee research:**\nFor organizations with international workforces, run employee experience research in employees' native languages. Non-English speakers are consistently underrepresented in HR, culture, and engagement research — with significant consequences for team health and retention.\n\n**Multilingual customer feedback programs:**\nBuild a continuous research pipeline that automatically interviews customers in their language after key touchpoints (onboarding, support interactions, renewals), then synthesizes findings across markets for product and strategy decisions.\n\n**Academic and institutional research:**\nFor research institutions studying cross-cultural phenomena, Koji enables standardized interview protocols to be deployed across multiple language markets simultaneously, with consistent methodology and comparable data quality.\n\n## Audio Quality Tips for Non-English Voice Interviews\n\nVoice interview quality varies somewhat by language. A few practical notes:\n\n- **Test each language before launch.** Run a self-test interview in the target language to verify naturalness of speech and appropriateness of the voice profile.\n- **Some accents and dialects may reduce transcription accuracy.** For highly regional accents, enabling text mode as a fallback is prudent.\n- **Encourage headphones.** This applies universally — headphones improve audio quality in any language.\n- **Consider time zones when monitoring.** If you are watching response rates in real time, remember that your participants across different markets may be active at very different hours.\n\n## Getting Started with Multilingual Research\n\nTo run your first non-English study in Koji:\n\n1. Create a new study and describe your research goal in the target language\n2. Set the language in **Customize → Interaction Mode → Default Language**\n3. Write your landing page headline and intake form in the target language\n4. Run a test interview to verify the AI sounds natural and covers your research topics\n5. Share your language-specific link with your target participants\n\nWith tools like Koji, multilingual research is no longer a specialized, resource-intensive effort reserved for large enterprise research teams. It is a standard capability available to any team that cares about understanding customers wherever they are in the world.\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [B2C User Research: How to Understand Consumer Behavior at Scale (2026)](/blog/b2c-user-research-guide-2026) — B2C user research is systematically underinvested at most consumer companies. While B2B teams run structured customer discovery as a matter \n- [Getting Started with Customer Research: A Beginner's Guide](/blog/getting-started-with-customer-research-a-beginner-s-guide) — A practical, step-by-step guide for Product Managers, UX Researchers, and Founders who want to start doing customer research today and build\n\n<!-- further-reading:blog -->\n","category":"guides","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Multi-Language User Research: How to Run Interviews in Any Language — Koji Docs","metaDescription":"How to configure Koji for multilingual user research. Covers supported languages, brief localization, single vs. multi-market study setup, cross-language synthesis, and best practices for international research.","keywords":["multilingual user research","international user interviews","multi-language research","user research in Spanish French German","global user research","non-English research"],"aiSummary":"Koji supports AI voice and text interviews in 15+ languages including Spanish, French, German, Japanese, Hindi, Portuguese, and more. This guide covers setting up single-language studies, running multi-market research with separate links, using URL language parameters for embedded research, localizing research briefs, and synthesizing findings across language groups. Platforms like Koji make multilingual research as easy and affordable as single-language research.","aiDifficulty":"intermediate","aiEstimatedTime":"12 minutes"},{"type":"documentation","id":"5ab65d0b-cc15-4e0b-8e9d-d2f51257614f","slug":"research-consent-form-templates","title":"Research Consent Form Templates: GDPR-Compliant Forms for Every Study","url":"https://www.koji.so/docs/research-consent-form-templates","summary":"Ready-to-use consent form templates for user research and AI-moderated interviews. Covers GDPR compliance, informed consent best practices, AI disclosure requirements, and how Koji collects consent automatically through intake forms before every interview.","content":"# Research Consent Form Templates: GDPR-Compliant Forms for Every Study\n\nEvery research study — whether it's a 5-minute product feedback session or a 60-minute in-depth interview — requires informed consent. Getting this right protects your participants, protects your organization, and produces data you can actually use.\n\nThis guide gives you copy-ready consent form templates, explains what each clause means, and shows how platforms like Koji automate consent collection so you never miss a signature before an interview begins.\n\n## What Is a Research Consent Form?\n\nA research consent form (also called an informed consent form or participant agreement) is a document that explains:\n\n- **What the study is about** — its purpose and how findings will be used\n- **What participation involves** — time commitment, activities, recording\n- **How data will be stored and used** — anonymization, retention period, sharing\n- **Participant rights** — the right to withdraw at any time without consequence\n- **Who to contact** — the researcher or organization responsible for the data\n\nInformed consent isn't just a legal checkbox. It's the ethical foundation of research. Participants who understand what they're agreeing to give better, more honest responses — because they trust you.\n\n## Why Consent Matters More Than Ever in 2026\n\nThree forces have made consent more important in recent years:\n\n**GDPR and global privacy law.** Under GDPR (and similar laws like CCPA, PIPEDA, and Brazil's LGPD), processing personal data without a lawful basis is a violation. For research, explicit informed consent is typically the appropriate lawful basis — especially when interviews are recorded or transcribed.\n\n**AI-powered research tools.** When an AI system conducts interviews, participants deserve to know that. Transparency about AI moderation is not just ethical — it's increasingly expected.\n\n**Rising participant sophistication.** People are more aware of their data rights. A clear, honest consent process builds trust and reduces dropout rates.\n\n## The Anatomy of a Research Consent Form\n\nEvery compliant consent form should include these seven elements:\n\n### 1. Study Title and Purpose\nA plain-language explanation of what you're researching — not the internal project name, but what you're trying to learn. Avoid jargon.\n\n*Example: \"We're researching how product teams discover and evaluate new research tools, so we can improve how Koji is explained and onboarded.\"*\n\n### 2. What Participation Involves\nDescribe the activities, time required, and format (text, voice, video, survey).\n\n*Example: \"You'll participate in a 15-20 minute AI-moderated voice or text interview. The conversation will be guided by an AI interviewer who may ask follow-up questions based on your answers.\"*\n\n### 3. Recording and Data Collection\nBe explicit about what data is collected: conversation transcripts, audio recordings, intake form responses, demographic data.\n\n*Example: \"This interview will be transcribed automatically. Audio recordings are deleted after transcription. Transcripts are stored securely and used only for research analysis.\"*\n\n### 4. How Data Will Be Used and Stored\nExplain anonymization, data retention, and who has access.\n\n*Example: \"Your responses will be anonymized before analysis. Individual responses will not be shared externally. Aggregated insights may be used in internal reports or published research.\"*\n\n### 5. Voluntary Participation and Right to Withdraw\nMake clear that participation is voluntary and withdrawal has no negative consequences.\n\n*Example: \"Your participation is entirely voluntary. You can stop at any time by closing the interview window. Partial responses will be deleted upon request.\"*\n\n### 6. Incentives (If Applicable)\nIf you're offering a gift card, raffle entry, or other incentive, state this clearly.\n\n*Example: \"As a thank-you for your time, you'll receive a €25 Amazon gift card within 5 business days of completing the interview.\"*\n\n### 7. Contact Information\nProvide a name and email for the research lead so participants can ask questions or withdraw consent after the fact.\n\n*Example: \"Questions? Contact [Name] at [email]. You can request deletion of your data at any time by emailing us.\"*\n\n## Ready-to-Use Consent Form Templates\n\n### Template 1: Standard User Research Consent Form\n\nSuitable for: product interviews, UX research, customer feedback sessions.\n\n---\n**Research Study: [Study Name]**\n\n**Conducted by:** [Your Name / Organization]\n\n**Purpose:** We are conducting research to [brief purpose, e.g., \"understand how teams make buying decisions for research tools\"]. Your insights will help us [outcome, e.g., \"improve our product and content\"].\n\n**What to expect:**\n- This session will take approximately [duration] minutes\n- [Format: e.g., \"An AI interviewer will guide the conversation via text or voice\"]\n- You may be asked follow-up questions based on your responses\n\n**Data and privacy:**\n- Your responses will be recorded and transcribed\n- All data is stored securely and accessed only by the research team\n- Your responses will be anonymized before any analysis or reporting\n- Data is retained for [retention period, e.g., \"12 months\"] and then securely deleted\n\n**Your rights:**\n- Participation is entirely voluntary\n- You may withdraw at any time without consequence\n- You may request deletion of your data by emailing [contact email]\n\n**Incentive:** [If applicable: \"You will receive [incentive] upon completing the study.\"]\n\n☐ I have read and understood the above information and agree to participate.\n\n**Name:** _______________\n**Date:** _______________\n\n---\n\n### Template 2: Short-Form Consent Statement (for Intake Pages)\n\nFor brief studies or when collecting consent via a checkbox on an interview landing page:\n\n---\n*By clicking \"Start Interview,\" you agree that your responses will be recorded and used for research purposes. All data is anonymized and stored securely. You can withdraw at any time. [Privacy Policy link]*\n\n---\n\n### Template 3: AI-Moderated Interview Consent (Full Disclosure)\n\nFor AI-powered interviews (like those run on Koji), best practice is to explicitly disclose AI moderation:\n\n---\n**[Study Name] — Participation Agreement**\n\nThis study uses AI-moderated interviews. An AI system (not a human) will conduct the conversation, ask follow-up questions, and generate a transcript.\n\n**What this means for you:**\n- No human is watching the conversation in real time\n- Your responses are analyzed by AI to identify themes and insights\n- The research team reviews anonymized summaries, not individual transcripts, by default\n\n**Your data:** Transcripts are stored on secure servers in [EU/US/etc.]. We do not sell or share individual responses. You may request data deletion at any time.\n\n**Consent:** By proceeding, you confirm that you are 18 or older and agree to participate in this AI-moderated research study.\n\n---\n\n### Template 4: GDPR-Specific Consent Form (EU Research)\n\nFor organizations subject to GDPR:\n\n---\n**Informed Consent and Data Processing Agreement**\n\n**Data Controller:** [Organization name and address]\n**Purpose of Processing:** Research and analysis to [stated purpose]\n**Legal Basis:** Explicit consent (Article 6(1)(a) GDPR)\n**Data Categories:** Conversational responses, transcript data, intake form responses\n**Recipients:** Research team members only. No third-party sharing without separate consent.\n**Retention:** [Period, e.g., \"Data retained for 12 months from collection date, then deleted\"]\n**International Transfers:** [If applicable — describe safeguards, e.g., Standard Contractual Clauses]\n\n**Your GDPR Rights:**\n- Right to access your data (Article 15)\n- Right to rectification (Article 16)\n- Right to erasure (Article 17)\n- Right to restrict processing (Article 18)\n- Right to data portability (Article 20)\n- Right to withdraw consent at any time (Article 7(3))\n- Right to lodge a complaint with your national supervisory authority\n\nTo exercise any of these rights, contact: [Data Protection contact email]\n\n☐ I give my explicit consent to participate in this research study and for my data to be processed as described above.\n\n---\n\n## How to Collect Consent Automatically with Koji\n\nManually distributing and tracking consent forms is one of the biggest administrative headaches in research. Koji's intake form system handles this automatically — no separate document, no manual tracking.\n\n**Here's how it works:**\n\n1. **Enable the intake form** in your study's interview settings\n2. **Add a consent checkbox field** — Koji supports checkboxes with custom label text, so you can paste your consent statement directly\n3. **Set the field as required** — participants cannot start the interview without checking the box\n4. **Consent is captured automatically** in the respondent record, timestamped and stored with the session data\n\nThis approach means consent is always collected before the interview begins — with a timestamp you can reference if questions ever arise. No chasing signatures, no separate forms.\n\nFor GDPR-compliant workflows, add your privacy policy link to the footer text of the intake page (configurable in the branding settings).\n\n## Consent in AI-Moderated Interviews: Special Considerations\n\nWhen an AI system conducts the interview, a few additional best practices apply:\n\n**Disclose AI moderation explicitly.** Don't bury this in fine print. Participants deserve to know that they're talking to an AI — and in most cases, this doesn't reduce participation rates. Honesty builds trust.\n\n**Clarify what \"recording\" means in AI context.** In AI interviews, there's no video recording — but there is a transcript. Make clear that the conversation is transcribed and analyzed by AI, and that the research team reviews aggregated insights.\n\n**Address voice data separately if using voice interviews.** Voice interviews generate audio data in addition to transcripts. Confirm whether audio is retained or deleted after transcription, and include this in your consent language.\n\n**Koji's voice interviews:** Audio processed during voice sessions is handled by the voice provider and is not retained by Koji after the session ends. Transcripts are retained per your data processing agreement.\n\n## Common Consent Mistakes to Avoid\n\n**Using consent as a legal shield rather than a trust builder.** Dense, legalistic consent forms increase drop-off rates and erode trust. Write for comprehension, not liability alone.\n\n**Forgetting to include withdrawal instructions.** GDPR and other regulations require you to explain how participants can withdraw consent and request data deletion. Always include a contact email.\n\n**Using blanket consent for future unrelated research.** Consent should be specific to the study. If you want to contact participants for future studies, get separate consent for that.\n\n**Not updating consent language when methods change.** If you switch from human-moderated to AI-moderated interviews, or add voice recording, update your consent form accordingly.\n\n**Collecting more data than you need.** Only collect data you'll actually use. Asking for name, company, phone number, and demographic data when you only need role and company size creates unnecessary privacy exposure.\n\n## A Quick Compliance Checklist\n\nBefore launching your next study, verify:\n\n- ☐ Purpose of research is clearly stated in plain language\n- ☐ Recording and transcription are disclosed\n- ☐ AI moderation is disclosed (if applicable)\n- ☐ Data storage, retention, and access are explained\n- ☐ Anonymization approach is described\n- ☐ Withdrawal process is included with contact information\n- ☐ Consent is captured before the interview begins\n- ☐ GDPR rights are listed (for EU participants)\n- ☐ Privacy policy link is accessible\n- ☐ Consent records are stored with timestamps\n\n## Related Resources\n\n- [Intake Forms and Consent](/docs/intake-forms-and-consent) — How to configure Koji's intake form settings\n- [Customizing Your Study](/docs/customizing-your-study) — Branding, landing pages, and intake customization\n- [Research Ethics Guide](/docs/research-ethics-guide) — Deeper dive into ethical research practices\n- [How to Find and Recruit Research Participants](/docs/finding-research-participants) — Building your participant pipeline\n- [Structured Questions Guide](/docs/structured-questions-guide) — Using Koji's 6 question types to collect structured data alongside consent\n- [Managing Research Participants](/docs/managing-research-participants) — Tracking participants and their status in Koji\n\n## Further reading on the blog\n\n- [Koji vs Microsoft Forms: AI-Powered Research vs Enterprise Form Builder (2026)](/blog/koji-vs-microsoft-forms-2026) — Microsoft Forms is free with every M365 subscription, which makes it the default for a lot of teams. But it's a form builder — not a researc\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n- [AI Agents for User Research in 2026: How Autonomous Research Is Reshaping Customer Insight](/blog/ai-agents-user-research-2026) — AI agents are taking over user research in 2026 — moderating interviews, synthesizing themes, and producing insight reports in hours. The fu\n\n<!-- further-reading:blog -->\n","category":"Participant Recruitment","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Research Consent Form Templates: GDPR-Compliant Forms for User Research","metaDescription":"Copy-ready consent form templates for user research, UX studies, and AI interviews. Covers GDPR compliance, informed consent best practices, and how Koji automates consent collection.","keywords":["research consent form template","user research consent form","informed consent template","GDPR user research","participant consent form","research participant agreement"],"aiSummary":"Ready-to-use consent form templates for user research and AI-moderated interviews. Covers GDPR compliance, informed consent best practices, AI disclosure requirements, and how Koji collects consent automatically through intake forms before every interview.","aiPrerequisites":["Understanding the basics of user research ethics"],"aiLearningOutcomes":["Write a GDPR-compliant consent form","Configure consent collection in Koji intake forms","Apply AI-specific consent best practices","Avoid the most common consent mistakes"],"aiDifficulty":"beginner","aiEstimatedTime":"10 minutes"},{"type":"documentation","id":"3f981bf3-4a5d-4624-81eb-ed0885b80538","slug":"b2b-user-research-guide","title":"B2B User Research: How to Interview Enterprise Customers Without the Scheduling Nightmare","url":"https://www.koji.so/docs/b2b-user-research-guide","summary":"B2B user research differs from consumer research because of limited access to enterprise users, multi-stakeholder complexity, constrained participant volume, and high scheduling overhead. AI-moderated interview platforms like Koji solve the scheduling problem by enabling asynchronous sessions — participants complete interviews on their own schedule without calendar coordination. Koji supports voice interviews, multilingual AI, intake form screening, and six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that help researchers capture both quantitative ratings and qualitative context from enterprise participants.","content":"## B2B User Research: How to Interview Enterprise Customers Without the Scheduling Nightmare\n\nB2B user research is harder than it should be. Enterprise customers are busy, access to them is often gated through account managers or customer success teams, scheduling a 45-minute interview can take two weeks of back-and-forth, and when a participant finally shows up, the conversation time is too short to go deep.\n\nTraditional research methods weren't designed for the realities of B2B. But a new generation of AI-native platforms like Koji is changing what's possible — making it feasible to run deep, qualitative B2B research at a scale that was impossible to staff five years ago.\n\n---\n\n## What Makes B2B User Research Different\n\n**1. Access is limited and gated.**\nIn consumer research, you can recruit from a panel of millions and schedule sessions in days. In B2B, your users are professionals — often senior ones — with packed calendars. Getting an enterprise buyer on a 45-minute call requires navigating procurement blockers, legal concerns, and the relationship dynamics between your sales and CS team and the customer.\n\n**2. Stakeholders are plural.**\nB2B products are rarely used by one person. The enterprise buyer, the daily user, the team lead, and the IT administrator may all have different experiences and needs. Good B2B research requires talking to multiple personas at the same account.\n\n**3. Context is complex.**\nEnterprise users embed your product in a workflow with dozens of other tools, internal processes, and legacy systems. Understanding their experience requires understanding their entire operational context — not just their experience with your product in isolation.\n\n**4. Volume is constrained.**\nConsumer researchers can talk to 100+ participants with modest effort. B2B researchers often struggle to hit 15–20 interviews, which limits qualitative generalization and statistical comparisons across segments.\n\n**5. Sensitivity is high.**\nEnterprise customers may be under NDA, reluctant to discuss competitors, or concerned about their own internal processes being shared externally. This calls for extra care in how questions are framed and how data is stored and distributed.\n\n---\n\n## The High-Value Questions in B2B Research\n\nThe best B2B research goes beyond \"what features do you need?\" It explores the organizational and behavioral context behind product usage.\n\n**Discovery and adoption questions:**\n- \"Walk me through how you first heard about and decided to adopt [product].\"\n- \"What were you using before, and what drove the switch?\"\n- \"Who else in your organization is involved in decisions about [product]?\"\n\n**Workflow and integration questions:**\n- \"How does [product] fit into your team's daily workflow?\"\n- \"What other tools does it connect to or replace?\"\n- \"What are the handoffs that involve [product]?\"\n\n**Pain and friction questions:**\n- \"When does [product] slow you down rather than speed you up?\"\n- \"Has your team built anything internally to work around a limitation in [product]?\"\n- \"What would need to be true for you to use [product] for [use case X] that you currently do elsewhere?\"\n\n**Decision and ROI questions:**\n- \"How do you measure whether [product] is delivering value?\"\n- \"What would you tell your manager if asked to justify the budget for [product]?\"\n- \"What would need to happen for you to consider switching to a competitor?\"\n\n**Structured check-in questions for quantitative anchoring:**\n- Scale (1–10): \"How central is [product] to your team's workflow, where 1 = barely use it and 10 = can't work without it?\"\n- Single choice: \"Which department uses [product] most heavily at your company?\"\n- Yes/No: \"Has your team considered replacing [product] in the past 12 months?\"\n\n---\n\n## The Scheduling Problem — and How AI Interviews Solve It\n\nThe biggest bottleneck in B2B research is scheduling. Even with motivated participants, getting 20 enterprise users on a 45-minute call takes weeks of coordination. AI-moderated research platforms like Koji eliminate this bottleneck by running interviews asynchronously.\n\n**How it works:**\n1. You create an interview study in Koji with your research questions\n2. You share a link with target participants — via email, Slack, or your account management team\n3. Participants complete the AI interview on their own schedule: 15 minutes when it's convenient, rather than 45 minutes on a specific day\n4. The AI follows up, probes for depth, and adapts to each participant's unique context\n5. You review transcripts and AI-generated reports — no note-taking, no scheduling coordination needed\n\nThis changes the economics of B2B research dramatically. Instead of 10–15 interviews over 4 weeks, teams using Koji regularly complete 40–80 interviews in the same timeframe — at a fraction of the cost.\n\n**Voice mode for on-the-go enterprise users:** Koji supports voice interviews that participants complete via phone or browser. For busy enterprise users who commute, travel, or work in field roles, a 10-minute voice interview is far more accessible than blocking time for a video call.\n\n**Multilingual support:** Enterprise research often spans global user bases. Koji's AI can conduct interviews in any language, making B2B research accessible to international teams without requiring a multilingual research staff or translation intermediaries.\n\n---\n\n## Structuring a B2B Research Study in Koji\n\nA well-structured B2B study balances open-ended qualitative depth with structured quantitative anchors.\n\n**Intake form (screening):**\n- Company size\n- Department / role\n- Tenure with the product\n- Primary use case\n\n**Opening questions (open-ended):**\n- \"How does your team primarily use [product]?\"\n- \"What challenges were you trying to solve when you adopted [product]?\"\n\n**Mid-study structured questions:**\n- Scale (1–10): \"How much of your team's workflow runs through [product]?\"\n- Single choice: \"Which best describes how your team's usage has changed over the past 6 months?\"\n- Yes/No: \"Has your team implemented integrations with [product] beyond the defaults?\"\n- Ranking: \"Rank the following capabilities by importance to your team's day-to-day use.\"\n\n**Deep-dive questions (open-ended):**\n- \"Tell me about a recent time [product] helped you accomplish something you couldn't have done otherwise.\"\n- \"Is there a type of workflow your team needs support for that [product] doesn't currently address?\"\n\n**Closing:**\n- \"Is there anything about your experience with [product] that would be helpful for us to know?\"\n\nKoji's AI will probe each response, uncovering organizational context, workarounds, and unmet needs automatically — giving you the kind of depth you'd normally only get from a skilled human moderator.\n\n---\n\n## Recruiting B2B Participants for Your Study\n\n**Customer Success team:** Your CS team is closest to enterprise customers and can warm participants up before a research invitation goes out. Coordinate to identify high-value research targets — don't cold-email important accounts without a CS introduction.\n\n**In-product prompts:** For B2B SaaS, in-product CTAs (\"Share your feedback — it takes 10 minutes\") reach daily users directly. Koji study links can be embedded within your product or sent through in-app messaging.\n\n**Slack and LinkedIn communities:** Many B2B users participate in professional communities. Posting research invitations with transparency about who you are and what you're studying can yield engaged, motivated participants.\n\n**Email outreach to your user database:** Segment by criteria relevant to your study (company size, feature usage, plan tier) and send targeted invitations. A 5–10% response rate is typical for research with a $25–75 incentive.\n\n**Research panels:** Platforms like Respondent.io specialize in B2B participant recruitment and can find matching professionals within 24–72 hours for specialized studies.\n\n---\n\n## Multi-Persona B2B Research: How to Cover All Stakeholders\n\nOne of the most underutilized strategies in B2B user research is running parallel studies for different personas at the same company type. Consider:\n\n- **The buyer** (VP, Director): Cares about ROI, vendor reliability, contract terms, and strategic fit\n- **The daily user** (Individual contributor): Cares about workflow fit, speed, and frustration points\n- **The team lead** (Manager): Cares about visibility, reporting, and team adoption\n- **The IT admin**: Cares about security, integrations, and maintenance overhead\n\nWith Koji, you can create four separate study briefs — each tuned for a different persona — and recruit participants into the right study via intake form segmentation. The AI interview adapts its probing to each participant's context without you having to moderate separately.\n\nThe insights from multi-persona B2B studies often reveal misalignments: what the buyer values most isn't always what the daily user values most. Surfacing these gaps is often the most valuable output of a B2B research program.\n\n---\n\n## Analyzing and Presenting B2B Research Findings\n\nB2B research data is rich but noisy — every enterprise customer has a slightly different context. To make sense of it:\n\n**Look for organizational patterns, not just individual ones.** Note which job titles, company sizes, or team structures correlate with specific pain points or needs.\n\n**Track structured question distributions.** When 60% of enterprise users rate workflow integration as 3/10 or below, that's a product prioritization signal even without reading every transcript.\n\n**Use AI-generated reports as a starting point.** Koji automatically synthesizes themes, clusters findings, and surfaces representative quotes. Start with the report, then dive into specific transcripts for additional color.\n\n**Build a research repository.** For ongoing B2B research programs, store insights in a shared tool (Notion, Confluence, or Koji's report archive) so findings accumulate over time and inform roadmap decisions quarter over quarter.\n\n**Present with specificity.** B2B research findings land better when grounded in real organizational context. Quotes from participants explaining their workflow or organizational constraints are far more persuasive to a product team than aggregated percentages.\n\n---\n\n## Related Resources\n\n- [How to Recruit B2B Participants for User Research](/docs/recruiting-b2b-participants) — Strategies for reaching hard-to-access enterprise users\n- [AI-Moderated Interviews: How Automated Research Works](/docs/ai-moderated-interviews) — Why AI moderation changes the economics of B2B research\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — How to use scale, choice, ranking, and yes/no questions to anchor B2B qualitative data\n- [Multi-Language User Research](/docs/multilingual-research-guide) — Running international B2B studies with Koji's multilingual AI\n- [How to Set Up AI Voice Interviews](/docs/setting-up-voice-interviews) — Voice mode for on-the-go enterprise participants\n- [Continuous Discovery: Weekly Customer Interviews Without Burning Out](/docs/continuous-discovery-user-research) — Building a sustainable B2B research cadence at scale\n\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [Best B2B Customer Research Tools in 2026: 10 Platforms Compared](/blog/best-b2b-customer-research-tools-2026) — B2B customer research is harder than B2C — smaller samples, busier participants, complex buying committees. We compared 10 platforms (Koji, \n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"B2B User Research: How to Interview Enterprise Customers | Koji Docs","metaDescription":"B2B user research guide: how to interview enterprise customers, handle scheduling constraints, recruit stakeholders, and run multi-persona studies with AI-moderated interviews.","keywords":["b2b user research","enterprise user research","b2b customer interviews","how to interview enterprise customers","b2b ux research","enterprise research methods","b2b qualitative research"],"aiSummary":"B2B user research differs from consumer research because of limited access to enterprise users, multi-stakeholder complexity, constrained participant volume, and high scheduling overhead. AI-moderated interview platforms like Koji solve the scheduling problem by enabling asynchronous sessions — participants complete interviews on their own schedule without calendar coordination. Koji supports voice interviews, multilingual AI, intake form screening, and six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that help researchers capture both quantitative ratings and qualitative context from enterprise participants.","aiPrerequisites":["Understanding of user research fundamentals","Access to enterprise customers or B2B user base"],"aiLearningOutcomes":["Understand what makes B2B user research uniquely challenging","Structure a B2B study brief that balances qualitative and quantitative questions","Use AI-moderated interviews to overcome B2B scheduling constraints","Run multi-persona B2B research across buyer, user, and admin roles"],"aiDifficulty":"intermediate","aiEstimatedTime":"16 minutes"},{"type":"documentation","id":"3117ab54-7f63-4c15-acb3-845ade3b7220","slug":"customer-retention-research","title":"Customer Retention Research: The Complete 2026 Playbook for Reducing Churn Before It Happens","url":"https://www.koji.so/docs/customer-retention-research","summary":"A practitioner's guide to customer retention research covering the four research streams (churn interviews, stay interviews, NPS follow-ups, continuous VoC), sample-size math, a 30-day starter plan, common mistakes, and how AI-moderated platforms like Koji compress 2–4 week retention research cycles into hours. Includes statistics from Bain/Reichheld, GrowSurf, CustomerGauge, and UXPA, plus an expert quote from Frederick Reichheld on loyalty.","content":"# Customer Retention Research: The Complete 2026 Playbook for Reducing Churn Before It Happens\n\n**Bottom line up front:** Customer retention research is the systematic study of why customers stay, why they leave, and what changes their loyalty over time. Done well, it lifts profits 25–95% per 5-point retention gain (Bain & Company, Reichheld) and cuts retention spend by 25% versus companies running ad-hoc feedback programs (CustomerGauge). The modern approach combines four research streams — churn interviews, stay interviews, NPS follow-ups, and continuous voice-of-customer interviews — and uses AI-moderated platforms like Koji to scale them across every customer segment, not just the ones a researcher has time to call.\n\nThis guide covers the full retention research stack: when to use each method, the exact questions to ask, sampling math, and how to turn findings into measurable churn reduction.\n\n---\n\n## Why customer retention research is the highest-ROI research you can run\n\nThree numbers explain why retention research deserves the largest slice of your research budget:\n\n- **A 5% increase in customer retention produces a 25–95% increase in profits.** This is the foundational finding from Frederick Reichheld at Bain & Company, replicated across SaaS, financial services, retail, and B2B services for three decades ([Bain & Company / HBR](https://hbr.org/1990/09/zero-defections-quality-comes-to-services)).\n- **Acquiring a new customer costs 5–25x more than retaining one.** Customer Acquisition Cost (CAC) in B2B SaaS now ranges from $750–$1,300, while Customer Retention Cost (CRC) averages $100–$500 ([GrowSurf, 2026 retention statistics](https://growsurf.com/statistics/customer-retention-statistics/)). CAC has surged 222% in the last five years — making retention research a defensive necessity, not a nice-to-have.\n- **85% of customer churn is preventable.** 73% of consumers name poor service or experience as the #1 reason they leave ([GrowSurf](https://growsurf.com/statistics/customer-retention-statistics/)). Yet most companies only find this out *after* the customer is gone.\n\nThe strategic implication: every dollar invested in understanding *why* customers stay or leave returns more than a dollar invested in acquiring new ones to replace them. Yet most product and growth teams over-index on acquisition research (concept tests, brand studies, ICP work) and under-invest in retention research — because retention listening has historically been slow, manual, and expensive.\n\n> \"Loyalty is, in many ways, a stamp of approval. It is the surest sign that a firm is delivering superior value.\" — Frederick Reichheld, founder of Bain's Loyalty practice and inventor of the Net Promoter Score\n\n---\n\n## The four research streams in a complete retention program\n\nA mature customer retention research practice runs all four of these continuously. Each answers a different question.\n\n### 1. Churn interviews (post-cancellation)\n**Question answered:** *Why did customers who left actually leave?*\n\nRun within 7–14 days of cancellation. The customer's reasons are still fresh, and the emotional charge is high enough to surface honest answers. Survey-only exit data lies — 25–40% of customers select \"too expensive\" because it's the socially acceptable answer, when the real reason is unmet expectations, poor onboarding, or a competitor switch.\n\nUse neutral interviewers (your churned customer is unlikely to be candid with their former CSM). Cover:\n- The *first* moment they considered leaving (the \"trigger event\")\n- The job they hired you to do and what you got wrong\n- Where they went and why\n- What would have changed their mind\n\nSee our deep dive on [churned customer interviews](/docs/churned-customer-interviews) for the full question battery.\n\n### 2. Stay interviews (active customers)\n**Question answered:** *Why do current customers stay, and what would make them leave?*\n\nThe most underused retention method. While churn interviews tell you why people left, stay interviews tell you what the next 20% of churners are about to do — *before* it happens. Run 8–12 per quarter across your top revenue tier, mid-tier, and at-risk accounts.\n\nStay interview questions surface latent dissatisfaction:\n- \"If we disappeared tomorrow, what would you actually use instead?\"\n- \"What were you doing before us, and what made you switch?\"\n- \"When was the last time you almost canceled?\"\n- \"What is the one thing that, if we changed it, would make you cancel?\"\n\n### 3. NPS detractor and passive follow-ups\n**Question answered:** *What's the specific service breakdown driving low scores?*\n\nThe score itself is noise. The *verbatim* follow-up is the signal. NPS at scale is only useful when paired with a structured interview workflow that interrogates every detractor and passive within 48 hours. See our [NPS follow-up interviews](/docs/nps-follow-up-interviews) playbook.\n\n### 4. Continuous voice-of-customer (VoC) interviews\n**Question answered:** *What is changing in our customers' world that will affect retention in 6–12 months?*\n\nRun weekly or bi-weekly with a rotating cohort. Discovers shifts in workflows, competitive moves, regulatory changes, and unmet jobs *before* they show up in churn data. Pair with the [continuous discovery handbook of weekly customer interviews](/docs/customer-research-30-day-habit) cadence.\n\n---\n\n## How many retention interviews do you need? (The math)\n\nA common mistake is running 5 churn interviews and declaring victory. Retention research has stricter sample requirements than discovery research because churn is heterogeneous — different segments leave for entirely different reasons.\n\nA defensible retention research plan:\n\n| Research stream | Cadence | Sample size per cycle | Why |\n|---|---|---|---|\n| Churn interviews | Weekly cohort | 10–15 per month, segmented by plan tier and tenure | You need 5+ per segment to identify recurring themes |\n| Stay interviews | Quarterly | 8–12 across top, mid, at-risk accounts | Catches early warning signals |\n| NPS detractor calls | Continuous | 100% of detractors, sampled passives | Service-recovery + insight |\n| Continuous VoC | Weekly | 1–2 interviews per week | Trend detection |\n\nFor a 500-customer SaaS company with a 15% annual churn rate, that's roughly **75 churn interviews + 32 stay interviews + 60+ NPS calls + 50 VoC sessions = ~217 conversations per year**. Few teams can run that volume manually — which is why AI-moderated platforms have become the default approach for sub-1,000-customer SaaS companies.\n\n---\n\n## How Koji automates the retention research stack\n\nTraditional retention listening requires a dedicated researcher, scheduling tools, recording infrastructure, transcription, manual coding, and a report-writing cycle. The average research cycle takes **2–4 weeks per study** — far too slow to act on churn signal before the next cohort leaves.\n\nKoji compresses this from weeks to hours:\n\n- **AI-moderated interviews.** Customers join a voice or web interview with a Koji AI moderator that asks open-ended questions, probes follow-ups in real time, and adapts to what the customer says. No researcher has to be on the call. Run 50 interviews in parallel.\n- **Structured + open-ended question mix.** Use Koji's six [structured question types](/docs/structured-questions-guide) — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — alongside the AI moderator's conversational probing. You get quantitative segmentation *and* qualitative depth in one interview.\n- **Automatic thematic analysis.** Themes, sentiment, and quality scores (1–5 scale) are extracted as interviews come in. By the time you've recruited the 30th churner, you already know the top three reasons they're leaving.\n- **Methodology presets.** Koji includes ready-to-launch templates for churn interviews, stay interviews, NPS follow-ups, and exit surveys — calibrated for retention research specifically.\n- **Real-time insights chat.** Ask the data questions in natural language: *\"Compare why Enterprise customers churned in Q1 vs. Q2.\"* No more 4-week reporting cycle.\n\nTeams using AI-assisted research tools report **60% faster time-to-insight** ([UXPA, 2025](https://uxpa.org/ux-research-in-2025-from-insights-to-action/)), and that compression is exactly what retention research needs — because every week you wait for findings, another cohort cancels for the same reason.\n\n> \"Companies that build a robust voice-of-customer program spend 25% less on customer retention than those who don't.\" — CustomerGauge B2B retention benchmarks ([source](https://customergauge.com/blog/voice-of-customer-methodologies))\n\n---\n\n## A 30-day retention research starter plan\n\nDay 1–5: Run 10 churn interviews on customers who cancelled in the last 30 days. Use Koji's churn template. Tag responses by plan tier and tenure.\n\nDay 6–10: Run 6 stay interviews — 2 power users, 2 mid-tier, 2 customers who downgraded but didn't cancel.\n\nDay 11–15: Wire NPS follow-up into your post-cancellation flow so every detractor automatically gets an interview invite. Target 100% coverage.\n\nDay 16–20: Start a weekly VoC interview cadence — 1 customer per week, rotating across your top three segments.\n\nDay 21–30: Synthesize. Identify the top 3 churn drivers (will be specific to your product). Build a prioritization tree using the [opportunity solution tree](/docs/opportunity-solution-tree) framework. Pass to product and CS for action.\n\nBy day 30 you will know more about your churn than 90% of SaaS teams — because most never run this research at all.\n\n---\n\n## Common mistakes that kill retention research\n\n1. **Asking the wrong people.** Surveying current customers about why other customers left produces speculation, not data. Talk to the actual churners.\n2. **Letting the CSM run the exit interview.** Power dynamic kills honesty. Use a neutral interviewer or an AI moderator.\n3. **Over-relying on NPS scores.** The score is a temperature reading; the conversation is the diagnosis.\n4. **Looking only at the cancellation reason field.** It's a forced-choice trap. Run a follow-up interview.\n5. **Treating retention research as a one-time project.** Churn is a moving target. Make it continuous.\n6. **Skipping stay interviews.** You only learn about problems from people who left — by then it's too late.\n\n---\n\n## Related Resources\n\n- [Churned Customer Interviews: The Complete Guide](/docs/churned-customer-interviews)\n- [Win-Back Customer Interviews](/docs/win-back-customer-interviews)\n- [NPS Follow-Up Interviews](/docs/nps-follow-up-interviews)\n- [Churn Survey Guide](/docs/churn-survey-guide)\n- [Structured Questions Guide](/docs/structured-questions-guide)\n- [How to Prioritize Customer Feedback](/docs/how-to-prioritize-customer-feedback)\n- [Employee Retention Research Guide](/docs/employee-retention-research-guide)\n- [Activating Research Insights](/docs/activating-research-insights)","category":"Research Methods","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Customer Retention Research: The 2026 Playbook to Cut Churn — Koji","metaDescription":"Run customer retention research that actually cuts churn. Combine churn interviews, stay interviews, NPS follow-ups, and continuous VoC with AI-moderated research that compresses 4-week cycles into hours.","keywords":["customer retention research","churn reduction","stay interviews","voice of customer","retention analytics","churn interviews","customer loyalty research","AI customer research","SaaS retention","customer feedback program"],"aiSummary":"A practitioner's guide to customer retention research covering the four research streams (churn interviews, stay interviews, NPS follow-ups, continuous VoC), sample-size math, a 30-day starter plan, common mistakes, and how AI-moderated platforms like Koji compress 2–4 week retention research cycles into hours. Includes statistics from Bain/Reichheld, GrowSurf, CustomerGauge, and UXPA, plus an expert quote from Frederick Reichheld on loyalty.","aiPrerequisites":["Basic understanding of SaaS or subscription business models","Familiarity with customer success workflows","Awareness of NPS or CSAT measurement"],"aiLearningOutcomes":["Understand the four research streams that make up a complete retention program","Choose between churn interviews, stay interviews, NPS follow-ups, and continuous VoC for each use case","Calculate the right sample size for retention research by segment","Avoid the six most common mistakes that kill retention research","Run a 30-day retention research starter plan from cold start","Use AI-moderated research to scale retention listening across every customer"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"b3008238-3b0c-4acb-8f17-257afc622dd9","slug":"brand-tracking-study-guide","title":"Brand Tracking Studies: How to Measure Brand Health Over Time (2026)","url":"https://www.koji.so/docs/brand-tracking-study-guide","summary":"A brand tracking study is a longitudinal research program that measures awareness, consideration, perception, and loyalty metrics at regular intervals to detect changes in brand health over time. Traditional quarterly trackers cost $80K-$300K per wave and arrive 4-8 weeks late; AI-native platforms collapse this to $8K-$20K with continuous always-on data and qualitative depth on every wave. Best practice: lock 6-12 KPIs, sample 400+ per wave for 5-point sensitivity, and pair quantitative metrics with conversational AI interviews for the why behind every metric movement.","content":"# Brand Tracking Studies: How to Measure Brand Health Over Time (2026)\n\n**A brand tracking study is a longitudinal research program that measures the same brand health metrics — awareness, consideration, perception, NPS, sentiment — at regular intervals so you can detect changes over time. Done right, brand tracking is your early warning system for marketing performance, competitive shifts, and category-level trends. Done wrong (the way most legacy trackers run), it is an expensive, slow, low-resolution survey that produces a deck nobody reads.**\n\nBrand tracking has been a foundational marketing research practice since the 1960s, but the methodology has barely evolved. A typical Fortune 500 brand tracker still costs $80K-$300K per quarter, takes 4-8 weeks per wave, and delivers a static PDF report that arrives after the marketing decisions it was meant to inform have already been made.\n\nAI-native research platforms are rewriting that playbook. This guide covers what a brand tracker should measure, how often to run it, how to design questions that detect real change instead of noise, and how to run a continuous brand tracking program at a fraction of traditional cost using AI-moderated interviews.\n\n---\n\n## What a Brand Tracking Study Measures\n\nA robust brand tracking program measures four layers of brand health, each with its own set of questions and KPIs.\n\n### 1. Awareness\n\n- **Unaided awareness:** \"When you think of [category], what brands come to mind?\" — captured open-ended, then coded\n- **Aided awareness:** \"Which of these brands have you heard of?\" with a list including yours and competitors\n- **Top-of-mind awareness (TOMA):** Percentage who name your brand first in unaided recall\n\nAwareness is the floor of the funnel. Without it, nothing else matters.\n\n### 2. Consideration & Funnel\n\n- **Familiarity:** Self-rated on a 5-point scale\n- **Consideration:** \"Would you consider [Brand] for [need/category]?\"\n- **Preference:** \"If you had to choose one, which brand would you pick?\"\n- **Purchase intent:** Forward-looking buying behavior\n\nThe funnel — Awareness → Familiarity → Consideration → Preference → Purchase — reveals where prospects are getting stuck.\n\n### 3. Brand Perception\n\n- **Brand attribute associations:** \"Which of these brands do you associate with [innovative / trustworthy / fast / expensive…]?\" — usually a matrix with 8-15 attributes across 3-5 brands\n- **Brand personality:** Open-ended descriptions or forced-choice between archetypes\n- **Brand promise alignment:** Does your audience associate the right benefits with you?\n\nThis is where brand investment shows up — or does not.\n\n### 4. Loyalty & Advocacy\n\n- **NPS** (Net Promoter Score)\n- **Repurchase intent**\n- **Word-of-mouth behavior:** \"Have you recommended [Brand] in the past 6 months?\"\n- **Switching intent:** Risk indicator for churn\n\nThese metrics signal whether the funnel converts to long-term value. For a deeper guide on NPS specifically, see our [NPS survey guide](/docs/nps-survey-guide).\n\n---\n\n## How Often to Run a Brand Tracker\n\nThe right cadence depends on your category dynamics and your budget. Most teams over-invest in expensive quarterly waves and under-invest in always-on signal.\n\n| Cadence | Best for | Cost (legacy) | Cost (AI-native) |\n|---|---|---|---|\n| **Annual** | Mature B2B, low ad spend | $40K-$80K | $2K-$5K |\n| **Semi-annual** | Established consumer brands | $80K-$160K | $4K-$10K |\n| **Quarterly** | Growth-stage SaaS, retail | $160K-$320K | $8K-$20K |\n| **Monthly / always-on** | High-velocity consumer, performance marketing | $400K+ | $15K-$40K |\n\nContinuous brand tracking — collecting a small sample every week or month — produces sharper signal than quarterly waves because trend lines are based on more data points and shorter detection windows. The challenge has always been cost. AI-native platforms have collapsed that.\n\n---\n\n## Designing a Brand Tracker That Detects Real Change\n\nThe biggest failure mode of brand trackers is waves that look identical for years. Usually this means the questions are too high-level to detect movement, or the sample is too small to find statistically significant change.\n\n### Sample size\n\nFor a single brand:\n- **n=200 per wave** is enough to detect 8-point shifts in metrics in the 30-70% range\n- **n=400 per wave** detects 5-point shifts\n- **n=800+ per wave** detects 3-point shifts and supports segment-level analysis\n\nFor competitive comparison, you typically need 200+ respondents per brand.\n\n### Question wave consistency\n\nOnce you commit to a tracker, never change the question wording. Even minor edits (\"brand X\" vs \"X brand\") break the time series. Pretest your battery exhaustively at the start, then lock it.\n\n### Sub-group cuts\n\nThe aggregate trend hides everything interesting. Plan your design so you can cut by:\n\n- Audience segment (current customer / lapsed / never used)\n- Demographic (age, region, role)\n- Awareness state (knows brand / does not)\n\nThis is where having individual-level data — not just toplines — pays off.\n\n### Open-ended verbatims\n\nNumbers tell you *what* changed. Verbatim responses tell you *why*. Every wave should include 2-3 open-ended questions (\"What is the first word that comes to mind when you think of [Brand]?\") plus AI-moderated probing on a sample of respondents.\n\nThis is where AI-native platforms have a structural advantage. Traditional trackers either skip open-ends (because coding is expensive) or include them but only deliver a word cloud months later. Koji AI runs open-ended conversations at every wave, codes them automatically using its [thematic analysis](/docs/thematic-analysis-guide) capabilities, and surfaces theme shifts wave-over-wave with no manual coding step.\n\n---\n\n## How AI-Moderated Interviews Replace the Legacy Tracker\n\nThe traditional brand tracker is a static survey. The modern brand tracker is a continuous AI-moderated interview program. The data you get is different — and more useful.\n\n**Traditional tracker output:**\n- 95% aided awareness ↑ 2 points\n- NPS 42 ↓ 3 points\n- \"Innovative\" attribute association 38% ↓ 1 point\n- 60-page deck, distributed 6 weeks after fieldwork\n\n**AI-native tracker output:**\n- Same KPIs, plus:\n- Top 12 themes shifting wave-over-wave\n- Verbatim quote evidence for every metric movement\n- Segment-level breakdown of *why* NPS dropped\n- Real-time dashboard, no deck delay\n\nKoji [insights chat](/docs/insights-chat-guide) lets brand managers query the data in plain English: \"What is driving the consideration drop in the 25-34 segment?\" returns an answer with quote evidence, sourced from actual customer conversations, in seconds.\n\n---\n\n## Setting Up Your First Brand Tracker\n\n### Step 1 — Define the brand health KPI tree\n\nPick 6-12 metrics that map to business decisions you actually make. Do not track 40 metrics — you will never act on them, and the noise drowns out signal.\n\nA good starter tree:\n- Unaided awareness\n- Aided awareness\n- Familiarity\n- Consideration\n- Preference vs top 2 competitors\n- 4-6 brand attribute associations\n- NPS\n- Brand promise statement (open-ended)\n\n### Step 2 — Lock the question wording\n\nPretest the full battery with 30-50 pilot respondents. Confirm every question is interpreted as intended. Make every wording edit *before* wave 1 — once locked, do not change.\n\n### Step 3 — Define the audience\n\nFor most B2B trackers: target buyers and influencers in your category. For B2C: nat-rep within category-relevant demographics. Document the screener carefully — small audience drift between waves looks like brand movement.\n\n### Step 4 — Choose a cadence and stick to it\n\nQuarterly is the standard. Consider monthly or always-on if your category moves fast or your marketing spend justifies tighter measurement.\n\n### Step 5 — Build the dashboard, not the deck\n\nThe deliverable should be a dashboard your marketing leadership reviews monthly — not a PDF buried in a shared drive. The metrics that matter are the deltas, not the levels.\n\n### Step 6 — Layer qualitative depth\n\nEvery wave, run 20-50 conversational AI interviews on top of the survey to capture the *why* behind any moving metric. This is where AI-native platforms unlock 10x value over legacy trackers.\n\n---\n\n## What a Brand Tracker Cannot Do\n\nBrand tracking is a measurement system, not a research method for discovery. Use it to detect change — not to explain it without supporting research.\n\nFor deep understanding of perception shifts, layer in:\n\n- [Brand research interviews](/docs/brand-research-interviews) when a metric moves significantly\n- [Customer journey mapping](/docs/customer-journey-mapping) when consideration drops\n- [Win-loss analysis](/docs/win-loss-analysis) when preference vs competitor erodes\n- [Switch interviews](/docs/switch-interviews-jtbd-method) when retention drops alongside NPS\n\nThe tracker tells you something changed. Generative research tells you why.\n\n---\n\n## Common Brand Tracker Pitfalls\n\n1. **Changing wording mid-program.** Breaks the time series. Pretest exhaustively at the start, then lock.\n2. **Sample drift.** If your screener accidentally shifts demographics between waves, \"brand movement\" is actually sample movement.\n3. **Tracking too many metrics.** 40 KPIs = nobody acts on any of them. 6-12 is the sweet spot.\n4. **Running waves too far apart.** Annual trackers detect catastrophic shifts; they miss campaigns.\n5. **Skipping the qualitative layer.** Without verbatim and conversational depth, you have a number with no narrative.\n6. **Over-investing in fieldwork; under-investing in dissemination.** A perfect tracker that nobody reads is wasted.\n\n---\n\n## The Cost Argument for AI-Native Brand Tracking\n\nA quarterly legacy tracker costs $160K-$320K annually for a single brand. The same coverage with a continuous AI-moderated program runs $20K-$40K — and produces sharper signal because trend detection is based on weekly data points, not quarterly averages.\n\nFor most companies under $100M ARR, traditional brand tracking has been priced out of reach entirely. AI-native platforms like Koji bring it within the marketing budget of growth-stage SaaS, DTC brands, and series-B startups for the first time.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — How Koji six question types support both quantitative tracking and qualitative depth in a single conversation\n- [Brand Research Interviews](/docs/brand-research-interviews) — Deep qualitative brand research to complement your tracker\n- [Brand Perception Survey Guide](/docs/brand-perception-survey-guide) — Survey templates for measuring brand perception\n- [NPS Survey Guide](/docs/nps-survey-guide) — How to build the loyalty metric inside your tracker\n- [Longitudinal Research Guide](/docs/longitudinal-research-guide) — Methods for any kind of repeated-measures research over time\n- [Voice of Customer Research Program](/docs/voice-of-customer-research-program) — How to build a continuous voice-of-customer program that complements brand tracking\n- [Insights Chat](/docs/insights-chat-guide) — Query your tracker data in plain English with AI\n\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [B2C User Research: How to Understand Consumer Behavior at Scale (2026)](/blog/b2c-user-research-guide-2026) — B2C user research is systematically underinvested at most consumer companies. While B2B teams run structured customer discovery as a matter \n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Brand Tracking Studies: The Complete 2026 Guide","metaDescription":"How to design a brand tracking study that detects real change — and how AI interviews make continuous brand tracking affordable for the first time.","keywords":["brand tracking study","brand tracker","brand health metrics","brand awareness measurement","brand perception tracking","brand equity research","continuous brand tracking","always-on brand tracking","brand KPI dashboard"],"aiSummary":"A brand tracking study is a longitudinal research program that measures awareness, consideration, perception, and loyalty metrics at regular intervals to detect changes in brand health over time. Traditional quarterly trackers cost $80K-$300K per wave and arrive 4-8 weeks late; AI-native platforms collapse this to $8K-$20K with continuous always-on data and qualitative depth on every wave. Best practice: lock 6-12 KPIs, sample 400+ per wave for 5-point sensitivity, and pair quantitative metrics with conversational AI interviews for the why behind every metric movement.","aiPrerequisites":["Familiarity with brand metrics (awareness, consideration, NPS)","Understanding of longitudinal/repeated-measures research"],"aiLearningOutcomes":["Identify the four layers of brand health a tracker should measure","Choose the right cadence (annual to always-on) for your category","Calculate sample size needed to detect statistically meaningful change","Design a brand tracker that combines quantitative KPIs with qualitative depth","Replace expensive legacy trackers with continuous AI-moderated programs"],"aiDifficulty":"intermediate","aiEstimatedTime":"17 minutes"},{"type":"documentation","id":"ac99592c-f785-45b7-9862-d2dbc167f315","slug":"win-loss-analysis-guide","title":"Win/Loss Analysis: How to Learn Why You Win and Lose Deals","url":"https://www.koji.so/docs/win-loss-analysis-guide","summary":"A comprehensive guide to win/loss analysis covering why CRM data is insufficient (85% inaccurate), how to design and conduct buyer interviews within 2-4 weeks of deal close, question frameworks across six categories, common mistakes to avoid, and how AI-powered platforms like Koji enable more honest responses and faster analysis. Includes B2B win rate benchmarks by segment and ROI data from Gartner and Clozd research.","content":"# Win/Loss Analysis: How to Learn Why You Win and Lose Deals\n\nWin/loss analysis is the practice of systematically interviewing buyers after a deal closes — whether you won or lost — to understand the real reasons behind their decision. Done well, it is one of the highest-ROI research activities a B2B company can run.\n\nMost sales teams think they know why they win and lose. The data says otherwise. Research from Clozd found that **buyer and seller reasons for lost deals only align 15% of the time** — meaning 85% of the \"lost reason\" data sitting in your CRM is inaccurate. Reps guess, rationalize, or pick the least embarrassing option. The truth lives with the buyer, and the only way to get it is to ask.\n\n## Why Win/Loss Analysis Matters\n\nThe business case is hard to ignore. Gartner research found that companies implementing comprehensive win/loss analysis see **up to a 50% improvement in win rates and a 15–30% increase in revenue**. Clozd's 2025 State of Win-Loss Report found that 63% of companies report win-rate increases from win/loss programs, jumping to 84% for programs running longer than two years.\n\nAs Todd Berkowitz, VP Analyst at Gartner, put it:\n\n> \"A formal and rigorous win/loss analysis program enables better segmentation, product strategy choices and sales enablement. Those that take a more comprehensive approach have seen a 15% to 30% increase in revenue and up to 50% improvement in win rates.\"\n\nDespite this, fewer than 20% of companies conduct formal win/loss interviews. The gap between knowing you should do it and actually doing it is where most teams stall.\n\n## What You Actually Learn\n\nWin/loss interviews reveal insights you cannot get any other way:\n\n- **Why you really lost.** Not the polite version your champion shared with the rep — the real competitive, pricing, product, or trust gaps that killed the deal.\n- **Why you really won.** Your actual differentiators, as buyers perceive them, not as your marketing team assumes them.\n- **How buyers evaluate vendors.** The criteria, internal politics, and emotional dynamics that shape purchase decisions.\n- **Where your sales process helps or hurts.** Specific moments where the buying experience built or eroded trust.\n- **Competitive positioning gaps.** How prospects perceive you relative to alternatives — in their own words.\n\n## How to Run Win/Loss Interviews\n\n### Timing\n\nInterview within **2–4 weeks** of the decision. After 30 days, memory degrades significantly. After 90 days, buyers struggle to recall relevant details, and recruitment becomes difficult. Won deals are especially time-sensitive — implementation experiences quickly overshadow evaluation memories.\n\n### Who to Interview\n\nTarget the **economic buyer or primary decision-maker** who drove vendor selection. Avoid interviewing only champions or influencers. Gather an equal number of won and lost interviews to avoid skewed feedback.\n\n### Who Should Conduct the Interview\n\nNever the sales rep who handled the deal. Buyers give polite, relationship-protecting answers to the people who sold them. Companies that partner with a third party or use neutral interviewers are **over 2x more likely to be satisfied** with feedback quality. Buyers share up to **40% more critical feedback** with neutral or AI-powered interviewers than with human researchers from the selling organization.\n\nThis is where AI-powered research platforms like Koji change the game. Koji conducts structured win/loss interviews at scale — without the scheduling overhead, without interviewer bias, and with consistent question delivery across every conversation. Buyers can respond on their own time, in their own words, without the social pressure of a live call with someone from the company that just sold to (or lost) them.\n\n### Question Design\n\nThe best win/loss interviews follow a hybrid approach: a structured core of consistent questions (for trend data) combined with open-ended follow-ups that adapt to what the buyer reveals (for depth).\n\n**Core question categories:**\n\n1. **Decision context** — What triggered the search? What problem were you solving?\n2. **Evaluation process** — Who was involved? What criteria mattered most? How did you shortlist?\n3. **Competitive comparison** — How did vendors compare on key criteria? What stood out?\n4. **Sales experience** — How was the sales process? What helped or hindered?\n5. **Decision drivers** — What ultimately tipped the decision? What was the #1 factor?\n6. **Perception** — How did you perceive each vendor's strengths and weaknesses?\n\nStart broad and open-ended: \"Walk me through your evaluation process from the beginning.\" Use \"what\" and \"why\" questions rather than yes/no — they introduce less bias and keep conversations flowing. Use the Five Whys technique to dig past surface-level answers.\n\n### Getting Honest Answers\n\nHonesty is the central challenge of win/loss research. Buyers default to diplomatic responses, especially on losses.\n\n- **Use a neutral interviewer.** Buyers are dramatically more candid with someone who has no stake in the outcome.\n- **Normalize critique.** Frame it: \"Many buyers we talk to share concerns about X. Did you experience anything similar?\"\n- **Embrace silence.** After a buyer answers, pause. They often reveal their true thoughts after a moment of reflection.\n- **Probe emotional factors.** B2B buyers make emotional decisions as much as rational ones. Ask about personal risk, career impact, and trust dynamics.\n\nKoji's AI interviewer is particularly effective here. It maintains a warm, neutral tone without the subtle cues — facial expressions, tone shifts, defensive body language — that make buyers self-censor in live interviews. The result is richer, more honest data.\n\n### How Many Interviews You Need\n\nStart with **5–8 interviews per month** (roughly 2 per week). Wait for 10–15 interviews before making structural changes. When a factor appears in 40%+ of losses but less than 10% of wins, that is a systemic signal, not anecdotal.\n\nEqual representation of wins and losses is essential for valid comparisons.\n\n## Common Mistakes\n\n1. **Relying on CRM \"Lost Reason\" fields.** This is the #1 mistake. The data is 85% inaccurate.\n2. **Having the sales rep conduct the interview.** Buyers protect the relationship rather than revealing truth.\n3. **Only interviewing losses.** Win interviews tell you what is working. Without them, you fixate on weaknesses and neglect your actual differentiators.\n4. **Interviewing too late.** Beyond 30 days, memory reconstruction omits critical details.\n5. **Running a project instead of a program.** One-time studies produce a point-in-time snapshot. Ongoing programs (85% positive ROI) vastly outperform project-based approaches (55% positive ROI).\n6. **Poor distribution of findings.** Insights die in a spreadsheet if marketing runs the program and never tells sales, product, or leadership.\n\n## Analyzing and Acting on Findings\n\nTag every interview against consistent themes: pricing, product gaps, sales experience, competitive strengths, decision process. Track theme frequency across wins vs. losses. Segment by deal size, vertical, competitor, and buyer persona.\n\nLook for **asymmetric patterns**: factors that appear frequently in losses but rarely in wins (or vice versa). These are your highest-leverage improvement opportunities.\n\nDistribute findings to every team that touches revenue:\n- **Sales enablement** — battle cards, objection handling, competitive positioning\n- **Product** — feature gaps, integration needs, usability issues\n- **Marketing** — messaging, positioning, content gaps in the buyer journey\n- **Leadership** — strategic pricing, market positioning, competitive strategy\n\n## Win/Loss Analysis with Koji\n\nTraditional win/loss programs require scheduling calls, hiring third-party firms, and waiting weeks for reports. Koji collapses this cycle.\n\n**How it works:**\n\n1. **Start from a template.** Koji's GTM Win/Loss Analysis template gives you a research-backed interview plan designed specifically for post-decision buyer interviews.\n2. **The AI Consultant refines your brief.** Tell it about your product, market, and what you want to learn. It builds a tailored interview plan with the right question types, probing depth, and guardrails.\n3. **Share the link with buyers.** No scheduling, no calendar coordination. Buyers complete the interview on their own time — voice or text.\n4. **AI conducts the interview.** Koji's AI interviewer follows your structured plan while adapting follow-up questions in real time based on what the buyer reveals. It probes deeper on competitive mentions, emotional decision factors, and sales experience gaps.\n5. **Get analysis immediately.** Koji synthesizes themes, extracts quotes, and surfaces patterns across interviews — no manual coding required.\n\nThe result: more interviews, more honest responses, faster time-to-insight, and a fraction of the cost of traditional programs.\n\n## Win Rate Benchmarks\n\nFor context, here is where typical B2B win rates land:\n\n| Segment | Typical Win Rate | Top Performers |\n|---|---|---|\n| SMB | 30–40% | 45%+ |\n| Mid-Market | 25–35% | 40%+ |\n| Enterprise | 20–25% | 30%+ |\n| Large deals (>$100K) | 15–25% | — |\n\nThe overall average B2B win rate is roughly 20–21% — meaning 4 out of 5 qualified opportunities are lost or end in no-decision. Even a 10% relative improvement (e.g., moving from 20% to 22%) on a $10M pipeline yields an incremental $1M in revenue.\n\n## Key Takeaways\n\n- **CRM data is not win/loss analysis.** The reasons in your CRM are wrong 85% of the time.\n- **Interview within 2–4 weeks.** Memory degrades fast.\n- **Use a neutral interviewer.** AI-powered platforms like Koji remove bias and social pressure.\n- **Interview both wins and losses.** You need both sides to see the full picture.\n- **Make it a program, not a project.** Ongoing programs deliver 85% positive ROI vs. 55% for one-time studies.\n- **Distribute findings broadly.** Win/loss insights are only valuable if they reach the teams that can act on them.\n\n---\n\n## Related Resources\n\n- [Competitive Intelligence Guide](/docs/competitive-intelligence-survey-guide) — Broader competitive research\n- [Customer Discovery Interviews](/docs/customer-discovery-interviews) — Pre-deal customer research\n- [Churn Survey Guide](/docs/churn-survey-guide) — Post-churn analysis\n- [Koji for Product Managers](/docs/koji-for-product-managers) — Product team workflows\n- [Post-Demo Feedback Guide](/docs/post-demo-feedback-survey-guide) — Demo-stage feedback\n\n*Explore [structured questions](/docs/structured-questions-guide) for combining win/loss scales with AI-powered buyer interviews.*\n\n## Further reading on the blog\n\n- [50+ Win/Loss Interview Questions That Reveal Why You Really Win and Lose Deals (2026)](/blog/win-loss-interview-questions-2026) — Your sales team has one version of why you lost that deal. Your buyer has a completely different one. Here are 50+ win/loss interview questi\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Win/Loss Analysis Guide: How to Learn Why You Win and Lose Deals — Koji","metaDescription":"Learn how to run win/loss analysis interviews that improve win rates, sharpen positioning, and give your team real competitive intelligence. Includes question frameworks, benchmarks, and best practices.","keywords":["win loss analysis","win loss interviews","competitive intelligence","sales win rate","B2B sales research","deal analysis","buyer interviews","post-decision interviews","competitive analysis","sales enablement"],"aiSummary":"A comprehensive guide to win/loss analysis covering why CRM data is insufficient (85% inaccurate), how to design and conduct buyer interviews within 2-4 weeks of deal close, question frameworks across six categories, common mistakes to avoid, and how AI-powered platforms like Koji enable more honest responses and faster analysis. Includes B2B win rate benchmarks by segment and ROI data from Gartner and Clozd research.","aiPrerequisites":["Basic understanding of B2B sales processes","Familiarity with CRM deal tracking"],"aiLearningOutcomes":["Understand why CRM lost-reason data is unreliable and how interviews fix it","Design a structured win/loss interview with six core question categories","Know when to interview, who to target, and how to get honest answers","Analyze and distribute findings across sales, product, and marketing teams","Set up an ongoing win/loss program using Koji templates"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"7a3c5017-9b32-4b71-997b-4970dbc4200e","slug":"employee-retention-research-guide","title":"Employee Retention Research: Stay Interviews, Exit Interviews, and What Actually Works","url":"https://www.koji.so/docs/employee-retention-research-guide","summary":"A comprehensive guide to employee retention research covering the trillion-dollar turnover problem, why 42% of turnover is preventable, stay interview methodology using Finnegan's 5-question SHRM framework, exit interview best practices, engagement pulse research, predictive attrition indicators, and how AI platforms like Koji enable honest feedback at scale. Includes statistics from Gallup, McKinsey, and SHRM, plus expert quotes from Dick Finnegan and McKinsey researchers.","content":"# Employee Retention Research: Stay Interviews, Exit Interviews, and What Actually Works\n\nVoluntary turnover costs U.S. businesses an estimated **$1 trillion per year** (Gallup). Replacing a single employee costs between 50% and 200% of their annual salary — and for senior or specialized roles, that figure can reach 400%. Yet **42% of employees who voluntarily leave** say their manager or organization could have done something to prevent it (Gallup, 2024).\n\nThe problem is not that retention is impossible. The problem is that most organizations do not systematically listen to the people who are still there — and by the time they conduct an exit interview, it is too late to change anything.\n\nThis guide covers how to use qualitative research to understand what keeps employees and what drives them away, before they hand in their notice.\n\n## The Listening Gap\n\nGallup's 2024 research found that **45% of employees who left** reported that no manager or leader proactively discussed their job satisfaction, performance, or future with the organization in the **three months before their departure**. Not one conversation.\n\nMeanwhile, global employee engagement fell from 23% to 21% in 2024 — only the second decline in 12 years — costing the world economy an estimated $438 billion in lost productivity. **70% of team engagement** is attributable to the manager, and low-engagement teams experience turnover rates 18% to 43% higher than highly engaged teams.\n\nAs Dick Finnegan, the world's leading authority on stay interviews, put it:\n\n> \"Everything I had been taught about turnover in HR was wrong. Most employees don't leave because of the actions of the CEO — they leave because of the actions of their direct supervisor.\"\n\nRetention is not an HR problem. It is a management problem. And it requires a management-level research practice to fix.\n\n## Stay Interviews: The Most Underused Retention Tool\n\nA stay interview is a structured, one-on-one conversation with a current employee designed to understand what keeps them at the organization and what might cause them to leave. Only **17.2% of employers** conduct stay interviews — yet companies that regularly seek employee feedback see a **14.9% reduction in turnover rates**.\n\n### When to Conduct Stay Interviews\n\n- **Every 6–12 months** for established employees, separate from performance review cycles\n- **At 90 days and 180 days** for new hires\n- **Quarterly** for employees in transition roles or after organizational changes\n- Never combine stay interviews with performance reviews — the power dynamic corrupts the data\n\n### Who to Prioritize\n\nStart with the people whose departure would hurt most:\n- **High performers and high-potential employees** — highest cost of loss\n- **New hires in their first year** — highest flight risk\n- **Employees in critical or hard-to-fill roles**\n\nThen expand organization-wide.\n\n### The Core Questions\n\nDick Finnegan's SHRM-validated 5-question framework:\n\n1. \"What do you look forward to when you come to work each day?\"\n2. \"What are you learning here?\"\n3. \"Why do you stay here?\"\n4. \"When was the last time you thought about leaving, and what prompted it?\"\n5. \"What can I do as your manager to make your experience at work better?\"\n\nQuestion 4 is the most diagnostically powerful. It surfaces real risk factors that employees rarely volunteer unprompted. The key is strong follow-up probing — \"What prompted it?\" and \"What happened next?\" — to reach specific, actionable detail.\n\n### Getting Honest Answers\n\nThe central challenge: employees fear that honesty will damage their standing. Research shows employees systematically underreport dissatisfaction with management and culture, instead citing safe answers like \"better opportunity\" or \"pay.\"\n\n**How to get past the filter:**\n- **Normalize the conversation.** Frame it as something every employee participates in, not a signal that something is wrong.\n- **Separate from performance reviews entirely.** If employees associate the conversation with evaluation, they will self-censor.\n- **Send questions in advance.** Give employees a week to reflect before the conversation.\n- **Start with positive questions.** \"What do you look forward to?\" builds rapport before the harder questions.\n- **Use a neutral interviewer for sensitive topics.** When exploring management quality or organizational culture, third-party or AI interviewers get dramatically more honest responses.\n\n## Exit Interviews: Getting Truth from Departing Employees\n\nExit interviews capture why employees left. But standard paper-based exit surveys achieve participation rates as low as **15–30%**, and even when employees participate, they systematically underreport the real reasons for leaving.\n\n### Timing\n\nConduct exit interviews during the **notice period but not the actual last day**. Too early and the employee is still in \"professional mode.\" Too late and they are mentally checked out.\n\nConsider a **follow-up survey 3–6 months post-departure**. With distance and perspective, former employees are often more candid about the real reasons they left.\n\n### The Honesty Problem\n\nDeparting employees protect relationships. They say \"I got a better offer\" when they mean \"My manager never once asked about my career goals.\" They say \"It was time for a change\" when they mean \"I watched three colleagues get promoted past me.\"\n\n**How to get real answers:**\n- **The interviewer must not be the employee's direct manager.** Ever.\n- **Third-party interviewers yield the most honest feedback** but require budget.\n- **AI-powered interviews remove the social pressure entirely.** Employees can be candid without worrying about burning bridges, giving a bad reference, or making the conversation awkward.\n- **Guarantee anonymity and demonstrate it through actions** — aggregate findings, never attribute specific feedback to individuals in reports.\n\n### Core Exit Interview Questions\n\n1. \"What prompted you to start looking for a new role?\"\n2. \"What could we have done differently to keep you?\"\n3. \"How would you describe your relationship with your direct manager?\"\n4. \"Did you feel you had opportunities for growth and development here?\"\n5. \"What would you tell our leadership team if you knew it would be taken seriously?\"\n6. \"On a scale of 1 to 10, how likely are you to recommend this organization as a place to work?\"\n\n## Engagement Pulse Research\n\nStay and exit interviews capture individual depth. Engagement pulse research captures organizational patterns.\n\nResearch identifies the strongest predictive indicators of employee attrition as: **job satisfaction** (single most influential factor), **years with current manager**, **years in current role**, and **years since last promotion**. Machine learning models analyzing engagement data can achieve 77.5% accuracy in predicting attrition.\n\nThe practical implication: if you collect structured engagement data regularly, you can identify retention risks before they become resignations.\n\n**What to measure:**\n- Manager relationship quality\n- Growth and development opportunities\n- Sense of purpose and meaning in work\n- Workload sustainability\n- Recognition frequency\n- Trust in leadership\n\n## Common Mistakes in Retention Research\n\n1. **Only doing exit interviews.** By then, the person has already left. Stay interviews are preventive; exit interviews are diagnostic.\n2. **Having the direct manager conduct stay interviews without training.** Untrained managers turn stay interviews into performance check-ins or defensive conversations.\n3. **Treating it as an HR initiative instead of a management practice.** Retention research works when managers own the conversations and act on them.\n4. **Asking but never acting.** Nothing destroys trust faster than soliciting feedback and visibly ignoring it.\n5. **Relying on annual engagement surveys alone.** Annual surveys measure sentiment at one point in time. Ongoing qualitative research captures the \"why\" behind the numbers.\n\nMcKinsey's Great Attrition research framed it clearly:\n\n> \"People keep quitting at record levels, yet companies are still trying to attract and retain them the same old ways. Organizations can treat this as the 'Great Attrition' or reframe it as the 'Great Attraction' by genuinely listening and adapting.\"\n\n## Retention Research with Koji\n\nTraditional stay and exit interviews are bottlenecked by the same constraint: someone has to schedule, conduct, and analyze every conversation. For a 500-person organization running annual stay interviews, that is 500 conversations — an impossible workload for most HR teams.\n\nKoji makes this scalable.\n\n**How it works:**\n\n1. **Start from a template.** Koji offers HR-specific templates for Stay Interviews, Exit Interviews, New Hire Check-ins, Manager Effectiveness research, and Employee Engagement studies — each built on validated frameworks like Finnegan's 5-question model.\n2. **Customize with the AI Consultant.** Tell it about your organization, your retention challenges, and what you want to learn. It builds a tailored interview plan with the right question types — open-ended for exploration, scale questions for measuring satisfaction, and yes/no questions for benchmarking.\n3. **Distribute the link.** Send it to employees via email, Slack, or your HRIS. No scheduling required. Employees complete the interview on their own time, voice or text.\n4. **AI conducts the interview.** Koji's AI interviewer follows the structured plan while adapting follow-ups in real time. When an employee mentions frustration with their manager, the AI probes deeper. When someone rates their growth opportunities as a 3 out of 10, it asks what would make it a 7. No human interviewer can maintain this consistency across 500 conversations.\n5. **Get actionable patterns.** Koji synthesizes themes across all interviews — surfacing the top retention risks, the strongest engagement drivers, and the specific actions that would make the biggest difference. No manual coding, no spreadsheet analysis.\n\nThe result: the depth of 1:1 stay interviews at the scale of an annual survey, with the honesty of a neutral third-party interviewer.\n\n## Predictive Retention: From Reactive to Proactive\n\nThe most advanced retention practices use qualitative data to predict attrition before it happens. Research shows that engagement data collected 6–12 months prior reliably surfaces turnover patterns across job roles, tenure levels, and business units.\n\nBy running ongoing stay interview programs through Koji, organizations build a living dataset of retention signals. Over time, patterns emerge: which teams have the highest flight risk, which manager behaviors correlate with retention, which career development gaps surface repeatedly.\n\nThis transforms retention from a reactive scramble (\"Why did our best engineer quit?\") into a proactive practice (\"Our engineering team's growth satisfaction scores dropped 20% this quarter — let's investigate before we lose anyone\").\n\n## Key Takeaways\n\n- **42% of turnover is preventable** — but only if someone asks before the employee decides to leave.\n- **Stay interviews are the highest-ROI retention practice.** Only 17% of employers use them.\n- **Exit interviews alone are insufficient.** Departing employees underreport real reasons. Follow up 3–6 months later for honest feedback.\n- **The manager is the #1 factor.** 70% of engagement variance is attributable to the direct manager.\n- **Neutral interviewers get more honest answers.** AI-powered platforms like Koji remove the social pressure that causes employees to self-censor.\n- **Make it ongoing, not annual.** Continuous qualitative research catches retention risks before they become resignations.\n\n---\n\n## Related Resources\n\n- [Exit Interview Guide](/docs/exit-interview-survey-guide) — Exit interview framework\n- [Stay Interview Guide](/docs/stay-interview-survey-guide) — Proactive retention\n- [Employee Engagement Guide](/docs/employee-engagement-survey-guide) — Engagement measurement\n- [Pulse Survey Guide](/docs/pulse-survey-guide) — Ongoing engagement tracking\n- [Employee Wellness Guide](/docs/employee-wellness-survey-guide) — Wellbeing research\n\n*Explore [structured questions](/docs/structured-questions-guide) for combining retention scales with AI-powered employee interviews.*\n\n## Further reading on the blog\n\n- [Why \"Price\" Is Never the Real Reason Customers Churn (And How AI Interviews Prove It)](/blog/why-price-is-never-the-real-churn-reason) — Exit surveys say 40% of churn is about price. AI interviews reveal that \"price\" is actually code for 5 different problems. Here is how to fi\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Employee Retention Research: Stay Interviews, Exit Interviews, and What Actually Works — Koji","metaDescription":"Learn how to use stay interviews, exit interviews, and engagement research to understand why employees leave and what makes them stay. Includes frameworks, question design, and how AI scales retention conversations.","keywords":["employee retention","stay interviews","exit interviews","employee engagement","turnover prevention","retention research","HR research","employee satisfaction","manager effectiveness","engagement surveys"],"aiSummary":"A comprehensive guide to employee retention research covering the trillion-dollar turnover problem, why 42% of turnover is preventable, stay interview methodology using Finnegan's 5-question SHRM framework, exit interview best practices, engagement pulse research, predictive attrition indicators, and how AI platforms like Koji enable honest feedback at scale. Includes statistics from Gallup, McKinsey, and SHRM, plus expert quotes from Dick Finnegan and McKinsey researchers.","aiPrerequisites":["Basic understanding of HR processes","Familiarity with employee lifecycle management"],"aiLearningOutcomes":["Understand why traditional exit interviews fail to capture real departure reasons","Design and conduct stay interviews using the Finnegan 5-question framework","Know when to conduct stay interviews vs. exit interviews vs. engagement pulses","Get more honest employee feedback using neutral and AI-powered interviewers","Build an ongoing retention research practice that predicts attrition before it happens","Use Koji HR templates to run stay and exit interviews at organizational scale"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"52c0afdb-319e-45e3-84f8-3b17d1e9be67","slug":"voice-of-employee-program","title":"Voice of the Employee: How to Build an Employee Listening Program That Drives Real Change","url":"https://www.koji.so/docs/voice-of-employee-program","summary":"A comprehensive guide to building a Voice of the Employee (VoE) program. Covers the limitations of annual surveys, 5 listening channels (pulse surveys, stay interviews, exit research, in-depth interviews, manager conversations), a 5-step implementation framework, and how AI-moderated employee interviews reduce social desirability bias and scale listening without headcount.","content":"# Voice of the Employee: How to Build an Employee Listening Program That Drives Real Change\n\n**The core reality:** In 2024, Gallup reported that employee engagement in the United States dropped to 31% — a 10-year low. Despite 92% of employees saying they are fully committed to their company's mission, 79% report experiencing burnout. The gap between stated commitment and actual experience is the problem that Voice of the Employee programs exist to close — and most organizations are failing to close it.\n\nA Voice of the Employee (VoE) program is a systematic approach to continuously collecting, analyzing, and acting on employee feedback across the full lifecycle of employment. It is not an annual engagement survey. It is not an HR compliance exercise. Done well, it is a competitive advantage: organizations that employees feel heard by are 4.6 times more likely to see employees deliver their best work, and companies with strong employee listening cultures are 12 times more likely to achieve high retention.\n\nThis guide explains what a VoE program actually requires, why most organizations fall short, and how to build one that produces genuine change rather than data that gets filed and forgotten.\n\n## What Is a Voice of the Employee Program?\n\nA Voice of the Employee program is the organizational infrastructure for hearing, understanding, and acting on what employees think, feel, and experience. It encompasses:\n\n- **Listening channels:** The mechanisms through which feedback is collected — surveys, interviews, focus groups, stay conversations, exit interviews, and digital sentiment monitoring\n- **Analysis systems:** The processes that transform raw feedback into patterns, themes, and actionable signals\n- **Action loops:** The organizational commitments and workflows that ensure feedback translates into visible change\n- **Trust infrastructure:** The psychological safety, confidentiality norms, and communication practices that make employees willing to share honestly\n\nMost organizations have some of these elements. Very few have all of them working together — and the missing piece is almost always the last one: the action loop that closes the feedback cycle and demonstrates that listening actually changes anything.\n\n## Why Annual Engagement Surveys Are Not Enough\n\nThe engagement survey is the default VoE tool in most organizations. HR sends one annually, results come back three months later, leadership reviews a dashboard of aggregate scores, and a few action items are logged into a project plan that may or may not get tracked.\n\nThis model fails for four structural reasons:\n\n**1. The lag problem.** By the time annual survey results are analyzed and shared, the experiences that generated the data are months old. The manager who was causing a team's disengagement may have already left. The process frustration that drove survey scores down may have been fixed already — or may have gotten significantly worse.\n\n**2. The aggregation problem.** Aggregate engagement scores obscure more than they reveal. A company-wide engagement score of 67% tells you almost nothing actionable. The signal is in the sub-group analysis — the specific teams, tenures, roles, and locations where engagement is dramatically different from the average. Most annual survey processes produce dashboards designed to be presented, not dashboards designed to drive investigation.\n\n**3. The action problem.** Research consistently shows that employees who give feedback and see no change become *more* disengaged than employees who never gave feedback at all. Annual surveys, with their slow feedback loops, frequently trigger exactly this dynamic: employees invest time and emotional energy in honest responses, then watch nothing change, then disengage more deeply.\n\n**4. The social desirability problem.** Annual surveys — even anonymous ones — often suffer from response bias. Employees who are most disengaged are least likely to participate. Employees who are concerned about anonymity (justified or not) self-censor. The result is data from the employees most willing to respond, which skews toward the moderately engaged middle, missing both the highly enthusiastic and the dangerously disengaged.\n\n> *\"The annual survey is a snapshot. But employees live in a continuous stream. The mismatch between the measurement model and the reality of employee experience is baked into the format.\"*  \n> — CIPD Employee Voice Factsheet, 2024\n\n## The 5 Listening Channels of a Mature VoE Program\n\nA robust employee listening program uses multiple complementary channels, each optimized for different types of signal:\n\n### 1. Continuous Pulse Surveys\n\nShort (3–7 question), high-frequency (weekly or biweekly) surveys that track key engagement indicators over time. Pulse surveys sacrifice depth for velocity — they are designed to catch trend movements early, not to diagnose root causes. They answer: *\"Is engagement moving in the right direction for this team?\"*\n\n**Best for:** Early warning signals, tracking the impact of organizational changes, identifying which teams need deeper investigation.\n\n### 2. In-Depth Employee Interviews\n\nStructured conversations — 20–45 minutes — designed to understand the *why* behind survey scores, the texture of employee experience, and the specific factors driving engagement or disengagement. Interviews produce the qualitative depth that surveys cannot.\n\nInterviews are especially valuable during:\n- Onboarding (30/60/90 day conversations)\n- Manager transitions\n- After organizational changes\n- For employees approaching the 1-year and 3-year tenure marks (common flight risk points)\n- As stay interviews for high-performers who show any disengagement signals\n\n### 3. Stay Interviews\n\nA stay interview is a proactive conversation with a current employee — specifically one you want to retain — focused on understanding what keeps them engaged and what risks might cause them to leave. Unlike exit interviews (which capture data too late), stay interviews give organizations the opportunity to act on retention risks before they become departures.\n\nResearch shows that employees who have stay conversations with their managers report significantly higher engagement — not just because the organization learns useful information, but because the act of being asked demonstrates that the employee's perspective is valued.\n\n### 4. Exit Interviews and Offboarding Research\n\nExit interviews capture the experiences of employees who have decided to leave. While this data cannot retain the individual, it builds an accurate picture of systemic issues driving turnover — issues that often cannot be surfaced any other way because they represent the accumulated frustrations that tipped someone from \"staying\" to \"leaving.\"\n\nThe caveat: traditional exit interviews with a direct HR representative notoriously suffer from social desirability bias. Departing employees fear burning bridges or harming references, so they often give sanitized answers. Exit interview data collected through neutral third parties or AI-moderated interviews tends to be significantly more candid.\n\n### 5. Manager-Team Conversations\n\nThe most frequent and highest-impact feedback channel in most organizations is also the least systematized: the regular 1:1 between managers and their teams. Organizations that train managers to conduct effective listening conversations — and who equip them with structured approaches — create a distributed VoE capability that reaches every employee, not just those who fill out surveys.\n\n## Building Your VoE Program: A 5-Step Framework\n\n### Step 1: Define What You're Listening For\n\nBefore deploying any tools, define the specific questions your VoE program needs to answer. Generic listening produces generic insights. Specific questions produce actionable data.\n\nGood VoE program questions:\n- What are the specific drivers of turnover among engineers with 18–36 months of tenure?\n- How does engagement differ between remote and in-office employees in the same role?\n- What aspects of the onboarding experience best predict 6-month engagement?\n- Which manager behaviors most strongly correlate with team retention?\n\n### Step 2: Design for Psychological Safety First\n\nNo listening program produces honest data from employees who don't believe it's safe to share. Psychological safety for employee feedback requires:\n\n- **Genuine anonymity** (not just promised anonymity — employees need to believe it)\n- **Proof that negative feedback doesn't result in negative consequences** — ideally, demonstrated by leadership acting positively on critical feedback\n- **Accessible channels** that don't require employees to navigate manager relationships to participate\n- **AI-moderated options** — employees consistently report greater willingness to share sensitive feedback with an AI than with an HR professional, because there is no human social relationship to protect\n\n### Step 3: Build Analysis Infrastructure Before You Start Collecting\n\nThe most common VoE failure is collecting more data than the organization can analyze. Before launching any listening initiative, define:\n\n- **Who owns analysis** and how often it runs\n- **How qualitative data (interview transcripts) gets synthesized** — manually, by team, or through AI thematic analysis\n- **What the output format looks like** — what does actionable look like vs. what gets filed\n\nModern AI research platforms like Koji transform this step. Instead of manually coding hundreds of interview transcripts, Koji's automatic thematic analysis surfaces patterns across all conversations, identifies the most frequently raised issues, and generates structured reports ready for leadership review — turning weeks of synthesis work into hours.\n\nKoji's [structured question types](/docs/structured-questions-guide) are particularly powerful in employee research: scale questions (e.g., \"Rate your sense of belonging on this team: 1–10\") aggregate automatically across all participants, while open-ended follow-up probes capture the nuanced reasoning behind every score.\n\n### Step 4: Create the Action Loop\n\nAction loops are the organizational infrastructure that ensures feedback translates into change. Without them, even the most sophisticated listening program becomes a cycle of data collection and filing that breeds employee cynicism.\n\nAn effective action loop includes:\n\n**Acknowledgment:** Communicate back to employees what you heard — within 30 days of a major listening cycle, share a summary of themes. Not conclusions. Not action plans yet. Just: *\"Here is what we heard.\"*\n\n**Prioritization:** Not every piece of feedback requires action. Be explicit about what you are and are not acting on, and why. Employees can accept that some things cannot change. What they cannot accept is silence.\n\n**Visible change:** For every listening cycle, commit to at least one visible, tangible change that is directly traceable to employee feedback. This creates the proof that listening leads to action.\n\n**Follow-through tracking:** Assign owners and timelines to every committed action. Share progress updates in the same channels where you shared the original summary.\n\n### Step 5: Measure the Program, Not Just the Employees\n\nA VoE program should track its own effectiveness alongside employee sentiment:\n\n- **Participation rates** by team, tenure, and role — low participation is a signal of low trust, not low opinions\n- **Response-to-action ratio** — what percentage of major themes raised in feedback cycles resulted in visible organizational action?\n- **Trend lines, not snapshots** — engagement measured over time shows whether the program is having impact\n\n## The AI Advantage in Employee Listening\n\nTraditional employee listening is constrained by the same cost and scale problems as customer research. Scheduling, conducting, and manually analyzing 200 stay interviews across a 1,000-person organization is a multi-month project for an HR team.\n\nAI-moderated employee interviews through platforms like Koji change this dynamic:\n\n- **Scale without headcount:** Run 200 stay interviews in the time it would take to schedule 20. The AI conducts conversations autonomously, at any time, without calendar bottlenecks.\n- **Reduced social desirability bias:** Employees share more candidly with an AI interviewer. Research consistently shows that the absence of human social dynamics removes much of the impression management pressure that distorts moderated interviews.\n- **Automatic synthesis:** Koji's thematic analysis processes all transcripts simultaneously, surfacing the themes that matter — by department, tenure, role, and location — without manual coding.\n- **Structured + qualitative data:** [Structured question types](/docs/structured-questions-guide) capture quantitative signals (NPS-style loyalty scores, ranking of concerns, yes/no on benefit satisfaction) alongside the qualitative depth of open-ended interview conversation — in a single session.\n- **Continuous, not episodic:** Because the cost per interview is dramatically lower, organizations can run listening at the frequency employee experience actually requires — not the frequency their analyst headcount permits.\n\n## The Business Case for Employee Listening\n\nThe ROI of a mature VoE program is well documented:\n\n- Organizations that listen and act on employee feedback achieve **21% higher profitability** than bottom-quartile peers (Gallup)\n- Employees who receive meaningful recognition and feel heard are **45% less likely to turn over** within two years (iSolvedHCM Voice of the Workforce 2024–2025)\n- Regular feedback reduces employee turnover by an estimated **15%** (contactmonkey research)\n- Organizations that actively listen to employees are **3.6 times more likely to innovate well** (Deloitte)\n- According to Deloitte's 2024 Global Human Capital Trends report, **86% of workers and 74% of leaders** say trust and transparency are critically important to organizational performance\n\nThe cost of *not* listening is even starker: replacing an employee costs an estimated 50–200% of their annual salary in recruiting, onboarding, and lost productivity costs. For an organization of 1,000 employees with 15% annual turnover, a 3-percentage-point improvement in retention is worth millions.\n\n## Common VoE Mistakes to Avoid\n\n**Listening without acting:** The fastest way to kill employee trust is to run a listening program that produces no visible change. Every silent follow-up to a feedback cycle teaches employees that honesty is pointless.\n\n**Over-surveying:** Survey fatigue is real. Organizations that deploy too many pulse surveys, too frequently, see response rates drop — and the remaining respondents are increasingly non-representative. Quality over quantity.\n\n**Anonymity theater:** Promising anonymity and then sharing verbatim quotes that identify the speaker. Or sharing results segmented so narrowly (3 people in a 5-person team) that attribution is obvious. True anonymity requires thoughtful data handling, not just a policy statement.\n\n**Centralizing feedback too tightly:** VoE programs that funnel all data to central HR without giving local managers visibility into their team's themes miss the most actionable level. Teams respond to their immediate manager's behavior — that is where most engagement levers actually sit.\n\n**Skipping the action communication:** Even when organizations act on feedback, they often fail to communicate that the action was driven by employee input. Employees cannot credit the program for changes they do not know are connected to their feedback.\n\n## Key Takeaways\n\n- Voice of the Employee is the organizational infrastructure for hearing, understanding, and acting on employee experience — it is not an annual survey\n- Employee engagement hit a 10-year low of 31% in 2024 (Gallup), creating urgent organizational need for more effective listening programs\n- Mature VoE programs combine multiple channels: pulse surveys, in-depth interviews, stay interviews, exit research, and manager-team conversations\n- The action loop — acknowledgment, prioritization, visible change, follow-through — is the most neglected component of most VoE programs, and its absence is the primary cause of program failure\n- AI-moderated employee interviews dramatically reduce social desirability bias, scale without headcount constraints, and transform weeks of synthesis work into hours\n- The ROI is significant: organizations with effective employee listening achieve 21% higher profitability, 45% lower turnover risk among recognized employees, and are 3.6x more likely to innovate\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — combine quantitative metrics and qualitative depth in a single employee research session\n- [Stay Interviews at Scale](/docs/stay-interviews-at-scale) — how AI makes the best retention tool actually usable\n- [Employee Retention Research: Stay Interviews, Exit Interviews, and What Actually Works](/docs/employee-retention-research-guide) — the complete retention research guide\n- [Koji for HR and People Teams](/docs/koji-for-hr-people-teams) — how HR teams use Koji for employee research at scale\n- [How to Build a Continuous Product Feedback Loop](/docs/product-feedback-loop-guide) — the VoC equivalent for product teams\n- [Semi-Structured Interviews: The Complete Guide](/docs/semi-structured-interview-guide) — the interview format used in most stay and exit conversations\n\n\n## Further reading on the blog\n\n- [How to Build a Voice of Customer Program in 2026: The Complete Guide](/blog/how-to-build-voice-of-customer-program-2026) — Companies with best-in-class Voice of Customer programs grow revenue 10x faster than those without. Here's a proven 8-step framework for bui\n- [Best Voice of Customer (VoC) Software in 2026: Top 10 Platforms Compared](/blog/best-voice-of-customer-software-2026) — The 10 best Voice of Customer platforms in 2026, ranked by AI sophistication, accessibility, and time-to-insight. Why Koji is the modern AI-\n- [Customer Journey Mapping Guide 2026: How to Build Maps That Actually Drive Decisions](/blog/customer-journey-mapping-guide-2026) — A modern, AI-native playbook for customer journey mapping in 2026 — including the 5-step process, the questions that surface real emotion at\n\n<!-- further-reading:blog -->\n","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Voice of the Employee Program: Complete Guide to Employee Listening","metaDescription":"Build a Voice of the Employee program that drives real change. Learn continuous listening methods, stay interviews, AI-powered employee research, and how to close the feedback loop.","keywords":["voice of the employee","employee listening program","VoE program","employee feedback","stay interviews","employee engagement research","HR research"],"aiSummary":"A comprehensive guide to building a Voice of the Employee (VoE) program. Covers the limitations of annual surveys, 5 listening channels (pulse surveys, stay interviews, exit research, in-depth interviews, manager conversations), a 5-step implementation framework, and how AI-moderated employee interviews reduce social desirability bias and scale listening without headcount.","aiDifficulty":"intermediate","aiEstimatedTime":"14 minutes"},{"type":"documentation","id":"ffcc33bf-339f-44d8-9fd7-962f2a28637e","slug":"ai-interview-data-privacy-security","title":"AI Interview Data Privacy & Security: A Buyer's Evaluation Guide","url":"https://www.koji.so/docs/ai-interview-data-privacy-security","summary":"Before choosing an AI customer research platform, evaluate it on five points: encryption (in transit and at rest), data residency and sub-processors (and whether your data trains third-party AI models — it should not), PII handling and anonymization, retention and deletion controls, and a signed Data Processing Agreement. AI interviews surface more PII than surveys because participants speak freely, so privacy matters more. Koji encrypts data in transit and at rest, offers a DPA to all business customers (not just enterprise), gates signup to verified business email (blocking consumer, disposable, and school domains), supports anonymized studies, does not train third-party foundation models on your data, and gives you retention and deletion control aligned with GDPR. Privacy-by-design practices: collect only what you need, get consent, prefer structured questions for sensitive attributes, and set retention windows.","content":"## The Short Answer\n\nWhen you run AI interviews, you are collecting **first-party customer data** — names, opinions, sometimes sensitive context — and processing it through AI models. Before you pick a platform, evaluate it on five things: **encryption**, **data residency and sub-processors**, **PII handling and anonymization**, **retention and deletion controls**, and **a signed Data Processing Agreement (DPA)**. Any vendor that cannot answer these clearly should not hold your customers' voices.\n\nKoji is built for business customers who care about this: data is encrypted in transit and at rest, a **DPA is available to all business customers**, signup is **gated to verified business email** (consumer, disposable, and school domains are blocked), participant data can be **anonymized**, and you control **retention and deletion**. This guide gives you a vendor-neutral checklist first, then shows how Koji answers each point.\n\n---\n\n## Why AI Research Raises the Stakes\n\nTraditional surveys collect short, structured answers. AI interviews — especially [voice conversations](/docs/voice-interview-experience) — collect rich, open-ended narratives where participants volunteer far more than a form ever captures. That depth is exactly why AI research is valuable, and exactly why privacy and security matter more, not less.\n\nThree things change with AI-moderated research:\n\n1. **More PII surfaces naturally.** People mention employers, health details, financial situations, and names of colleagues when they talk freely.\n2. **Transcripts are processed by AI models.** You need to know whether your data trains third-party models (it should not).\n3. **Recordings may exist.** Voice studies can produce audio; you need to know how it is stored and for how long.\n\n---\n\n## The 5-Point Evaluation Checklist\n\nUse these questions with **any** research vendor — Koji, SurveyMonkey, Qualtrics, Typeform, or a niche AI tool.\n\n### 1. Encryption\n- Is data encrypted **in transit** (TLS) and **at rest**?\n- Who can access raw transcripts internally?\n\n### 2. Data Residency & Sub-Processors\n- Where is data physically stored?\n- Which sub-processors (AI model providers, hosting, transcription) touch the data?\n- **Is your data used to train third-party AI models?** (The answer you want is *no*.)\n\n### 3. PII Handling & Anonymization\n- Can you **anonymize** or pseudonymize participant identities?\n- Can you redact PII from transcripts before sharing reports?\n- Are participant identifiers separated from response content?\n\n### 4. Retention & Deletion\n- Can you set a **retention window** and auto-delete after it?\n- Can a participant exercise a **right-to-erasure** request, and how fast?\n- Can you export everything for your own records before deletion?\n\n### 5. Contracts & Compliance\n- Will the vendor sign a **DPA**?\n- Do they support **GDPR** obligations (lawful basis, data-subject rights)?\n- For regulated data, can they support **HIPAA**-aligned workflows?\n\nIf a vendor dodges any of these, treat it as a red flag.\n\n---\n\n## How Koji Answers Each Point\n\n### Encryption\nCustomer and participant data is encrypted **in transit (TLS) and at rest**. Access to raw transcripts is restricted to the workspace that owns the study.\n\n### Data residency, sub-processors, and model training\nKoji uses vetted AI model providers to run interviews and analysis. **Your interview data is not used to train third-party foundation models.** Sub-processors are disclosed so your security team can review the chain before you commit.\n\n### PII handling & anonymization\nBecause AI interviews surface more personal detail than surveys, Koji supports [anonymizing customer interview data](/docs/anonymizing-customer-interview-data) — you can run studies without collecting real names, and reports can present themes and quotes without exposing identities. Structured questions also help: instead of free-typing sensitive data, you can capture it as a [scale or single-choice answer](/docs/structured-questions-guide) that is inherently easier to govern.\n\n### Retention & deletion\nYou control how long data lives. Studies can be exported for your records and then deleted, and participant erasure requests can be honored — the foundation of [GDPR-compliant research](/docs/gdpr-compliant-ai-user-research).\n\n### Contracts & compliance\nKoji provides a **Data Processing Agreement (DPA) to all business customers** — not just enterprise plans. The product is designed around **GDPR** principles, and for teams handling protected health information, see [HIPAA-compliant AI user research](/docs/hipaa-compliant-ai-user-research) for the right configuration.\n\n### Access control at the front door\nSignup is **gated to verified business email**. Consumer mailbox providers, disposable/temporary email domains, and school domains are blocked. That keeps workspaces tied to real organizations and reduces the risk of anonymous accounts hoarding customer data.\n\n---\n\n## Privacy-by-Design Research Practices\n\nTooling is half the story. These practices reduce risk regardless of platform:\n\n- **Collect only what you need.** If a study does not require names, do not ask for them. Koji studies run fine fully anonymous.\n- **Tell participants what happens to their data.** A one-line consent notice before the interview builds trust and satisfies lawful-basis requirements.\n- **Prefer structured questions for sensitive attributes.** A [single_choice or scale question](/docs/structured-questions-guide) about income band is easier to govern than a free-text field.\n- **Set a retention window up front** so old studies do not become liability.\n- **Restrict report sharing** to the people who need the insight.\n\nPlatforms like Koji make privacy-by-design the default: anonymous studies, automatic [thematic analysis](/docs/ai-transcript-analysis-guide) that summarizes without exposing every raw quote, and exportable, deletable data.\n\n---\n\n## Buyer Red Flags\n\n- \"We *might* use your data to improve our models.\" → walk away\n- No DPA, or DPA \"only on enterprise\" → governance gap\n- Cannot tell you where data is stored or who the sub-processors are\n- No deletion or export path\n- Free signups with personal email and no organizational control\n\n---\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — capture sensitive attributes as governable structured answers\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — lawful basis, data-subject rights, and retention\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) — configuring studies for protected health information\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — run studies without collecting real identities\n- [AI Transcript Analysis Guide](/docs/ai-transcript-analysis-guide) — summarize insight without over-exposing raw PII\n- [MCP Overview](/docs/mcp-overview) — how Koji connects to AI clients securely","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"AI Interview Data Privacy & Security: Buyer's Evaluation Guide","metaDescription":"Evaluate the privacy and security of an AI research platform: encryption, sub-processors, PII and anonymization, retention, deletion, and DPA. Plus how Koji handles each — DPA for all business customers, business-email-gated signup, and GDPR-aligned design.","keywords":["ai interview data privacy","ai research security","customer research data protection","ai survey privacy","research platform gdpr","data processing agreement research","anonymize interview data"],"aiSummary":"Before choosing an AI customer research platform, evaluate it on five points: encryption (in transit and at rest), data residency and sub-processors (and whether your data trains third-party AI models — it should not), PII handling and anonymization, retention and deletion controls, and a signed Data Processing Agreement. AI interviews surface more PII than surveys because participants speak freely, so privacy matters more. Koji encrypts data in transit and at rest, offers a DPA to all business customers (not just enterprise), gates signup to verified business email (blocking consumer, disposable, and school domains), supports anonymized studies, does not train third-party foundation models on your data, and gives you retention and deletion control aligned with GDPR. Privacy-by-design practices: collect only what you need, get consent, prefer structured questions for sensitive attributes, and set retention windows.","aiPrerequisites":["Basic understanding of data privacy concepts (PII, GDPR)","Knowing what customer data your research will collect"],"aiLearningOutcomes":["Evaluate any AI research vendor on five concrete security dimensions","Understand why AI interviews raise privacy stakes versus surveys","Know how Koji handles encryption, anonymization, retention, and DPAs","Apply privacy-by-design practices to your studies","Recognize buyer red flags before you commit"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 minutes"},{"type":"documentation","id":"711783ba-2d23-4136-bdbb-27fc1506d846","slug":"enterprise-security-ai-research-platforms","title":"Enterprise Security for AI Customer Research Platforms: SOC 2, SSO, and Vendor Review","url":"https://www.koji.so/docs/enterprise-security-ai-research-platforms","summary":"Enterprise buyers evaluate AI research platforms on encryption (AES-256 at rest, TLS 1.2+ in transit), SOC 2 Type II attestation or roadmap, SSO/SAML, transparent sub-processors and data residency, annual penetration testing, audit logging with configurable retention, and a signable DPA. Koji runs on SOC 2 Type II-attested cloud infrastructure, uses AES-256 and TLS 1.2+, runs annual third-party pen tests, and publishes its compliance posture, with its own SOC 2 and ISO 27001 on a dated roadmap.","content":"## The Bottom Line\n\nWhen you bring an AI customer research platform into an enterprise, the buying decision is rarely made by the research team alone — it passes through security review, legal, and procurement. The platforms that clear that gate quickly share five traits: encryption in transit and at rest, a SOC 2 Type II attestation (or a credible, dated roadmap to one), SSO/SAML for access control, transparent sub-processor and data-residency disclosure, and a Data Processing Agreement (DPA) ready to sign. Koji is built on SOC 2 Type II-attested cloud infrastructure (AWS and Google Cloud), encrypts data with AES-256 at rest and TLS 1.2+ in transit, commissions independent annual penetration testing, and publishes its compliance posture openly — so your security review moves in days, not quarters.\n\nThis guide gives you the exact checklist to run a vendor security assessment on any AI research tool, and shows where Koji stands on each line item.\n\n## Why AI research platforms get extra scrutiny\n\nA customer research platform is not a low-stakes tool. It collects first-party voice and text from your customers, employees, or prospects — often including names, opinions about your product, and sometimes regulated personal data. The moment a platform records, transcribes, and analyzes those conversations with AI, three risks land on your security team's desk:\n\n- **Data exposure**: interview transcripts and recordings are sensitive. A breach is both a privacy incident and a competitive one.\n- **Sub-processor sprawl**: AI features route data to model providers, transcription engines, and analytics vendors. Each is a sub-processor your legal team must vet.\n- **Access control**: research data often gets shared widely inside a company. Without SSO and role-based permissions, that sharing becomes a liability.\n\nTraditional survey tools were never designed for this level of qualitative depth. AI-native platforms like Koji are — which means security is engineered in, not retrofitted.\n\n## The enterprise security checklist\n\nUse this checklist to evaluate any AI research vendor. Send it verbatim to your security team.\n\n### 1. Encryption\nConfirm encryption **in transit** (TLS 1.2 or higher) and **at rest** (AES-256). Ask whether certificate management is automatic and whether any data is ever stored unencrypted, even temporarily. *Koji: TLS 1.2+ in transit, AES-256 at rest, with automatic certificate management handled by the underlying cloud platform.*\n\n### 2. SOC 2 Type II\nThis is the single most common gate. Ask for the attestation report under NDA, or — if the vendor is earlier-stage — a dated roadmap with a defined audit period. Be wary of vendors who claim compliance with no report and no timeline. *Koji: runs on two SOC 2 Type II-attested cloud platforms (AWS and Google Cloud); Koji's own SOC 2 Type II and ISO/IEC 27001 attestations are on the published compliance roadmap with a defined target audit period.*\n\n### 3. Penetration testing\nAsk how often independent third-party penetration tests run and whether a summary letter is available. Annual cadence is the baseline. *Koji: independent third-party penetration testing is scoped on an annual cadence alongside its audit engagement.*\n\n### 4. SSO and access control\nFor any team over a handful of seats, SSO/SAML is non-negotiable — it lets you enforce your own password and MFA policy and deprovision instantly. Confirm role-based permissions so a viewer cannot edit studies or export raw data. *Koji: supports SSO/SAML and role-based access.*\n\n### 5. Sub-processors and data residency\nRequest the current sub-processor list and where data is stored and processed (region matters for GDPR and data-residency requirements). *Koji: publishes its sub-processor list and data-residency information on its compliance pages.*\n\n### 6. Audit logging and retention\nAsk whether administrative and authentication events are logged, how long logs are retained, and whether you control data-retention windows. *Koji: maintains database and authentication audit logs with configurable retention (up to six-year retention available by contract).*\n\n### 7. DPA and privacy framework\nConfirm a signable DPA, GDPR alignment, and support for data subject access and deletion requests. *Koji offers a DPA to business customers and is built for GDPR-aligned workflows including anonymization and deletion.*\n\n## How Koji is architected for enterprise trust\n\nKoji's approach to security follows a simple principle: collect rich qualitative data without becoming a liability. A few design choices matter here.\n\n**Async, link-based interviews reduce recording risk.** Because Koji interviews are conducted through a shareable link rather than a live, recorded video call, there is no third-party meeting recorder in the loop and consent is captured in the interview flow itself. Fewer moving parts means a smaller attack surface and a cleaner consent trail.\n\n**The quality gate limits unnecessary data processing.** Koji only counts conversations that score 3 or higher on its quality scale toward your plan — low-effort or junk sessions are filtered. That same gate means your analysis (and the data you retain) focuses on genuine signal.\n\n**Structured questions keep data predictable.** Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — so you decide exactly what is collected. Quantitative fields stay quantitative and predictable, while open-ended answers get AI follow-up probing. Knowing your data schema up front makes retention and anonymization policies far easier to enforce. See the [structured questions guide](/docs/structured-questions-guide) for the full breakdown.\n\n**Transparent posture, not vague assurances.** Koji publishes its security, sub-processor, incident-response, and certification-status pages openly. For a security reviewer, public documentation that names specifics (AES-256, TLS 1.2+, annual pen testing, six-year log retention) is worth more than a marketing claim of being enterprise-grade.\n\n## Running the vendor review efficiently\n\nA few practical tips to get research tools approved fast:\n\n1. **Loop security in during the trial, not after.** Send the checklist above the moment a tool reaches your shortlist. Security review is the longest pole — start it early.\n2. **Ask for documentation links, not promises.** A vendor that can point you to a live security page and a DPA template is one that has done this before.\n3. **Scope data minimization into your study design.** Use Koji's structured questions and screeners to collect only what you need. The less personal data you gather, the lighter your compliance burden.\n4. **Set retention deliberately.** Decide how long transcripts should live and configure retention accordingly rather than defaulting to forever.\n\n## Where this leaves you\n\nThe modern, AI-native research platforms win enterprise deals precisely because they treat security as a feature. With AES-256 encryption, SSO/SAML, annual penetration testing, transparent sub-processor disclosure, a signable DPA, and a published path to its own SOC 2 Type II and ISO 27001 attestations, Koji gives your security team the artifacts they need to say yes — while your research team gets AI voice and text interviews, automatic analysis, and real-time reports that traditional survey tools cannot match.\n\n## Red flags in a vendor security review\n\nA few warning signs should slow a purchase until they are resolved:\n\n- **Compliance claims with no artifact.** A vendor that says it is SOC 2 compliant but cannot share a report, a roadmap, or a status page is asserting something you cannot verify. Credible vendors point to documentation.\n- **No DPA, or a take-it-or-leave-it contract.** A platform handling customer conversations should expect to sign a DPA. Resistance here is a signal about how they treat data obligations generally.\n- **Vague sub-processor disclosure.** AI features route data to model and transcription providers. If a vendor cannot name its sub-processors, your legal team cannot assess the chain of custody.\n- **No SSO on business plans.** If single sign-on is locked away or unavailable, centralized access control and instant deprovisioning become manual and error-prone.\n- **Recorded live calls with no consent trail.** Tools that depend on third-party meeting recorders add a sub-processor and a consent burden. Koji's async, link-based interviews avoid both by capturing consent in the flow.\n\nScoring a shortlist against these red flags — alongside the seven-point checklist above — turns a subjective security conversation into a comparable, defensible evaluation you can document for procurement.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how they keep collected data predictable\n- [AI Interview Data Privacy & Security](/docs/ai-interview-data-privacy-security) — how interview data is protected end to end\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — running research under GDPR\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) — research in regulated healthcare settings\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — de-identification best practices\n- [Exporting Research Data](/docs/exporting-research-data) — getting data out securely","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Enterprise Security for AI Customer Research Platforms (SOC 2, SSO & Vendor Review)","metaDescription":"How to evaluate the security of an AI customer research platform: SOC 2, AES-256 encryption, SSO/SAML, data residency, sub-processors, and a procurement-ready vendor checklist.","keywords":["enterprise security","soc 2 research platform","ai research security","vendor security review","data residency","sso saml research tool","dpa research platform","secure customer research","penetration testing","customer research compliance"],"aiSummary":"Enterprise buyers evaluate AI research platforms on encryption (AES-256 at rest, TLS 1.2+ in transit), SOC 2 Type II attestation or roadmap, SSO/SAML, transparent sub-processors and data residency, annual penetration testing, audit logging with configurable retention, and a signable DPA. Koji runs on SOC 2 Type II-attested cloud infrastructure, uses AES-256 and TLS 1.2+, runs annual third-party pen tests, and publishes its compliance posture, with its own SOC 2 and ISO 27001 on a dated roadmap.","aiPrerequisites":["Basic familiarity with vendor security review","Understanding of your organization's compliance requirements"],"aiLearningOutcomes":["Run a structured security assessment on any AI research vendor","Distinguish credible security claims from marketing language","Understand Koji's encryption, SOC 2 posture, and access controls","Design studies that minimize data and ease compliance"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 minutes"},{"type":"documentation","id":"4ebf0f84-5517-4032-9cb9-3d6f6c8c957c","slug":"gdpr-compliant-ai-user-research","title":"GDPR-Compliant AI User Research: A Practical Guide","url":"https://www.koji.so/docs/gdpr-compliant-ai-user-research","summary":"GDPR-compliant AI user research requires six things on paper: a lawful basis (usually consent), purpose limitation, data minimization, storage limitation, participant rights, and sub-processor transparency. For AI interviews, the privacy notice must name the LLM vendor and disclose cross-border transfers. Koji handles each requirement with a built-in consent intake form, per-study retention controls, EU-region routing, BYOK for direct LLM contracting, DPA on request, and SAR-ready data export. Best practices: pseudonymize participants, minimize demographic fields, anonymize transcripts in reports, document retention in your Record of Processing Activities. Most simple research doesn't need a DPIA; sensitive categories or vulnerable groups do.","content":"## What GDPR-compliant AI user research means\n\nGDPR-compliant AI user research is a research practice where every participant interaction — recruitment, consent, interview, transcript, analysis, and storage — satisfies the EU General Data Protection Regulation. The two things that matter most are (1) a lawful basis for processing each participant's data, and (2) clear participant control over their own data, including the right to withdraw at any time.\n\nWhen the moderator is an AI rather than a human, GDPR still applies — sometimes more strictly, because LLM providers may be sub-processors located outside the EU. This guide explains how to run GDPR-compliant AI user research end to end, the questions your DPO will ask, and how Koji is built so EU teams can deploy AI interviews without a legal sprint.\n\nNothing in this guide is legal advice. Run your specific use case past counsel before processing data from EU residents.\n\n## The six GDPR essentials for AI research\n\nEvery GDPR-compliant AI user research program needs to answer these six questions on paper:\n\n1. **Lawful basis** — usually consent (Art. 6(1)(a)) for research, occasionally legitimate interest (Art. 6(1)(f)) for existing customers.\n2. **Purpose limitation** — research participants are told exactly what their data will be used for and you don't silently use it for something else (training a model, marketing, etc.).\n3. **Data minimization** — collect only what the research goal requires. Don't ask for date of birth if age band is enough.\n4. **Storage limitation** — define a retention period, document it, and delete after.\n5. **Participant rights** — provide a clear way to access, rectify, port, and delete their data.\n6. **Sub-processor transparency** — disclose every third party that touches participant data (LLM provider, transcription service, hosting region).\n\nThe rest of this guide walks each one through the lens of running an AI-moderated interview study.\n\n## Lawful basis: when consent is required\n\nFor most user research, consent is the cleanest lawful basis because participation is voluntary, the data is sensitive (qualitative answers often reveal personal opinions), and you want unambiguous proof of agreement.\n\nConsent under GDPR has to be:\n\n- **Freely given** — no dark patterns, no penalty for declining.\n- **Specific** — for this study, not all future research.\n- **Informed** — the participant knows what data is collected, by whom, for how long, and who else will see it.\n- **Unambiguous** — affirmative action, not a pre-checked box.\n\nKoji handles this with the built-in intake form ([intake forms and consent](/docs/intake-forms-and-consent)). You can require participants to read your privacy notice and tick a consent box before the AI moderator starts. For each study, the consent record is timestamped and retrievable.\n\nIf you're researching existing customers and the research is closely related to the service you already provide, legitimate interest may apply — but you still owe participants a clear notice and an easy opt-out. Document the balancing test.\n\n## The participant-facing privacy notice\n\nEvery AI research study processing EU data needs a privacy notice at the start of the interview. It should cover, in plain language:\n\n- **Who you are** (data controller) and contact info.\n- **What you'll ask** and roughly how long the interview takes.\n- **Whether the interview is recorded** (voice mode) or transcript-only (text mode).\n- **Which AI provider transcribes / moderates** (OpenAI, Anthropic, Google — name the LLM vendor).\n- **Where data is stored** and for how long.\n- **Whether data leaves the EU** and what safeguards apply (SCCs, adequacy decisions).\n- **How to withdraw consent** and request deletion.\n- **Whether any decisions affecting the participant are made automatically** (under Art. 22).\n\nKoji ships customizable notice fields in the intake step, and the [research consent form templates](/docs/research-consent-form-templates) include EU-ready language you can adapt.\n\n## Data minimization for AI interviews\n\nThe AI moderator doesn't need a lot of personal data to do its job. Best practice:\n\n- **Use pseudonymous IDs.** Pass `participant_id=abc-123` instead of `email=jane@example.com` where possible. See [personalized interview links](/docs/personalized-interview-links).\n- **Skip demographic questions you won't analyze.** Don't ask for nationality, exact age, or income if cohort-level data is enough.\n- **Anonymize transcripts.** Koji can strip names, employers, and email addresses from analysis exports — useful when sharing reports beyond the research team.\n- **Aggregate, don't identify.** When publishing findings, summarize at the theme level. Verbatim quotes need separate consent.\n\nThe fewer columns of personal data you store, the smaller the GDPR surface area and the simpler your DPIA becomes.\n\n## Retention: how long is \"as long as necessary\"?\n\nGDPR says you can keep personal data only as long as you need it for the stated purpose. For research, common retention bands are:\n\n- **30 days** for raw audio recordings (used only for transcription verification).\n- **6–12 months** for transcripts (long enough for follow-up analysis and report iteration).\n- **12–24 months** for de-identified themes and aggregated insights (those usually don't qualify as personal data once anonymized).\n\nKoji lets you set per-study retention. Configure it in the study settings, and Koji automatically purges raw conversations on schedule while keeping the aggregated report intact.\n\n## Right to withdraw, access, port, and delete\n\nEvery participant must be able to:\n\n- Withdraw consent at any time, including mid-study.\n- Access the personal data you hold about them.\n- Receive a portable copy in a common format.\n- Request deletion (right to erasure).\n\nOperationally:\n\n- Provide a single email address (`privacy@yourcompany.com`) in the intake notice.\n- Train CS or the research team to action these requests within 30 days.\n- Use Koji's [exporting research data](/docs/exporting-research-data) feature to produce a participant-specific export when a Subject Access Request comes in.\n- Use the delete-interview action in the study admin to remove a participant's session and transcript.\n\nDocument each request and the response date for audit purposes.\n\n## Sub-processors and cross-border transfers\n\nThe biggest GDPR question with AI research is: which third parties touch the data, and where are they?\n\nKoji discloses every sub-processor on its public sub-processor page (cloud host, LLM provider, transcription provider, email delivery, etc.). For EU customers, key points:\n\n- **Data residency**: studies can be configured to keep transcripts within the EU; LLM inference may happen in a US region under Standard Contractual Clauses with additional safeguards.\n- **Bring Your Own Key (BYOK)**: [Enterprise plan](/enterprise) customers can route LLM calls through their own contracted OpenAI / Anthropic / Google / Azure accounts so the LLM relationship is direct. Self-serve users can also enable BYOK per-user. See [bring your own key](/docs/bring-your-own-key).\n- **DPA on file**: Koji has a pre-signed Data Processing Agreement available at [/compliance/dpa](/compliance/dpa); Enterprise customers counter-sign within 1 business day. EU SCCs (2021) and UK Addendum incorporated.\n- **No training on customer data**: explicit contractual commitment covered in [/compliance/ai-governance](/compliance/ai-governance). Koji's LLM contracts disable training on customer prompts and outputs.\n- **Full GDPR program**: see [/compliance/gdpr](/compliance/gdpr) for the complete EU GDPR + UK GDPR positioning, EU member state nuances, lawful bases, data-subject rights, retention, and transfer mechanisms.\n\nIf your organization has strict residency rules (financial services, public sector), discuss BYOK and EU-region routing during procurement.\n\n## DPIA: do you need one?\n\nA Data Protection Impact Assessment is mandatory under Art. 35 when processing is \"likely to result in a high risk\" to participants. Most simple AI user research (voluntary, no sensitive categories, anonymized output) doesn't cross that threshold. You should run a DPIA if:\n\n- You're collecting health data, sexual orientation, political views, or other [Art. 9 special categories](https://gdpr-info.eu/art-9-gdpr/).\n- You're researching vulnerable groups (children, patients, employees in power-imbalanced contexts).\n- The interview is mandatory (employees in mandatory feedback, customers tied to service access).\n- The AI makes any consequential decision automatically.\n\nFor everything else, document the lawful basis, consent flow, and retention in a lightweight Record of Processing Activities (Art. 30) and you're typically covered.\n\n## How Koji compares to running AI research with raw ChatGPT\n\nSome teams paste customer interview transcripts into raw ChatGPT to analyze them. Under GDPR, this is risky:\n\n- Pasting personal data into a general-purpose ChatGPT account is a transfer to a sub-processor you may not have authorized.\n- Free-tier ChatGPT trains on inputs by default.\n- There's no DPA, no consent record, no retention policy, no deletion path.\n\nKoji, in contrast, is purpose-built for compliant research: contracted LLM use with training disabled, consent records, per-study retention, deletion workflow, EU residency option, and a DPA available. See [can I paste user interviews into ChatGPT](/docs/can-i-paste-user-interviews-into-chatgpt-a-guide-to-gdpr-and-llms) for the deeper comparison.\n\n## Practical setup checklist for EU teams\n\nBefore publishing your first study:\n\n1. Draft a participant-facing privacy notice covering the seven elements above.\n2. Configure the Koji intake form to require explicit consent.\n3. Set per-study retention (audio, transcript, themes).\n4. Confirm your DPA is signed.\n5. If your DPO requires EU residency, request EU-region routing or enable BYOK.\n6. Document the lawful basis and retention in your Record of Processing Activities.\n7. Train CS / research on handling Subject Access Requests within 30 days.\n\nWith those seven steps, your AI user research program meets the GDPR bar without slowing discovery to a crawl.\n\n## Related Resources\n\n- [Intake forms and consent](/docs/intake-forms-and-consent) — configure GDPR-ready consent screens\n- [Research consent form templates](/docs/research-consent-form-templates) — EU-ready notice language\n- [Personalized interview links](/docs/personalized-interview-links) — pseudonymize participants without losing context\n- [Bring your own key](/docs/bring-your-own-key) — route LLM calls through your own contracted account\n- [Exporting research data](/docs/exporting-research-data) — produce Subject Access Request exports\n- [Can I paste user interviews into ChatGPT? A GDPR guide](/docs/can-i-paste-user-interviews-into-chatgpt-a-guide-to-gdpr-and-llms)\n- [Structured questions guide](/docs/structured-questions-guide) — design briefs that minimize personal data collection","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"GDPR-Compliant AI User Research: A Practical Guide | Koji","metaDescription":"Run AI-moderated customer interviews under GDPR. Lawful basis, consent flows, data minimization, retention, sub-processors — and how Koji handles each requirement.","keywords":["GDPR user research","GDPR AI research","GDPR-compliant survey","EU customer research compliance","data protection user research","research consent EU","AI research privacy","research DPIA","research sub-processors"],"aiSummary":"GDPR-compliant AI user research requires six things on paper: a lawful basis (usually consent), purpose limitation, data minimization, storage limitation, participant rights, and sub-processor transparency. For AI interviews, the privacy notice must name the LLM vendor and disclose cross-border transfers. Koji handles each requirement with a built-in consent intake form, per-study retention controls, EU-region routing, BYOK for direct LLM contracting, DPA on request, and SAR-ready data export. Best practices: pseudonymize participants, minimize demographic fields, anonymize transcripts in reports, document retention in your Record of Processing Activities. Most simple research doesn't need a DPIA; sensitive categories or vulnerable groups do.","aiPrerequisites":["Familiarity with GDPR basics","Access to your organization DPO or counsel","Awareness of your data residency requirements"],"aiLearningOutcomes":["Identify the correct lawful basis for an AI research study","Draft a GDPR-ready participant privacy notice","Configure consent, retention, and sub-processor disclosure in Koji","Decide when a DPIA is required","Handle Subject Access and erasure requests within 30 days"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"203dbc40-d0f1-4b51-925c-8e619d75efaa","slug":"ccpa-user-research-compliance","title":"CCPA/CPRA Compliance for Customer Research: The 2026 Practitioner Guide","url":"https://www.koji.so/docs/ccpa-user-research-compliance","summary":"CCPA/CPRA applies to for-profit businesses doing business in California meeting one threshold: $26.625M annual revenue, 100,000+ California consumers, or 50%+ revenue from selling/sharing personal information. B2B contacts and employees are in scope since the 2023 exemption sunset. The highest-risk research gap is vendor contracts: without six specific service provider terms, disclosing participant data to a research platform, panel, or transcription vendor is legally a sale or sharing that triggers opt-out rights. Voice recordings are biometric information only if an identifier template can be extracted, and sensitive personal information only when actually used to identify a consumer, so ordinary voice research is not automatically sensitive. Enforcement in 2025-26 targeted broken opt-outs, excessive verification, asymmetrical choices, missing vendor terms, and over-collection: Honda $632,500, Healthline $1.55M, a youth sports platform $1.10M. Statutory penalties are $2,663 per unintentional and $7,988 per intentional violation.","content":"**Answer first:** If your research vendor contract lacks the specific CCPA service provider clauses, handing them your participant list is legally a **\"sale\" or \"sharing\"** of personal information — which triggers opt-out rights you almost certainly are not honouring. That contract language, not your consent form, is the single highest-risk gap in most US research programmes.\n\nThe second thing to know: California enforcement in 2025–26 has not targeted exotic edge cases. It has targeted **broken opt-outs, excessive verification, and missing vendor contract terms** — all three of which are routine failures in research operations.\n\nIf your research reaches European participants too, read this alongside our [GDPR-compliant AI user research guide](/docs/gdpr-compliant-ai-user-research). The regimes overlap but their mechanics differ sharply, and California's are less forgiving in one specific place: the paperwork with your vendors.\n\n## Does CCPA even apply to you?\n\nThe CCPA applies to for-profit businesses doing business in California that meet at least one threshold. Adjusted for inflation, the revenue threshold now stands at **$26.625 million** in annual gross revenue. The alternatives: buying, selling, or sharing the personal information of **100,000+** California consumers or households annually, or deriving **50% or more** of annual revenue from selling or sharing personal information.\n\nTwo traps worth naming:\n\n- **\"Consumer\" includes B2B contacts and employees.** California abandoned the B2B and HR exemptions in 2023. Your enterprise buyer interviews and your employee research are both in scope. This surprises people constantly.\n- **You do not need a California entity.** Doing business in California is the test, and serving California customers over the internet counts.\n\nIf you are under every threshold, CCPA does not bind you today — but write your research programme to comply anyway. Crossing $26.625M in revenue should not require rebuilding your consent architecture.\n\n## The vendor contract that decides everything\n\nThis is the part of CCPA that research teams get wrong most often, and it is worth understanding precisely because the consequence is severe and invisible.\n\nWhen you disclose participant personal information to an outside company — a research platform, a recruiting panel, a transcription service, an incentive fulfilment vendor — California asks what that recipient is. There are three answers:\n\n| Classification | What it means | Opt-out consequence |\n| --- | --- | --- |\n| **Service provider / contractor** | Processes data only for your specified purposes, under a written contract with required terms | No opt-out right triggered |\n| **Third party** | Anyone who does not meet the service provider criteria | Disclosure is a \"sale\" or \"sharing\" — consumers can opt out |\n\nHere is the mechanism that catches people: **the classification is created by the contract, not by the relationship.** To qualify as a service provider, the transfer must be pursuant to a written contract that prohibits the recipient from retaining, using, or disclosing the personal information for any purpose other than the specific purposes named in that contract.\n\nWithout those terms, the disclosure is a sale or sharing to a third party by default — even if your vendor never does anything commercially inappropriate with the data. Good behaviour does not cure a missing clause.\n\nYour CCPA vendor contract must:\n\n1. Prohibit use of the personal information beyond the specified purposes\n2. Require the vendor to provide **the same level of privacy protection** the CPRA requires of you\n3. Grant you rights to take reasonable steps to verify appropriate use\n4. Require the vendor to **notify you** if it can no longer meet its obligations\n5. Grant you rights to stop and remediate unauthorised use\n6. Bind subcontractors to equivalent terms\n\n**Your action item:** pull every research vendor contract you have and check for these six terms. Recruiting panels and transcription services are the usual offenders — research platforms tend to have their paper in order, incentive and panel vendors frequently do not. A vendor who cannot produce a compliant DPA is not a procurement inconvenience; they are converting your research operations into an unreported sale of personal information.\n\n## Voice recordings, biometrics, and where the line sits\n\nVoice research raises a question worth answering carefully, because the common assumption — \"audio is biometric, therefore sensitive\" — is wrong in a way that leads to unnecessary compliance theatre.\n\nUnder California law, a voice recording can qualify as biometric information **if an identifier template can be extracted from it**. But it constitutes *sensitive* personal information only when the recording is actually **used to identify a consumer**. Recording an interview to understand what a customer thinks about your onboarding flow is not identification. Running voiceprint matching against that audio is.\n\nSo ordinary voice research does not automatically pull you into the sensitive-personal-information regime with its additional right to limit use. What does pull you in: collecting health information, precise geolocation, racial or ethnic origin, or — added by SB 1223 and effective January 2025 — **neural data**.\n\nNote also that AB 1008, effective January 2025, clarified that personal information includes data **embedded in AI models and other abstract digital systems**. If participant data was used to fine-tune a model, that model may itself hold personal information. This is a strong argument for research tooling that does not train on your data, and for asking vendors the question explicitly.\n\nWhere research *does* commonly touch sensitive categories is in screeners. A health-condition screener or an ethnicity quota question collects sensitive personal information directly — see [research screener questions](/docs/research-screener-questions) for how to collect only what your quotas genuinely require.\n\n## What California actually enforces\n\nThe enforcement record is the most useful compliance document available, because it shows what regulators care about rather than what commentators speculate about. Recent actions:\n\n- **Honda — $632,500.** For requiring excessive personal information to verify privacy rights requests, presenting asymmetrical privacy choices, making authorised-agent requests unnecessarily difficult, and — directly relevant here — **sharing personal information with vendors without the required contract terms.**\n- **Healthline — $1.55 million**, the largest CCPA penalty to date and the first data-minimisation enforcement action. The consent banner logged rejections while tracking continued regardless.\n- **A youth sports media platform — $1.10 million**, for opt-out and consumer notice failures.\n\nStatutory penalties in 2026 run to **$2,663 per violation** for unintentional violations and **$7,988** for intentional violations or those involving minors. Per violation means per consumer — the arithmetic across a participant database escalates quickly.\n\nThe pattern is unmistakable: **asymmetrical choices, excessive verification, broken opt-outs, missing vendor contract terms, and collecting more than you need.** Four of those five are ordinary research-operations failure modes.\n\n## Your compliance checklist\n\n1. **Notice at collection, before you collect.** At or before the point of collection, tell participants what categories you collect, why, how long you keep it, and whether it is sold or shared. On the recruitment screen — not in a policy they reach afterwards.\n2. **Fix the vendor contracts.** All six terms, every vendor, including panels and transcription.\n3. **Make opt-outs real.** If any disclosure is a sale or sharing, honour Global Privacy Control signals and provide a working opt-out. Honda's fine says asymmetrical choices are independently punishable.\n4. **Verify proportionately.** Do not demand a government ID to process a deletion request. Excessive verification is itself a violation.\n5. **Minimise deliberately.** Collect what the research question requires. Do not retain a full participant profile because it might be useful later — that is precisely the Healthline theory.\n6. **Build deletion you can execute.** A deletion request must reach transcripts, recordings, analysis artefacts, and your repository. If deletion cannot reach your insight repository, you cannot comply. See [research repository guide](/docs/research-repository-guide).\n7. **Honour the 45-day clock.** Respond to requests within 45 days, extendable to 90 with notice.\n8. **Keep a data inventory.** Categories collected, sources, purposes, disclosure recipients, retention. Everything above depends on it.\n\n## How Koji helps\n\nCompliance gets dramatically easier when the platform is designed so the compliant path is the default one:\n\n- **Service provider by contract.** Koji operates as a service provider under CCPA, with the required terms in place — your disclosure to Koji does not become a sale.\n- **No training on your research data.** Which keeps AB 1008's model-embedded-personal-information problem from arising in the first place.\n- **Consent and notice at the front door.** Koji's [intake forms and consent](/docs/intake-forms-and-consent) present notice at collection before the first question, which is exactly where California requires it.\n- **Per-study retention with automatic purge.** Set retention when you design the study and let raw conversations expire on schedule while aggregated themes survive. Retention becomes a property of the study rather than a quarterly cleanup project.\n- **Deletion that reaches everything.** Because transcripts, analysis, and reports live in one system, a deletion request is one operation — not a hunt across a transcription vendor, a spreadsheet, a Notion page, and three stakeholders' downloads. That fragmentation is the real reason deletion requests go unfulfilled, and consolidation is the fix.\n- **Minimisation through better instrumentation.** This is where methodology and compliance align. Koji's six [structured question types](/docs/structured-questions-guide) — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — let you capture exactly the data point you need rather than an open field you later mine. A `scale` question asking satisfaction directly collects one number. A free-text field that participants fill with health details, employer names, and family circumstances collects sensitive personal information you never intended to hold and now must protect, disclose, and delete. Precise instrumentation is data minimisation implemented in the study design.\n\nThe structural contrast: legacy survey tools optimise for collecting as much as possible and sorting it out later. That instinct was harmless in 2015 and is now a liability with a per-consumer price tag.\n\n## Related Resources\n\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the European regime, and where it diverges from California's\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) — AI-specific obligations layered on top of privacy law\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types, and why precise instrumentation is data minimisation\n- [AI Interview Data Privacy & Security](/docs/ai-interview-data-privacy-security) — storage, encryption, and retention mechanics\n- [Intake Forms and Consent](/docs/intake-forms-and-consent) — presenting notice at collection correctly\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — copy-ready participant notices\n- [Research Screener Questions](/docs/research-screener-questions) — qualifying participants without over-collecting\n\n*Regulatory information current as of July 2026. Practitioner orientation, not legal advice — confirm your obligations with qualified privacy counsel.*\n","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"CCPA/CPRA Compliance for Customer Research (2026 Guide)","metaDescription":"What CCPA/CPRA actually requires for interviews, surveys, and voice research: the six service provider contract terms that keep vendor disclosures from becoming a \"sale\", when voice recordings count as biometric data, and the 2026 enforcement record.","keywords":["ccpa user research","cpra customer research","ccpa compliance interviews","california privacy research","ccpa service provider contract","ccpa voice recording biometric","ccpa research participants","cpra sensitive personal information","ccpa penalties 2026","ccpa data minimization research"],"aiSummary":"CCPA/CPRA applies to for-profit businesses doing business in California meeting one threshold: $26.625M annual revenue, 100,000+ California consumers, or 50%+ revenue from selling/sharing personal information. B2B contacts and employees are in scope since the 2023 exemption sunset. The highest-risk research gap is vendor contracts: without six specific service provider terms, disclosing participant data to a research platform, panel, or transcription vendor is legally a sale or sharing that triggers opt-out rights. Voice recordings are biometric information only if an identifier template can be extracted, and sensitive personal information only when actually used to identify a consumer, so ordinary voice research is not automatically sensitive. Enforcement in 2025-26 targeted broken opt-outs, excessive verification, asymmetrical choices, missing vendor terms, and over-collection: Honda $632,500, Healthline $1.55M, a youth sports platform $1.10M. Statutory penalties are $2,663 per unintentional and $7,988 per intentional violation.","aiPrerequisites":["Basic understanding of user research operations","Familiarity with research participant recruitment"],"aiLearningOutcomes":["Determine whether CCPA/CPRA applies to your research programme","Audit vendor contracts for the six required service provider terms","Distinguish ordinary voice research from regulated biometric processing","Present a compliant notice at collection during recruitment","Build deletion workflows that reach transcripts, analysis, and repositories","Apply data minimisation through study design rather than post-hoc cleanup"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"61fd3b49-d9f2-49db-9b67-fde0865802cb","slug":"interview-recording-consent-laws","title":"Interview Recording Consent Laws: One-Party, All-Party, and Biometric Rules (2026)","url":"https://www.koji.so/docs/interview-recording-consent-laws","summary":"US federal law (18 U.S.C. 2511) permits recording with one-party consent, but roughly twelve states require all-party consent: California, Connecticut, Delaware, Florida, Illinois, Maryland, Massachusetts, Montana, New Hampshire, Oregon, Pennsylvania, and Washington - with Connecticut and Oregon applying different rules to phone versus in-person communications. Because participant location cannot be reliably determined, research teams should apply all-party consent universally, captured before recording starts. Biometric laws are separate: Illinois BIPA requires written notice and a written release before collecting voiceprints, with damages of $1,000 negligent and $5,000 intentional per violation. An audio recording is not automatically a voiceprint - the trigger is deriving a template to identify a person.","content":"## The short answer\n\n**Get affirmative, on-the-record consent from every participant before recording starts. Always. Everywhere.**\n\nThat one rule resolves nearly every legal question on this page, because it satisfies the strictest standard that could apply to you. The reason it matters:\n\n- **Federal law** (the Electronic Communications Privacy Act, [18 U.S.C. § 2511](https://www.law.cornell.edu/uscode/text/18/2511), originally the Wiretap Act of 1968) sets a **one-party consent** baseline — recording is permitted if at least one party to the conversation consents. As the person recording, you are that party.\n- **Roughly a dozen states** require **all-party consent** instead, and violations there can be criminal, not merely civil.\n- **Biometric privacy laws** are a completely separate regime. Illinois BIPA requires **written notice and a written release** before collecting a voiceprint — and consent to *record* is not the same as consent to create a **biometric identifier**.\n\n> This guide is written for research and product teams and is not legal advice. State law changes, and several of the distinctions below are genuinely unsettled. Confirm current requirements with counsel before rolling out a recording program — especially a multi-state one.\n\n## One-party vs all-party consent\n\nA **one-party consent** jurisdiction requires only one participant in the conversation to agree to the recording. A **two-party** — more accurately **all-party** — jurisdiction requires everyone.\n\nAs of 2026, the states commonly listed as all-party consent are:\n\n**California, Connecticut, Delaware, Florida, Illinois, Maryland, Massachusetts, Montana, New Hampshire, Oregon, Pennsylvania, and Washington.**\n\nTreat that list as a starting point rather than gospel. Published lists disagree at the margins because several statutes distinguish between *types* of communication, and courts keep refining the boundaries. Two documented examples of exactly that:\n\n- **Connecticut** requires all-party consent for recording **phone calls**, but follows a one-party rule for in-person conversations under its criminal statute.\n- **Oregon** requires all-party consent for **in-person oral** communications, but applies a one-party rule to electronic communications.\n\nThis is precisely why \"which list is right?\" is the wrong question for a research team to spend time on.\n\n## The rule that makes the map irrelevant\n\nIn research you almost never know a participant's physical location with legal certainty. Someone recruited as a New York customer may take the session from a hotel in Seattle. Remote panels cross state lines constantly, and cross-border sessions can pull in foreign law entirely.\n\nSo do not build a compliance program that depends on geolocating participants. **Apply the strictest applicable standard to everyone:**\n\n1. State the recording intent in the invitation, before anyone joins\n2. Capture explicit consent as a discrete step before the session begins\n3. Restate it on the record at the start of the session\n4. Give a genuine way to decline that still lets the person participate — or to withdraw partway through\n5. Log the consent with a timestamp alongside the recording\n\nStep 4 carries more weight than teams expect. Consent that is a precondition to getting paid is not obviously \"freely given,\" which matters a great deal under GDPR and is a fair criticism of many incentive-driven panels. Offer a text-only or notes-only path for people who decline recording.\n\n## Biometrics: the duty most teams miss\n\nHere is the distinction that catches sophisticated teams: **an audio recording is not automatically a voiceprint.**\n\nIllinois BIPA governs **biometric identifiers** — fingerprints, iris scans, face geometry, and **voiceprints**. A voiceprint is a template derived from a voice for the purpose of **identifying** a person. Simply storing an audio file of someone answering questions is generally not that. Running that audio through a system that builds a speaker template to recognize who is talking is.\n\nBIPA's requirements, when it applies:\n\n- **Written notice** that a biometric identifier is being collected, why, and for how long it will be retained\n- A **written release** — obtained before collection\n- A published retention and destruction schedule\n\nThe exposure is significant: statutory damages of **$1,000 per negligent violation** and **$5,000 per intentional or reckless violation**, plus attorney fees and injunctive relief. A 2024 amendment (Public Act 103-769) softened the arithmetic considerably by treating repeated collections of the same biometric from the same person as a **single** violation rather than one per scan — but the per-person exposure remains real.\n\nTwo 2026 developments worth knowing. First, litigation has expanded sharply: in **May 2026**, Illinois voice actors, narrators, podcasters, and journalists filed suits against a long list of major technology and AI companies, framing the claim around the **extraction of voiceprints** rather than copyright — a theory that travels further for class treatment. Second, practitioners have flagged that **passively** building voiceprints from ordinary call audio, with no enrollment step, is arguably *higher* risk than an explicit enrollment flow, because there is no consent moment at all.\n\nA spoken disclosure at the top of a call does not satisfy BIPA. The statute asks for a **written** release — an SMS link, an email confirmation, or a signed portal step.\n\nIllinois is the sharpest instrument, but not the only one: Texas and Washington have their own biometric statutes, and most of the 20 US state comprehensive privacy laws now treat biometric data as **sensitive** personal information requiring opt-in consent.\n\n## Outside the US\n\n**GDPR** does not have a one-party/all-party concept. A recording of an identifiable person is personal data, so you need a lawful basis (usually consent for research), purpose limitation, minimization, a retention limit, and a route for participants to exercise their rights. Voice data processed **to uniquely identify** someone is special category data under Article 9 and requires a stronger basis. See [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research).\n\nFrom **2 August 2026**, the EU AI Act adds a separate transparency duty when people interact with an AI system — a disclosure obligation about the AI itself, independent of anything recording law requires.\n\n## A practical policy you can adopt\n\n| Element | What to do |\n|---|---|\n| Default standard | All-party consent, applied to every participant regardless of location |\n| Timing | Consent captured **before** recording begins, never retroactively |\n| Record of consent | Timestamped, stored with the session |\n| Decline path | Text-based or notes-only participation, no loss of incentive |\n| Voiceprints | Do not create them for research. If a vendor does, demand written notice and release |\n| Speaker identification | Off unless there is a documented need |\n| Retention | Published schedule with automatic deletion — see [research data retention](/docs/research-data-retention-deletion) |\n| Minors | Parental consent plus child assent — see [research with children](/docs/user-research-with-children-teens) |\n| Sensitive topics | Extra care on withdrawal rights — see [trauma-informed research](/docs/trauma-informed-user-research) |\n\n## How Koji is structured around this\n\nThe mechanics of AI-moderated research remove several of the sharpest edges — not because the law differs, but because the workflow does.\n\n- **Participants start the session themselves.** There is no bot silently joining a meeting already in progress, which is the fact pattern generating most of the current recording-consent litigation.\n- **Consent is a structured step before the interview begins**, not a verbal aside a moderator may forget under time pressure. See [intake forms and consent](/docs/intake-forms-and-consent).\n- **Text mode sidesteps voice capture entirely.** For sensitive studies, multi-state panels, or Illinois participants, running text interviews means no audio and therefore no voiceprint question at all. [Voice vs text interviews](/docs/voice-vs-text-interviews) covers the tradeoffs — voice yields richer affect and detail, text yields cleaner compliance and faster review.\n- **Analysis works from the transcript, not from the voice.** Koji derives themes, quality scores, and a sentiment label from what the participant **said** — not from vocal timbre, prosody, or a speaker template. No voiceprints are created for identification and no biometric categorisation is performed.\n- **Structured questions reduce your reliance on inferred signal.** With six question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` — you can measure how strongly someone feels by **asking them** on a scale and letting the AI probe why, rather than inferring it from their voice. The [structured questions guide](/docs/structured-questions-guide) shows how to build that instrument.\n- **Every session produces a verbatim transcript** with a consent record attached, which is exactly the artifact you want if a participant later asks what you hold about them.\n\nConfirm current configuration, sub-processors, and retention settings during procurement — see [enterprise security for AI research platforms](/docs/enterprise-security-ai-research-platforms).\n\n## Common mistakes\n\n1. **Geolocating participants to pick a legal standard.** Fragile, and it fails the moment someone travels. Apply the strictest rule universally.\n2. **Treating a spoken \"is it okay if I record?\" as sufficient for biometrics.** BIPA wants written notice and a written release.\n3. **Assuming a recording is a voiceprint, or that it never is.** The trigger is whether a template is derived to identify someone.\n4. **Making consent a condition of payment.** Undermines the \"freely given\" requirement and reads badly to regulators.\n5. **Leaving speaker identification enabled by default** because a vendor ships it that way.\n6. **No published retention schedule.** BIPA explicitly asks for one, and GDPR storage limitation expects it.\n7. **Recording first and asking later.** In an all-party state this can be a criminal exposure, not a paperwork problem.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — measure intensity by asking rather than inferring from voice\n- [How to Record Customer Interviews](/docs/how-to-record-customer-interviews) — the practical capture-and-transcribe workflow\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the EU lawful-basis picture\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — consent language you can adapt\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — publish the schedule these laws expect\n- [Voice vs Text Interviews](/docs/voice-vs-text-interviews) — when dropping audio is the cleaner choice\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — reduce what you hold in the first place","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Interview Recording Consent Laws: One-Party vs All-Party (2026)","metaDescription":"Federal law permits one-party consent recording, but about a dozen states require all-party consent and Illinois BIPA adds written-consent duties for voiceprints. A practical recording policy for research teams.","keywords":["interview recording consent laws","two party consent states 2026","all party consent recording research","is it legal to record a customer interview","bipa voiceprint consent","wiretap law research recording","recording consent research participants","biometric privacy voice research","one party consent states","call recording laws research"],"aiSummary":"US federal law (18 U.S.C. 2511) permits recording with one-party consent, but roughly twelve states require all-party consent: California, Connecticut, Delaware, Florida, Illinois, Maryland, Massachusetts, Montana, New Hampshire, Oregon, Pennsylvania, and Washington - with Connecticut and Oregon applying different rules to phone versus in-person communications. Because participant location cannot be reliably determined, research teams should apply all-party consent universally, captured before recording starts. Biometric laws are separate: Illinois BIPA requires written notice and a written release before collecting voiceprints, with damages of $1,000 negligent and $5,000 intentional per violation. An audio recording is not automatically a voiceprint - the trigger is deriving a template to identify a person.","aiPrerequisites":["Familiarity with running recorded customer or user interviews","Basic understanding of participant consent workflows"],"aiLearningOutcomes":["Distinguish one-party from all-party consent jurisdictions and know which states require which","Apply a single strictest-standard recording policy instead of geolocating participants","Recognize when audio collection becomes biometric voiceprint collection under BIPA","Build a consent workflow that captures agreement before recording begins","Decide when text-based interviews are the cleaner compliance choice"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"bceabd4c-aede-4355-b9e8-e1e753cd8fde","slug":"research-data-retention-deletion","title":"Research Data Retention and Deletion: How Long Should You Keep Interview Data?","url":"https://www.koji.so/docs/research-data-retention-deletion","summary":"There is no universal legal retention period for research data - GDPR Article 5(1)(e) storage limitation requires keeping personal data only as long as necessary for the collection purpose, which means the absence of a documented schedule is itself the compliance failure. Retention should be decided per artifact, not per study: raw audio 30-90 days after transcription, identifiable transcripts 6-12 months, de-identified transcripts and quotes 24-36 months or indefinitely if genuinely anonymized, analysis outputs indefinitely, consent records as long as the data plus the limitation period. Genuinely anonymized data falls outside GDPR so the retention clock stops, while pseudonymized data remains personal data. Key regimes: COPPA prohibits indefinite retention of children data, Illinois BIPA requires a publicly available destruction schedule for biometric identifiers.","content":"## The short answer\n\n**There is no universal legal retention period for research data. The law requires you to decide, document, and enforce one — so the absence of a schedule is itself the compliance failure.**\n\nUnder GDPR's storage limitation principle (Article 5(1)(e)), personal data may be kept only as long as necessary for the purpose it was collected for. That is a *relative* standard: it obliges you to define the purpose, derive a period from it, and then actually delete on schedule.\n\nThe second key idea, which most teams miss:\n\n**Retention is a per-artifact decision, not a per-study decision.** A single 45-minute interview generates a recording, a transcript, a set of coded themes, a handful of verbatim quotes, a consent record, and a contact record for paying the incentive. Those six artifacts have wildly different risk profiles and wildly different useful lifespans. One blanket rule gets it wrong in both directions at once — holding raw audio far too long while deleting the analysis you actually needed.\n\n> Not legal advice. Retention periods interact with your jurisdiction, sector, and contracts; validate your schedule with counsel.\n\n## The recommended tiered schedule\n\nAdapt the periods, keep the structure. These are defensible defaults, not legal minimums.\n\n| Artifact | Suggested retention | Why |\n|---|---|---|\n| **Raw audio / video** | 30–90 days after transcription | Highest-risk, lowest marginal value once transcribed. The single biggest easy win. |\n| **Identifiable transcripts** | 6–12 months | Long enough to re-analyze and verify; short enough to limit exposure. |\n| **De-identified transcripts and quotes** | 24–36 months, or indefinitely if genuinely anonymized | Where nearly all analytical value lives. |\n| **Analysis outputs — themes, reports, aggregate scores** | Indefinite | Business records containing no personal data. |\n| **Consent records** | As long as you hold the data, plus your limitation period | This is your evidence of lawful basis. Do not delete it with the data. |\n| **Contact details and incentive/payment records** | Per finance and tax rules, in a separate system | Different purpose, different clock — never in the research corpus. |\n| **Voiceprints or other biometric identifiers** | Do not collect for research | If collected, a published destruction schedule is mandatory — see below. |\n\nThe recording tier deserves emphasis. Raw audio is the most sensitive artifact you hold and, once you have an accurate transcript, usually the least useful. Teams keep it out of vague anxiety about needing to \"go back to the tape.\" In practice they almost never do — and meanwhile the recording is the thing that turns a routine security incident into a serious one.\n\n## Anonymization is the lever that lets you keep insight forever\n\nThis is the most valuable mechanic in retention design, and it is widely misunderstood.\n\n- **Pseudonymized** data — real identifiers swapped for codes, with a key that still exists somewhere — is **still personal data**. The retention clock keeps running.\n- **Genuinely anonymized** data — where re-identification is no longer reasonably possible, and no key exists — falls **outside** GDPR's scope. The clock stops.\n\nSo the way to retain research value indefinitely without indefinite risk is not to argue for longer retention periods. It is to **de-identify at the point of synthesis**, so that your durable artifacts — reports, theme libraries, quote banks — never contain personal data in the first place.\n\nBe honest about the bar, though. Qualitative data resists anonymization more than survey data, because narrative detail identifies people. \"The VP of Engineering at a 40-person Berlin fintech who joined last March\" is identifiable no matter what you call them. Strip role-plus-company-plus-timeline combinations, not just names. See [anonymizing customer interview data](/docs/anonymizing-customer-interview-data) for the mechanics.\n\n## What specific regimes require\n\n**GDPR** — storage limitation (Art. 5(1)(e)) plus the right to erasure (Art. 17). You must be able to find and delete one participant's data on request. That is an architecture requirement, not a policy statement: if you cannot locate everything about one person, you cannot comply. See [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research).\n\n**COPPA** — the amended Rule explicitly prohibits retaining a child's personal information indefinitely, limits retention to what is reasonably necessary for the collection purpose, and expects a written, published retention policy. See [research with children and teens](/docs/user-research-with-children-teens).\n\n**Illinois BIPA** — requires a **publicly available** written retention schedule and destruction guidelines for biometric identifiers, with destruction when the initial purpose is satisfied or within three years of the individual's last interaction, whichever comes first. The simplest compliance posture for a research team is to not create voiceprints at all — see [interview recording consent laws](/docs/interview-recording-consent-laws).\n\n**US state privacy laws** — the comprehensive laws now in effect across roughly 20 states generally require disclosing retention practices and honoring deletion requests, with sensitive categories attracting opt-in consent.\n\n**Sector and contract terms** frequently override all of the above. Healthcare, financial services, and enterprise DPAs often specify periods directly; your customer's contract may be stricter than any statute.\n\n## Handling a deletion request without losing the finding\n\nThe hard case: a participant asks you to delete their data six months after their quote landed in a report that shaped your roadmap.\n\nWhat must go: the recording, the transcript, identifiers, the linkage between person and response.\n\nWhat generally survives: **aggregate findings and genuinely de-identified insights.** Deleting a source does not obligate you to unlearn a conclusion. If four of twelve participants struggled with the same step, the finding \"a third of participants struggled here\" remains valid and contains no personal data.\n\nWhere teams get stuck is verbatim quotes in published decks. A distinctive quote can identify its speaker. Two ways to avoid the problem entirely:\n\n1. **De-identify quotes at synthesis time**, so nothing downstream needs revisiting.\n2. **Keep a quote-to-source map in one place only** — the research platform — so a deletion request has a single point of execution rather than a scavenger hunt across Notion, Slack, decks, and someone's laptop.\n\nThat second point is the real argument for keeping research data in one system. Sprawl is what makes deletion requests unanswerable.\n\n## Writing the policy: six elements\n\n1. **Scope** — which artifacts, which systems, which studies\n2. **A period per artifact tier**, with the reasoning recorded\n3. **The trigger** — is the clock from collection date, study close, or last participant interaction?\n4. **The mechanism** — automatic deletion beats manual cleanup, which never happens\n5. **Exceptions** — legal hold, active dispute, contractual requirement, and who authorizes them\n6. **An owner and a review cadence** — annually, with a named accountable person\n\nThen do the thing almost nobody does: **run a deletion drill.** Pick one past participant and try to delete everything about them. Whatever you cannot find is your actual compliance gap, and it is invariably a spreadsheet or a slide deck rather than the research platform.\n\n## How this works with Koji\n\n- **The transcript is the durable artifact, and the recording does not have to be.** Because every session yields a complete verbatim transcript, deleting raw audio early costs you very little analytically — which makes the shortest, highest-value retention tier genuinely practical.\n- **Structured questions produce findings that survive deletion.** With the six question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` — quantitative results aggregate into distributions and rankings that carry no personal data. When a participant exercises erasure, your scale distributions and choice frequencies remain intact because they were never personal data to begin with. That is a real structural advantage over a pile of recordings, and the [structured questions guide](/docs/structured-questions-guide) covers building studies that way.\n- **Reports are synthesized outputs**, so the insight layer you keep long-term is separable from the identifiable layer you delete on schedule.\n- **Collect less up front.** The cheapest retention policy is a study scoped so it never gathers identifiers — no full names, no employer, no free-text field inviting people to volunteer personal details.\n- **Text mode removes the audio tier entirely** for sensitive studies. See [voice vs text interviews](/docs/voice-vs-text-interviews).\n- **Export before deletion.** Take the de-identified analysis into your repository so the schedule never costs you institutional memory — see the [research repository guide](/docs/research-repository-guide).\n\nConfirm current storage locations, sub-processors, and configurable retention settings during procurement, and record the answers in your policy — see [enterprise security for AI research platforms](/docs/enterprise-security-ai-research-platforms) and [AI interview data privacy and security](/docs/ai-interview-data-privacy-security).\n\n## Common mistakes\n\n1. **No written schedule.** The most common failure, and the one that turns any incident or audit into an unnecessary problem.\n2. **One blanket period for everything.** Different artifacts, different risk, different value.\n3. **Keeping raw recordings forever** because deleting feels irreversible. Transcribe, verify, then delete.\n4. **Confusing pseudonymization with anonymization.** Only genuine anonymization stops the clock.\n5. **Deleting consent records along with the data.** You need them to show your lawful basis.\n6. **Manual cleanup.** If deletion is a calendar reminder, it will not happen. Automate it.\n7. **Data sprawl across decks, spreadsheets, and Slack**, making erasure requests impossible to fulfill honestly.\n8. **Mixing incentive and payment records into the research corpus.** Different purpose, different clock, different system.\n9. **Never testing the policy.** Run the deletion drill.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — build studies whose findings survive deletion\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — the de-identification mechanics that stop the clock\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — storage limitation and erasure in practice\n- [Interview Recording Consent Laws](/docs/interview-recording-consent-laws) — why biometric artifacts need a published schedule\n- [User Research With Children and Teens](/docs/user-research-with-children-teens) — the stricter COPPA retention rule\n- [Research Repository Guide](/docs/research-repository-guide) — keep the insight after the data is gone\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — what to verify during procurement\n- [IRB Approval for User Research](/docs/irb-approval-user-research) — the data-management plan reviewers expect","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Research Data Retention and Deletion: How Long to Keep Interview Data","metaDescription":"No universal legal retention period exists - so having no schedule is the compliance failure. A tiered per-artifact schedule for recordings, transcripts, quotes and reports, plus how to handle deletion requests without losing insights.","keywords":["research data retention policy","how long to keep interview data","research data deletion schedule","gdpr storage limitation research","interview recording retention","participant data deletion request","research data management plan","anonymization retention clock","ux research data policy","deletion request research data"],"aiSummary":"There is no universal legal retention period for research data - GDPR Article 5(1)(e) storage limitation requires keeping personal data only as long as necessary for the collection purpose, which means the absence of a documented schedule is itself the compliance failure. Retention should be decided per artifact, not per study: raw audio 30-90 days after transcription, identifiable transcripts 6-12 months, de-identified transcripts and quotes 24-36 months or indefinitely if genuinely anonymized, analysis outputs indefinitely, consent records as long as the data plus the limitation period. Genuinely anonymized data falls outside GDPR so the retention clock stops, while pseudonymized data remains personal data. Key regimes: COPPA prohibits indefinite retention of children data, Illinois BIPA requires a publicly available destruction schedule for biometric identifiers.","aiPrerequisites":["An existing research practice that collects interviews, recordings, or survey responses","Basic familiarity with GDPR or comparable privacy obligations"],"aiLearningOutcomes":["Build a tiered per-artifact retention schedule instead of one blanket period","Understand why genuine anonymization stops the GDPR retention clock and pseudonymization does not","Handle a participant deletion request without discarding valid aggregate findings","Write a retention policy with the six required elements and test it with a deletion drill","Scope studies to collect fewer identifiers so retention obligations shrink"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"faecdc6f-bdc0-4312-911a-4ef3c7be9a55","slug":"research-data-residency-international-transfers","title":"Research Data Residency and International Transfers: Where Your Interview Data Actually Lives","url":"https://www.koji.so/docs/research-data-residency-international-transfers","summary":"Data residency is where data is stored, data sovereignty is whose law governs it and who can compel access, and data localisation is a statutory requirement to keep data in country. Under GDPR, any transfer outside the EEA is restricted and needs a Chapter V mechanism: adequacy decision, the EU-US Data Privacy Framework for certified US recipients, Standard Contractual Clauses under Implementing Decision 2021/914, Binding Corporate Rules, or Article 49 derogations (occasional transfers only). Since Schrems II, SCCs require a transfer impact assessment plus supplementary measures; pseudonymisation before transfer is the highest-value measure for research. The DPF remains valid but the Latombe appeal is pending at the CJEU and PCLOB lost quorum in 2025, so pair DPF reliance with SCCs. Research is exposed because transcripts contain incidental identifiers and the processing chain spans recruitment, interviewing, transcription, LLM analysis, and storage. Ask vendors about storage and processing countries, sub-processors, transfer mechanism, model training, deletion SLA, and pseudonymisation support. Koji shortens the chain, supports anonymisation before sharing, allows deletion at any time, and its six structured question types reduce how much narrative data crosses any border.","content":"## The short answer\n\n**\"Where is our data stored?\" is the wrong first question. The right one is \"who can lawfully compel access to it, and under which country's law?\"** Storage location is one input to that answer, not the answer itself. A dataset sitting in a Frankfurt data centre operated by a company subject to another country's disclosure laws is not automatically protected by its postcode.\n\nFor research teams, three distinct concepts get collapsed into one word:\n\n- **Data residency** — the geographic location where data is stored at rest. Usually a contractual commitment you choose.\n- **Data sovereignty** — whose laws govern that data, including who can compel disclosure. Follows the operator and the legal entity, not just the server.\n- **Data localisation** — a legal requirement that data must stay in a country. Imposed by statute, not chosen.\n\nUnder GDPR, moving personal data outside the EEA is a **restricted transfer** and needs a Chapter V mechanism regardless of where you store it. Getting this right takes an afternoon. Getting it wrong surfaces during a security review, three weeks before you needed the research.\n\n## Why this lands on research teams\n\nInterview data is unusually exposed. Survey responses are short and often anonymous; interview transcripts are long, narrative, and full of incidental identifiers — a person naming their employer, their manager, a medical condition, a customer account. Recorded voice adds another layer, since in some jurisdictions voiceprints are biometric data with their own rules.\n\nMeanwhile the processing chain is long. A single study can touch a recruitment tool, an interview platform, a transcription service, an LLM analysis provider, and a repository — potentially in four countries. **Your transfer analysis has to cover the whole chain, not just the platform you bought.** The most common finding in a research security review is not that the vendor is in the wrong country; it is that nobody mapped the sub-processors.\n\n## GDPR Chapter V: the transfer mechanisms\n\n| Mechanism | When to use | Practical notes |\n|---|---|---|\n| **Adequacy decision** (Art 45) | Destination country recognised by the European Commission | Simplest route. Covers the UK, Switzerland, Canada (commercial), Japan, South Korea, New Zealand, and others |\n| **EU-US Data Privacy Framework** | US recipient that has self-certified | An adequacy decision adopted 10 July 2023. Only covers **certified** recipients — verify the certification, do not assume it |\n| **Standard Contractual Clauses** (Art 46) | The default for most vendors | Commission Implementing Decision 2021/914; four modules for different controller/processor relationships. Requires a transfer impact assessment |\n| **Binding Corporate Rules** | Intra-group transfers in large organisations | Slow to approve; rarely relevant to a research purchase |\n| **Article 49 derogations** | Occasional, non-repetitive transfers | Explicit consent or contractual necessity. **Not** a basis for systematic, ongoing research operations |\n\n### A note on the Data Privacy Framework's stability\n\nThe DPF remains valid law and is the simplest route for certified US vendors, but it is under active challenge. The EU General Court dismissed the *Latombe* action on 3 September 2025, upholding the framework; that ruling was appealed to the Court of Justice on 31 October 2025 and the appeal is still pending. Separately, the US Privacy and Civil Liberties Oversight Board lost its quorum in January 2025 when three of its five members were removed — and PCLOB oversight was one of the safeguards the Commission relied on in finding US protections adequate.\n\nNone of this means you should avoid the DPF. It means you should **not build a research programme whose only transfer mechanism is the DPF.** The resilient pattern is DPF certification *plus* SCCs as fallback in the same contract. That way an adverse ruling is a paperwork event rather than a fieldwork stoppage. Ask your vendor directly whether their DPA includes SCCs alongside any DPF reliance.\n\n### Transfer impact assessments\n\nSince *Schrems II* (2020), relying on SCCs is not enough on its own. You must assess whether the destination country's law undermines the protection the clauses promise, and add supplementary measures where it does — encryption with keys held in the EEA, pseudonymisation before transfer, or contractual commitments to challenge and report access requests.\n\nFor research specifically, the highest-value supplementary measure is **pseudonymisation before transfer**. If names, employers, and direct identifiers are stripped and held separately in the EEA, what crosses the border is far less sensitive and your TIA becomes considerably easier to write.\n\n## Beyond the EU\n\n- **UK GDPR** — separate regime. Transfers out of the UK use the International Data Transfer Agreement or the UK Addendum to the EU SCCs. The UK's own adequacy status under EU law has been extended rather than made permanent and is subject to periodic review, so confirm the current position rather than assuming it holds.\n- **Switzerland** — its own framework and a Swiss-US arrangement; the EU SCCs need Swiss-specific amendments.\n- **Brazil (LGPD)** — transfers require adequacy, contractual clauses, or specific consent; the ANPD has issued its own standard clauses.\n- **Canada (PIPEDA)** — no localisation requirement federally, but transfers require comparable protection and transparency to individuals about offshore processing. Some provincial public-sector rules are stricter.\n- **India (DPDP Act)** — a blocklist model rather than an allowlist: transfers permitted except to countries the government restricts.\n- **China (PIPL)** — genuinely restrictive, with security assessments, standard contracts, or certification required, plus localisation duties for some operators. Treat China research as a separate project with local counsel.\n\n**Practical implication:** if you run global research, do not try to satisfy every regime with one policy. Segment by participant region, apply the strictest applicable rule per segment, and document the segmentation.\n\n## Six questions to ask any research vendor\n\nSend these before the security review, not during it:\n\n1. **In which countries is interview data stored at rest, and in which is it processed?** These are different answers.\n2. **List every sub-processor** that touches participant data — transcription, LLM inference, storage, email. Which are on your DPA's sub-processor list?\n3. **Which transfer mechanism applies**, and does the DPA include SCCs *in addition to* any adequacy or DPF reliance?\n4. **Is data used to train models?** For a research platform, the answer should be no for customer content. Get it in the contract, not the FAQ.\n5. **What is the deletion path and its SLA**, including backups?\n6. **Can you support pseudonymisation before data leaves our control?**\n\nQuestion 4 is the one that changed most in 2026. Model-training terms move faster than security pages; verify against the current DPA rather than a blog post.\n\n## How Koji reduces the surface area\n\nMost transfer risk in research is created by the *shape* of the data, not the route it travels. Koji is built so that less sensitive data exists in the first place.\n\n**Structured questions collect bounded values.** Koji supports six question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. A `scale`, `single_choice`, or `yes_no` answer is a value, not a narrative that might contain a name, a diagnosis, or a customer's identity. Every question you convert reduces what crosses any border. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Anonymisation before sharing.** Transcripts can be anonymised before they are shared or exported, which is exactly the supplementary measure a transfer impact assessment asks for. See [anonymizing customer interview data](/docs/anonymizing-customer-interview-data).\n\n**A short chain.** Recruitment link, AI-moderated interview, automatic transcription, automatic analysis, report — inside one platform. Compare that with the usual stack of a scheduler, a video tool, a separate transcription vendor, an LLM analysis add-on, and a repository, each with its own sub-processors and its own transfer analysis. Fewer hops means a shorter list to assess.\n\n**Deletion you control.** Interview data, transcripts, and analysis belong to your account, and studies and their associated data can be deleted at any time — the concrete evidence a deletion SLA needs.\n\n**Text-only mode as a mitigation.** Where voice recording raises biometric questions in a given jurisdiction, running the study as a text interview removes the issue entirely while keeping the conversational depth and AI follow-up probing. That option does not exist in a video-based research stack.\n\nFor teams with contractual residency requirements or a specific enterprise data-handling posture, contact the Koji team to discuss enterprise options rather than inferring the answer from a docs page.\n\n## Five common mistakes\n\n1. **Treating storage location as the whole answer.** Sovereignty follows the operator and applicable law too.\n2. **Mapping the platform but not the sub-processors.** The transcription and inference layers are where surprises live.\n3. **Relying on Article 49 consent for ongoing operations.** The derogations are for occasional transfers, not a running research programme.\n4. **Relying on the DPF alone.** Pair it with SCCs so a court ruling is a paperwork problem, not an outage.\n5. **Doing the analysis once.** Vendors change sub-processors and regions. Re-check at renewal.\n\n## Related Resources\n\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — SOC 2, SSO, and the vendor review process\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — lawful basis and participant rights\n- [DPIA for User Research](/docs/dpia-user-research) — the risk assessment that references your transfer analysis\n- [AI Interview Data Privacy and Security](/docs/ai-interview-data-privacy-security) — the buyer's evaluation checklist\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — pseudonymisation as a supplementary measure\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — how long data should exist anywhere\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and data minimisation by design","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Research Data Residency & International Transfers Guide (2026)","metaDescription":"Data residency vs sovereignty vs localisation for research teams: GDPR Chapter V mechanisms, SCCs, transfer impact assessments, the DPF's current status, and six questions to ask any vendor.","keywords":["research data residency","international data transfers gdpr","standard contractual clauses research","transfer impact assessment","eu us data privacy framework 2026","data sovereignty vs residency","where is interview data stored","gdpr chapter v transfers","cross border research data","vendor sub-processor review"],"aiSummary":"Data residency is where data is stored, data sovereignty is whose law governs it and who can compel access, and data localisation is a statutory requirement to keep data in country. Under GDPR, any transfer outside the EEA is restricted and needs a Chapter V mechanism: adequacy decision, the EU-US Data Privacy Framework for certified US recipients, Standard Contractual Clauses under Implementing Decision 2021/914, Binding Corporate Rules, or Article 49 derogations (occasional transfers only). Since Schrems II, SCCs require a transfer impact assessment plus supplementary measures; pseudonymisation before transfer is the highest-value measure for research. The DPF remains valid but the Latombe appeal is pending at the CJEU and PCLOB lost quorum in 2025, so pair DPF reliance with SCCs. Research is exposed because transcripts contain incidental identifiers and the processing chain spans recruitment, interviewing, transcription, LLM analysis, and storage. Ask vendors about storage and processing countries, sub-processors, transfer mechanism, model training, deletion SLA, and pseudonymisation support. Koji shortens the chain, supports anonymisation before sharing, allows deletion at any time, and its six structured question types reduce how much narrative data crosses any border.","aiPrerequisites":["Familiarity with GDPR basics","Knowledge of your current research tool stack"],"aiLearningOutcomes":["Distinguish data residency, sovereignty, and localisation","Select the correct GDPR Chapter V transfer mechanism","Write a transfer impact assessment with proportionate supplementary measures","Map sub-processors across the full research chain","Evaluate a research vendor with six targeted questions"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"64943f27-30cf-4e04-b33a-701411ada6ac","slug":"ai-governance-frameworks-research","title":"AI Governance for Customer Research: ISO 42001, the NIST AI RMF, and What Procurement Actually Asks","url":"https://www.koji.so/docs/ai-governance-frameworks-research","summary":"Procurement conflates three different things: the EU AI Act (binding law), ISO/IEC 42001 (a voluntary certifiable AI management system standard), and the NIST AI RMF (a voluntary framework). The key correction is that ISO 42001 is not a harmonised standard under the AI Act, so certification carries no presumption of conformity — prEN 18286 is the deliverable intended for that role. A second correction: NIST AI RMF was never legally mandatory, and guidance citing Executive Order 14110 is stale because that order was revoked in January 2025. Includes five overlapping control areas to build once, and seven checkable questions to ask an AI research vendor.","content":"## The short answer\n\n**ISO/IEC 42001 certification does not make you compliant with the EU AI Act, and the NIST AI RMF has never been a law.** Both are worth having. Neither is what most procurement questionnaires assume it is.\n\nIf you are buying or defending an AI-moderated research tool in 2026, three different things get conflated in the same email thread:\n\n| | What it is | Legal force | What it gives you |\n|---|---|---|---|\n| **EU AI Act** | Regulation | Binding law in the EU | Obligations you must meet, with penalties |\n| **ISO/IEC 42001:2023** | Certifiable management system standard | Voluntary | A third-party audited certificate that you govern AI systematically |\n| **NIST AI RMF 1.0** | Voluntary framework | None | A shared vocabulary and structure for AI risk work |\n\nGetting this hierarchy right saves a research team weeks. It also stops you paying for a certificate to solve a problem the certificate does not solve.\n\n## ISO/IEC 42001: the certifiable one\n\nPublished in December 2023, ISO/IEC 42001 is the first international standard for an **AI management system** (AIMS). It is structured like ISO 27001 — the same Annex SL management-system shape — but its subject is AI: how you inventory AI systems, assess their impact on the people they touch, assign accountability, and monitor behaviour across the lifecycle.\n\nCertification works the way ISO certification always works. An accredited certification body runs a Stage 1 (documentation readiness) and Stage 2 (implementation) audit. The certificate is valid for three years, with annual surveillance audits in between. Published cost and timeline estimates vary widely by scope and organisation size; consultancies commonly quote figures in the tens of thousands of dollars and four to nine months of preparation. Treat those as directional, not as a quote.\n\nTwo supporting standards arrived in 2025 and are worth knowing by name, because sophisticated buyers cite them:\n\n- **ISO/IEC 42005:2025** — guidance for conducting AI system impact assessments. It is guidance, not a certifiable requirement, and it complements 42001 with lifecycle-continuous assessment practice.\n- **ISO/IEC 42006:2025** — requirements for the bodies that audit and certify AIMS. This one matters indirectly: it is what makes one vendor's 42001 certificate comparable to another's.\n\n## The presumption-of-conformity trap\n\nHere is the point that most vendor questionnaires get wrong, and the single most useful thing to know in this whole area.\n\nUnder Article 40 of the EU AI Act, applying a **harmonised standard** — one cited in the Official Journal of the EU — gives you a presumption of conformity with the corresponding AI Act requirements. That is a real, valuable legal shortcut.\n\n**ISO/IEC 42001 is not a harmonised standard.** A 42001 certificate therefore carries no presumption of conformity with the AI Act. The European deliverable being developed to serve that role is **prEN 18286**; until it is cited in the Official Journal, nobody has the shortcut. Neither 42001 nor 42005 was designed to operationalise the full set of AI Act obligations in the first place.\n\nSo when a security reviewer writes \"we require ISO 42001 certification to satisfy the EU AI Act,\" the honest answer is that the two are complementary but not substitutes: the certificate is credible evidence of governance maturity, and the AI Act obligations still have to be met on their own terms.\n\nFor customer research specifically, those obligations are usually lighter than people fear. AI-moderated interviews sit in the AI Act's limited-risk transparency tier, where the core duty is to tell participants they are talking to an AI before the conversation starts. Our [EU AI Act guide](/docs/eu-ai-act-user-research-compliance) walks through the tiering and the narrow cases that escalate it.\n\n## NIST AI RMF: useful framework, frequently mis-cited\n\nNIST released the AI Risk Management Framework 1.0 in January 2023, organised around four functions — **GOVERN, MAP, MEASURE, MANAGE** — plus a companion Generative AI Profile (NIST AI 600-1) published in July 2024.\n\nIt is genuinely good structure, and it is free. But be careful with claims about its legal status. Much of the guidance still circulating online states that federal agencies are directed to align with the AI RMF under Executive Order 14110. **EO 14110 was revoked on 20 January 2025**, and replaced days later by a different executive order with a different posture. The framework itself is unaffected — NIST still publishes it, and enterprises still use it — but it remains what it always was: voluntary. If a vendor tells you AI RMF alignment is legally mandatory, they are working from stale material.\n\nWhat is true is that AI RMF vocabulary has become the lingua franca of enterprise AI questionnaires. Answering in its terms (here is our GOVERN evidence, here is our MEASURE evidence) makes a review go faster whether or not anyone is certified.\n\n## The overlap you can build once\n\nAcross the AI Act, ISO 42001, and the NIST AI RMF, the same five control areas keep appearing:\n\n1. **Risk and impact documentation** — what the system does, who it affects, what could go wrong\n2. **Data governance** — provenance, minimisation, retention, quality\n3. **Human oversight** — who can intervene, and how\n4. **Incident monitoring** — detection, logging, escalation\n5. **Transparency documentation** — disclosure to affected people, and system documentation for reviewers\n\nBuild the evidence once, tag each artefact against all three frameworks, and most questionnaires answer themselves.\n\n## What to actually ask an AI research vendor\n\nSkip the framework name-dropping and ask questions whose answers are checkable:\n\n- **Which model providers process interview data, and are they named as sub-processors?** AI interview tools route text and audio to model and transcription providers. Unnamed sub-processors are the real risk, not the absence of a certificate.\n- **Is participant data used to train models?** Get this in the contract, not in a sales email.\n- **How is AI disclosure handled in the participant flow?** Under the AI Act this is the operative obligation for research. Ask to see the actual screen.\n- **What human oversight exists over the AI interviewer?** Can a researcher review, correct, and re-run analysis?\n- **What is logged, and for how long?** This is where AI governance and your retention schedule meet — see [research data retention and deletion](/docs/research-data-retention-deletion).\n- **Do you hold ISO 42001, or have a dated roadmap?** A credible roadmap beats a vague yes. Ask which certification body and what scope — scope is where certificates get thin.\n- **What is your AI incident process?** Ask for one worked example.\n\nAsk the same seven of every vendor on the shortlist, including Koji. Answers that vary in specificity tell you more than answers that vary in confidence.\n\n## How Koji's architecture changes the governance conversation\n\nTwo structural properties of Koji make several of these questions easier to answer than they are for the tools research teams are migrating from.\n\n**Interviews are asynchronous, link-based, and transcript-first.** There is no third-party meeting recorder in the chain, which removes a sub-processor and the consent trail that comes with it. Participants receive the AI disclosure in the flow before the conversation begins, so the AI Act transparency duty is satisfied by the product's default path rather than by researcher discipline.\n\n**Analysis runs on transcripts, not on biometric signals.** This distinction does real work under the AI Act: emotion recognition obligations attach to *biometric* inference, so analysing what someone said sits in a materially different place from inferring affect from vocal tone. A transcript-first architecture keeps you out of the harder tier by design rather than by configuration.\n\nOn the research-quality side, Koji's six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — matter more for governance than they first appear. Structured questions produce predictable, typed data with a stable question ID from interview plan through to report. When a reviewer asks what data you collect and how it is processed, \"these six typed fields plus transcript\" is a far easier answer than \"free-text, variably\" — and it is the same property that makes deletion and access requests tractable.\n\n## Common mistakes\n\n- **Buying a certificate to answer a regulation.** ISO 42001 is evidence of governance maturity, not a conformity shortcut under the AI Act.\n- **Citing EO 14110.** It was revoked in January 2025. Cite the framework, not the executive order.\n- **Accepting \"we're SOC 2\" as an AI governance answer.** SOC 2 addresses information security controls. It says nothing about model behaviour, training use, or human oversight. Both matter; they are not interchangeable.\n- **Ignoring certificate scope.** A 42001 certificate covering a parent company's internal tooling tells you little about the product you are buying.\n- **Treating AI governance as a one-time procurement gate.** All three frameworks assume ongoing monitoring, and the EDPB expects controllers to re-verify processing chains over time.\n\n## Frequently asked questions\n\n**Does ISO 42001 certification make me EU AI Act compliant?**\nNo. ISO/IEC 42001 is not a harmonised standard under the Act, so a certificate carries no presumption of conformity. The European deliverable intended for that role is prEN 18286. The certificate is still strong evidence of governance maturity in a vendor review.\n\n**Is the NIST AI RMF mandatory?**\nNo. It is and always has been a voluntary framework. Guidance claiming a federal mandate typically traces to Executive Order 14110, which was revoked on 20 January 2025.\n\n**Do we need ISO 42001 to run AI-moderated customer research?**\nNo. Customer research generally sits in the AI Act's limited-risk transparency tier, where the operative duty is disclosing to participants that they are interacting with an AI. Certification becomes relevant when your own enterprise buyers demand it of you.\n\n**What is the difference between ISO 42001 and ISO 42005?**\n42001 is the certifiable management system standard. 42005 is guidance for conducting AI system impact assessments and is not certifiable; it complements 42001.\n\n**What is the difference between SOC 2 and ISO 42001 for an AI vendor?**\nSOC 2 attests to information security controls — access, encryption, availability. ISO 42001 attests to how the organisation governs AI systems specifically. Enterprise buyers increasingly want both, for different reviewers.\n\n**How should a small research team answer an AI governance questionnaire?**\nAnswer in NIST AI RMF terms even without certification: document what the AI does, who it affects, what human oversight exists, what is logged, and what you disclose to participants. That covers most of what the questionnaire is reaching for.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and the typed, predictable data they produce\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) — the binding obligations, and which tier research falls into\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — SOC 2, SSO, and the wider vendor security review\n- [AI Interview Data Privacy & Security](/docs/ai-interview-data-privacy-security) — how interview data is protected end to end\n- [DPIA for User Research](/docs/dpia-user-research) — when an impact assessment is legally required, and how to write it\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — the logging and retention half of AI governance\n- [User Research for AI Products](/docs/user-research-for-ai-products) — researching AI features, as distinct from governing them","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"AI Governance for Research: ISO 42001, NIST AI RMF, EU AI Act (2026)","metaDescription":"ISO 42001 is not a harmonised standard, so it buys no EU AI Act presumption of conformity. What each AI governance framework covers, and the seven questions to ask a research vendor.","keywords":["ai governance customer research","iso 42001 certification","nist ai rmf","eu ai act presumption of conformity","pren 18286","iso 42005","ai vendor questionnaire","ai management system","ai governance framework comparison","responsible ai research platform"],"aiSummary":"Procurement conflates three different things: the EU AI Act (binding law), ISO/IEC 42001 (a voluntary certifiable AI management system standard), and the NIST AI RMF (a voluntary framework). The key correction is that ISO 42001 is not a harmonised standard under the AI Act, so certification carries no presumption of conformity — prEN 18286 is the deliverable intended for that role. A second correction: NIST AI RMF was never legally mandatory, and guidance citing Executive Order 14110 is stale because that order was revoked in January 2025. Includes five overlapping control areas to build once, and seven checkable questions to ask an AI research vendor.","aiDifficulty":"intermediate","aiEstimatedTime":"12 min"},{"type":"documentation","id":"f52c73f4-b762-45d7-bad9-f20cd8a18891","slug":"us-state-privacy-laws-research","title":"US State Privacy Laws for Customer Research: The Multi-State Compliance Guide (2026)","url":"https://www.koji.so/docs/us-state-privacy-laws-research","summary":"Twenty US states have comprehensive consumer privacy laws on the books as of early 2026. The decisive scope question for research is whether a participant is a consumer: every comprehensive state law except California defines a consumer as a resident acting in an individual or household context and excludes people acting in a commercial or employment context, so B2B research participants and employees fall outside those laws. California is the exception because the CCPA covers B2B contacts and employees after the partial exemptions expired on 1 January 2023. Key 2026 dates: Maryland MODPA in effect 1 October 2025 and enforceable April 2026; Indiana, Kentucky and Rhode Island effective 1 January 2026 along with Oregon amendments, expanded California data broker and health data rules, the Nebraska Age-Appropriate Design Code and the Texas Responsible AI Governance Act; further changes in Connecticut, Arkansas and Utah on 1 July 2026; new California data broker registration obligations on 1 August 2026. Rhode Island and Maryland apply at 35,000 consumers against Oregon 100,000. Most laws require opt-in consent for sensitive data including health, race, religion, sexual orientation, immigration status, precise geolocation and biometric data processed to identify someone, and open-ended interviews collect such data accidentally. Maryland bans the sale of sensitive personal data outright regardless of consent and requires collection to be reasonably necessary and proportionate, which consent cannot cure. Voice recordings analysed for content are not ordinarily biometric data because identification is not the purpose. Disclosures to a processor or service provider under contract terms covering documented instructions, no independent use, no onward sale, confidentiality, deletion or return and sub-processor flow-down are not sales. Universal opt-out mechanisms such as Global Privacy Control are honoured in around a dozen states but are a website and advertising obligation rather than a research one. The practical answer is a single research process built to the strictest provision in each dimension rather than twenty state variants.","content":"Two facts reshape most research teams' compliance work once they are understood. First, **outside California, your B2B research participants are almost certainly not \"consumers\" under any state privacy law** — every other comprehensive state law excludes individuals acting in a commercial or employment context. Second, when the laws do apply, the provisions that bite in research are narrow and predictable: opt-in consent for sensitive data, deletion and access rights over transcripts, data minimisation, and whether handing recordings to a vendor counts as a sale.\n\nEverything else in the state privacy landscape — universal opt-out signals, targeted advertising opt-outs, data broker registration — is a website and adtech problem, not a research problem. Knowing which is which is what keeps a research programme from being reviewed as if it were a marketing pixel.\n\n## The 2026 landscape\n\n**Twenty states have comprehensive consumer privacy laws on the books** as of early 2026, and the count keeps rising as each legislative session closes; some trackers already count higher after the mid-2026 wave. The dates that matter for anyone re-papering their process this year:\n\n| Date | What changed |\n|---|---|\n| 1 Oct 2025 | Maryland Online Data Privacy Act in effect (enforceable from April 2026) |\n| 1 Jan 2026 | Indiana, Kentucky and Rhode Island comprehensive laws take effect |\n| 1 Jan 2026 | Oregon amendments (HB 2008); California expanded data broker and health data rules; Nebraska Age-Appropriate Design Code |\n| 1 Jan 2026 | Texas Responsible AI Governance Act (HB 149) takes effect |\n| 1 Jul 2026 | Further changes land in Connecticut, Arkansas and Utah |\n| 1 Aug 2026 | New California data broker registration obligations |\n\nRhode Island is worth flagging because its applicability threshold is unusually low — processing the data of at least 35,000 consumers, or 10,000 if more than 20% of revenue comes from selling personal data. Maryland matches the 35,000 figure, against Oregon's 100,000. **Threshold shopping is not a strategy**; if you operate nationally you will cross a threshold somewhere.\n\n## Step 1 — Is your research participant a \"consumer\" at all?\n\nThis is the question to answer first, because it disposes of most B2B research entirely.\n\nEvery comprehensive state law except California's defines \"consumer\" as a resident **acting in an individual or household context**, and expressly excludes people acting in a commercial or employment context. Colorado, Connecticut, Utah and Virginia all take this approach, and the newer laws follow it.\n\nCalifornia is the exception, and it is a total one. The CCPA as amended defines a consumer as any California resident — which includes employees, job applicants and B2B contacts. The partial exemptions for employee and B2B data expired on **1 January 2023** and have not returned.\n\nWhat this means concretely:\n\n- Interviewing a procurement manager at a customer account about your product, in her professional capacity: outside the scope of every state law except California's.\n- Interviewing that same person about her personal banking app: a consumer, everywhere.\n- Interviewing your own employees about an internal tool: outside every state law except California's.\n- Interviewing a sole trader about the tools she uses to run her business: genuinely ambiguous, and worth treating as in-scope.\n\nTwo cautions before you relax. This analysis governs the **state privacy laws only** — recording consent laws, contractual obligations to your customers, sectoral rules like HIPAA, and GDPR for anyone in Europe all apply on their own terms. And if any of your participants are California residents, you are in scope regardless, which for most US-national research means designing to the California standard anyway.\n\n## Step 2 — Recognise the sensitive data you did not mean to collect\n\nNearly every state law requires **opt-in consent before processing sensitive data**, and Virginia, Connecticut, Colorado, Indiana, Kentucky and Rhode Island all take that approach. Sensitive categories typically include health conditions, racial or ethnic origin, religious beliefs, sexual orientation, citizenship or immigration status, precise geolocation, and biometric data processed to identify someone.\n\nThe research problem is not that teams deliberately collect sensitive data. It is that **open-ended interviews collect it accidentally.** Ask a customer why she cancelled and she may tell you about a cancer diagnosis. Ask about a missed payment and you may hear about a divorce and an immigration status. None of that was on your discussion guide, and all of it is now in your transcript.\n\nThree practical controls:\n\n1. **Say so in the consent.** State that the interview may touch on personal circumstances, that the participant should share only what they are comfortable with, and what happens to the recording.\n2. **Redact on ingest, not at report time.** Anonymisation applied when you write the summary leaves the raw transcript sitting in your system.\n3. **Never let sensitive data leave in a \"sale\".** Maryland goes furthest here: MODPA **bans the sale of sensitive personal data outright, regardless of consent** — the strictest sensitive-data rule in any US state law. If your process cannot guarantee that, design it so sensitive data never enters the flow that could be characterised as a sale.\n\n**On voice recordings specifically:** most state definitions treat biometric data as data processed *for the purpose of uniquely identifying* an individual. A voice recording captured to be transcribed and analysed for content is not ordinarily biometric processing, because identification is not the purpose. That is a meaningful distinction for any team running voice interviews — but write the purpose down explicitly, keep voiceprint-style matching out of your stack, and check Maryland separately, since its biometric and consumer health data definitions are stricter than most.\n\n## Step 3 — Data minimisation is now a design constraint\n\nOlder state laws tied minimisation to disclosed purposes. Maryland changed the shape of the obligation: collection must be **reasonably necessary and proportionate**, and consent does not cure over-collection.\n\nFor research that argues against several common habits:\n\n- Recording video when the analysis only ever uses audio and transcript.\n- Retaining full recordings indefinitely because \"we might re-analyse later.\"\n- Importing an entire CRM export to personalise a study that needed three fields.\n- Capturing demographics you never cross-tabulate.\n\nThe defensible pattern is the boring one: collect the fields the analysis actually uses, keep raw recordings for a defined window, and keep the de-identified transcript and structured answers for the longer term. That also happens to be a better research archive, because structured answers stay comparable across studies while recordings rot.\n\n## Step 4 — Consent that meets the statutory definition\n\nState laws converge on the same consent standard: a clear affirmative act that is freely given, specific, informed and unambiguous. Consent obtained through dark patterns is not valid consent, and several laws say so explicitly.\n\nResearch consent is usually easier to get right than product consent because the context is transparent. Present the purpose, the recording, the retention period, who sees the data, and the withdrawal route on one screen before the interview begins, with a single affirmative action. Avoid pre-ticked boxes, bundled consents that mix research with marketing, and any design where declining is harder than accepting.\n\nWhere sensitive data is genuinely part of the study — health research, financial hardship research — take a separate, specific opt-in for that category rather than folding it into a general consent.\n\n## Step 5 — Rights requests reach into your transcripts\n\nAccess, correction, deletion and portability rights apply to research data held about an in-scope consumer, and a deletion request is the one that finds the weak point in most research operations. A transcript typically lives in more than one place: the platform, an export, a slide, a repository, a shared drive.\n\nBefore you need it, you should be able to answer: where does participant data live, how is a participant located across those stores, what is the deletion runbook, and what do you retain in de-identified form afterwards? De-identified data generally falls outside these laws — but only if it is genuinely de-identified, you commit not to re-identify it, and you bind recipients to the same.\n\n## Step 6 — Is sending transcripts to a vendor a \"sale\"?\n\n\"Sale\" is defined broadly in most states — the exchange of personal data for monetary or other valuable consideration — and this is the provision that most often surprises research teams.\n\nDisclosures to a **processor or service provider** acting on your documented instructions are not sales, provided the contract carries the required terms: process only on instructions, no use for the vendor's own purposes, no selling on, confidentiality, deletion or return at the end, and flow-down to sub-processors. Get those contract terms in place with every platform, transcription service and analysis tool that touches interview data, and the sale question resolves cleanly.\n\nTwo arrangements to look at carefully: any tool that trains its own models on your interview content for its own benefit, and any panel or data partner that receives participant data as part of a commercial exchange.\n\nUniversal opt-out mechanisms such as Global Privacy Control — now honoured under around a dozen state laws including California, Colorado, Connecticut, Delaware, Maryland, Minnesota, Montana, New Jersey, New Hampshire, Oregon and Texas — are a website obligation about sales and targeted advertising. They rarely touch a research programme directly, but they do matter if you recruit participants through advertising.\n\n## The pragmatic answer: build to the strictest standard once\n\nMaintaining twenty variants of a research process is not viable. Build one, tuned to the strictest provision in each dimension:\n\n| Dimension | Build to |\n|---|---|\n| Scope | Assume California applies (B2B and employment included) |\n| Sensitive data | Opt-in consent, and never in a sale — Maryland standard |\n| Minimisation | Reasonably necessary and proportionate — Maryland standard |\n| Consent quality | Freely given, specific, informed, unambiguous; no dark patterns |\n| Retention | Defined window for raw recordings; de-identified archive after |\n| Vendors | Processor terms with every tool touching interview data |\n| Rights | A documented, tested deletion runbook across every store |\n\nA single process at that level satisfies all twenty and most of GDPR besides, and it costs far less than tracking divergence state by state.\n\n## Common mistakes\n\n- Applying consumer privacy analysis to B2B interviews that are out of scope everywhere except California, and drowning the programme in unnecessary process.\n- Assuming the reverse — that B2B is always exempt — and forgetting California entirely.\n- Treating consent as the whole obligation, when minimisation, retention and rights carry equal weight.\n- Recording video by default when the analysis never uses it.\n- Sending transcripts to tools without processor terms in place.\n- Calling data anonymised when it still contains names, employers and identifiable circumstances in the transcript body.\n- Having no deletion runbook until a request arrives.\n\n## Frequently asked questions\n\n**Do US state privacy laws apply to B2B user research?**\nGenerally not, outside California. Every other comprehensive state law defines a consumer as a resident acting in an individual or household context and excludes people acting in a commercial or employment context. California is the exception: the CCPA covers B2B contacts and employees, and the partial exemptions for that data expired on 1 January 2023. Since most US-national research includes California residents, many teams design to the California standard regardless.\n\n**How many US states have comprehensive privacy laws in 2026?**\nTwenty are on the books as of early 2026, with more added as each legislative session closes. Indiana, Kentucky and Rhode Island took effect on 1 January 2026, Maryland took effect on 1 October 2025 and became enforceable in April 2026, and further changes land in Connecticut, Arkansas and Utah on 1 July 2026.\n\n**Is a recorded research interview biometric data?**\nUsually not. Most state definitions treat biometric data as data processed for the purpose of uniquely identifying an individual, and a recording captured to be transcribed and analysed for content is not being processed for identification. Write that purpose down explicitly, keep voiceprint matching out of your stack, and check Maryland separately because its biometric and consumer health data definitions are stricter than most states.\n\n**Does sharing interview transcripts with a research platform count as a sale of personal data?**\nNot when the platform acts as a processor or service provider under a contract with the required terms: processing only on your documented instructions, no use for its own purposes, no onward selling, confidentiality, deletion or return at the end, and flow-down to sub-processors. Look carefully at any tool that trains its own models on your interview content, and at panel or data partners receiving participant data as part of a commercial exchange.\n\n**What is different about the Maryland Online Data Privacy Act?**\nThree things matter for research. Data collection must be reasonably necessary and proportionate, and consent does not cure over-collection. The sale of sensitive personal data is banned outright regardless of consent, which is the strictest such rule in the country. And its definitions of biometric data, consumer health data and sensitive personal data are broader than most states, alongside a low 35,000-consumer applicability threshold.\n\n**Do we need a separate research process for every state?**\nNo, and it is not sustainable to try. Build one process to the strictest provision in each dimension — California scope, Maryland minimisation and sensitive-data handling, statutory consent quality, defined retention, processor terms with every vendor, and a tested deletion runbook. A single process at that level satisfies all of them and most of GDPR as well.\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — collecting countable answers without over-collecting personal data\n- [CCPA/CPRA Compliance for Customer Research](/docs/ccpa-user-research-compliance) — the California-specific detail behind the scope rule\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the European equivalent of this analysis\n- [DSARs for Research Data](/docs/dsar-research-data) — handling access, deletion and portability requests\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — what genuine de-identification requires\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — setting defensible retention windows\n- [Interview Recording Consent Laws](/docs/interview-recording-consent-laws) — the separate one-party and two-party consent regime","category":"Research Operations","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"US State Privacy Laws for Customer Research: 2026 Compliance Guide","metaDescription":"What twenty US state privacy laws require of research interviews: the B2B and employment scope rule, sensitive data opt-in, Maryland minimisation, deletion rights over transcripts, and processor terms.","keywords":["US state privacy laws customer research","VCDPA research compliance","Colorado Privacy Act research","Maryland Online Data Privacy Act","state privacy law research consent","sensitive data opt-in consent","B2B exemption state privacy law","research data deletion rights"],"aiSummary":"Twenty US states have comprehensive consumer privacy laws on the books as of early 2026. The decisive scope question for research is whether a participant is a consumer: every comprehensive state law except California defines a consumer as a resident acting in an individual or household context and excludes people acting in a commercial or employment context, so B2B research participants and employees fall outside those laws. California is the exception because the CCPA covers B2B contacts and employees after the partial exemptions expired on 1 January 2023. Key 2026 dates: Maryland MODPA in effect 1 October 2025 and enforceable April 2026; Indiana, Kentucky and Rhode Island effective 1 January 2026 along with Oregon amendments, expanded California data broker and health data rules, the Nebraska Age-Appropriate Design Code and the Texas Responsible AI Governance Act; further changes in Connecticut, Arkansas and Utah on 1 July 2026; new California data broker registration obligations on 1 August 2026. Rhode Island and Maryland apply at 35,000 consumers against Oregon 100,000. Most laws require opt-in consent for sensitive data including health, race, religion, sexual orientation, immigration status, precise geolocation and biometric data processed to identify someone, and open-ended interviews collect such data accidentally. Maryland bans the sale of sensitive personal data outright regardless of consent and requires collection to be reasonably necessary and proportionate, which consent cannot cure. Voice recordings analysed for content are not ordinarily biometric data because identification is not the purpose. Disclosures to a processor or service provider under contract terms covering documented instructions, no independent use, no onward sale, confidentiality, deletion or return and sub-processor flow-down are not sales. Universal opt-out mechanisms such as Global Privacy Control are honoured in around a dozen states but are a website and advertising obligation rather than a research one. The practical answer is a single research process built to the strictest provision in each dimension rather than twenty state variants.","aiPrerequisites":["Research participants who are US residents","A view of where interview recordings, transcripts and exports are stored","Vendor contracts for every tool that processes interview data"],"aiLearningOutcomes":["Determine whether a research participant is a consumer under state privacy law","Apply the B2B and employment scope rule correctly, including the California exception","Handle sensitive data that interviews collect unintentionally","Meet the Maryland data minimisation and sensitive-data sale standards","Distinguish processor disclosures from a sale of personal data","Build one multi-state research process instead of twenty state variants"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"7c598815-e2b9-4841-87b8-aead41a53990","slug":"employee-engagement-survey-guide","title":"How to Build an Employee Engagement Survey That People Actually Answer Honestly","url":"https://www.koji.so/docs/employee-engagement-survey-guide","summary":"Definitive guide to employee engagement surveys. Explains why traditional surveys fail (social desirability bias, survey fatigue, action gap), how to structure 7-theme conversational engagement studies with Koji, and best practices for frequency, anonymity, and acting on results.","content":"# How to Build an Employee Engagement Survey That People Actually Answer Honestly\n\nEmployee engagement surveys are broken. Companies spend millions on annual engagement programs, yet Gallup reports that only 23% of employees worldwide are engaged at work. The problem isn't that employees don't care. It's that they don't trust the survey.\n\nWhen an employee sees a 50-question SurveyMonkey form from HR, they think: \"Is this really anonymous? Will my manager see my answers? Will anything actually change?\" So they give safe, middle-of-the-road answers. The resulting data looks fine. The organization misses the warning signs. Then people quit, and leadership is blindsided.\n\nKoji solves this by replacing the intimidating survey form with a natural, conversational AI interview that feels more like talking to a trusted colleague than filling out a corporate questionnaire.\n\n## Why Traditional Engagement Surveys Fail\n\n### Social desirability bias\nEmployees say what they think is expected, not what they actually feel. Even with \"anonymous\" surveys, people self-censor because they don't believe anonymity is real. In a conversation with an AI, employees share 40-60% more candid feedback because there's no human on the other end who might judge them.\n\n### Survey fatigue\nThe average engagement survey is 40-80 questions. Completion rates drop after question 15. By question 40, employees are clicking randomly to finish. The data from the back half of long surveys is statistically unreliable.\n\n### Action gap\n87% of engagement surveys result in no visible action. Employees learn that feedback goes into a black hole. Participation drops each subsequent year. The survey becomes a compliance exercise, not a listening tool.\n\n### Timing problems\nAnnual surveys capture a single snapshot. Employee sentiment fluctuates throughout the year based on project cycles, org changes, market conditions, and personal circumstances. Annual data is inherently stale.\n\n## Building Engagement Surveys with Koji\n\n### Study Architecture\n\nDesign your engagement study around 5-7 core themes, each with one quantitative anchor and conversational follow-up:\n\n**Theme 1: Overall Engagement**\n- Scale (1-10): \"On a scale of 1 to 10, how energized do you feel about coming to work each day?\"\n- Open-ended: \"What contributes most to that feeling?\"\n- Probing: AI explores specific drivers (role, team, projects, culture)\n\n**Theme 2: Manager Relationship**\n- Scale (1-5): \"How supported do you feel by your direct manager?\"\n- Open-ended: \"Can you give me an example of when your manager did or didn't support you well?\"\n- Probing: AI digs into communication, feedback, development\n\n**Theme 3: Growth and Development**\n- Single choice: \"Which best describes your career growth here?\" (Options: Growing rapidly / Steady growth / Stagnant / Declining)\n- Open-ended: \"What would accelerate your development?\"\n- Probing: AI explores specific skills, opportunities, barriers\n\n**Theme 4: Team and Collaboration**\n- Scale (1-5): \"How effective is collaboration within your team?\"\n- Open-ended: \"What's one thing that would make your team work together better?\"\n\n**Theme 5: Recognition and Compensation**\n- Yes/No: \"Do you feel fairly compensated for your work?\"\n- Open-ended: \"Tell me more about that.\"\n- Probing: AI separates compensation concerns from recognition needs\n\n**Theme 6: Work-Life Balance**\n- Scale (1-5): \"How sustainable is your current workload?\"\n- Open-ended: \"What would improve your day-to-day experience?\"\n\n**Theme 7: Belonging and Culture**\n- Scale (1-10): \"How strongly do you feel you belong here?\"\n- Open-ended: \"What makes you feel included or excluded?\"\n- Probing: AI explores DEI dimensions sensitively\n\n### Why This Structure Works\n\nTraditional surveys ask all 40+ questions in a flat list. Koji's conversational approach means:\n- Each theme gets 1 quantitative question (chartable, benchmarkable) plus AI-driven conversation\n- Total explicit questions: 7-10. Total depth: equivalent to a 60-minute focus group.\n- The AI adapts its probing based on the quantitative answer. A 2/10 on belonging triggers very different follow-ups than a 9/10.\n\n### Distribution Strategy\n\n- **All-hands rollout:** Send during a team meeting. Frame it as a conversation, not a survey.\n- **Staggered by team:** Helps identify team-level patterns and prevents the \"survey week\" crush.\n- **Anonymous links:** No login required. No tracking cookies. True anonymity builds trust.\n- **Voice option:** Some employees (especially in operational roles) prefer speaking to typing. Koji supports both.\n\n## Best Practices\n\n### Frequency\n- **Annual deep dive:** Full 7-theme study once per year\n- **Quarterly pulse:** 2-3 themes per quarter, rotating focus areas\n- **Event-triggered:** After org changes, layoffs, new leadership, acquisitions\n\n### Anonymity and trust\n- Use Koji's anonymous mode where no personally identifiable information is collected\n- Communicate the anonymity clearly before and during the study\n- Share aggregate results openly. Transparency builds trust for future surveys.\n- Never attempt to identify respondents from their qualitative answers.\n\n### Acting on results\n- Share results within 2 weeks. Longer delays signal that feedback doesn't matter.\n- Commit to 2-3 specific actions, not 20 vague promises.\n- Follow up in 90 days with progress updates.\n- Run a pulse survey to check if the actions are working.\n\n### Question design\n- Use first-person language: \"How energized do you feel?\" not \"How engaged are employees?\"\n- Ask about specific behaviors and experiences, not abstract concepts\n- Avoid double-barreled questions: \"Are you satisfied with your role and compensation?\" asks two things\n- Include both strengths and weaknesses: don't only ask what's wrong\n\n## Koji vs Traditional Engagement Platforms\n\n| Feature | Traditional (Culture Amp, Lattice, Glint) | Koji |\n|---------|-------------------------------------------|------|\n| Format | 40-80 question form | 7-10 questions + AI conversation |\n| Depth | Surface-level (Likert scales) | Deep qualitative + quantitative |\n| Candor | Social desirability bias | AI reduces bias significantly |\n| Completion time | 15-25 minutes of clicking | 8-12 minutes of natural conversation |\n| Completion rate | 60-75% | 80-90% (conversations are more engaging) |\n| Analysis | Pre-built dashboards | AI-synthesized themes with quotes |\n| Cost | $5-15 per employee per month | Credit-based, starting at less than $1 per conversation |\n| Languages | Limited | 30+ languages natively |\n| Voice option | Never | Built-in AI voice interviews |\n\n## Common Engagement Survey Mistakes\n\n- **Too many questions:** More than 15 explicit questions kills completion rates. Let Koji's AI do the probing.\n- **Manager-visible results:** If managers can see individual responses, no one will be honest.\n- **No action plan:** The fastest way to kill future participation is to ignore this round's feedback.\n- **Survey-only approach:** Combine engagement data with exit interviews, stay interviews, and pulse surveys for a complete picture.\n- **Benchmarking obsession:** Your engagement score relative to \"industry average\" matters less than your trend. Are you improving?\n\n## What Koji Reports Reveal\n\nAfter running an engagement study, Koji's automated report includes:\n- **Engagement index** with distribution charts per theme\n- **Theme-level heatmap** showing strongest and weakest areas\n- **Verbatim analysis** with sentiment tagging and key quotes\n- **Cross-theme correlations** (e.g., low manager scores predict low engagement scores)\n- **Department/team comparison** (if demographic data is collected anonymously)\n- **Recommended actions** prioritized by impact and frequency\n\nEvery insight is traced back to specific (anonymous) quotes, so leadership can see the human stories behind the numbers.\n\n---\n\n## Related Survey Guides\n\n- [Pulse Survey Guide](/docs/pulse-survey-guide) — Frequent lightweight engagement checks\n- [Employee Wellness Guide](/docs/employee-wellness-survey-guide) — Measure holistic wellbeing\n- [Stay Interview Guide](/docs/stay-interview-survey-guide) — Prevent attrition proactively\n- [Exit Interview Guide](/docs/exit-interview-survey-guide) — Learn why people leave\n- [DEI Survey Guide](/docs/dei-survey-guide) — Measure inclusion and belonging\n\n*Use [structured questions](/docs/structured-questions-guide) to combine engagement scales with AI-powered honest conversations.*\n\n## Further reading on the blog\n\n- [Customer Journey Mapping Guide 2026: How to Build Maps That Actually Drive Decisions](/blog/customer-journey-mapping-guide-2026) — A modern, AI-native playbook for customer journey mapping in 2026 — including the 5-step process, the questions that surface real emotion at\n- [Best Online Survey Software in 2026: The Complete Buyer's Guide](/blog/best-survey-software-2026) — From SurveyMonkey to Koji, we compare the top survey tools of 2026 across features, pricing, and use case fit — and explain when traditional\n- [How to Build a Voice of Customer Program in 2026: The Complete Guide](/blog/how-to-build-voice-of-customer-program-2026) — Companies with best-in-class Voice of Customer programs grow revenue 10x faster than those without. Here's a proven 8-step framework for bui\n\n<!-- further-reading:blog -->\n","category":"Survey & Study Templates","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Employee Engagement Survey Guide: Build Surveys People Actually Answer Honestly | Koji","metaDescription":"Learn how to design employee engagement surveys that surface real sentiment. Conversational AI eliminates social desirability bias and delivers 40-60% more candid feedback than traditional forms.","keywords":["employee engagement survey","employee satisfaction survey","engagement survey template","workplace survey","employee feedback survey","HR survey","employee pulse survey"],"aiSummary":"Definitive guide to employee engagement surveys. Explains why traditional surveys fail (social desirability bias, survey fatigue, action gap), how to structure 7-theme conversational engagement studies with Koji, and best practices for frequency, anonymity, and acting on results."},{"type":"documentation","id":"97511a5d-b067-40a3-90a8-ba3394fde135","slug":"commercial-due-diligence-customer-interviews","title":"Commercial Due Diligence Customer Interviews: An AI-Powered Playbook","url":"https://www.koji.so/docs/commercial-due-diligence-customer-interviews","summary":"Koji lets deal teams run commercial due diligence customer interviews at scale: instead of a handful of slow, cherry-picked manual reference calls, an AI interviewer runs the same structured battery — retention, switching cost, expansion, and competitive win-loss — across a target company entire customer base in days. Structured questions produce comparable, chartable evidence, AI follow-up probing surfaces the real reasons behind loyalty or churn, and a live report drops straight into the investment memo.","content":"## The Bottom Line\n\nCommercial due diligence lives or dies on one question: do this company customers actually value what it sells, and will they keep paying for it? The fastest way to answer that is to talk to those customers directly — but traditional reference calls are slow, expensive, and cover a tiny, cherry-picked sample. Koji lets a deal team run structured, AI-moderated customer interviews across a target company entire reference base in days: every conversation captures the same retention, satisfaction, and switching-cost signals, the AI probes for the real reasons behind loyalty or discontent, and the results land in a live report you can drop straight into the investment memo.\n\nIf you are a private equity or venture investor, a corporate development lead, or a founder preparing for a raise or an acquisition, this playbook shows how to replace a handful of manual reference calls with rigorous, scalable customer evidence.\n\n## Why customer interviews are the heart of commercial due diligence\n\nFinancials tell you what happened; customer interviews tell you whether it will continue. In a diligence process, primary customer research answers the questions a data room cannot:\n\n- **Is revenue durable?** Reported net revenue retention can hide concentration risk and quiet dissatisfaction. Customers will tell you whether they are renewing out of loyalty or inertia.\n- **How deep is the moat?** Switching costs, integration depth, and the strength of alternatives only surface when you ask customers what they would do if the product disappeared tomorrow.\n- **Is the growth story real?** Customers reveal whether they are expanding usage, holding flat, or planning to churn — the leading indicator behind every projection.\n- **What is the competitive reality?** Win-loss patterns from real buyers validate or puncture management claims about differentiation.\n\nThe problem has always been execution. A diligence window is short, and manual reference calls are the bottleneck: scheduling across time zones, a partner spending an hour per call, and a sample so small it is statistically meaningless. That is exactly the constraint Koji removes.\n\n## The old way vs. the Koji way\n\nTraditional commercial due diligence customer work looks like this: an analyst emails ten or fifteen customers the target introduces, a handful respond, and a senior person runs 45-minute calls over two weeks. The sample is tiny, self-selected by management, and inconsistent — every call covers different ground, so you cannot compare answers.\n\nWith Koji, the same effort covers far more ground:\n\n- Send a personalized AI interview link to dozens or hundreds of customers at once.\n- Every respondent answers the **same** structured questions, so results are directly comparable and chartable.\n- Koji AI conducts each interview — voice or text — and automatically asks follow-up questions when an answer is vague, capturing the depth of a live call without a human on the line.\n- Customers respond asynchronously, on their own schedule, across any time zone or language.\n- The analysis is automatic: Koji clusters themes, flags churn-risk accounts, and surfaces representative quotes in a report you can review the same day responses arrive.\n\nThe result is a customer evidence base an order of magnitude larger than manual reference calls, produced in a fraction of the time — turning a two-week bottleneck into a two-day workstream.\n\n## What to ask: a diligence interview structure\n\nKoji six [structured question types](/docs/structured-questions-guide) — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — let you capture hard metrics and soft narrative in a single conversation. A strong commercial diligence interview typically includes:\n\n- **scale** — overall satisfaction and likelihood to recommend (a clean NPS input across the whole sample)\n- **scale** — likelihood to renew or continue purchasing at the next cycle\n- **single_choice** — the primary reason they chose this vendor over alternatives\n- **ranking** — how they weigh price, product capability, support, and switching cost\n- **yes_no** — whether they evaluated or currently use a competitor\n- **open_ended** — what would make them leave, and what they would do if the product disappeared, with Koji AI probing automatically for specifics\n\nBecause every customer answers the same battery, you get a distribution — not three anecdotes — behind every claim in the memo. You can segment by contract size, tenure, or industry to expose concentration and cohort risk.\n\n## Reading the signals that matter\n\nWhen you review the Koji report, focus on the patterns that predict post-close performance:\n\n- **Retention conviction:** Do customers describe the product as essential infrastructure or a nice-to-have? Essential-infrastructure language is the strongest durability signal there is.\n- **Switching-cost reality:** If customers struggle to name a viable alternative or describe painful migration, the moat is real. If they rattle off three competitors, model higher churn.\n- **Expansion appetite:** Customers planning to add seats, modules, or use cases validate the upside case; flat or shrinking usage is a yellow flag.\n- **Detractor concentration:** A few unhappy large accounts can outweigh many happy small ones. Koji lets you weight and segment so you see revenue-weighted sentiment, not a raw average.\n\n## Beyond the deal: post-close and founder use\n\nCommercial due diligence interviews are not only for buyers. Founders preparing for a raise or sale run the same Koji study proactively to walk into the process with independent, third-party-style customer evidence — a credibility advantage over management assertions. And after close, the same interview cadence becomes an always-on voice-of-customer program, so the investment thesis is monitored continuously rather than revisited only at the next transaction. For the founder-facing version of this work, see our guides on [market validation](/docs/market-validation-ai-research) and [product-market-fit interviews](/docs/product-market-fit-interviews).\n\n## Getting started\n\n1. **Define the thesis questions** the interviews must answer — durability, moat, expansion, competitive position.\n2. **Draft the brief.** Koji AI converts your diligence questions into a structured interview with the right scale and open-ended mix.\n3. **Load the customer list** the target provides, or a broader sample where available.\n4. **Launch voice or text interviews** with personalized links; customers respond asynchronously.\n5. **Review the live report** as responses arrive, segment by account size and tenure, and export the evidence into your memo.\n\nA workstream that once meant a fortnight of scheduling becomes a same-week deliverable — with a bigger, cleaner, more defensible customer sample behind every conclusion.\n\n## Common pitfalls in customer diligence — and how to avoid them\n\nEven with the right tool, customer diligence goes wrong in predictable ways. Watch for these:\n\n- **Only interviewing management-selected references.** A curated shortlist is self-serving by design. Because Koji makes interviewing cheap and fast, widen the sample as far as the customer list allows — a broad, representative read is more defensible than five hand-picked advocates.\n- **Averaging sentiment instead of revenue-weighting it.** A raw satisfaction average buries concentration risk. Segment the Koji results by contract size so a few unhappy large accounts do not hide behind many happy small ones.\n- **Confusing satisfaction with switching cost.** Happy customers still churn when a better, cheaper alternative appears. Always ask what they would do if the product disappeared and how hard migration would be — that, not satisfaction, is the durability signal.\n- **Skipping former and churned customers.** The customers who already left often hold the most important lessons about the moat. Include them in the study and let Koji AI probe why they moved on.\n- **Leading the witness.** Manual reference calls run by a deal advocate tend to fish for confirmation. Koji AI asks every respondent the same neutral questions, so the evidence reflects the customer base, not the interviewer hopes.\n\nAvoiding these keeps the customer workstream honest — and makes its conclusions hold up under scrutiny from an investment committee.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types behind comparable, chartable diligence data\n- [Win-Loss Analysis Guide](/docs/win-loss-analysis-guide) — validate competitive positioning with real buyers\n- [Customer Research for Investors](/docs/customer-research-for-investors) — how funds use AI interviews across the portfolio\n- [Market Validation with AI Research](/docs/market-validation-ai-research) — pressure-test demand before you commit\n- [Product-Market-Fit Interviews](/docs/product-market-fit-interviews) — measure how essential a product really is\n- [Churned Customer Interviews](/docs/churned-customer-interviews) — learn why customers leave before it shows up in the numbers","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Commercial Due Diligence Customer Interviews with AI | Koji","metaDescription":"Run commercial due diligence customer interviews faster with Koji AI voice interviews. Capture retention, moat, and win-loss evidence across a target company customer base in days — not weeks of manual reference calls.","keywords":["commercial due diligence","customer reference interviews","due diligence customer research","private equity customer interviews","investor customer research","AI reference calls","deal diligence research","customer retention diligence","win-loss diligence","vendor due diligence interviews"],"aiSummary":"Koji lets deal teams run commercial due diligence customer interviews at scale: instead of a handful of slow, cherry-picked manual reference calls, an AI interviewer runs the same structured battery — retention, switching cost, expansion, and competitive win-loss — across a target company entire customer base in days. Structured questions produce comparable, chartable evidence, AI follow-up probing surfaces the real reasons behind loyalty or churn, and a live report drops straight into the investment memo.","aiPrerequisites":["Access to a target or portfolio company customer list","A defined diligence thesis or set of questions to validate"],"aiLearningOutcomes":["Replace slow manual reference calls with scalable AI-moderated customer interviews","Structure a diligence interview that captures retention, moat, and competitive signals","Read customer evidence for durability, switching cost, and expansion risk","Segment results to expose concentration and cohort risk","Turn customer interviews into memo-ready, defensible evidence in days"],"aiDifficulty":"advanced","aiEstimatedTime":"13 minutes"},{"type":"documentation","id":"109cb05e-d119-451d-95df-56aa1b79cd83","slug":"customer-research-for-investors","title":"Customer Research for Investors: Customer Due Diligence with AI Interviews","url":"https://www.koji.so/docs/customer-research-for-investors","summary":"Investors use Koji's AI-moderated interviews to run customer due diligence at scale — fielding 20 to 30 candid customer and churned-customer interviews in days instead of relying on a handful of founder-arranged reference calls. Koji probes retention, value, competitive position, and expansion automatically, supports structured questions for comparable metrics like likelihood-to-renew, and synthesizes real-time reports for the investment committee. Useful pre-investment, for win-loss validation, and for always-on portfolio voice-of-customer programs.","content":"## The Bottom Line\n\nThe single most predictive signal in a deal is what real customers think — yet customer due diligence is often the most rushed part of the process, squeezed into a handful of reference calls that founders pre-select. Koji lets investors run independent customer due diligence at scale: AI-moderated voice and text interviews with a target company's customers (or churned customers), fielded in days, with automatic analysis that surfaces retention drivers, unmet needs, and competitive risk. Instead of three friendly reference calls, you get 30 candid conversations and a synthesized report — turning customer DD from a gut check into a defensible, data-backed input to the investment decision.\n\nThis guide covers how investors use AI interviews across the deal lifecycle, what to ask, and how Koji makes it fast enough to fit a live process.\n\n## Why customer due diligence is broken\n\nTraditional commercial and customer due diligence has three weaknesses:\n\n- **Selection bias.** Reference calls are usually arranged through the founder, so you talk to the happiest customers. The dissatisfied and the churned — where the real risk lives — rarely make the list.\n- **Tiny samples.** Five or six calls cannot tell you whether a retention problem is an anomaly or a pattern.\n- **Speed pressure.** In a competitive process, you have days, not weeks. Scheduling moderated calls across busy buyers is the bottleneck.\n\nThe consequence is that investors often underweight the most important question — *do customers actually love this product, and will they keep paying?* — because the method to answer it well is too slow.\n\n## Where AI interviews fit across the deal lifecycle\n\n**Pre-investment (diligence).** Run independent interviews with a sample of the target's customers to validate the retention and expansion story, probe the real reasons for adoption, and stress-test the moat. Pair this with churned-customer interviews to understand why accounts leave.\n\n**Win-loss validation.** Interview prospects who chose the company and those who chose a competitor to map the true competitive dynamics behind the revenue.\n\n**Portfolio value creation (post-investment).** After the deal closes, stand up an always-on voice-of-customer program across portfolio companies to track NPS drivers, surface product gaps, and feed the board real qualitative signal between updates.\n\n**Thesis and market research.** Before you even have a target, use AI interviews to test a market thesis — talking to buyers in a category to understand budgets, pain, and willingness to switch.\n\n## Why Koji fits an investor's timeline\n\n**Speed.** Because interviews are AI-moderated and asynchronous, a customer completes one on their own schedule from a link — no calendar coordination. A study that would take weeks of reference-call scheduling fields in days, which is what a live process demands.\n\n**Candor at scale.** An AI interviewer with no stake in the outcome often gets more candid answers than a founder-arranged call. And because Koji scales to dozens of conversations cheaply, you move from anecdote to pattern.\n\n**Automatic follow-up probing.** When a customer says \"it does most of what we need,\" Koji probes what is missing, how painful the gap is, and what would make them switch — the diligence questions that matter, asked consistently every time.\n\n**Structured questions for comparable metrics.** Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Use a scale question for likelihood-to-renew or NPS, a ranking question for purchase drivers, and open-ended questions for the narrative — then read distributions and frequency charts in the report. This gives you quantitative benchmarks you can compare across targets and across a portfolio. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Real-time synthesis.** Koji clusters themes, pulls representative quotes, and quantifies structured answers as interviews arrive, so you can brief the deal team with findings the same week — not after a transcription cycle.\n\n**Quality gate.** Only conversations scoring 3 or higher count toward your plan, so junk responses do not pollute your diligence dataset.\n\n## A customer due diligence interview blueprint\n\nA strong customer DD study usually probes five areas. Build them as a mix of open-ended and structured questions:\n\n1. **Adoption and use case** — What problem does this product solve for you? How central is it to your workflow? (open-ended + a scale on criticality)\n2. **Value and ROI** — What measurable value have you gotten? Would you say it is essential or nice-to-have? (open-ended + single_choice)\n3. **Retention and renewal** — How likely are you to renew, and why? What would make you leave? (scale + open-ended probing)\n4. **Competitive position** — What did you use before, and what would you switch to? (open-ended + ranking of decision factors)\n5. **Expansion** — Would you buy more seats or modules? What would justify it? (yes_no + open-ended)\n\nRun the same blueprint with churned customers, reframed to the past tense, and the contrast between current and churned cohorts becomes one of the most revealing artifacts in your diligence pack.\n\n## Practical and ethical notes\n\n- **Source your own sample where possible.** Independent recruiting reduces selection bias. When you must use a founder-provided list, interview enough of it that patterns — not curation — drive the read.\n- **Be transparent and compliant.** Capture consent in the interview flow, anonymize where appropriate, and handle data under a DPA. Koji encrypts data in transit and at rest and supports anonymization and retention controls.\n- **Triangulate.** Customer interviews are strongest alongside usage data and financials. Treat the qualitative read as the *why* behind the numbers your deal team already has.\n\n## Getting started\n\nPick the highest-risk assumption in the deal — usually retention or competitive durability — and design a 20-to-30 interview study around it with the blueprint above. Field it the week you get access to a customer list, read the real-time report, and bring a data-backed customer view to the investment committee. Post-close, convert the same study into an always-on program across the portfolio. That is the difference between hoping customers love the product and knowing they do.\n\n## What good looks like: reading the signal\n\nOnce interviews are in, a few patterns separate a strong investment from a risky one:\n\n- **Criticality, not just satisfaction.** A high satisfaction score paired with low criticality (\"nice-to-have\") is a churn risk. The pairing of your scale and open-ended answers tells you whether the product is embedded in the workflow or merely liked.\n- **Consistent, specific value stories.** When customers independently describe the same concrete ROI in their own words, the value proposition is real. Vague praise that never names a measurable outcome is a yellow flag.\n- **Switching cost in the language.** Listen for how hard customers say it would be to leave. Genuine lock-in shows up as specific dependencies, not generic loyalty.\n- **The churned-cohort contrast.** If churned customers cite a fixable onboarding gap, that is a value-creation lever. If they cite a structural product limitation echoed by current customers, that is a thesis risk.\n\nReading these signals across 20 to 30 interviews — rather than inferring them from three calls — is what makes AI-powered customer due diligence a genuine edge in a competitive process.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — building comparable metrics into diligence interviews\n- [Customer Discovery Interview Guide](/docs/customer-discovery-interviews) — interview technique fundamentals\n- [Win-Loss Analysis Guide](/docs/win-loss-analysis-guide) — mapping competitive dynamics\n- [Product-Market Fit Interviews](/docs/product-market-fit-interviews) — measuring how much customers need the product\n- [Voice of Customer Research Program](/docs/voice-of-customer-research-program) — always-on portfolio listening\n- [Generating Research Reports](/docs/generating-research-reports) — synthesizing interviews into a diligence pack","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Customer Research for Investors: Customer Due Diligence with AI Interviews | Koji","metaDescription":"How VCs, growth equity, and PE firms run customer due diligence with AI interviews — validating retention, demand, and competitive position in days, with synthesized reports.","keywords":["customer due diligence","customer research for investors","vc due diligence","commercial due diligence","private equity customer research","reference calls","retention diligence","investor customer interviews","growth equity research","portfolio voice of customer"],"aiSummary":"Investors use Koji's AI-moderated interviews to run customer due diligence at scale — fielding 20 to 30 candid customer and churned-customer interviews in days instead of relying on a handful of founder-arranged reference calls. Koji probes retention, value, competitive position, and expansion automatically, supports structured questions for comparable metrics like likelihood-to-renew, and synthesizes real-time reports for the investment committee. Useful pre-investment, for win-loss validation, and for always-on portfolio voice-of-customer programs.","aiPrerequisites":["Familiarity with deal diligence processes","Access to a target company's customer or churned-customer list"],"aiLearningOutcomes":["Run independent customer due diligence at scale","Reduce selection bias versus founder-arranged reference calls","Design a five-part customer DD interview blueprint","Stand up always-on voice-of-customer programs across a portfolio"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 minutes"},{"type":"documentation","id":"ab1362aa-077a-4e6d-a2d7-5ee1d96ffe85","slug":"anonymous-employee-research-ai-interviews","title":"Anonymous Employee Research with AI Interviews: Get the Honest Feedback Surveys Miss","url":"https://www.koji.so/docs/anonymous-employee-research-ai-interviews","summary":"Anonymous employee research with Koji uses AI voice and text interviews to capture honest, in-depth feedback on culture, leadership, and retention without storing names, emails, or identifiers. Disable the lead form and enable the Anonymous badge in the study Customize tab; share a single link via email, Slack, or QR code; and watch themes, quote clusters, and engagement scores update in real-time. Combine 6 structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) for both qualitative depth and quantitative tracking. Available on every Koji plan including Free; only conversations scoring 3+ on quality consume credits.","content":"## Anonymous employee research, in one sentence\n\nAnonymous employee research is a study where employees share open, candid feedback through interviews or surveys without any identifier — name, email, manager, team — being attached to their response. With Koji, you switch on Anonymous Mode in your study settings, share a single interview link, and the AI moderator runs a 5–15 minute voice or text conversation that surfaces the same depth as a 1:1 — minus the awkwardness, the calendar tetris, and the fear of retaliation.\n\nIf you have ever sent an \"anonymous\" survey through a tool that asks for an email at the start, you already know the problem: people do not believe it. The response rate plummets, and the people who do respond water down their answers. AI-moderated interviews flip that dynamic. Employees talk to a neutral AI that cannot identify them, will not gossip, and probes their answers gently — the same way a skilled qualitative researcher would.\n\n## Why anonymous matters more than ever\n\nAccording to Gallup's 2025 State of the Global Workplace report, only 23% of employees worldwide report being engaged at work, and disengagement is now estimated to cost the global economy roughly $8.8 trillion in lost productivity. The harder data point: research from MIT Sloan and Glassdoor consistently shows that employees who fear identification withhold their most valuable feedback — exactly the signals leaders need before someone walks out the door.\n\nThe traditional fix has been engagement surveys. They are easy to send, but they have three structural problems:\n\n- **Multiple choice cannot capture nuance.** \"Somewhat agree\" tells you nothing about why.\n- **Open-text fields get one-line answers.** No follow-up means no depth.\n- **People assume the data is not actually anonymous.** Email links, IP logging, small team filters — they all leak identity.\n\nKoji solves all three. The AI conducts a real conversation, asks follow-up questions automatically, and runs in fully anonymous mode with no intake form, no email capture, and no participant ID stored.\n\n## How Koji's Anonymous Mode actually works\n\nWhen you create an employee study in Koji, you have full control over identity capture. Two settings determine anonymity:\n\n1. **Lead form disabled** — In your study's Customize tab, switch off the pre-interview lead form. No name, email, or department field is shown to the participant before they start.\n2. **Anonymous badge on** — Display a visible \"Anonymous\" or \"Confidential\" badge on the interview landing page so employees see the guarantee before they speak.\n\nUnder the hood, each anonymous response is stored as an opaque session token. There is no email, no employee ID, no LDAP lookup. Two employees from the same team look identical in the dataset — except for the content of their answers, which Koji clusters into themes automatically.\n\n## A 6-step playbook for running anonymous employee interviews\n\n### Step 1 — Pick the right research question\n\nAnonymous interviews shine for sensitive topics:\n\n- \"What would make you stay another two years?\" (stay interviews)\n- \"What is the one thing leadership keeps getting wrong?\" (culture audit)\n- \"Why do you think people on this team are burning out?\" (retention risk)\n- \"If you could change one part of your manager relationship, what would it be?\" (manager 360s)\n- \"What is your honest take on our DEI commitments — promises kept or just slogans?\" (DEI pulse)\n\nWrite a single sharp research question first. Koji's AI Consultant will turn it into a full interview brief in about 90 seconds.\n\n### Step 2 — Build the study with structured questions\n\nKoji's 6 structured question types let you mix candid open-ended exploration with quantitative tracking — all in the same anonymous interview:\n\n- `open_ended` for the deep \"why\" questions\n- `scale` for engagement ratings (1–10 NPS-style or 1–5 satisfaction)\n- `single_choice` for forced trade-offs (\"If we could only fix one thing this quarter…\")\n- `multiple_choice` for blockers (\"Which of these gets in the way most?\")\n- `ranking` for priority (\"Rank what matters most about your job\")\n- `yes_no` for clear binary signals (\"Would you recommend this team to a friend?\")\n\nThis means you can run an anonymous study that produces both a CEO-friendly engagement number AND the quotes that explain it. Most engagement survey tools force you to choose. See the [Structured Questions in AI Interviews](/docs/structured-questions-guide) guide for the full pattern library.\n\n### Step 3 — Turn off identity capture\n\nGo to your study → Customize → Lead form. Disable it. In the Badges section, enable \"Anonymous\" or set custom text like \"Confidential — no identifiers stored.\" Optionally lock the data export to omit any metadata that could indirectly identify someone (timestamps, device hints, etc.).\n\n### Step 4 — Distribute through low-friction channels\n\nShare the single interview link via:\n\n- An all-hands email from the CEO (not HR — leadership endorsement matters)\n- A Slack `#general` post pinned for two weeks\n- A link in Lattice / Culture Amp / 15Five next to the regular pulse\n- A QR code on break room posters for hourly workforces\n\nKoji links never require a login. Employees click, talk to the AI for 5–15 minutes, and leave.\n\n### Step 5 — Watch insights build in real time\n\nAs anonymous interviews come in, Koji's Insights Dashboard updates live. You see emerging themes, quote clusters, sentiment shifts, and quality scores without waiting for a researcher to code transcripts. The dashboard groups feedback by topic — manager trust, workload, growth, comp — even though no participant is identified. See [Real-Time Research Insights](/docs/real-time-research-insights) for what to look for in the first 24 hours.\n\n### Step 6 — Report up without breaking anonymity\n\nKoji's auto-generated report aggregates 30, 100, or 500 interviews into a sharable narrative with verbatim quotes, themes, scale distributions, and recommendations. A small-team safeguard: if a theme is reported by fewer than 5 unique respondents, you can suppress it from the published report so a single voice cannot be triangulated.\n\nShare the report via Koji's public link or export it as PDF/JSON for your HRIS. Your CHRO sees the patterns. Individual employees stay anonymous. See [Publishing & Sharing Reports](/docs/publishing-sharing-reports).\n\n## Where anonymous AI interviews beat traditional employee surveys\n\n| Capability | Annual Engagement Survey | Anonymous AI Interview (Koji) |\n|---|---|---|\n| Captures the \"why\" | Rare — only in optional comments | Always — AI probes 1–3 follow-ups per question |\n| Time to insight | 4–8 weeks (ship → close → analyze) | Real-time, theme detection as responses come in |\n| Response rate | 30–50% industry average | 60–80% in our customer studies (lower friction) |\n| Identity protection | Often nominal — emails frequently logged | True anonymous mode — no PII captured at all |\n| Depth per response | 1–3 sentences in open text | 800–2,500 words of conversational data |\n| Cost per response | $4–$8 (Qualtrics, Culture Amp) | 1 credit (text) / 3 credits (voice) on Koji |\n\nPlatforms like Koji automate the parts that traditional engagement tools leave to humans: probing follow-ups, theme clustering, quote extraction, and report writing. A People team of one can run a full anonymous engagement study end-to-end in a single afternoon.\n\n## Designing questions that produce honest answers\n\nThree principles, drawn from *The Mom Test* and decades of qualitative HR research:\n\n1. **Ask about the past, not the future.** \"Tell me about the last time you considered leaving\" beats \"Are you a flight risk?\" — the second is a yes/no nobody answers truthfully.\n2. **Ground every question in a specific story.** \"Walk me through your last 1:1 with your manager\" gets you twenty times more signal than \"How is your manager?\"\n3. **Let the AI probe.** Set `maxFollowUps: 2` on important open-ended questions. Koji's AI will gently dig into the surface answer and surface the real story. See the [AI Follow-Up Probing](/docs/ai-probing-guide) guide.\n\n## Common pitfalls to avoid\n\n- **Don't fake anonymity.** If your tool captures email, do not call the study anonymous — employees will eventually find out and trust collapses.\n- **Don't over-segment.** Filtering to \"engineers in Berlin who joined after 2023\" can re-identify people in small samples. Set a minimum group size (e.g., 5) before showing demographic breakdowns.\n- **Don't skip the close-the-loop step.** The quickest way to kill future participation is to collect feedback and never report back. Koji's public report links make this easy — share the same dashboard internally.\n- **Don't over-rely on one annual study.** Anonymous AI interviews are cheap enough to run quarterly. Continuous listening reveals trends an annual snapshot misses.\n\n## When to escalate beyond anonymous interviews\n\nAnonymous interviews are the right starting point, but for some sensitive topics — harassment, discrimination, ethics violations — you need a formal investigation channel with legal protections. Use Koji to identify *patterns* and *prevalence*, then route specific allegations to your ethics hotline or legal team. Koji's research debrief flow makes this handoff explicit. See [Research Debrief Guide](/docs/research-debrief-guide).\n\n## Pricing and access\n\nKoji's Anonymous Mode is available on every plan, including Free. Each text interview costs 1 credit, voice interview 3 credits. The Insights plan (€29/mo, 29 credits) is enough for ~30 anonymous interviews per month — roughly the size of a department-level study. The Interviews plan (€79/mo, 79 credits) supports company-wide programs at growth-stage SaaS scale.\n\nQuality gate: only conversations scoring 3+ on completion quality consume credits. Half-finished or low-effort responses are free, so noise does not eat your budget.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the 6 question types that mix qualitative depth with quantitative tracking\n- [Voice of the Employee Program](/docs/voice-of-employee-program) — building a continuous employee listening program\n- [Stay Interviews at Scale](/docs/stay-interviews-at-scale) — running retention interviews with AI moderation\n- [How the Quality Gate Works](/docs/how-the-quality-gate-works) — why you only pay for high-quality responses\n- [Avoiding Bias in Research Interviews](/docs/avoiding-bias-in-interviews) — keeping anonymous studies methodologically sound\n- [Koji for HR and People Teams](/docs/koji-for-hr-people-teams) — the full HR research playbook\n\n## Further reading on the blog\n\n- [The Death of Static Surveys: Why AI Interviews Are the Future of User Research](/blog/death-of-static-surveys) — Static surveys are suffering from <2% response rates. Learn why \"Survey Fatigue\" is killing your data and how AI-moderated interviews are ac\n- [Best Delighted Alternatives in 2026: AI Research Interviews Beat NPS Surveys](/blog/delighted-alternatives-2026) — Delighted is shutting down June 30, 2026. Here are the 6 best alternatives — and why the smartest migration isn't a replacement, it's an upg\n- [Focus Groups vs Interviews: Which Research Method Gets You Better Data? (2026)](/blog/focus-groups-vs-interviews-2026) — Should you run focus groups or one-on-one interviews? This guide compares both methods on depth, cost, bias, and speed — and shows why AI-mo\n\n<!-- further-reading:blog -->\n","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Anonymous Employee Research with AI Interviews | Koji","metaDescription":"Run truly anonymous employee research with AI voice & text interviews. Capture honest feedback on culture, retention, and leadership — no PII, real-time insights.","keywords":["anonymous employee research","anonymous employee survey","anonymous feedback tool","employee engagement interviews","AI employee research","confidential employee feedback","employee voice platform","HR research tool","stay interview tool","anonymous engagement survey"],"aiSummary":"Anonymous employee research with Koji uses AI voice and text interviews to capture honest, in-depth feedback on culture, leadership, and retention without storing names, emails, or identifiers. Disable the lead form and enable the Anonymous badge in the study Customize tab; share a single link via email, Slack, or QR code; and watch themes, quote clusters, and engagement scores update in real-time. Combine 6 structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) for both qualitative depth and quantitative tracking. Available on every Koji plan including Free; only conversations scoring 3+ on quality consume credits.","aiPrerequisites":["Basic understanding of employee surveys","Koji account (Free or paid plan)"],"aiLearningOutcomes":["Configure a fully anonymous Koji study with no PII capture","Write open-ended and structured questions that get honest answers","Distribute the interview link through low-friction channels","Read aggregate themes and quotes without re-identifying employees","Avoid common anonymity pitfalls (small-team triangulation, fake anonymity, no close-the-loop)"],"aiDifficulty":"beginner","aiEstimatedTime":"9 min read"},{"type":"documentation","id":"fd97ec88-f7d9-4cb1-a7ae-1c87b842bedb","slug":"koji-for-hr-people-teams","title":"Koji for HR and People Teams: Run Employee Research at Scale","url":"https://www.koji.so/docs/koji-for-hr-people-teams","summary":"Koji helps HR and People teams run scalable employee listening programs including stay interviews, exit interviews, onboarding feedback, and DEI research. No scheduling required — employees click a personalized link and complete an AI-moderated conversation on their own schedule. Reports are automatically segmented by department, tenure, and location. Works best with personalized CSV imports and a mix of open-ended and structured questions for qualitative depth plus trackable metrics.","content":"\n# Koji for HR and People Teams: Run Employee Research at Scale\n\nEvery People leader knows they should be running more employee listening programs. The data is clear: organizations with robust employee listening cultures see 25% lower turnover, stronger engagement scores, and faster culture change. But the reality of executing employee research at scale is painful.\n\nScheduling 1:1 interviews across a distributed workforce is a logistics nightmare. Focus groups require a skilled facilitator and produce group-think. Annual engagement surveys take months to analyze and arrive too late to act on. Short pulse surveys generate data fatigue without generating genuine insight.\n\nKoji changes this equation entirely. By deploying AI-moderated conversations that employees take on their own schedule — anytime, anywhere, in any language — People teams can run the kind of deep qualitative research that was previously only possible with a dedicated research function and months of calendar wrangling.\n\n## What Koji Does for People Teams\n\nKoji is an AI-powered interview platform that runs fully automated, conversational interviews with employees. Instead of filling out a static form, participants have a real conversation — the AI asks contextually relevant follow-up questions, probes for detail, and surfaces themes that structured surveys consistently miss.\n\nKey capabilities for HR and People teams:\n\n**No scheduling required.** Share a link via Slack, email, or your HRIS. Employees complete the interview on their own time — during lunch, between meetings, or from home at 10pm.\n\n**Voice or text mode.** Employees who struggle to express nuance in writing can choose a voice interview instead. Koji transcribes and analyzes both modalities with equal depth.\n\n**Always-on global availability.** International teams spanning multiple time zones complete interviews without coordinator overhead. Koji is available 24/7.\n\n**AI follow-up questions.** The AI probes for depth without leading: \"You mentioned the onboarding process felt rushed — what specifically felt most rushed?\" That contextual follow-up is what makes qualitative research valuable. Koji does it automatically, for every participant.\n\n**Automatic analysis.** After each batch of interviews, Koji generates a report showing themes, representative quotes, sentiment patterns, and — if you included structured questions — quantitative distributions that give you trackable metrics over time.\n\n## Use Case 1: Stay Interviews at Scale\n\nStay interviews — conversations with current employees to understand what keeps them engaged and what might cause them to leave — are one of the highest-ROI retention tools available. Research consistently shows that employees who have had a recent stay interview are meaningfully more likely to remain at the organization.\n\nThe problem: stay interviews are time-intensive. A 30-minute interview with 100 employees represents 50 hours of manager or HR time, before analysis.\n\nWith Koji, you can run 100 stay interviews in a week with minimal overhead. Design the study once:\n\n- What keeps you energized and motivated in your work?\n- What is the most frustrating thing about working here right now?\n- What would need to change for you to see yourself here in two years?\n- On a scale of 1–10, how likely are you to still be working here in 12 months? *(structured scale question)*\n- Is there anything leadership could do that would immediately improve your experience?\n\nKoji handles the rest: distributing personalized links to employees, conducting the interviews, and generating a report showing patterns across the workforce — automatically segmented by department, tenure, or location if you imported those attributes via CSV.\n\nThis is a fundamentally different ROI equation than traditional stay interview programs that require months of scheduling and facilitation.\n\n## Use Case 2: Exit Interviews That Actually Surface the Truth\n\nExit interviews are notoriously unreliable when conducted by HR staff — departing employees soften feedback to protect relationships and future references. AI-moderated interviews change this dynamic in a well-documented way.\n\nResearch on AI-assisted interviewing consistently shows that participants disclose more candidly to AI interviewers than to human ones — particularly on sensitive topics. When employees know their responses are analyzed as aggregated themes (not attributed to them individually), they are more likely to share the real reasons for their departure.\n\nKoji exit interviews integrate directly into offboarding workflows. When an employee submits resignation, a personalized Koji interview link is automatically sent as part of the offboarding checklist.\n\nEffective exit interview questions:\n- What was the primary factor in your decision to leave?\n- What did you enjoy most about working here?\n- What would have made you stay? *(single-choice + open-ended follow-up)*\n- How would you rate your relationship with your direct manager? *(scale 1–10)*\n- What is one thing you wish you had said while you were here?\n\nThe resulting report aggregates exit themes across cohorts, helping you identify systemic issues — management problems in specific departments, compensation gaps for certain roles, culture friction that only appears after enough candid departures.\n\n## Use Case 3: Longitudinal Onboarding Research\n\nThe first 90 days determine whether a new hire reaches their potential or exits within the year. Most organizations rely on a single survey at day 30 — too structured to surface real friction, and too early to reflect the complete picture.\n\nKoji enables a continuous onboarding listening program:\n\n**Day 7 check-in**: \"Tell me about your first week. What has been clearest, and what has felt most confusing?\" *(open-ended only — this early, you want exploration, not structure)*\n\n**Day 30 reflection**: \"Now that you have been here a month, what is working and what is not?\" *(add structured questions: onboarding satisfaction scale, clarity of role expectations yes/no, likelihood to recommend company scale)*\n\n**Day 90 debrief**: \"Looking back, what would have made your first 90 days more effective? What surprised you most about working here?\"\n\nBecause each interview is AI-moderated, you do not need a dedicated HR coordinator to run 30 new-hire conversations per month. The program scales automatically with your hiring velocity.\n\nOver time, you accumulate a longitudinal dataset that shows whether onboarding quality is improving, which departments generate the most friction, and which cohorts retain best — all from qualitative conversations that no pulse survey could capture.\n\n## Use Case 4: DEI and Culture Research\n\nUnderstanding how different employee groups experience your culture requires psychological safety — the genuine belief that it is safe to tell the truth without career consequences. Anonymous AI interviews create conditions for more honest disclosure than manager-led conversations or named surveys.\n\nKoji for culture and belonging research:\n\n- **Belonging assessments**: Do employees across backgrounds feel included and valued in day-to-day work?\n- **Inclusive communication audits**: Are there communication patterns that certain groups experience differently?\n- **Advancement perception research**: Do employees believe that promotions and recognition are equitable?\n- **Manager effectiveness**: How do teams experience their managers' communication, feedback, and decision-making?\n\nKoji's [structured questions](/docs/structured-questions-guide) enable mixed-methods culture research. A belonging scale (1–10) combined with an open-ended question — \"Tell me about a moment that made you feel especially included or excluded\" — gives you both a trackable metric and the qualitative texture to understand it.\n\nTrack belonging scores by department quarter over quarter. The structured data tells you where to focus; the qualitative interviews tell you why.\n\n## Setting Up Employee Research: Personalized Links for HR\n\nKoji's personalized link feature is particularly powerful for People teams. By [importing your employee list via CSV](/docs/importing-participants-csv) with attributes like `department`, `tenure`, `role_level`, `location`, and `manager`, you can:\n\n- Send personalized interview links that address each employee by name\n- Automatically segment analysis by department, cohort, or tenure without manual tagging\n- Monitor completion rates by group to ensure representative participation\n- Follow up with non-responders in specific departments without knowing individual responses\n\nFor detailed setup instructions, see [Personalized Interview Links](/docs/personalized-interview-links).\n\nExample CSV for a stay interview program:\n\n```csv\nemail,name,department,tenure_years,location,manager\nalice@company.com,Alice,Engineering,3,Berlin,Marcus\nbob@company.com,Bob,Sales,1,New York,Jennifer\n```\n\nWhen Alice opens her personalized link, the AI knows she has been in Engineering for 3 years based in Berlin. Context it can reference: \"As someone with three years on the engineering team, what has kept you here this long?\" versus a cold \"Why do you like working here?\"\n\n## Privacy and Trust: The Foundation of Employee Research\n\nEmployee research only works if participants trust the process. Koji supports privacy-forward research design at every step:\n\n**Anonymous mode**: Remove the name field from your personalized links — participants know their responses will not be individually attributed to them.\n\n**Aggregate-only reporting**: Koji reports show themes across all respondents, not individual-level breakdowns by default. You see \"40% of participants mentioned career development concerns\" not \"Alice said she is thinking of leaving.\"\n\n**Custom consent language**: Edit the intake form to explain exactly how data is used, who can access it, and how long it is retained.\n\n**GDPR compliance**: Koji's data processing infrastructure is compliant with European privacy regulations.\n\nCommunicating these protections explicitly in your invitation message is one of the single most effective things you can do to improve participation rates and response candor.\n\n## Analyzing Employee Research with Insights Chat\n\nAfter collecting employee interviews, Koji's Insights Dashboard gives you an aggregate view of themes, sentiment patterns, and quantitative distributions. The [Insights Chat](/docs/insights-chat-guide) lets you query the data conversationally:\n\n- \"What do employees in the Engineering department say about their career development opportunities?\"\n- \"Which themes are more prevalent in employees with less than one year of tenure?\"\n- \"What are the top three reasons employees said they considered leaving?\"\n\nThese are questions that would take hours to answer by reading transcripts manually. With Insights Chat, they take seconds.\n\n## Getting Started with Your First Employee Study\n\n1. **Define your research question**: What decision are you trying to make? (\"Why is voluntary turnover highest in Customer Success?\")\n2. **Design your study**: Mix open-ended questions for qualitative texture with [structured questions](/docs/structured-questions-guide) for trackable metrics\n3. **Import your employee list**: Include segmentation attributes in your CSV (department, tenure, location)\n4. **Distribute via Slack or email**: Personalized links can be sent through any communication channel your team uses\n5. **Review your report**: Koji generates structured analysis after each batch of interviews\n6. **Share findings**: Use [Publishing and Sharing Reports](/docs/publishing-sharing-reports) to distribute findings to leadership\n\nFor voice interviews, see [Setting Up Voice Interviews](/docs/setting-up-voice-interviews). Voice is especially effective for sensitive HR topics where employees may find speaking more natural and candid than typing.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — Add quantitative metrics like satisfaction scales, yes/no questions, and feature rankings to your employee research\n- [Personalized Interview Links](/docs/personalized-interview-links) — Send named, contextualized invitations to individual employees with department and tenure pre-filled\n- [Importing Participants via CSV](/docs/importing-participants-csv) — Set up your employee list with segmentation attributes\n- [Stay Interviews at Scale](/docs/stay-interviews-at-scale) — Deep dive on AI-powered stay interview programs\n- [Insights Chat Guide](/docs/insights-chat-guide) — Query your employee research data with AI after collection\n- [Setting Up Voice Interviews](/docs/setting-up-voice-interviews) — Enable voice mode for more candid employee responses\n\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [B2C User Research: How to Understand Consumer Behavior at Scale (2026)](/blog/b2c-user-research-guide-2026) — B2C user research is systematically underinvested at most consumer companies. While B2B teams run structured customer discovery as a matter \n- [Customer Research Done Right: A Complete Guide for Product Teams](/blog/customer-research-done-right-a-complete-guide-for-product-teams) — Customer research is the foundation of every successful product decision. Learn the types, methods, and best practices that help product tea\n\n<!-- further-reading:blog -->\n","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Koji for HR and People Teams: Employee Research at Scale | Koji Docs","metaDescription":"Run stay interviews, exit interviews, onboarding feedback, and DEI research at scale — no scheduling required. Koji AI moderates every conversation automatically and generates instant reports segmented by department, tenure, and location.","keywords":["employee research tool AI","AI employee interviews","HR research platform","stay interviews AI","exit interview automation","people team research","employee listening program AI"],"aiSummary":"Koji helps HR and People teams run scalable employee listening programs including stay interviews, exit interviews, onboarding feedback, and DEI research. No scheduling required — employees click a personalized link and complete an AI-moderated conversation on their own schedule. Reports are automatically segmented by department, tenure, and location. Works best with personalized CSV imports and a mix of open-ended and structured questions for qualitative depth plus trackable metrics.","aiPrerequisites":["Koji account on any plan","Employee list exportable as CSV from your HRIS"],"aiLearningOutcomes":["Set up a stay interview program with AI moderation","Automate exit interview collection in your offboarding flow","Run longitudinal onboarding research at scale","Design DEI and culture research with anonymous AI interviews","Analyze employee research by department and tenure cohort"],"aiDifficulty":"beginner","aiEstimatedTime":"10 minutes"},{"type":"documentation","id":"6abe5408-a45b-46cd-b044-c5a9c70c3b49","slug":"employee-ai-adoption-research","title":"Employee AI Adoption Research: How to Find Out How Your Team Actually Uses AI (2026)","url":"https://www.koji.so/docs/employee-ai-adoption-research","summary":"Employee AI adoption cannot be measured with licence telemetry or named HR surveys: seat data captures access rather than value, and attributable surveys under-report usage because a large share of employees conceal AI use or use unapproved tools. The reliable method is a short, genuinely anonymous conversational study that reconstructs specific episodes — which tool, which task, what was pasted in, what happened next — and answers four questions: what work is handed over, whether the pattern is substitution or augmentation or rework, what the verification burden costs, and what blocks or hides usage. Sample everyone rather than licence holders only, run 12-18 minute asynchronous interviews in text or voice, promise and honour amnesty, set a minimum segment cell size before launch, and repeat quarterly. Koji fits because an AI moderator removes the colleague from the room, probes each answer up to three times automatically, and mixes all six structured question types with open-ended probing in one session, producing distributions and themed quotes together.","content":"# Employee AI Adoption Research: How to Find Out How Your Team Actually Uses AI (2026)\n\n**The bottom line:** Seat licences and login dashboards measure access, not adoption. The interesting behaviour — which tasks people actually hand to AI, which outputs they quietly rewrite, and which unapproved tools they use because the sanctioned one is slower — is invisible to telemetry and under-reported in named HR surveys. The fastest way to see it is a short, genuinely anonymous conversational study that probes specifics: which tool, which task, what got pasted in, what happened next. Platforms like Koji run that study as an AI-moderated interview, which removes the manager from the room and, with it, most of the incentive to give the safe answer.\n\n---\n\n## Why your AI adoption numbers are wrong\n\nMost organisations measure AI adoption three ways, and all three mislead:\n\n1. **Licence utilisation.** Counts activations and weekly active users. A person who opens the assistant, asks one question, dislikes the answer, and returns to their old workflow looks identical to a power user.\n2. **Prompt volume.** Counts messages. It cannot distinguish a genuine work task from experimentation, and it rises fastest during the honeymoon period — the [novelty effect](/docs/novelty-effect) in its purest form.\n3. **A named engagement survey question.** Asks \"do you use approved AI tools?\" of people who know their answer is attributable and who suspect there is a policy they may have broken.\n\nThe gap between these measures and reality is not small. Public workplace research through 2025 and 2026 has consistently found that roughly half of US workers now use AI in their jobs, while a comparable share say they have used AI at work without telling their employer, and a striking proportion of professionals report using tools they believed were not permitted under company policy. Some surveys put the share of employees who conceal AI use — often out of fear of looking replaceable or of looking lazy — at close to half. Whatever the precise figure in your organisation, the direction is reliable: **self-reported, attributable AI usage understates real usage, and the underestimate is largest exactly where the risk sits.**\n\nThat has three practical consequences:\n\n- **Governance gaps look smaller than they are.** If a third of your data-handling exposure lives in personal accounts, a security review built on licence data will not find it.\n- **Enablement money goes to the wrong place.** Teams buy more seats when the binding constraint is that nobody trusts the output enough to skip the manual re-check.\n- **Productivity claims cannot be defended.** \"AI saved us 4,000 hours\" collapses the first time a CFO asks how the number was derived.\n\n---\n\n## The four questions worth researching\n\nA useful employee AI study does not ask \"do you use AI?\" It reconstructs specific episodes. Four questions carry most of the value:\n\n### 1. What work is actually being handed over?\nNot \"do you use AI for writing\" but \"walk me through the last document you produced with AI help — what did you type, what came back, what did you change?\" Task-level detail is what turns into enablement content and workflow redesign.\n\n### 2. Is it substitution, augmentation, or theatre?\nThree very different patterns hide behind the same usage metric: AI **replaces** a task, AI **accelerates** a task the person still owns, or AI produces something the person then rebuilds by hand. The third is common and expensive, and it never shows up as a negative signal in a dashboard.\n\n### 3. What is the verification burden?\nEvery AI output carries a checking cost. The honest question is whether checking takes less time than doing the task unaided. Where verification exceeds the saving, people abandon the tool quietly — and describe the tool as \"fine\" in surveys.\n\n### 4. What is blocking or hiding usage?\nUnclear policy, fear of judgement, a sanctioned tool that lacks the model or the integration people want, data-sensitivity worries, or simply no time to learn. Blockers and concealment are the same research question asked from two directions.\n\n---\n\n## Why anonymous AI-moderated interviews work here\n\nThis is a topic where the method determines whether the data is worth anything. Three properties matter:\n\n**No human on the other side.** Admitting that you pasted a customer list into a personal AI account is not something most people will say to their manager, an HR business partner, or an internal researcher whose name they recognise. [Social desirability bias](/docs/social-desirability-bias) is at its strongest when the behaviour is both common and technically against the rules. An AI moderator is not a colleague, does not have a promotion decision, and does not react — and participants consistently disclose more to it in exactly these sensitive domains.\n\n**Depth without scheduling.** Reconstructing a real episode needs follow-up questions: which tool, which task, what did you do with the output. A static form cannot ask them; a human interviewer can, but you will not get 200 half-hour slots across engineering, finance, and support. Koji's AI moderator probes each answer automatically — up to three follow-ups per question, tuned per question — so a 15-minute asynchronous conversation produces the texture of a moderated interview at survey scale.\n\n**Structure where you need to count.** Adoption research has to produce both narrative and numbers. Koji's six [structured question types](/docs/structured-questions-guide) — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — let you mix them in a single conversation, so you get a clean distribution for \"how often do you use AI for this task\" and a probed explanation of *why* in the same session. Traditional survey platforms make you pick one or the other; running a SurveyMonkey form for the counts and a separate interview round for the depth doubles the timeline and halves the sample.\n\n**One caveat worth stating plainly:** anonymity has to be real, and it has to be operationally true. If you collect department, tenure, seniority, and location, some respondents become uniquely identifiable in a 400-person company. Decide your minimum cell size before you launch and drop segment cuts that fall below it.\n\n---\n\n## Designing the study\n\n**Population.** Everyone, not just the people with licences. Non-users and abandoners are the most informative segment, and they are systematically missing from tool-based sampling — a textbook case of [survivorship bias](/docs/survivorship-bias-customer-research).\n\n**Length.** 12 to 18 minutes. Long enough for two probed episodes, short enough to hold completion rates.\n\n**Mode.** Text for sensitive disclosure and for anyone at a shared desk; voice where you want richer narrative from field or frontline staff. Koji supports both from the same interview link, and participants choose.\n\n**Amnesty framing.** Say, in the invitation and in the first screen, that the purpose is to fix tooling and policy rather than to identify individuals, and that nothing said will be used in performance decisions. Then honour it. One breach of that promise ends your ability to measure this topic for years.\n\n**Cadence.** Quarterly waves with a stable core question set. AI capability changes fast enough that annual measurement is useless, and the trend line is more decision-useful than any single wave.\n\n---\n\n## The question bank\n\nAdapt the wording; keep the structure. Types map directly to Koji question types.\n\n**Baseline (structured)**\n- `single_choice` — In a typical week, how often do you use any AI tool for work? (Daily / A few times a week / A few times a month / Rarely / Never)\n- `multiple_choice` — Which of these have you used for work in the last month? (Approved assistant / ChatGPT or similar personal account / AI features inside another tool / Coding assistant / AI notetaker / None)\n- `yes_no` — Have you ever avoided mentioning AI use in a work context? *(Probe on yes: what made you hold back?)*\n\n**Episode reconstruction (open-ended with probing)**\n- Walk me through the last real work task where AI helped. What did you ask for, and what did you do with what came back?\n- Tell me about a time an AI output was wrong or unusable. How did you notice, and what did it cost you?\n- Is there a task you tried to hand to AI and then took back? What went wrong?\n\n**Value and verification**\n- `scale` (1–7) — For that task, how much time did AI actually save you, net of checking the output? *(Anchor probe: what would have to change to move that up two points?)*\n- `scale` (0–10) — How confident are you that you can tell when an AI output is wrong in your domain?\n- `open_ended` — What do you always check before you use an AI output in front of a customer or an executive?\n\n**Barriers and governance**\n- `ranking` — Rank what most limits your AI use: unclear policy, output quality, data sensitivity, no time to learn, missing integrations, no need.\n- `open_ended` — If policy allowed anything, what would you want to use AI for that you cannot today?\n- `yes_no` — Do you know where to find your organisation's AI policy? *(A brutally clarifying question. The yes rate is usually well below what the policy owner expects.)*\n\n**Enablement**\n- `open_ended` — Who on your team is best at this, and what do they do differently?\n- `single_choice` — What would help most: examples for my role, a better tool, clearer rules, or protected time to learn?\n\n---\n\n## Turning it into decisions\n\nFour outputs justify the study:\n\n1. **A task map, not a tool map.** Rank tasks by frequency × reported net time saved. Enablement content should follow that ranking, not the vendor's feature list.\n2. **A verification-cost list.** Tasks where checking eats the saving are candidates for workflow redesign, better context grounding, or removal from the AI story entirely.\n3. **A shadow-AI exposure picture.** Which unapproved tools, for which data classes, driven by which missing capability. This is the input to procurement, and it is usually cheaper to close the capability gap than to police the behaviour. Pair it with your [data retention and deletion](/docs/research-data-retention-deletion) rules.\n4. **A defensible productivity narrative.** Self-reported net time saved, segmented, with sample sizes and the anonymity caveat attached — far more credible in front of a board than a vendor-supplied multiplier.\n\nKoji generates the analysis layer automatically: per-question distributions for the structured items, themes with supporting quotes for the open-ended ones, and a report you can refresh as later waves land — so quarter-over-quarter comparison is a re-read of the same report rather than a fresh analysis project.\n\n---\n\n## Common mistakes\n\n- **Asking only licence holders.** Guarantees an inflated adoption rate and no barrier data.\n- **Naming the respondent.** Halves your disclosure rate on precisely the questions that matter.\n- **Leading with policy.** Opening with \"our policy states…\" turns the interview into a compliance quiz. Ask about behaviour first, policy last.\n- **Treating counts as census.** A 200-person voluntary study is a strong signal, not a population statistic. Report it as such.\n- **Measuring once.** A single wave during a rollout captures the novelty peak and nothing else.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types and when to use each\n- [Anonymous Employee Research with AI Interviews](/docs/anonymous-employee-research-ai-interviews) — running internal studies where honesty depends on anonymity\n- [Social Desirability Bias](/docs/social-desirability-bias) — why attributable surveys under-report sensitive behaviour\n- [Stay Interviews at Scale](/docs/stay-interviews-at-scale) — the same anonymous-interview pattern applied to retention\n- [User Research for AI Products](/docs/user-research-for-ai-products) — the external-facing counterpart: researching trust in AI features you ship\n- [Scale Questions in AI Interviews](/docs/scale-questions-guide) — building the quantitative backbone of a tracking study\n","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Employee AI Adoption Research: Measure Real Usage (2026)","metaDescription":"Licence dashboards measure access, not adoption. How to research real employee AI usage — including shadow AI — with anonymous AI-moderated interviews, a full question bank, and a quarterly tracking model.","keywords":["employee ai adoption research","ai adoption survey questions","employee ai usage survey","shadow ai","measure ai adoption at work","internal ai rollout research","ai adoption metrics","anonymous employee ai survey","ai enablement research","ai policy compliance survey"],"aiSummary":"Employee AI adoption cannot be measured with licence telemetry or named HR surveys: seat data captures access rather than value, and attributable surveys under-report usage because a large share of employees conceal AI use or use unapproved tools. The reliable method is a short, genuinely anonymous conversational study that reconstructs specific episodes — which tool, which task, what was pasted in, what happened next — and answers four questions: what work is handed over, whether the pattern is substitution or augmentation or rework, what the verification burden costs, and what blocks or hides usage. Sample everyone rather than licence holders only, run 12-18 minute asynchronous interviews in text or voice, promise and honour amnesty, set a minimum segment cell size before launch, and repeat quarterly. Koji fits because an AI moderator removes the colleague from the room, probes each answer up to three times automatically, and mixes all six structured question types with open-ended probing in one session, producing distributions and themed quotes together.","aiPrerequisites":["A defined internal AI rollout or AI policy","Ability to invite employees to an anonymous study"],"aiLearningOutcomes":["Explain why licence and prompt-volume metrics overstate real adoption","Design an anonymous employee AI study that surfaces shadow AI","Write episode-reconstruction questions that produce task-level detail","Separate substitution, augmentation, and rework in your usage data","Build a quarterly tracking cadence with a stable core question set"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"b72806a8-9a9b-493b-b644-6b68c5524025","slug":"churn-interview-questions","title":"Churn Interviews: 20 Questions to Uncover Why Customers Really Leave (2026)","url":"https://www.koji.so/docs/churn-interview-questions","summary":"A guide to qualitative churn interviews: how they differ from churn surveys, which churned segments to interview, a push-pull (Jobs to be Done) question framework with 20 non-leading example questions, how to code and rank churn reasons, and how Koji runs churn interviews automatically with an AI moderator.","content":"## What is a churn interview, and why run one? (Answer first)\n\n**A churn interview is a qualitative conversation with a customer who has cancelled, downgraded, or gone dormant — designed to uncover the real reason they left, not just the reason they ticked on a cancellation form.** A one-click exit survey tells you *what* category the churn fell into (\"too expensive\"). A churn interview tells you the story behind it: what changed, what they switched to, and what would have kept them. That story is where your retention roadmap actually comes from.\n\nThe catch is that churned customers are the hardest people to get on a call — they have already left, and goodwill is low. That is why most teams settle for a survey and never learn the \"why.\" Platforms like Koji close that gap: you drop a single conversational link into your cancellation flow or win-back email, and Koji's AI interviewer runs the full probing conversation 24/7 — no scheduling, no moderator, no awkward call. You get interview-grade depth at survey-grade reach.\n\n> **Bottom line:** Surveys quantify churn reasons; interviews explain them. Run interviews when you need to *fix* churn, not just measure it. Combine the two — a structured cancellation survey that branches into an AI-led interview — and you get both the number and the narrative in one flow.\n\n## Churn interview vs. churn survey: which do you need?\n\n| | Churn survey | Churn interview |\n|---|---|---|\n| Output | Categorized reasons, % breakdown | Root cause, narrative, switching story |\n| Depth | Shallow | Deep — probes the \"why behind the why\" |\n| Best for | Trend tracking, dashboards | Roadmap decisions, save-flow design |\n| Effort (traditional) | Low | High (scheduling, moderating) |\n| Effort (with Koji) | Low | Low — AI moderates automatically |\n\nIf you only have a survey today, start with [How to Build Churn Surveys That Actually Save Customers](/docs/churn-survey-guide), then layer interviews on top for the segments that matter most.\n\n## Who to interview (and when)\n\nNot all churn is the same, and the segments tell different stories:\n\n- **Voluntary churn** — they actively cancelled. Interview within 3–7 days, while the decision is fresh.\n- **Involuntary churn** — payment failed, they did not return. Interview to learn whether they even meant to leave.\n- **Dormant / silent churn** — still subscribed but inactive. The most valuable and most overlooked group — catch them *before* they cancel. See [Dormant User Reactivation Research](/docs/dormant-user-reactivation-research).\n- **Downgrades** — they stayed but spent less. The interview reveals which value they stopped seeing.\n\nAim for 8–12 interviews per segment. Qualitative saturation — the point where new interviews stop surfacing new reasons — usually arrives around 10–15 conversations per coherent segment.\n\n## The framework: borrow the \"push and pull\" from Jobs to be Done\n\nThe most useful churn interviews are structured around the forces that drove the switch, an idea from the [Jobs to Be Done Framework](/docs/jobs-to-be-done-framework):\n\n- **Push** — what about your product (or their situation) pushed them to look elsewhere?\n- **Pull** — what attracted them to the alternative (which may be a competitor, a workaround, or \"doing nothing\")?\n- **Anxiety** — what almost stopped them from leaving? (This is your retention lever.)\n- **Habit** — what made staying feel costly or pointless?\n\nMapping answers to these four forces turns a vague \"they left because of price\" into \"the product stopped delivering weekly value (push), a free tool covered their shrunken use case (pull), and nothing in our experience reminded them what they would lose (no anxiety).\"\n\n## 20 churn interview questions that work\n\n**Opening (context, not blame)**\n1. Take me back to when you first signed up — what were you hoping it would do for you?\n2. Walk me through how you were using it in a typical week.\n\n**The trigger (the push)**\n3. When did you first start thinking about cancelling?\n4. What was happening at that point — what changed?\n5. What was the most frustrating part of using us toward the end?\n\n**Alternatives (the pull)**\n6. What are you using now instead? (Including \"nothing.\")\n7. What does that do better for you?\n8. How did you find or decide on it?\n\n**The decision (anxiety and habit)**\n9. What almost made you stay?\n10. Was there a moment we could have done something differently?\n11. How hard was the decision to leave — quick, or drawn out?\n\n**Value and price**\n12. When you think about what you paid, what were you really paying for?\n13. At what point did it stop feeling worth it?\n\n**Experience**\n14. How was getting started in the first place?\n15. Did you ever reach out to support? How did that go?\n16. Was there a feature you expected that we did not have?\n\n**The counterfactual (your roadmap)**\n17. What would have had to be true for you to stay?\n18. If we fixed one thing, what should it be?\n\n**Closing**\n19. Would you consider coming back? Under what conditions?\n20. Is there anything I should have asked but did not?\n\nNotice none of these are leading — see [How to Avoid Leading Questions](/docs/avoiding-leading-questions). \"What frustrated you?\" assumes nothing; \"Was it too expensive?\" plants the answer.\n\n## How to analyze churn interviews\n\nRun every transcript through a consistent coding pass: tag each mentioned reason, map it to push/pull/anxiety/habit, and cluster near-duplicate reasons into a canonical list. Then weight by segment — a reason cited by ten dormant power users matters more than one cited by a single trial tire-kicker. The output you want is a ranked list of *fixable* churn drivers, each backed by verbatim quotes.\n\nThis is exactly the synthesis Koji does for you. Every interview is transcribed, the open-ended answers are coded into themes, and near-identical reasons are merged into a real-time **churn-reason report** with the supporting quotes attached — no spreadsheet, no manual tagging.\n\n## How Koji automates churn interviews\n\n1. **Reach customers who would never take a call.** Drop a Koji link into your cancellation flow, win-back email, or downgrade confirmation. The AI interviewer runs the conversation whenever they choose to engage.\n2. **Mix structured and open questions.** Koji supports six **structured question types** (open_ended, scale, single_choice, multiple_choice, ranking, yes_no). Ask \"primary reason for leaving\" as a `single_choice`, satisfaction as a `scale`, and \"tell me more about that\" as an `open_ended` question with adaptive AI follow-ups.\n3. **Probe automatically.** When a customer says \"it got too expensive,\" Koji follows up — \"expensive relative to what?\" — surfacing the real story instead of stopping at the label.\n4. **Voice or text.** Some churned users will talk; others will type. Koji handles both.\n5. **Real-time reporting.** Watch churn reasons cluster and rank as responses arrive, then share the report with one link.\n\nThe result: the depth of a moderated exit interview, at the scale and cost of a survey.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types behind hybrid churn studies\n- [How to Build Churn Surveys That Actually Save Customers](/docs/churn-survey-guide)\n- [The B2B SaaS Retention Interview Template](/docs/b2b-saas-retention-interview-template)\n- [Dormant User Reactivation Research](/docs/dormant-user-reactivation-research)\n- [Jobs to Be Done Framework: The Complete Guide](/docs/jobs-to-be-done-framework)\n- [Closing the Loop on Customer Feedback](/docs/closing-the-loop-customer-feedback)\n","category":"Use Cases","lastModified":"2026-08-01T03:21:55.717966+00:00","metaTitle":"Churn Interviews: 20 Questions to Find Why Customers Leave (2026)","metaDescription":"Run customer churn interviews that reveal the real reason customers leave: survey vs interview, who to talk to, 20 non-leading questions in a push-pull framework, and how to automate churn interviews with AI on Koji.","keywords":["churn interview","churn interview questions","customer churn interview","exit interview questions","why customers leave","churn analysis","retention research","cancellation interview","customer retention interviews"],"aiSummary":"A guide to qualitative churn interviews: how they differ from churn surveys, which churned segments to interview, a push-pull (Jobs to be Done) question framework with 20 non-leading example questions, how to code and rank churn reasons, and how Koji runs churn interviews automatically with an AI moderator.","aiPrerequisites":["A product with customers who have churned or downgraded","Access to cancellation or win-back touchpoints","Basic familiarity with customer interviews"],"aiLearningOutcomes":["Decide when to run a churn interview vs a churn survey","Identify which churned segments to prioritize","Use the push-pull framework to structure churn questions","Ask 20 non-leading churn questions","Code and rank fixable churn drivers","Automate churn interviews with Koji AI moderation"],"aiDifficulty":"beginner","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"6a4bc130-fb28-4460-8dac-80e693fca038","slug":"post-merger-integration-customer-research","title":"Post-Merger Integration Customer Research: Keeping the Customers You Just Paid For","url":"https://www.koji.so/docs/post-merger-integration-customer-research","summary":"Commercial due diligence assesses a customer base that has not yet been asked to change anything; the events that actually drive post-acquisition churn — the announcement, account team changes, platform migration and commercial harmonisation — all occur after close. Bain notes that mergers can trigger rapid customer attrition and that the remedy is taking the customer view when making integration decisions, while large deals take months to close and two to three years to realise run-rate synergies, leaving a long window in which relationships erode. Post-close research must answer five questions: relationship continuity risk, honestly measured switching cost, migration tolerance, brand and naming reaction, and whether the modelled cross-sell synergy is real. A workable cadence runs an announcement pulse in days 0-15, a relationship and commitment audit at day 30, a migration tolerance study at day 45-60, cross-sell validation at day 60, a brand check at day 90, then a quarterly combined-base tracker. AI-moderated interviews make the cadence deliverable, and Koji structured question types provide an identical quantitative spine across waves so acquirer and acquired bases can be compared directly. Results must be segmented by base, coverage change, migration exposure, contract timing and size band.","content":"**Answer first: the customer research that determines whether a deal works happens after it closes, not before it — and almost nobody runs it.** Commercial due diligence answers \"should we buy this?\" It is a point-in-time assessment of a customer base that has not yet been asked to change billing systems, learn a new product name, accept a new price, or say goodbye to the account manager they trusted. Every one of those events happens during integration, and each one is a fresh opportunity for a customer to reconsider. Bain's work on merger integration makes the point plainly: mergers can trigger customer attrition quickly, and the way to avoid it is to adopt the customer's view of the merger when making integration decisions. Adopting the customer's view requires actually asking them, repeatedly, while there is still time to change course.\n\nThis guide covers what post-close research has to answer that diligence cannot, the four integration events that generate churn, a cadence you can actually run, and how to structure the studies so the answers are comparable over time.\n\n## Why diligence evidence expires at close\n\nDiligence interviews are conducted with a static proposition: how do you feel about this product, this vendor, this relationship, as things stand. They are also, structurally, a small sample — a handful of reference calls, often arranged through the seller, weighted toward customers willing to speak well of the company.\n\nThe moment the deal closes, the proposition changes. The questions that now decide revenue are ones diligence never asked:\n\n- Does the customer know the acquisition happened, and what did they conclude from how they found out?\n- Which specific commitment made by their old account team do they believe still stands?\n- What would they do if the product name changed? If the price rose 8%? If their integration needed re-authenticating?\n- Is there a competitor already in the account because of the news?\n\nNone of these are answerable from a pre-close data room. All of them are answerable in a week of interviews — see [commercial due diligence customer interviews](/docs/commercial-due-diligence-customer-interviews) for the pre-close counterpart to this work.\n\n## The four integration events that generate churn\n\nAttrition after an acquisition is rarely a slow drift. It clusters around discrete events, which is convenient, because it means you can research ahead of each one.\n\n| Event | The customer's actual question | Research window |\n|---|---|---|\n| **The announcement** | \"Am I still a priority to anyone?\" | Days 1–30 |\n| **Account team changes** | \"Who do I call now, and do they know my history?\" | Whenever coverage is remapped |\n| **Product and platform migration** | \"How much work is this going to cost me?\" | 60–90 days before any forced migration |\n| **Commercial harmonisation** | \"Am I being repriced, and is it still worth it?\" | Before, not after, the renewal notice |\n\nThe migration window is the expensive one. Truist, formed by the 2019 BB&T and SunTrust merger, is the standard cautionary example: moving customers onto a different digital platform alongside branch rebranding and cost decisions took longer than anticipated and contributed to significant customer-service problems. The lesson is not that migrations are avoidable — it is that the customer's tolerance for one has to be measured before it is scheduled, not discovered afterwards in the support queue.\n\nTiming matters more than volume. Large deals typically take months from announcement to close and then a further two to three years to realise the bulk of run-rate cost synergies, which means the integration period is long enough for a customer relationship to erode quietly and long enough for research to catch it.\n\n## The five questions post-close research must answer\n\n**1. Relationship continuity risk.** Which relationships were with a *person* rather than with the *product*? When that person leaves — and in integrations, many do — the account's retention probability changes. Ask customers to describe who they rely on and what would happen if that person were gone.\n\n**2. Switching cost, honestly measured.** Diligence estimates switching cost from the outside: data volume, contract length, integration depth. Only the customer knows whether they have already scoped a migration, whether a competitor has offered to fund it, and whether an internal champion has left. High measured switching cost with low perceived value is the most dangerous combination in the book, and it is invisible in usage data.\n\n**3. Migration tolerance.** Not \"will you accept a migration\" — everyone says yes in the abstract — but what specifically breaks: which integration, which workflow, which compliance approval would have to be re-obtained. This is the finding that most often changes an integration plan.\n\n**4. Brand and naming reaction.** If you are retiring a brand your customers chose deliberately, find out what they think it stood for before you replace it. Brand migration failures are usually failures to understand what the old name signalled — see [brand tracking studies](/docs/brand-tracking-study-guide) for the measurement approach.\n\n**5. Cross-sell reality.** Deal models routinely carry revenue synergies from selling the acquirer's products to the acquired base. That assumption is testable in week two: do these customers have the problem the other product solves, do they already buy a competing solution, and would they buy it from you *now*, during integration, or only later?\n\n## The cadence\n\nThe failure mode is a single \"customer sentiment study\" at day 90, by which point the migration plan is locked and the churn is priced in. Run a rhythm instead.\n\n| Timing | Study | Population | Purpose |\n|---|---|---|---|\n| Day 0–15 | Announcement pulse | Top accounts by revenue plus a random sample of the long tail | Detect immediate flight risk and mis-set expectations |\n| Day 30 | Relationship and commitment audit | Accounts whose coverage changed | Surface promises made pre-close that nobody has recorded |\n| Day 45–60 | Migration tolerance study | Anyone facing a platform, contract or workflow change | Set the migration sequence and the exception list |\n| Day 60 | Cross-sell validation | Acquired base, segmented | Test the revenue synergy before the sales team is retargeted |\n| Day 90 | Brand and positioning check | Both bases | Decide on brand retirement, timing and messaging |\n| Quarterly | Combined-base tracker | Both bases, matched questions | Compare the two customer bases on identical measures |\n| Post-migration | Experience debrief | Everyone who migrated | Fix the next wave before it happens |\n\nThe long tail is not optional. Integration research defaults to the top twenty accounts because they are easy to reach, but the long tail is where attrition is silent, unmanaged and, in aggregate, frequently larger.\n\n## Why this is a research automation problem\n\nThe reason almost nobody runs this cadence is arithmetic. Seven studies in ninety days, across two customer bases, in multiple languages, while the integration team is also doing the integration, is simply not deliverable with scheduled one-hour calls. A traditional programme means recruiters, calendars, moderators, transcription and a synthesis backlog — and the finding arrives after the decision it was meant to inform.\n\nThis is exactly what AI-moderated interviews are for. With a platform like Koji, each study in the table is a link sent to a segment; the AI conducts a real conversation, asks follow-up questions when an answer is thin, and the analysis is generated as responses arrive. A migration tolerance study can go from \"we need to know\" to \"here is what 200 customers said\" inside a week, with no moderator in the loop and no scheduling at all. Interviews run in voice or text, so a busy operations manager answers at 11pm and a CTO talks through it hands-free.\n\nTwo capabilities matter specifically for integration work.\n\n**Comparability across studies and bases.** Koji's six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — let you carry an identical quantitative spine across every wave. Migration tolerance as a scale, blocking factors as a ranking, competitive contact as `yes_no`, and open-ended probing on the reasoning behind each. That means you can say \"migration tolerance fell from 7.2 to 5.8 in the accounts whose account manager changed\" — a sentence a survey of free-text responses can never produce, and exactly the sentence an integration steering committee can act on. Running the same instrument across the acquirer's base and the acquired base is what turns two customer sets into one comparable picture.\n\n**Depth without a moderator.** The reason integration teams distrust surveys is that surveys return \"pricing\" when the truth is \"pricing, because the person who justified it internally left and nobody else can defend the renewal.\" Koji's AI probes that second layer automatically, which is the difference between a number that describes churn and a finding that prevents it.\n\n## Segments that must be reported separately\n\nAggregate integration research hides the risk. Always split by:\n\n- **Base** — acquirer versus acquired. Merging them into a single average is the most common analytical error in integration research.\n- **Coverage change** — accounts whose account manager or CSM changed versus those unaffected. This is usually the sharpest signal in the dataset.\n- **Migration exposure** — facing a forced change versus not.\n- **Contract timing** — renewal inside 6 months versus later. Attrition is only visible at renewal, so this segment is your early warning system.\n- **Size band** — because the long tail behaves differently and is managed less.\n\n## Turning findings into integration decisions\n\nResearch that does not change a plan is overhead. Wire each study to a specific decision owner and a specific reversible action:\n\n| Finding | Decision it should change |\n|---|---|\n| Migration tolerance low in a segment | Sequencing — migrate willing cohorts first, build an exception path for the rest |\n| Commitments made pre-close and now unrecorded | An explicit honour-or-renegotiate list, handled deliberately rather than by surprise |\n| Relationship risk concentrated on departing staff | Retention packages, or a structured handover with the customer present |\n| Cross-sell hypothesis not supported | Reforecast the revenue synergy now, not at the year-end review |\n| Brand equity in the retiring name | Extend a dual-brand period or change the messaging |\n\nThe reason to run this early is that all five of those actions are cheap in month one and expensive in month nine.\n\n## Frequently asked questions\n\n**How is this different from commercial due diligence customer interviews?**\nDiligence is pre-close and answers whether to buy at what price, using a small, often seller-arranged sample of customers facing no change. Post-merger integration research is post-close, larger-sample, repeated, and answers whether specific integration decisions — migration timing, rebranding, repricing, coverage changes — will cost you customers while those decisions can still be changed.\n\n**When should the first post-close study run?**\nWithin the first fifteen days. The announcement itself is an event customers react to, and the reaction sets the tone for everything after. Waiting until day 90 means measuring damage rather than preventing it.\n\n**Should we tell customers we are researching because of the acquisition?**\nYes. They already know the acquisition happened, and pretending otherwise reads as evasive at precisely the moment you need to look reliable. Being asked \"what worries you about this change\" is one of the few reassuring things that happens to a customer during an integration.\n\n**How many customers do we need to interview?**\nMore than diligence used, and stratified rather than cherry-picked. Cover your top accounts comprehensively and take a genuine random sample of the long tail — the tail is where unmanaged churn lives. Because AI-moderated interviews remove the scheduling constraint, sample size stops being the binding limit; segment coverage becomes the thing to plan around.\n\n**Can we research the acquired base before the deal closes?**\nGenerally not directly, and not without counsel — pre-close contact with a target's customers raises confidentiality, gun-jumping and competition-law issues. Run the design work pre-close so studies launch on day one, and keep pre-close customer contact inside whatever the deal documents and your lawyers permit.\n\n**What if the two customer bases have completely different products?**\nThen the comparable spine is not product satisfaction, it is relationship and change tolerance: continuity of contact, perceived commitment, willingness to accept change, and competitive exposure. Keep those items identical across both bases and let the product-specific questions differ.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make waves comparable over time\n- [Commercial Due Diligence Customer Interviews](/docs/commercial-due-diligence-customer-interviews) — the pre-close counterpart to this playbook\n- [Customer Research for Investors](/docs/customer-research-for-investors) — customer diligence from the investor's side\n- [Customer Retention Research](/docs/customer-retention-research) — the underlying retention methodology\n- [Churn Interviews: 20 Questions](/docs/churn-interview-questions) — what to ask customers who are leaving or considering it\n- [Brand Tracking Study Guide](/docs/brand-tracking-study-guide) — measuring brand equity before you retire a name\n- [B2B User Research](/docs/b2b-user-research-guide) — interviewing enterprise accounts without the scheduling nightmare","category":"Use Cases","lastModified":"2026-08-01T03:21:21.209322+00:00","metaTitle":"Post-Merger Integration Customer Research: The Post-Close Playbook","metaDescription":"Diligence ends at close; the churn starts after. The seven-study cadence that catches integration attrition — migration tolerance, relationship risk and cross-sell reality — while it is still reversible.","keywords":["post-merger integration customer research","post-acquisition customer interviews","M&A customer retention","integration customer churn","migration tolerance research","brand migration research","PMI customer experience","post-close customer research"],"aiSummary":"Commercial due diligence assesses a customer base that has not yet been asked to change anything; the events that actually drive post-acquisition churn — the announcement, account team changes, platform migration and commercial harmonisation — all occur after close. Bain notes that mergers can trigger rapid customer attrition and that the remedy is taking the customer view when making integration decisions, while large deals take months to close and two to three years to realise run-rate synergies, leaving a long window in which relationships erode. Post-close research must answer five questions: relationship continuity risk, honestly measured switching cost, migration tolerance, brand and naming reaction, and whether the modelled cross-sell synergy is real. A workable cadence runs an announcement pulse in days 0-15, a relationship and commitment audit at day 30, a migration tolerance study at day 45-60, cross-sell validation at day 60, a brand check at day 90, then a quarterly combined-base tracker. AI-moderated interviews make the cadence deliverable, and Koji structured question types provide an identical quantitative spine across waves so acquirer and acquired bases can be compared directly. Results must be segmented by base, coverage change, migration exposure, contract timing and size band.","aiPrerequisites":["A closed or imminent acquisition with an identifiable customer base","Access to customer contacts for both the acquirer and the acquired company","An integration plan with dated migration, rebrand or repricing milestones"],"aiLearningOutcomes":["Distinguish post-close customer research from pre-close commercial due diligence","Map the four integration events that cluster customer churn","Measure migration tolerance before scheduling a forced migration","Test modelled cross-sell revenue synergies within 60 days of close","Run a seven-study post-close cadence with AI-moderated interviews","Segment integration findings so silent long-tail attrition becomes visible"],"aiDifficulty":"advanced","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"82492cda-47b7-4049-a54a-24d53671d6ba","slug":"ai-auto-tagging-customer-interviews","title":"AI Auto-Tagging for Customer Interviews: Code 100 Interviews in Minutes","url":"https://www.koji.so/docs/ai-auto-tagging-customer-interviews","summary":"AI auto-tagging compresses 40-146 hours of manual qualitative coding into under 30 minutes. Koji runs a two-cycle pipeline: cycle-1 generates 1-3 descriptive codes per open-ended answer (2-5 word labels grounded in verbatim supporting quotes and message indices), then cycle-2 axial clustering at report time merges near-duplicate codes into a canonical codebook per question across all interviews. Structured question types (scale, choice, ranking, yes/no) get pre-coded automatically during the interview. Two modes: emergent (codebook derived from data, default for discovery) vs codebook-guided (predefined codes for longitudinal or cohort-comparison studies). Quality controls: per-answer confidence scores, verbatim supporting quotes, transcript traceability. Limits: sarcasm, niche jargon, very subtle theme distinctions, single-interview outliers.","content":"## The 30-Second Version\n\nAI auto-tagging is the automated application of qualitative codes to interview transcripts — and it has fundamentally changed how customer research scales. A single researcher coding 25 hour-long interviews by hand takes 40-100 hours and produces inconsistent results across the corpus. Koji's AI auto-tagging completes the same work in minutes, applies the same codebook consistently across every interview, and traces every code back to the verbatim respondent quote that justified it.\n\nThis is not generic AI summarization. It is a research-grade pipeline that performs **two-cycle coding** — descriptive cycle-1 codes per answer, then axial cycle-2 clustering across all interviews into a canonical codebook. The output is a coded dataset you can query, filter, and report on, not a free-text summary.\n\nThis guide explains what auto-tagging is, how Koji does it specifically, when to trust the output, and how to validate AI-generated codes against your own standards.\n\n## What Auto-Tagging Is — And Is Not\n\nA few terms get used interchangeably, but they mean different things:\n\n- **Tagging / coding** — applying a short atomic label to a segment of text (a sentence, paragraph, or message). Examples: \"Onboarding friction\", \"Pricing surprise\", \"Integration request\".\n- **Thematic analysis** — grouping codes into higher-level themes that answer the research question. See the [thematic analysis guide](/docs/thematic-analysis-guide) for the methodology.\n- **Auto-tagging** — the automation of the tagging step using AI.\n- **Summarization** — producing a free-text paragraph summary. This is not auto-tagging and is much weaker for research because the output is not structured or queryable.\n\nAuto-tagging is the structured input layer. Thematic analysis builds on top of it. [Insight repositories](/docs/atomic-research-nuggets-guide) store the tagged segments as atoms you can reuse across studies.\n\n## How Koji's Auto-Tagging Actually Works\n\nKoji performs auto-tagging in two passes, mirroring how a human qualitative researcher would code at scale.\n\n### Cycle-1: Descriptive Coding Per Answer\n\nFor every [open-ended question](/docs/structured-questions-guide) in every interview, Koji generates a small set of cycle-1 codes (typically 1-3 per answer). Each code includes:\n\n- A **label** — 2-5 words, in the study language (English), sentence-case. The label codes the meaning, not the verbatim words. Example: \"Convenience preference\" rather than \"they like that it's easy\".\n- A **kind** — either `descriptive` (analyst-paraphrased topic label, the default) or `in_vivo` (captures the participant's specific framing, translated to English; used sparingly when a topic label would lose nuance).\n- **Message indices** — exact pointers into the transcript so you can navigate from the code back to the source.\n- A **supporting quote** — the verbatim respondent words from the message that justified the code, kept in the participant's original language so the highlighted transcript span matches their voice.\n\nThis grounding step is what separates Koji's auto-tagging from generic AI summarization. Every code is anchored in a specific quote and a specific message, which makes it auditable and citable.\n\n### Cycle-2: Axial Clustering Across Interviews\n\nAfter every interview is cycle-1 coded, Koji performs cycle-2 axial coding during report aggregation. The job here is to **cluster near-duplicate codes into a canonical codebook** for each question across all interviews in the study.\n\nExample: across 25 interviews, cycle-1 might produce these labels for the same underlying concept:\n\n- \"Onboarding too long\"\n- \"Setup friction\"\n- \"Took too long to start\"\n- \"Slow first value\"\n\nCycle-2 clusters these into a single canonical code (e.g., \"Slow time-to-value\") and updates the report so the underlying respondent quotes are grouped, ranked, and chartable. The result is a coded dataset where you can ask \"how often does 'slow time-to-value' come up across the cohort\" and get a real answer with quotes attached.\n\n### Structured Questions Get Tags For Free\n\nThe other half of auto-tagging is that Koji's [structured question types](/docs/structured-questions-guide) — scale, single choice, multiple choice, ranking, yes/no — produce pre-coded answers automatically. There is no coding step. The AI moderator extracts the structured value (e.g., NPS = 8, ranked preferences = [Search, Filters, Settings], yes/no = yes) from natural conversation as the interview happens.\n\nSo a typical 10-question Koji interview ends up with:\n\n- 4-5 open-ended questions → cycle-1 coded automatically, then cycle-2 clustered in the report\n- 4-5 structured questions → pre-coded structured values ready to aggregate\n\nYour analysis is done by the time the interview ends.\n\n## Manual vs Auto-Tagging: The Math\n\nFor a typical mid-size qualitative study, the time savings are dramatic.\n\n| Step | Manual | Koji Auto-Tagging |\n|---|---|---|\n| Transcribe | 4-8 hr/interview | 0 — automatic during the interview |\n| Build initial codebook | 6-10 hr (sample read-through) | 0 — emerges from cycle-1 |\n| Code 25 interviews | 25-100 hr | ~10 minutes total |\n| Cluster into themes | 8-16 hr | ~minutes (axial pass) |\n| Build report | 6-12 hr | 0 — automatic |\n| **Total for 25 interviews** | **49-146 hr** | **Under 30 minutes** |\n\nA full-time qualitative researcher costs $80,000-$140,000 annually loaded. The cost of a single 25-interview manual coding pass is roughly $4,000-$10,000 in labor. Koji runs the same pass for 5 credits on the [report refresh](/docs/understanding-usage-limits), or roughly €5.\n\nThis is what makes weekly research cadences feasible. Manual coding makes you choose between depth and frequency; auto-tagging removes the trade-off.\n\n## Two Modes: Emergent vs Codebook-Guided\n\nKoji supports two modes of auto-tagging depending on how structured your research is.\n\n### Emergent mode (default)\n\nThe AI generates codes from the data without a predefined codebook. This is the right mode for:\n\n- Exploratory studies where you do not know what categories will emerge.\n- First-time research in a new domain.\n- Studies where you want to be open to surprises.\n- Most [customer discovery interviews](/docs/customer-discovery-interviews).\n\nCycle-2 clustering will still produce a clean canonical codebook in the report, but it is derived from the data rather than imposed.\n\n### Codebook-guided mode\n\nFor longitudinal studies, regulated research, or programs where you need codes to be comparable across waves, you can pre-define the codebook by:\n\n1. Specifying expected codes in your [research brief](/docs/how-to-write-research-brief) or as part of the question probing instructions.\n2. Running the study with the codebook hint included in the AI's coding prompt.\n3. Reviewing the cycle-1 codes after the first 3-5 interviews to confirm fit.\n\nThis mode trades some openness for comparability across studies — useful when you are running a quarterly customer health study or comparing cohorts over time.\n\n## Quality Controls: Trust But Verify\n\nAI auto-tagging is fast, but it is not infallible. Three quality controls let you trust the output.\n\n### Confidence scores\n\nEvery [structured answer](/docs/analyzing-ai-moderated-interview-results) carries a confidence rating (high / medium / low). Low-confidence extractions are flagged for human review. Filter the report to show only high-confidence answers when you need certainty.\n\n### Supporting quote anchoring\n\nEvery code links to the verbatim respondent quote that justified it. You can navigate from any code in the report back to the message in the transcript that produced it. This is the difference between trustable auto-tagging and a black-box summary.\n\n### Transcript traceability\n\nThe `messageIndices` field on every code points to the exact messages in the conversation. When you spot-check a theme, you can read the surrounding context, not just the highlighted snippet. This is essential for catching cases where the AI tagged correctly at the sentence level but missed the surrounding nuance.\n\nA reasonable validation cadence:\n\n- For the first 5 interviews in a new study, manually spot-check 20% of cycle-1 codes.\n- For ongoing studies, spot-check 10% per wave.\n- For high-stakes decisions (board-level reports, pricing changes), validate the top 5 themes by reading the supporting quotes directly.\n\n## Building a Codebook the AI Respects\n\nIf you want codebook-guided auto-tagging, here is what works:\n\n- **Short, conceptual labels**: 2-5 words. \"Pricing surprise\" beats \"the prospect was surprised by our pricing\".\n- **One concept per code**: \"Onboarding friction\" or \"Pricing surprise\", not \"Onboarding friction OR pricing surprise\".\n- **Define the boundary**: a one-line description of what is in scope vs. out of scope for each code. Example: \"Onboarding friction = anything in the first 7 days of product use that slowed activation. Does NOT include sales-cycle friction.\"\n- **Mix descriptive and in-vivo codes**: most codes should be descriptive (analyst-paraphrased), with a few in-vivo codes that capture distinctive participant framings.\n\nThis is the same approach you would use for [manual qualitative coding](/docs/coding-qualitative-data) — the AI just applies the codebook at scale.\n\n## What Auto-Tagging Does Not Do Well\n\nThe honest limits, because trust matters more than hype:\n\n- **Heavy sarcasm and irony** — the AI sometimes misreads sarcastic responses. Voice mode helps a little here because tone disambiguates.\n- **Domain-specific jargon** — niche industry terms that the model has not seen often will be coded generically. The fix is to include a short glossary in the [research brief](/docs/how-to-write-research-brief) context.\n- **Very subtle distinctions** — auto-tagging is excellent at the top 80% of insights. The last 20% — where two themes are subtly different in ways only a domain expert would catch — still benefits from human review.\n- **Single-interview outliers** — if a unique insight appears in only one interview, cycle-2 clustering will sometimes fold it into a nearby theme rather than preserve it as a singleton. Use [Insights Chat](/docs/chat-with-interview-transcripts-ai) to surface single-interview signals on demand.\n\nNone of these are reasons to avoid auto-tagging. They are reasons to keep a human in the loop for the highest-stakes interpretations.\n\n## When to Use Auto-Tagging Across Your Research Program\n\nThree patterns:\n\n- **Always-on customer discovery** — auto-tag every interview as it completes. Pair with a [continuous discovery cadence](/docs/continuous-discovery-tools-2026) for weekly synthesis without analyst burnout.\n- **Cohort comparison studies** — use codebook-guided auto-tagging to compare segments (enterprise vs SMB, North America vs Europe, new vs churned).\n- **Longitudinal tracking** — apply the same codebook to a quarterly customer health study and watch theme frequencies move over time.\n\nFor one-off, high-stakes interpretive studies (e.g., pre-IPO board research), auto-tagging is still useful as a first pass — but human qualitative researchers should review and re-code the highest-stakes themes.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the 6 question types that auto-extract structured answers in parallel with auto-tagging.\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) — the qualitative analysis framework auto-tagging accelerates.\n- [How to Code Qualitative Data](/docs/coding-qualitative-data) — the manual coding method auto-tagging automates.\n- [How to Analyze Interview Results](/docs/analyzing-interview-results) — the broader analysis workflow.\n- [Chat With Your Interview Transcripts](/docs/chat-with-interview-transcripts-ai) — querying the auto-tagged dataset.\n- [Atomic Research Nuggets Guide](/docs/atomic-research-nuggets-guide) — how to store and reuse auto-tagged segments across studies.\n- [How AI Interviewers Work](/docs/how-ai-interviewers-work) — what happens during the interview that produces the tags.","category":"Reports & Analysis","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"AI Auto-Tagging for Customer Interviews: Code 100 Interviews in Minutes","metaDescription":"How AI auto-tagging compresses 40+ hours of manual qualitative coding into minutes. Covers Koji's two-cycle coding (descriptive + axial), emergent vs codebook-guided modes, confidence scores, supporting-quote anchoring, and when to keep a human in the loop.","keywords":["AI auto-tagging customer interviews","automated qualitative coding","AI interview tagging","automated thematic coding","two-cycle coding AI","axial coding automation","interview transcript auto-tagging","AI codebook clustering","automated qualitative analysis","Koji auto-tagging"],"aiSummary":"AI auto-tagging compresses 40-146 hours of manual qualitative coding into under 30 minutes. Koji runs a two-cycle pipeline: cycle-1 generates 1-3 descriptive codes per open-ended answer (2-5 word labels grounded in verbatim supporting quotes and message indices), then cycle-2 axial clustering at report time merges near-duplicate codes into a canonical codebook per question across all interviews. Structured question types (scale, choice, ranking, yes/no) get pre-coded automatically during the interview. Two modes: emergent (codebook derived from data, default for discovery) vs codebook-guided (predefined codes for longitudinal or cohort-comparison studies). Quality controls: per-answer confidence scores, verbatim supporting quotes, transcript traceability. Limits: sarcasm, niche jargon, very subtle theme distinctions, single-interview outliers.","aiPrerequisites":["Basic familiarity with qualitative research and coding","Understanding of thematic analysis fundamentals","At least one completed interview in Koji to inspect codes against"],"aiLearningOutcomes":["Understand the difference between auto-tagging, thematic analysis, and AI summarization","See how Koji's two-cycle coding works end to end","Compare time and cost of manual coding vs auto-tagging at study scale","Choose between emergent and codebook-guided auto-tagging modes","Validate AI-generated codes using confidence scores, supporting quotes, and transcript traceability","Build a codebook the AI respects"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min read"},{"type":"documentation","id":"8b9c56bb-112a-4f39-8218-71d687e063ca","slug":"ux-research-report-template","title":"How to Create Effective UX Research Reports (+ Free Template)","url":"https://www.koji.so/docs/ux-research-report-template","summary":"A UX research report is a structured document communicating findings, insights, and recommendations from a research study to stakeholders. It should include an executive summary, methodology, findings with supporting quotes, and actionable recommendations tied to each insight. AI-native platforms like Koji auto-generate research reports automatically after each interview, reducing reporting time from days to minutes.","content":"# How to Create Effective UX Research Reports (+ Free Template)\n\nA great UX research report does one thing: it moves people to act. It transforms raw qualitative data — hours of interviews, transcripts, and notes — into a clear narrative that tells stakeholders exactly what users need and what the team should do next. Yet more than **51% of UX researchers say they wish they had more time for analysis and socializing findings** (Dscout, 2024), meaning the reporting step is one of the most underinvested parts of the research process.\n\nThis guide gives you a complete, reusable UX research report template and shows you how modern AI-native research platforms like Koji can auto-generate these reports in minutes — not days.\n\n---\n\n## What Is a UX Research Report?\n\nA UX research report is a structured document that communicates the findings, insights, and recommendations from a research study to stakeholders. It bridges the gap between raw user data and product decisions.\n\nA strong report answers three questions:\n1. **What did we learn?** (Findings)\n2. **What does it mean?** (Insights)\n3. **What should we do?** (Recommendations)\n\nThe report format varies depending on your audience — a concise executive summary for C-suite, a detailed findings document for product and design teams — but the structure remains consistent.\n\n---\n\n## Why Most UX Research Reports Fail\n\nThe most common reason research reports fail to drive action is a broken chain between data, insight, and recommendation. Researchers often:\n\n- **Overload with data** — including every metric and observation rather than prioritizing what matters\n- **Bury the insights** — leading with methodology instead of the most important finding\n- **Skip actionable next steps** — leaving stakeholders with interesting data but no clear path forward\n- **Use research jargon** — making reports inaccessible to non-researchers on the product team\n- **Create static documents** — reports that get filed and forgotten rather than shared and iterated on\n\nThe fix is a clear structure that puts the most important finding first and connects every insight to a concrete recommendation.\n\n---\n\n## The 6-Section UX Research Report Template\n\nUse this structure for any research study — usability tests, user interviews, surveys, or diary studies.\n\n### Section 1: Executive Summary\n\nThe executive summary is the most important section. Write it last but place it first. Keep it to 2-3 paragraphs covering:\n\n- **The research question** — What were you trying to learn?\n- **Key findings** — The 3-5 most important insights (stated as clear, specific claims)\n- **Top recommendations** — What should the team do next?\n\nMost stakeholders will only read the executive summary. If your key finding is not here, it will not get actioned.\n\n**Template:**\n```\nWe conducted [X interviews / survey with X participants] to understand [research question].\n\nKey findings:\n1. [Most important finding — stated as a specific insight, not just an observation]\n2. [Second finding]\n3. [Third finding]\n\nWe recommend: [Top 2-3 actionable recommendations]\n```\n\n### Section 2: Research Background & Objectives\n\nProvide context for why the research was conducted:\n\n- **Business context** — What product decision or challenge prompted this research?\n- **Research goals** — What specific questions were you trying to answer?\n- **Hypotheses** — What did the team believe going into the research (to be validated or invalidated)?\n- **Scope** — What was explicitly out of scope?\n\nThis section helps stakeholders understand why certain topics were explored and others were not.\n\n### Section 3: Methodology\n\nDescribe how the research was conducted:\n\n- **Method chosen** — User interviews, usability testing, surveys, diary study, etc.\n- **Why this method** — Brief rationale for why this method best answered the research question\n- **Participants** — Number of participants, screener criteria, key characteristics\n- **Session details** — Duration, moderated vs. unmoderated, in-person vs. remote\n- **Analysis approach** — How you synthesized the data (thematic analysis, affinity mapping, etc.)\n\nKeep this section concise — 1 page maximum. Stakeholders need enough context to trust the methodology, not a dissertation.\n\n### Section 4: Findings & Insights (The Core)\n\nThis is the heart of the report. Structure findings in one of three ways depending on your study:\n\n**Option A: By Research Goal**\nList each research goal and present all evidence supporting or contradicting it. Best for evaluative studies (usability tests, concept validation).\n\n**Option B: By Theme**\nOrganize by the patterns that emerged most strongly across participants. Best for generative/discovery research.\n\n**Option C: By Affinity Category**\nGroup insights by the natural categories that emerged during synthesis. Best for large datasets with many participants.\n\n**For each insight, include:**\n- **The insight statement** — A clear, specific claim (e.g., \"Users cannot find the export function because it is nested 3 levels deep in settings\")\n- **Supporting evidence** — 2-3 direct quotes from participants\n- **Frequency** — How many participants experienced this (e.g., \"7 of 8 participants\")\n- **Severity** — Critical / Major / Minor\n- **Visual evidence** — Screenshots, video clips, annotated UI\n\n**Example insight format:**\n\n> **Finding:** Users frequently abandon the checkout flow when asked to create an account.\n>\n> *\"I just wanted to buy one thing — I do not want to sign up for another account.\"* — Participant 4\n>\n> *\"Why do I need an account? I am never going to come back.\"* — Participant 7\n>\n> **Frequency:** 6 of 8 participants | **Severity:** Critical\n\n### Section 5: Recommendations\n\nTranslate every key insight into a specific, actionable recommendation. The most effective recommendations include:\n\n- **What to do** — The specific change or action\n- **Why** — Tied directly to the insight\n- **Priority** — P0 (immediate), P1 (next sprint), P2 (backlog)\n- **Owner** — Who should take this action (design, engineering, product)\n\n**Template:**\n```\nRecommendation: Allow guest checkout without account creation.\nWhy: 75% of participants abandoned checkout when required to create an account.\nPriority: P0 — blocks conversion\nOwner: Product + Engineering\n```\n\n### Section 6: Appendix\n\nInclude supporting materials for researchers and designers who want to dig deeper:\n\n- Full participant demographics\n- Interview guide / discussion guide\n- Raw data tables or session recordings\n- Affinity map or synthesis artifacts\n- Methodology limitations and caveats\n\nThe appendix keeps the main report clean while providing depth for those who need it.\n\n---\n\n## UX Research Report Best Practices\n\n### Lead With the Most Important Finding\n\nStructure your report like a newspaper article — the most important information first. UX researchers often make the mistake of building to a conclusion. Instead, state the conclusion upfront and then support it with evidence.\n\nThis is called the Minto Pyramid Principle: start with the top-level insight, then support it with evidence below.\n\n### Use Direct Quotes Strategically\n\nDirect quotes from participants are the most persuasive evidence in a research report. They create empathy and overcome stakeholder skepticism in a way that statistics cannot. Nielsen Norman Group emphasizes that \"video evidence is a strong UX storytelling tool that helps you improve comprehension, build empathy, and overcome skepticism when communicating research findings to stakeholders.\"\n\nChoose quotes that are:\n- Specific (not vague or abstract)\n- Representative of a pattern (not cherry-picked outliers)\n- Human and relatable\n\n### Quantify What You Can\n\nEven in qualitative research, numbers add credibility. \"7 of 8 participants struggled with X\" is more compelling than \"most participants struggled with X.\" Always specify how many participants experienced each finding.\n\n### Make It Visual\n\nAnnotated screenshots, journey maps, and comparison charts reduce cognitive load and make reports scannable. Use a consistent visual hierarchy:\n- **Bold** for finding statements\n- Quoted text for participant quotes\n- Tables for prioritized recommendations\n\n### Tailor for Your Audience\n\nCreate different versions of the same report for different audiences:\n- **Executive audience:** 1-page summary with top 3 findings and recommendations\n- **Product team:** Full findings document with supporting evidence\n- **Design team:** Detailed findings with UI annotations and specific design recommendations\n\n---\n\n## Research Reporting Timeline Reality\n\nThe average research project takes **42 days from start to finish** (Dscout, 2024), with analysis and reporting consuming a significant portion. Specifically:\n- Discovery research averages 60 days\n- Evaluative research averages 28 days\n\nNearly **60% of researchers report that reduced project time negatively affects the rigor of their methodology** and their creative approach — which means reporting quality suffers when timelines compress.\n\nOrganizations that invest in research are increasingly seeing results: from 8% in 2025 to **22% in 2026**, companies now view research as essential to their core business strategy — nearly tripling in one year (Maze Future of User Research Report, 2026).\n\n---\n\n## How AI Is Transforming Research Reporting\n\nThe traditional research reporting process involves manual transcription, time-consuming affinity mapping, hours of synthesis, and then writing the report from scratch. In 2026, **nearly 69% of researchers now use AI in at least some of their projects** (Maze, 2026), and the results are significant:\n\n- **63% report faster research turnaround**\n- **60% experience better team efficiency**\n- **56% achieve more optimized workflows**\n\nAI-native research platforms are fundamentally changing what is possible.\n\n### Traditional Approach vs. Koji AI-Native Approach\n\n| Step | Traditional | With Koji AI |\n|------|-------------|-------------|\n| Transcription | 2-4 hours per interview | Automatic, real-time |\n| Thematic analysis | 1-2 days for 10 interviews | Minutes |\n| Report generation | 4-8 hours per study | Auto-generated after each interview |\n| Cross-study synthesis | Days to weeks | Real-time dashboard |\n| Sharing findings | Static PDF or slide deck | Live, shareable research portal |\n\n### How Koji Auto-Generates Research Reports\n\nKoji is an AI-native research platform that conducts interviews autonomously — via text or voice — and automatically generates structured research reports. Here is how the reporting workflow works:\n\n1. **Set up your study** — Define your research questions using any of Koji's 6 structured question types: open-ended, scale, single choice, multiple choice, ranking, or yes/no.\n2. **Collect responses** — Koji's AI interviewer conducts interviews with your participants, probing for depth on open-ended questions.\n3. **Auto-analysis** — After each interview, Koji scores response quality (1-5 scale) and extracts structured answers.\n4. **Report generation** — Koji aggregates all responses into a comprehensive research report with themes, quotes, distributions for quantitative questions, and actionable insights — all organized by your research questions.\n5. **Share and publish** — Publish your report as a shareable link or export to CSV/JSON for further analysis.\n\nThe result: research teams can run studies with dozens or hundreds of participants and have a full analysis in hours rather than weeks.\n\n---\n\n## Downloadable UX Research Report Template\n\nHere is a complete, copy-paste-ready template for your next research report:\n\n```\n# [Study Title] Research Report\nDate: [Date]\nResearcher(s): [Names]\nStudy Type: [User interviews / Usability test / Survey]\nParticipants: [N participants, key characteristics]\n\n---\n\nExecutive Summary\n[2-3 paragraphs: context, key findings, top recommendations]\n\nKey Findings:\n1. [Finding 1]\n2. [Finding 2]\n3. [Finding 3]\n\nRecommendations:\n1. [Recommendation 1]\n2. [Recommendation 2]\n\n---\n\nBackground & Objectives\nBusiness context: [Why this research was needed]\nResearch goals: [What we were trying to learn]\nHypothesis: [What we believed going in]\nOut of scope: [What we intentionally did not explore]\n\n---\n\nMethodology\nMethod: [Research method]\nParticipants: [N participants, screener criteria]\nSessions: [Duration, moderated/unmoderated, remote/in-person]\nAnalysis: [How we synthesized the data]\n\n---\n\nFindings\n\nFinding 1: [Specific insight statement]\nEvidence:\n- \"[Quote from participant]\" — P[#]\n- \"[Quote from participant]\" — P[#]\nFrequency: [X of N participants]\nSeverity: [Critical / Major / Minor]\n\n---\n\nRecommendations\n\nRecommendation | Rationale | Priority | Owner\n[Action] | [Tied to finding] | P0/P1/P2 | [Team]\n\n---\n\nAppendix\n- Participant demographics\n- Interview guide\n- Raw data / session recordings\n- Methodology limitations\n```\n\n---\n\n## Related Resources\n\n- [Thematic Analysis: How to Find Patterns in Qualitative Data](/docs/thematic-analysis-guide)\n- [How to Write a Research Brief](/docs/research-brief-template)\n- [Turning Interviews Into Insights: Koji's Analysis Engine](/docs/turning-interviews-into-insights)\n- [Generating Research Reports with Koji](/docs/generating-research-reports)\n- [Structured Questions Guide: Mixing Qualitative and Quantitative Research](/docs/structured-questions-guide)\n- [Publishing and Sharing Research Reports](/docs/publishing-sharing-reports)\n\n## Further reading on the blog\n\n- [User Research Budget Template: How to Plan and Justify Research Spending in 2026](/blog/user-research-budget-template-2026) — Build a research budget that actually gets approved. Real benchmarks, line-item templates, ROI arguments, and stage-appropriate guidance — f\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n- [AI Agents for User Research in 2026: How Autonomous Research Is Reshaping Customer Insight](/blog/ai-agents-user-research-2026) — AI agents are taking over user research in 2026 — moderating interviews, synthesizing themes, and producing insight reports in hours. The fu\n\n<!-- further-reading:blog -->\n","category":"Analysis & Synthesis","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"UX Research Report Template: How to Write Reports That Drive Action (2026)","metaDescription":"Get a complete UX research report template plus best practices for writing reports that actually drive product decisions. Includes how Koji auto-generates research reports in minutes.","keywords":["ux research report template","user research report","how to write research report","research findings report","ux research findings presentation","research report structure"],"aiSummary":"A UX research report is a structured document communicating findings, insights, and recommendations from a research study to stakeholders. It should include an executive summary, methodology, findings with supporting quotes, and actionable recommendations tied to each insight. AI-native platforms like Koji auto-generate research reports automatically after each interview, reducing reporting time from days to minutes."},{"type":"documentation","id":"218d0287-ecca-4b8a-9278-8bb10c8065bf","slug":"qualitative-research-codebook","title":"How to Build a Qualitative Research Codebook (With Examples and Templates)","url":"https://www.koji.so/docs/qualitative-research-codebook","summary":"A qualitative codebook is a standalone document that defines every code in an analysis with name, definition, inclusion criteria, exclusion criteria, and example excerpt. Codebooks can be inductive (built bottom-up from data), deductive (built top-down from theory), or hybrid. The standard quality measure is Cohen's kappa: >0.80 is almost perfect, 0.60-0.80 is substantial, below 0.40 needs major revision. The 6-phase development process is immersion, open coding, define and consolidate, pilot, reach agreement, apply at scale with revision history. Koji uses the research brief as a high-level codebook and applies it via AI moderation and thematic analysis automatically, with structured questions providing a deductive coding layer.","content":"A codebook is the single most underused artifact in qualitative research. It is the rulebook that defines what each code means, when to apply it, and when *not* to. Without one, a team of three researchers coding the same set of interview transcripts will produce three different sets of themes — not because they disagree about what the data shows, but because they never aligned on what the codes mean in the first place.\n\nA well-built codebook is what separates a defensible qualitative analysis from a glorified set of personal impressions. It is also the artifact that AI-assisted coding tools depend on to produce consistent output, which makes codebook craft more relevant in 2026 than it was a decade ago, not less.\n\n## What a Codebook Actually Is\n\nA qualitative codebook is a standalone document — usually a structured table or spreadsheet — that lists every code used in an analysis along with the rules for applying it. It is not a list of themes. It is not the coded data itself. It is the *operational definition* of your coding scheme, designed so that a second analyst could pick it up and code new data the same way you did.\n\nThe classic codebook from a thematic analysis contains, at minimum, five columns per code:\n\n| Column | What it captures |\n|--------|------------------|\n| **Code name** | A short label (1–4 words) |\n| **Definition** | A precise sentence explaining what the code captures |\n| **Inclusion criteria** | Specific signals in the data that *should* be coded with this code |\n| **Exclusion criteria** | Signals that *should not* be coded with this code, even if they look similar |\n| **Example excerpt** | An actual quote from the data that exemplifies the code |\n\nMore comprehensive codebooks add a parent theme column (for hierarchical schemes), a frequency count, a coder's notes column for atypical cases, and a revision history.\n\nJohnny Saldaña, whose *Coding Manual for Qualitative Researchers* is the standard reference, defines a codebook plainly: it is \"a code-description-data\" reference document, distinct from an index of the corpus. The codebook tells you how to code; the index tells you what has been coded.\n\n## Inductive vs Deductive Codebooks\n\nThere are two fundamentally different ways to build a codebook, and the difference shapes everything that follows.\n\n**Inductive (bottom-up).** You build the codebook as you code. The codes emerge from the data itself rather than from prior theory. You start with no codes, code your first transcript, generate codes as you go, then continue refining and merging as you encounter more data. This is the dominant approach in exploratory and grounded-theory research.\n\n**Deductive (top-down).** You build the codebook *before* you code, drawing from existing theory, a research framework, or prior literature. The codes are predetermined and the analyst's job is to apply them consistently. This is the dominant approach in confirmatory research, evaluation studies, and any context where you're testing a specific framework.\n\n**Hybrid.** Most real-world projects mix both. A skeleton codebook from theory provides the initial structure; inductive coding fills in the gaps. The hybrid approach is recommended by Springer's 2019 case study on codebook development as the most practical model for applied research because it combines theoretical grounding with empirical openness.\n\nThe choice affects how strict the codebook needs to be. Inductive codebooks are living documents that should *expect* revision. Deductive codebooks need to be locked early, with very explicit inclusion and exclusion criteria, because the goal is consistency rather than discovery.\n\n## A Worked Codebook Example\n\nImagine an analysis of 25 customer interviews about a project management tool. A condensed slice of the codebook might look like:\n\n| Code | Definition | Inclusion | Exclusion | Example |\n|------|------------|-----------|-----------|---------|\n| **Onboarding friction** | Difficulty experienced during the first 7 days of product use | Statements about setup confusion, missing guidance, abandoning during trial, struggling to invite a team | Generic complaints about the product (those go under \"general dissatisfaction\"); friction after week 1 (use \"ongoing friction\") | \"I signed up Thursday and by Tuesday I still hadn't figured out how to add my team — I just gave up.\" |\n| **Notification fatigue** | Feeling overwhelmed by the volume or frequency of notifications | Mentions of \"too many,\" \"spam,\" \"noisy,\" \"muting\"; descriptions of disabling notifications entirely | Complaints about *missing* notifications (use \"missed alerts\") | \"It pinged me 40 times in an hour. I turned them all off and now I don't check the app at all.\" |\n| **Power-user frustration** | Frustration from a user who has mastered the product and now wants more advanced behavior | Statements implying long tenure (\"I've used this for 2 years…\"); requests for keyboard shortcuts, bulk actions, API access | New-user struggles (use \"onboarding friction\" or \"discoverability\"); general feature requests from non-power users | \"I've been here 18 months and there's still no way to bulk-archive completed projects. It's the only thing keeping me on Trello.\" |\n\nNote what the columns force the analyst to do: precisely scope the code, explicitly enumerate what *doesn't* count, and ground the definition in an actual quote. That discipline is what makes the codebook usable by someone other than its author.\n\n## How to Build a Codebook From Scratch\n\nThe canonical process for inductive codebook development, drawn from Braun & Clarke's six-phase thematic analysis and refined by the Springer 2019 codebook case study:\n\n### Phase 1 — Immersion\nRead 3–5 transcripts in full without coding anything. Take notes on impressions and recurring patterns. Resist the urge to label.\n\n### Phase 2 — Open coding\nRe-read the same transcripts. Generate short codes for anything that seems meaningful — a behavior, an emotion, a constraint, a recurring phrase. Aim for 30–60 candidate codes from the first batch.\n\n### Phase 3 — Define and consolidate\nReview the candidate codes. Merge near-duplicates. Split codes that have become umbrellas for too many distinct ideas. Write a precise definition, inclusion rule, and exclusion rule for each remaining code. This is when the codebook is born.\n\n### Phase 4 — Pilot\nApply the codebook to a transcript you have not coded yet. Track where the rules break down. Refine codes that produced ambiguous decisions. Document atypical cases in a coder's notes column.\n\n### Phase 5 — Reach agreement (if multiple coders)\nHave a second analyst independently code 2–3 transcripts. Compare results. Where disagreements cluster, the codebook is unclear — sharpen the definitions and inclusion criteria until two analysts produce substantively the same coding on new transcripts.\n\n### Phase 6 — Apply at scale, with revision history\nCode the rest of the corpus. When new patterns emerge that don't fit existing codes, add codes — and *log the revision* with a date. Late additions to the codebook should trigger re-coding of earlier transcripts under the new code, which is tedious but necessary for consistency.\n\n## Measuring Codebook Quality\n\nThe most common quantitative measure of codebook reliability is **Cohen's kappa** — a statistic that captures agreement between two coders while correcting for agreement that would happen by chance.\n\nCohen's kappa ranges from −1 (complete disagreement) to 1 (perfect agreement). 0 means agreement is no better than chance. Widely-used interpretation thresholds:\n- **< 0.40** — poor agreement; codebook needs major revision\n- **0.40–0.60** — moderate; refine ambiguous codes\n- **0.60–0.80** — substantial agreement; usable but worth sharpening\n- **> 0.80** — almost perfect; ready for analysis\n\nKappa's appropriateness for qualitative research is contested — some argue it imports a positivist frame onto interpretive work. A pragmatic position: kappa is useful as a *diagnostic* for where the codebook is unclear, not as a stamp of validity. If two coders disagree on a code 40% of the time, that disagreement points to ambiguous criteria — fix the criteria, not the coders.\n\nFor more than two coders, **Fleiss's kappa** or **Krippendorff's alpha** are the appropriate generalizations.\n\n## Common Codebook Mistakes\n\n- **Codes that are actually themes.** A code is a granular label applied to a passage; a theme is a higher-level pattern that organizes codes. A codebook entry called \"User experience problems\" is too broad to apply consistently — break it into specific codes.\n- **No exclusion criteria.** Inclusion criteria alone produce codebook entries that look complete but in practice swallow everything. Every code needs an explicit \"this does not count as X\" clause.\n- **No example quotes.** A definition without an example forces every coder to interpret it differently. A real excerpt anchors the meaning.\n- **Single-coder development for high-stakes work.** A codebook built by one researcher reflects one researcher's assumptions. For work that needs to be defensible, have a second analyst pressure-test the codebook before applying it at scale.\n- **No revision history.** Codebooks evolve. Without tracked changes, a stakeholder later cannot tell whether a code meant the same thing in transcript 1 as it did in transcript 25.\n\n## How AI Changes Codebook Work\n\nFor the first 30 years of qualitative software (NVivo, Atlas.ti, Dedoose), the codebook was a manual artifact and coding was a manual process. AI doesn't change the codebook itself — but it dramatically changes how it gets applied.\n\nWith a well-defined codebook, modern LLMs can apply codes consistently across hundreds of transcripts in minutes, with kappa scores that often match or exceed human inter-coder reliability when the codebook is sharp. The bottleneck shifts: the limiting factor is no longer how fast you can code, but how precisely you can articulate the codes.\n\nThat is exactly what a codebook does.\n\n## How Koji Handles Codebook-Driven Analysis\n\nKoji's analysis pipeline is, in effect, a codebook applied at machine speed. When you create a study, the research brief functions as a high-level codebook: it specifies the themes you're investigating, the structured questions, and the methodology framework (mom_test, jtbd, discovery, exploratory, or lead_magnet). The AI moderator applies that codebook during interviews — probing for evidence of each theme — and the analysis layer applies it again when consolidating findings across all responses.\n\nFor researchers who want explicit control, Koji supports [structured questions](/docs/structured-questions-guide) across six types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that effectively act as a deductive codebook for the quantitative slice of the study, while open-ended themes are coded inductively by the AI. The thematic analysis output names the codes, counts their frequency across participants, and surfaces verbatim quotes — the same artifacts a human-built codebook would produce, in roughly 1% of the time.\n\nTeams using AI-assisted thematic coding report dramatically faster time-to-insight on what was historically the slowest stage of qualitative work. The codebook craft still matters — clearer briefs produce sharper themes — but the manual labor of applying it has effectively collapsed.\n\n## A Codebook Template You Can Use\n\nFor a basic project, copy this structure into a spreadsheet:\n\n```\n| Code Name | Parent Theme | Definition | Inclusion Criteria | Exclusion Criteria | Example Excerpt | Coder Notes | Date Added | Last Revised |\n```\n\nKeep it in version control or shared cloud storage. Append revisions; don't overwrite. When a stakeholder later asks \"what did you mean by 'onboarding friction' on this date?\" you'll have the answer.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six structured question types that function as a deductive coding layer\n- [Coding Qualitative Data](/docs/coding-qualitative-data) — the broader process of applying codes\n- [Open, Axial, and Selective Coding](/docs/open-axial-selective-coding) — the grounded-theory coding sequence\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) — the larger framework codebooks support\n- [Research Synthesis Guide](/docs/research-synthesis-guide) — moving from codes to shareable insight\n- [Affinity Mapping](/docs/affinity-mapping) — a complementary technique for clustering codes into themes\n\n## Sources\n\n- Saldaña, J. (2021). *The Coding Manual for Qualitative Researchers* (4th ed.). SAGE.\n- Roberts, K., Dowell, A., & Nie, J. B. (2019). *Attempting rigour and replicability in thematic analysis of qualitative research data; a case study of codebook development.* BMC Medical Research Methodology, 19(66).\n- Braun, V., & Clarke, V. (2006). *Using thematic analysis in psychology.* Qualitative Research in Psychology, 3(2).\n- Cohen, J. (1960). *A coefficient of agreement for nominal scales.* Educational and Psychological Measurement, 20(1).","category":"Analysis & Synthesis","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Qualitative Research Codebook: Complete Guide with Templates (2026)","metaDescription":"A qualitative codebook defines how you code data: code names, definitions, inclusion criteria, examples. Learn how to build one inductively or deductively, measure quality with Cohen's kappa, and apply it at scale with AI.","keywords":["qualitative codebook","codebook","qualitative coding","thematic analysis","inductive coding","deductive coding","inter-rater reliability","Cohen kappa","qualitative research methodology","coding scheme"],"aiSummary":"A qualitative codebook is a standalone document that defines every code in an analysis with name, definition, inclusion criteria, exclusion criteria, and example excerpt. Codebooks can be inductive (built bottom-up from data), deductive (built top-down from theory), or hybrid. The standard quality measure is Cohen's kappa: >0.80 is almost perfect, 0.60-0.80 is substantial, below 0.40 needs major revision. The 6-phase development process is immersion, open coding, define and consolidate, pilot, reach agreement, apply at scale with revision history. Koji uses the research brief as a high-level codebook and applies it via AI moderation and thematic analysis automatically, with structured questions providing a deductive coding layer.","aiPrerequisites":["coding-qualitative-data"],"aiLearningOutcomes":["Build an inductive, deductive, or hybrid codebook from scratch using a 6-phase process","Write code definitions with explicit inclusion and exclusion criteria that hold up across multiple analysts","Measure codebook reliability with Cohen's kappa and interpret the result correctly","Avoid the five most common codebook mistakes that produce indefensible analyses","Apply a codebook at scale using AI-assisted thematic coding"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"644eafeb-740f-456f-a007-643e4130de44","slug":"how-to-prioritize-customer-feedback","title":"How to Prioritize Customer Feedback: A Framework for Product Teams","url":"https://www.koji.so/docs/how-to-prioritize-customer-feedback","summary":"A practitioner's guide to prioritizing customer feedback in 2026. Walks through consolidating channels, tagging by outcome (not feature), choosing the right scoring framework — RICE, MoSCoW, Kano, or Opportunity Solution Tree — and validating priorities with structured customer input. Includes a 30-day rollout plan and shows how AI-native research platforms like Koji collapse the workflow from 20–40 hours per quarter to minutes.","content":"## The short answer\n\n**The fastest way to prioritize customer feedback is to (1) consolidate every feedback channel into a single repository, (2) tag each piece of feedback with the customer segment, the problem theme, and the desired outcome, (3) score themes — not individual requests — using a framework like RICE or the Opportunity Solution Tree, and (4) re-run the scoring every two weeks.** Manual triage takes most product teams 8–12 hours per sprint. AI-native research platforms like Koji collapse the same workflow into minutes by clustering verbatim feedback into themes, surfacing the top opportunities, and feeding scored insights directly to your backlog.\n\nIf you only remember one thing: never prioritize individual feature requests. Prioritize the underlying *problem* (or \"opportunity\"). One problem usually has five viable solutions; ranking solutions before you rank problems is how teams build the wrong thing fast.\n\n## Why feedback prioritization is the bottleneck of modern product work\n\nProduct teams in 2026 do not have a feedback shortage — they have a feedback overload. A typical SaaS team now ingests requests from in-app surveys, support tickets, sales call recordings, NPS comments, community forums, customer success notes, and AI-generated interview transcripts. Productboard documents that mature teams centralize feedback from Intercom, Zendesk, Slack, Salesforce, and email into a single repository — and even then, \"product managers who are overwhelmed by a multitude of feedback sources need to capture, organize, and process user input more effectively.\"\n\nThe result is what Marty Cagan, in *Inspired*, calls the opportunity-versus-solution trap: teams ranking individual feature requests rather than the strategic problems behind them. Cagan's repeated counsel — \"prioritize business results rather than product ideas\" — is the foundation of every framework below.\n\nThree consequences when teams skip the prioritization step:\n\n1. **Roadmap whiplash.** The loudest customer wins. Whoever escalated last week sets the next sprint.\n2. **Feature bloat.** Saying yes to one-off requests inflates the product surface and the maintenance bill.\n3. **Strategy drift.** Without a tie-back to outcomes, the roadmap stops looking like a strategy and starts looking like a list.\n\nResearch-driven prioritization is the cure. Nielsen Norman Group recommends: \"Research findings should be funneled back into your backlog as tasks on existing backlog items, bugs to fix, or new backlog items,\" and notes that those findings \"should be brought up at the sprint review so items can be prioritized and added to the backlog immediately.\"\n\n## Step 1 — Consolidate every feedback channel\n\nBefore any framework, you need one source of truth. The HubSpot 2026 State of Marketing Report found that 69% of marketers cite data unification as a complex pain point — and the same is true on the product side.\n\nA minimum viable feedback repository captures, for every record:\n\n- **Source** — support ticket, NPS comment, interview, sales call\n- **Customer segment** — plan tier, industry, ARR band, persona\n- **Verbatim quote** — never a summary, always the raw words\n- **Sentiment** — positive, neutral, negative\n- **Problem theme** — the *underlying* pain, not the requested feature\n- **Date and recency** — feedback decays\n\n### How Koji helps\n\nKoji unifies every feedback channel into one searchable repository. AI-moderated voice and text interviews, structured surveys, and imported transcripts all land in the same workspace. Koji's automatic thematic analysis — built on Braun and Clarke's six-phase framework — clusters thousands of verbatim quotes into themes in minutes rather than the 60–120 hours typical of manual coding for a 10-interview study. Every theme links back to the verbatim quotes and the source interview, so you never lose traceability when a stakeholder challenges a finding.\n\n## Step 2 — Tag feedback against outcomes, not features\n\nTeresa Torres, who introduced the Opportunity Solution Tree in 2016, is unambiguous on this point: prioritize opportunities, not solutions. \"She doesn't recommend prioritizing solutions because teams end up comparing apples to oranges.\"\n\nAn opportunity is a customer problem, need, pain point, or desire — phrased in the customer's own words. It sits between your product outcome and the candidate solutions. The structure looks like:\n\n```\nOutcome (e.g., increase trial-to-paid conversion)\n  └─ Opportunity (e.g., new users can't tell if their setup is correct)\n      ├─ Solution A (in-app setup checklist)\n      ├─ Solution B (post-signup email sequence)\n      └─ Solution C (AI onboarding consultant)\n```\n\nWhen feedback arrives — \"I wish there was a tutorial\" — resist the urge to log it as a feature request. Instead, log it under the opportunity it implies (\"new users can't tell if their setup is correct\") so that ten different requests collapse into one prioritizable problem.\n\n## Step 3 — Apply the right scoring framework\n\nThere is no single best framework. The right choice depends on whether you are planning a roadmap, working under a deadline, or weighing emotional impact. The four most-used frameworks in product teams in 2026:\n\n### RICE — when you have data and need a roadmap\n\nDeveloped by Intercom, RICE scores each opportunity on four factors:\n\n- **Reach** — how many users this affects per quarter\n- **Impact** — qualitative score (0.25 minimal → 3 massive) of the per-user effect\n- **Confidence** — your confidence in the estimates (50% / 80% / 100%)\n- **Effort** — person-months to ship\n\n`RICE Score = (Reach × Impact × Confidence) / Effort`\n\nUse RICE for quarterly planning where you need to compare diverse opportunities on a common scale. Its strength is that the *Confidence* multiplier punishes bold claims that aren't backed by evidence — a built-in nudge to do the research before scoring.\n\n### MoSCoW — when you have a deadline\n\nMust have / Should have / Could have / Won't have. Originated in DSDM/Agile, MoSCoW is a categorization, not a calculation. It works when scope is fluid but the deadline isn't. Avoid the failure mode of more than 60% of items being labeled \"Must\" — that's a signal you haven't actually prioritized.\n\n### Kano Model — when emotional reaction matters\n\nNoriaki Kano's 1984 model classifies features into:\n\n- **Must-be** — basics; their absence causes dissatisfaction, presence is taken for granted\n- **Performance** — linear relationship between investment and satisfaction\n- **Attractive** (delighters) — unexpected; their presence creates strong positive reaction\n- **Indifferent** — users don't care\n- **Reverse** — actively annoy a segment\n\nNielsen Norman Group calls the Kano model \"a good approach for teams who have difficulty prioritizing based on the user, as it introduces user research directly into the prioritization process and mandates discussion around user expectations.\" Run Kano via a structured survey: for each feature ask one functional and one dysfunctional question, then map responses.\n\n### Opportunity Solution Tree — when you have outcomes, not features\n\nTeresa Torres's framework prioritizes the *opportunity space* before the solution space. Score opportunities on dimensions like reach, value, importance, and frequency; only then generate three or more solutions per top opportunity and run assumption tests to pick the winner.\n\n## Step 4 — Calibrate with structured customer input\n\nScoring without recent customer input is just opinion. Before you finalize a quarter's priorities, validate the top opportunities with a fresh, structured study. Three high-leverage techniques:\n\n1. **A ranking question.** Ask 100+ customers to drag the top opportunities into priority order. The aggregate ranking exposes consensus and disagreement.\n2. **Van Westendorp price sensitivity meter.** For paid features, four price questions surface the optimal price point.\n3. **Kano survey.** Two questions per feature (functional / dysfunctional) classify each feature into Kano categories.\n\nKoji's structured questions guide covers the six built-in question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — that make these techniques fast to set up. Each is auto-aggregated into a chart in the report; you don't have to clean spreadsheets to see the answer.\n\n## Step 5 — Close the loop\n\nThe most common failure of prioritization isn't the framework — it's that customers never hear back. A 2024 survey by Pendo found that customers who never see a response to feedback are 2.4× less likely to give feedback again. Two practices fix this:\n\n- **Public roadmap with status.** Tag each item Planned / In Progress / Shipped / Considered.\n- **Direct follow-ups.** When you ship something a customer requested, email them. Customers who get a personal \"we shipped what you asked for\" follow-up have 38% higher retention 12 months later, according to Pendo benchmarks.\n\n## The modern, AI-native approach with Koji\n\nTraditional feedback prioritization workflows look like: PM exports support tickets to a spreadsheet → manually tags each ticket → opens the spreadsheet at sprint planning → argues with engineering about reach estimates. Total time per quarter: 20–40 hours. Average freshness: 8 weeks stale.\n\nKoji collapses this into three actions:\n\n1. **Run an AI-moderated research study** — a 15-minute setup creates a voice or text interview that adapts probing to each respondent. Hundreds of customers complete in parallel.\n2. **Read the auto-generated insight report** — themes are clustered, frequencies counted, supporting verbatim quotes attached. The AI consultant can answer follow-up questions like \"what do enterprise customers want most?\" against the entire dataset.\n3. **Score and ship** — export the prioritized opportunities directly to Linear or Jira, or pipe them via the Koji MCP integration into your AI coding agent for continuous discovery.\n\nTeams using AI-assisted research tools report 60% faster time-to-insight and a measurable increase in research velocity — the difference between running one strategic study per quarter and running one per sprint.\n\nThe goal is not to replace the product manager's judgment. It's to remove the eight hours of tagging that stand between the manager and the judgment.\n\n## A 30-day rollout plan\n\n- **Days 1–7** — Pick one repository (Koji, Productboard, Notion, Airtable). Move the last 90 days of feedback into it. Tag by source, segment, and theme.\n- **Days 8–14** — Pick one framework (RICE for roadmap, OST for outcome work). Score the top 30 themes. Resist scoring more — the long tail rarely matters.\n- **Days 15–21** — Run one structured validation study on the top 5 themes. Use a ranking question and at least one open-ended probe.\n- **Days 22–30** — Publish the prioritized list. Close the loop with the customers whose verbatims drove the top three themes.\n\nRepeat every quarter. Continuous discovery only works if the cadence is non-negotiable.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types that drive accurate prioritization data\n- [The Complete Guide to Thematic Analysis](/docs/thematic-analysis-guide) — how to cluster verbatim feedback into themes\n- [Continuous Discovery: Weekly Customer Interviews](/docs/continuous-discovery-user-research) — keep your priorities fresh\n- [Customer Feedback Analysis](/docs/customer-feedback-analysis) — turning raw input into actionable insights\n- [Opportunity Solution Tree](/docs/opportunity-solution-tree) — Teresa Torres's framework for outcome-driven prioritization\n- [Kano Model](/docs/kano-model) — classify features by emotional impact\n\n## Further reading on the blog\n\n- [Best Product Discovery Tools in 2026: The Complete Buyer's Guide](/blog/best-product-discovery-tools-2026) — Product discovery is no longer a one-off pre-launch phase — it is a continuous loop. The best teams in 2026 are running weekly customer inte\n- [The Continuous Discovery Handbook: How Product Teams Run Weekly Customer Interviews (2026)](/blog/continuous-discovery-handbook-weekly-customer-interviews) — 64% of software features are rarely or never used. Continuous discovery — weekly customer interviews baked into your product workflow — is t\n- [Koji vs Canny: AI Customer Research vs Public Feedback Voting (2026)](/blog/koji-vs-canny-2026) — Canny tallies upvotes on a feature board. Koji actually understands the why behind requests by running AI-moderated voice and text interview\n\n<!-- further-reading:blog -->\n","category":"Analysis & Synthesis","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"How to Prioritize Customer Feedback: RICE, MoSCoW, Kano & OST (2026)","metaDescription":"Master customer feedback prioritization with RICE, MoSCoW, Kano, and the Opportunity Solution Tree. Step-by-step framework with examples, tools, and a 30-day rollout plan.","keywords":["prioritize customer feedback","feedback prioritization framework","RICE framework","MoSCoW method","Kano model","opportunity solution tree","product feedback management","customer feedback management"],"aiSummary":"A practitioner's guide to prioritizing customer feedback in 2026. Walks through consolidating channels, tagging by outcome (not feature), choosing the right scoring framework — RICE, MoSCoW, Kano, or Opportunity Solution Tree — and validating priorities with structured customer input. Includes a 30-day rollout plan and shows how AI-native research platforms like Koji collapse the workflow from 20–40 hours per quarter to minutes.","aiPrerequisites":["customer-feedback-analysis","thematic-analysis-guide"],"aiLearningOutcomes":["Consolidate feedback from multiple channels into a single repository","Tag feedback against outcomes and opportunities rather than feature requests","Apply the right scoring framework (RICE, MoSCoW, Kano, OST) to a given decision","Validate priorities with structured customer input using ranking, scale, and Kano questions","Close the feedback loop with customers and stakeholders"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"525dbb9a-ac6a-4532-a94d-a1356d517f00","slug":"moderated-usability-testing-guide","title":"Moderated Usability Testing: How to Run Sessions That Surface Real Problems (2026 Guide)","url":"https://www.koji.so/docs/moderated-usability-testing-guide","summary":"A practical reference on moderated usability testing: what it is, when to choose it over unmoderated, how to write task scenarios, run think-aloud sessions, measure task success/SEQ/SUS, choose sample size, avoid common mistakes, and scale moderation with Koji's AI interviewer.","content":"## What is moderated usability testing? (Answer first)\n\n**Moderated usability testing is a research method in which a facilitator guides a participant through realistic tasks on a product — live — while observing where they hesitate, struggle, or fail.** The moderator can ask \"why did you do that?\" in the moment, probe confusion as it happens, and adapt the session to what the participant reveals. That real-time probing is exactly what distinguishes moderation from a static survey or a hands-off recording: you do not just see *that* someone failed a task, you learn *why*.\n\nThe trade-off has always been cost. A traditional moderated study means scheduling 8–15 calls, sitting through every one, taking notes, and then spending days synthesizing recordings. Platforms like Koji change that economics: an AI moderator runs the think-aloud session, asks adaptive follow-up questions, and clusters the findings automatically — so you get moderated-quality depth at unmoderated-style scale.\n\n> **Bottom line:** Use moderated usability testing when you need to understand the *reasoning* behind behavior — early-stage designs, complex flows, or any time a metric alone will not tell you what to fix. Use unmoderated testing when you only need to confirm a known hypothesis at volume.\n\n## Moderated vs. unmoderated: when to choose which\n\n| Dimension | Moderated | Unmoderated |\n|---|---|---|\n| Depth of insight | High — probe the \"why\" live | Lower — behavior only |\n| Best for | New flows, ambiguous problems, B2B/expert users | Validated flows, A/B comparisons, large samples |\n| Speed per session | Slower (live) | Faster (self-serve) |\n| Cost to scale | Traditionally high | Low |\n\nThe historical rule was \"moderate for discovery, go unmoderated for validation.\" Koji collapses that divide: its AI moderator conducts a guided, probing session *and* runs many of them in parallel, so you no longer have to trade depth for sample size. For a deeper comparison, see [Unmoderated vs Moderated User Research](/docs/unmoderated-vs-moderated-research).\n\n## How to run a moderated usability test (step by step)\n\n**1. Define the research question, not the feature.** Write down what decision the test will inform. \"Can a new user complete checkout without help?\" is testable; \"Is the design good?\" is not.\n\n**2. Write task scenarios, not instructions.** A good task gives context and a goal but never names the UI element. Bad: \"Click the blue Filter button.\" Good: \"You want to find a jacket under €100 in your size — show me how you would do that.\" Naming the button tells the participant the answer and destroys the test.\n\n**3. Recruit the right participants.** Five users will surface roughly 85% of the usability problems in a single design (Nielsen Norman Group), which is why 5–8 participants per distinct user segment is the workhorse sample size for formative moderated tests. Add a screener so you talk to real target users, not whoever is available.\n\n**4. Run a think-aloud session.** Ask the participant to narrate their thoughts continuously: \"Tell me what you are looking at, what you expect to happen, and what you are trying to do.\" Stay quiet while they work. Resist the urge to help — a silence that feels painful to you is data.\n\n**5. Probe at the right moments.** When someone hesitates, hovers, or backtracks, that is your cue to ask a non-leading follow-up: \"What did you expect to happen there?\" or \"What are you looking for right now?\" This adaptive probing is the entire value of moderation — and it is exactly what Koji's AI interviewer automates with configurable follow-up depth (1–3 probes per question).\n\n**6. Capture both behavior and metrics.** Note task success/failure, where errors cluster, and the verbatim quotes that explain them.\n\n## The metrics that make moderated tests defensible\n\nQualitative observation is the heart of moderated testing, but pairing it with a few standard metrics makes findings far easier to defend to stakeholders:\n\n- **Task success rate** — % of participants who complete each task. The single most important usability metric.\n- **Time on task** — how long completion takes; spikes flag friction.\n- **Single Ease Question (SEQ)** — a 7-point post-task rating of difficulty. See the [Single Ease Question (SEQ) guide](/docs/single-ease-question-seq-guide).\n- **System Usability Scale (SUS)** — a validated 0–100 score for the whole experience. See the [System Usability Scale (SUS) guide](/docs/system-usability-scale-guide).\n\nIn Koji, you capture these with **structured questions** — Koji supports six types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no). You add SEQ as a `scale` question and the post-task \"what was confusing?\" as an `open_ended` question with AI probing. Because every scale answer is captured as a ground-truth structured value, Koji aggregates task-level difficulty automatically while still clustering the open-ended explanations into themed friction findings.\n\n## Common mistakes that ruin moderated sessions\n\n- **Leading the witness.** \"Was that easy?\" invites a polite yes. Ask \"How did that feel?\" instead. See [How to Avoid Leading Questions](/docs/avoiding-leading-questions).\n- **Helping too soon.** The moment you rescue a struggling user, you lose the finding.\n- **Testing the participant, not the product.** If someone fails, the design failed — never imply otherwise, or social-desirability bias will distort everything that follows.\n- **Skipping the pilot.** Always run one practice session to catch broken tasks before they cost you real participants.\n- **Synthesizing from memory.** Notes taken during a live call are lossy; a verbatim transcript with coded themes is not.\n\n## How Koji makes moderated usability testing faster\n\nTraditional moderation is bottlenecked by *you* — one researcher can only sit in so many calls. Koji removes that bottleneck without removing the depth:\n\n1. **AI moderator runs the think-aloud session** in voice or text, asking your tasks and probing hesitation with adaptive follow-ups — no scheduler, no calendar, available 24/7.\n2. **Voice mode** captures natural think-aloud narration; text mode renders interactive widgets for SEQ and choice questions.\n3. **Automatic analysis** transcribes every session, codes open-ended answers into themes, and aggregates task-level metrics into a **real-time report** you can share with one link.\n4. **Scale without losing nuance** — run 5 sessions or 50 in parallel; the per-question synthesis holds either way.\n\nA study that used to take two weeks of scheduling, moderating, and synthesizing becomes an afternoon. You bring the tasks and the judgment; Koji handles the moderation and the math.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types that power task metrics and probes\n- [Unmoderated vs Moderated User Research: How to Choose](/docs/unmoderated-vs-moderated-research)\n- [How to Conduct Usability Testing: The Complete Guide](/docs/usability-testing-guide)\n- [Usability Testing Script Template](/docs/usability-testing-script-template)\n- [Single Ease Question (SEQ): The 7-Point UX Metric](/docs/single-ease-question-seq-guide)\n- [System Usability Scale (SUS): Complete Guide](/docs/system-usability-scale-guide)\n","category":"Interview Techniques","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Moderated Usability Testing: Tasks, Think-Aloud & Metrics (2026 Guide)","metaDescription":"How to run moderated usability testing in 2026: write non-leading tasks, run think-aloud sessions, measure task success and SEQ, pick sample size, and scale moderation with AI on Koji.","keywords":["moderated usability testing","usability testing","think aloud testing","task success rate","moderated vs unmoderated","usability test tasks","user testing sessions","SEQ","remote usability testing"],"aiSummary":"A practical reference on moderated usability testing: what it is, when to choose it over unmoderated, how to write task scenarios, run think-aloud sessions, measure task success/SEQ/SUS, choose sample size, avoid common mistakes, and scale moderation with Koji's AI interviewer.","aiPrerequisites":["A product or prototype to test","Basic familiarity with usability concepts","Access to target users for recruiting"],"aiLearningOutcomes":["Decide when moderated testing beats unmoderated","Write task scenarios that do not lead participants","Run a think-aloud session and probe at the right moments","Capture task success, time on task, and SEQ correctly","Scale moderated sessions with an AI moderator on Koji"],"aiDifficulty":"beginner","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"743b89a8-6b9f-4f70-b6d5-74f585b405be","slug":"product-feedback-triage-guide","title":"Product Feedback Triage: A Framework for Turning Noise Into a Prioritized Backlog","url":"https://www.koji.so/docs/product-feedback-triage-guide","summary":"A practitioner framework for triaging product feedback before it reaches prioritization. Covers the five-stage loop — capture, dedupe, tag, assess severity, route — a frequency-by-impact severity matrix, and how AI-native platforms like Koji automate deduping and tagging while structured questions make triage metadata deterministic at the source.","content":"## The short answer\n\n**Product feedback triage is the upstream step before prioritization: you capture every incoming signal, deduplicate it, tag it by theme and segment, judge its severity, and route it to the right owner — so that only clean, structured insight ever reaches your scoring framework.** Triage answers \"what is this and who should see it?\" Prioritization answers \"what do we build next?\" Conflating the two is why most backlogs are bloated with 300 raw feature requests nobody can act on.\n\nA reliable triage loop has five stages: **Capture → Dedupe → Tag → Assess severity → Route.** Run it continuously, not in a quarterly cleanup sprint. AI-native research platforms like Koji collapse the slowest stages — deduping and tagging — into seconds by clustering verbatim feedback into themes automatically, so a job that used to eat a full day each sprint becomes a background process.\n\nIf you remember one rule: **triage the problem, not the words.** Ten customers asking for \"a dark mode,\" \"less eye strain,\" and \"a night setting\" are one theme, not three tickets.\n\n## Why triage is a distinct discipline from prioritization\n\nTeams routinely skip triage and jump straight to a RICE spreadsheet. The result is predictable: the spreadsheet fills with duplicates, vague one-liners, and requests that were never validated with a real customer. Prioritization frameworks assume the input is already clean. Triage is what makes it clean.\n\nProductboard and Pendo both report that mature product teams ingest feedback from eight or more channels — in-app surveys, support tickets, sales calls, NPS comments, community forums, churn interviews, customer success notes, and AI interview transcripts. Without a triage layer, every one of those channels dumps raw text directly onto the product manager. The PM becomes a human router, spending 8–12 hours per sprint copy-pasting and re-categorizing. That is the bottleneck triage removes.\n\nThe distinction matters because the two activities have different owners, cadences, and outputs:\n\n- **Triage** runs daily-to-weekly, is often owned by a PM, support lead, or research ops, and outputs a tagged, deduplicated stream of themes.\n- **Prioritization** runs every two weeks (theme-level) to quarterly (roadmap-level), is owned by the product trio, and outputs a ranked backlog.\n\n## The five-stage triage workflow\n\n### 1. Capture: one inbox, every channel\nThe first failure mode is fragmentation. Feedback scattered across Zendesk, Slack, Gong, and a spreadsheet cannot be triaged because nobody can see all of it at once. Consolidate into a single repository. The capture rule is simple: **if it isn't in the repository, it doesn't exist.** Train support, sales, and success teams to forward signals to one destination, or wire integrations that do it automatically.\n\n### 2. Dedupe: collapse variants into themes\nThis is the most time-consuming manual step and the one AI changes most dramatically. The goal is to recognize that \"the export keeps timing out,\" \"downloads fail on big files,\" and \"CSV never finishes\" are the same problem expressed three ways. Manual deduping requires a human to read every item and remember everything they've already read — which is exactly what humans are worst at. Koji clusters verbatim feedback into canonical themes automatically using the same axial-coding logic its analysis engine applies to interview transcripts, so duplicates merge without a person reading each line.\n\n### 3. Tag: attach the metadata that makes routing and scoring possible\nEvery triaged item needs at least four tags:\n\n- **Theme / opportunity** — the underlying problem, not the requested feature.\n- **Customer segment** — ARR tier, persona, or lifecycle stage. Volume of feedback correlates with how vocal a customer is, not how important their problem is, so segment weighting is essential.\n- **Type** — bug, usability friction, feature request, or strategic signal.\n- **Source channel** — so you can tell whether a theme is broad or just loud in one place.\n\nKoji's **structured questions** make this tagging deterministic at the source. Instead of free-text feedback you have to interpret after the fact, you can collect a `single_choice` segment, a `scale` severity rating, and a `ranking` of competing priorities directly inside the AI interview — so the metadata arrives pre-structured. (See the [six structured question types](/docs/structured-questions-guide): open_ended, scale, single_choice, multiple_choice, ranking, and yes_no.)\n\n### 4. Assess severity: a simple two-axis matrix\nNot every item deserves equal attention even before prioritization. Use a quick **frequency × impact** matrix:\n\n| | Low impact | High impact |\n|---|---|---|\n| **Low frequency** | Backlog / monitor | Investigate (could be a top-account risk) |\n| **High frequency** | Quick-win candidate | Escalate immediately |\n\nSeverity is a triage judgment, not a final score. A single high-impact request from a top-ARR account belongs in \"Escalate\" even if only one customer raised it.\n\n### 5. Route: send each theme to its owner with context\nThe last stage is routing. Bugs go to engineering with reproduction context. Strategic signals go to the product trio's discovery board. Validated, high-frequency opportunities go to the prioritization queue. Routing should carry the evidence with it — a verbatim quote and the segment — so the receiving team never has to re-investigate.\n\n## Validating triaged themes with structured customer input\n\nTriage tells you what people are asking for. It does not tell you how much they care, or which of three solutions they'd actually use. Before a theme graduates to the roadmap, validate it:\n\n- A **scale** question (1–10) measures how painful the problem really is.\n- A **ranking** question forces customers to trade competing priorities against each other — far more honest than asking \"would you like this?\" about each in isolation.\n- A **yes_no** question with an AI follow-up confirms whether your proposed solution actually addresses the underlying need.\n\nBecause Koji's AI interviewer asks adaptive follow-up questions automatically, a validation study that once required a moderator and two weeks of scheduling runs asynchronously over a weekend — voice or text, no moderator, with the analysis written the moment the last interview closes.\n\n## The AI-native triage loop\n\nPutting it together, a modern triage loop looks like this:\n\n1. Feedback lands in one repository from every channel.\n2. AI clusters it into themes and merges duplicates in real time.\n3. Each theme inherits structured tags — segment, severity, type — partly from structured questions captured at the source.\n4. High-severity themes route to owners; ambiguous ones spawn a quick validation interview.\n5. Clean, validated themes flow into prioritization, where RICE or an Opportunity Solution Tree does its job on trustworthy input.\n\nThe payoff is not just speed. Continuous triage means feedback never piles up into an unmanageable backlog, patterns surface while they're still actionable, and prioritization meetings argue about evidence instead of anecdotes.\n\n## Common triage mistakes\n\n- **Triaging words instead of problems.** Ten feature requests are usually two or three opportunities.\n- **Letting volume decide.** Weight by segment value, not ticket count.\n- **Batch triage.** A quarterly cleanup guarantees stale, decayed feedback. Triage continuously.\n- **Skipping validation.** A well-tagged theme is still a hypothesis until a customer confirms the pain and the fit.\n- **No routing context.** A ticket without a quote and a segment forces the receiver to redo the investigation.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make triage metadata deterministic\n- [How to Prioritize Customer Feedback](/docs/how-to-prioritize-customer-feedback) — the prioritization step that triage feeds\n- [Feature Request Management](/docs/feature-request-management) — managing the request pipeline end to end\n- [AI Auto-Tagging for Customer Interviews](/docs/ai-auto-tagging-customer-interviews) — how automatic theme clustering works\n- [Customer Feedback Analysis](/docs/customer-feedback-analysis) — turning tagged feedback into insight\n- [Closing the Loop on Customer Feedback](/docs/closing-the-loop-customer-feedback) — telling customers what you did with their input","category":"product-management","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Product Feedback Triage: Framework + AI Workflow (2026)","metaDescription":"Learn how to triage product feedback at scale: capture, dedupe, tag, assess severity, and route every request before prioritization. Includes a triage matrix and an AI-native workflow.","keywords":["product feedback triage","how to triage customer feedback","feedback triage framework","product feedback workflow","triage feature requests","customer feedback management"],"aiSummary":"A practitioner framework for triaging product feedback before it reaches prioritization. Covers the five-stage loop — capture, dedupe, tag, assess severity, route — a frequency-by-impact severity matrix, and how AI-native platforms like Koji automate deduping and tagging while structured questions make triage metadata deterministic at the source.","aiPrerequisites":["how-to-prioritize-customer-feedback","customer-feedback-analysis"],"aiLearningOutcomes":["Distinguish triage from prioritization and run them on the right cadence","Operate a five-stage triage loop: capture, dedupe, tag, assess severity, route","Apply a frequency-by-impact severity matrix to incoming feedback","Use structured questions to capture triage metadata at the source","Validate triaged themes with scale, ranking, and yes/no questions"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"6d0fe858-d470-4e2f-aa59-c4ec8d5bac59","slug":"inter-rater-reliability-qualitative-research","title":"Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement","url":"https://www.koji.so/docs/inter-rater-reliability-qualitative-research","summary":"Inter-rater (intercoder) reliability measures how consistently independent coders apply the same codes to qualitative data. Report it with a chance-corrected statistic — Cohen's kappa (two coders, nominal data) or Krippendorff's alpha (more flexible). Thresholds: 0.80+ is reliable, 0.667–0.80 supports tentative conclusions, below 0.667 is insufficient. Percent agreement alone is misleading because it ignores chance. AI-native platforms like Koji make consistent coding the default through automatic thematic analysis and structured question types.","content":"# Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement\n\n**Bottom line up front:** Inter-rater reliability (IRR) — also called intercoder reliability — measures how consistently two or more researchers apply the same codes to the same qualitative data. The most defensible way to report it is with a chance-corrected statistic such as Cohen's kappa or Krippendorff's alpha, where a value of **0.80 or higher is generally accepted as reliable**, 0.667–0.80 supports tentative conclusions, and anything below 0.667 is considered insufficient for drawing inferences. If two trained coders read the same interview and disagree on what it means, your themes aren't findings — they're opinions. This guide shows you how to measure agreement, which statistic to choose, and how AI-native platforms like Koji make consistent coding the default rather than an afterthought.\n\n## What Is Inter-Rater Reliability?\n\nInter-rater reliability is the degree to which independent coders assign the same codes, categories, or ratings to the same units of qualitative data. In practice, two researchers each read a set of [interview transcripts](/docs/coding-qualitative-data), apply a shared [codebook](/docs/qualitative-research-codebook), and then you compare how often they agreed.\n\nThe term \"inter-rater reliability\" is used interchangeably with \"intercoder reliability\" and \"intercoder agreement.\" Whatever you call it, the goal is the same: to demonstrate that your coding scheme is reproducible and not simply a reflection of one researcher's idiosyncratic interpretation. When a study reports strong IRR, a reader can trust that the themes would hold up if a different qualified researcher analyzed the same data.\n\nThis is distinct from broader [study-level validity and reliability](/docs/qualitative-research-validity), which concerns whether your entire research design produces trustworthy conclusions. IRR is narrower and more measurable: it is specifically about agreement at the point of coding.\n\n## Why Inter-Rater Reliability Matters\n\nQualitative analysis is interpretive by nature, and that is its strength — but interpretation without verification is where bias creeps in. Without a reliability check, you have no way to distinguish a genuine pattern in your data from a pattern that exists only in the analyst's head.\n\nAs Cliodhna O'Connor and Helene Joffe argue in their widely cited 2020 methodological review in the *International Journal of Qualitative Methods*, intercoder reliability \"can enhance the systematicity, communicability, and transparency of the coding process; prompt reflection and discussion among the research team; and help safeguard against the imposition of a single researcher's assumptions on the data.\" In other words, the act of measuring agreement improves the research itself, not just the credibility score you report.\n\nThe stakes are practical. Product and research teams routinely make roadmap, pricing, and positioning decisions on the back of a handful of coded interviews. If the coding is unreliable, every downstream decision inherits that error.\n\n## Percent Agreement Is Not Enough\n\nThe simplest measure of agreement is **percent agreement**: the proportion of coding decisions where coders matched. It is intuitive, but it has a fatal flaw — it ignores agreement that would happen by chance alone.\n\nImagine two coders deciding whether each quote expresses \"frustration.\" If 90% of quotes don't express frustration, two coders randomly guessing \"not frustrated\" most of the time would agree roughly 80% of the time without reading anything. A raw 80% agreement number sounds impressive but may reflect almost nothing.\n\nThat is why methodologists insist on **chance-corrected coefficients**. These statistics subtract out the agreement you would expect from random chance and report only the agreement beyond it.\n\n## Cohen's Kappa vs. Krippendorff's Alpha\n\nThe two most common chance-corrected statistics are Cohen's kappa and Krippendorff's alpha.\n\n**Cohen's kappa** is the most widely used coefficient because of its relative simplicity and because it accounts for chance agreement. Its main limitations: it handles only two coders and assumes nominal categories. It also behaves erratically when codes are highly imbalanced — the so-called kappa paradox, where high agreement can produce a low kappa.\n\n**Krippendorff's alpha** is considered more robust and flexible. It accommodates any number of coders, different levels of measurement (nominal, ordinal, interval, ratio), and missing data. For these reasons many measurement specialists, including the team behind the ATLAS.ti research hub, recommend Krippendorff's alpha over Cohen's kappa for most qualitative coding projects.\n\nA practical rule of thumb: if you have exactly two coders applying simple categorical codes, Cohen's kappa is fine and easy to explain. If you have three or more coders, ordinal scales, or incomplete coding, reach for Krippendorff's alpha.\n\n## What Counts as \"Reliable\"? Interpreting the Thresholds\n\nThe most cited benchmark comes from Landis and Koch (1977), who proposed the following gradient for kappa-type statistics:\n\n- **0.81–1.00** — almost perfect agreement\n- **0.61–0.80** — substantial agreement\n- **0.41–0.60** — moderate agreement\n- **0.21–0.40** — fair agreement\n- **0.00–0.20** — slight agreement\n\nFor publication-grade work, the conventional standard is stricter. Krippendorff recommends treating **α ≥ 0.80 as satisfactory**, **0.667–0.80 as adequate only for tentative conclusions**, and **below 0.667 as insufficient** for drawing reliable inferences. Miles and Huberman's influential guidance suggests aiming for agreement of around 0.80 across roughly 95% of your codes.\n\nDon't fetishize a single number. A high coefficient on a trivially easy coding scheme proves little, and a slightly lower coefficient on a nuanced interpretive scheme may still represent rigorous work — as long as you are transparent about how you got there.\n\n## How to Calculate Inter-Rater Reliability: Step by Step\n\n1. **Develop a clear codebook.** Each code needs a name, a definition, inclusion and exclusion criteria, and an example. Ambiguous definitions are the single biggest driver of low reliability. See our [codebook guide](/docs/qualitative-research-codebook).\n2. **Train your coders.** Walk through the codebook together and code a few practice transcripts as a group before going independent.\n3. **Code independently.** Two or more coders apply the codebook to the same subset of data — commonly 10–25% of the full dataset — without conferring.\n4. **Build an agreement matrix.** For each coded unit, record what each coder assigned.\n5. **Calculate the coefficient.** Compute Cohen's kappa or Krippendorff's alpha. Tools like ATLAS.ti, NVivo, Dedoose, and open-source R and Python packages do this automatically.\n6. **Resolve disagreements.** Where coders diverge, discuss, refine ambiguous code definitions, and re-code. This step often improves the codebook itself.\n7. **Report transparently.** State the statistic used, the value achieved, the proportion of data double-coded, and how disagreements were resolved.\n\n## Common Pitfalls That Sink Reliability\n\n- **Vague code definitions.** If two smart people can read the same definition differently, your kappa will suffer.\n- **Too many codes.** Bloated codebooks with overlapping categories invite disagreement.\n- **Coding the whole dataset before checking.** Catch reliability problems early on a sample, not after 40 hours of work.\n- **Reporting only percent agreement.** Reviewers and savvy stakeholders will discount it.\n- **Treating IRR as a one-time gate.** Reliability can drift as coders fatigue. Spot-check throughout.\n\n## The Modern Approach: Consistent Coding With AI\n\nHere is the uncomfortable truth about traditional IRR: it exists largely to compensate for the fact that humans are inconsistent. Two researchers get tired, bring different assumptions, and drift over a long coding session. Inter-rater reliability is the patch we apply to a fundamentally manual, error-prone process.\n\nAI-native research changes the equation. A well-tuned AI coder applies the same definitions to the first transcript and the five-hundredth with no fatigue and no drift — the consistency that IRR is designed to verify becomes the baseline. Recent research bears this out: a 2025 comparative study on arXiv evaluating large language models for deductive qualitative coding found that LLMs can achieve substantial-to-strong agreement with expert human coders on well-defined schemes, positioning AI as a powerful complement to human judgment rather than a replacement for it.\n\nThis is exactly how [Koji](/docs/structured-questions-guide) is built. Koji runs AI-moderated interviews and then applies **automatic thematic analysis** with a consistent coding logic across every conversation — so the \"second coder\" is effectively built in. Where you want quantifiable consistency, Koji's six **structured question types** (open_ended, scale, single_choice, multiple_choice, ranking, and yes_no) capture responses in pre-defined categories that need no subjective coding at all, eliminating inter-rater disagreement at the source for those items. For the open-ended responses that do require interpretation, Koji's [auto-tagging](/docs/ai-auto-tagging-customer-interviews) produces a transparent, reproducible code structure you can audit — and a human researcher stays in the loop to validate and refine themes.\n\nThe result: instead of spending 40 hours coding and then a reliability ritual to prove you were consistent, you start from a consistent, auditable analysis and spend your time on interpretation and decisions. Teams using AI-assisted analysis routinely report cutting time-to-insight dramatically while preserving — and arguably improving — coding consistency.\n\nYou don't need a PhD in measurement theory to produce trustworthy qualitative findings. You need clear definitions, a transparent process, and tooling that makes consistency the default.\n\n## Related Resources\n\n- [Qualitative Coding: How to Code Interview Data](/docs/coding-qualitative-data)\n- [How to Build a Qualitative Research Codebook](/docs/qualitative-research-codebook)\n- [Qualitative Research Validity and Reliability](/docs/qualitative-research-validity)\n- [The Complete Guide to Thematic Analysis](/docs/thematic-analysis-guide)\n- [AI Auto-Tagging for Customer Interviews](/docs/ai-auto-tagging-customer-interviews)\n- [Structured Questions Guide: The 6 Question Types](/docs/structured-questions-guide)","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Inter-Rater Reliability in Qualitative Research: Cohen's Kappa & Krippendorff's Alpha Guide","metaDescription":"How to measure inter-rater (intercoder) reliability in qualitative research — Cohen's kappa vs Krippendorff's alpha, reliable thresholds (0.80+), step-by-step calculation, and AI-native consistent coding.","keywords":["inter-rater reliability","intercoder reliability","Cohen's kappa","Krippendorff's alpha","intercoder agreement","qualitative coding reliability","coding agreement","qualitative research"],"aiSummary":"Inter-rater (intercoder) reliability measures how consistently independent coders apply the same codes to qualitative data. Report it with a chance-corrected statistic — Cohen's kappa (two coders, nominal data) or Krippendorff's alpha (more flexible). Thresholds: 0.80+ is reliable, 0.667–0.80 supports tentative conclusions, below 0.667 is insufficient. Percent agreement alone is misleading because it ignores chance. AI-native platforms like Koji make consistent coding the default through automatic thematic analysis and structured question types.","aiPrerequisites":["Basic understanding of qualitative coding","Familiarity with thematic analysis"],"aiLearningOutcomes":["Define inter-rater and intercoder reliability","Choose between Cohen's kappa and Krippendorff's alpha","Interpret reliability thresholds correctly","Calculate IRR step by step","Use AI to make coding consistent and auditable"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"c7f2064e-d82c-48d2-9cf0-c4f00826e902","slug":"cognitive-walkthrough-guide","title":"Cognitive Walkthrough: The Complete Guide to Learnability Inspection (2026)","url":"https://www.koji.so/docs/cognitive-walkthrough-guide","summary":"A definitive guide to the cognitive walkthrough method: the four original Wharton questions, Spencer’s streamlined two-question version, the difference between cognitive walkthrough and heuristic evaluation, a step-by-step workshop protocol, and how Koji turns inspection findings into validated insights via AI-moderated user interviews.","content":"## What Is a Cognitive Walkthrough?\n\nA **cognitive walkthrough** is a task-based usability inspection method in which a small group of evaluators steps through every individual action required to complete a user task and asks a fixed set of questions about whether a *first-time, learning-by-doing* user would know what to do at each step. It is one of the two foundational discount-usability inspection methods (alongside [heuristic evaluation](/docs/heuristic-evaluation-guide)), and is specifically engineered to evaluate **learnability** rather than overall interface quality.\n\nThe method was developed in the early 1990s by Cathleen Wharton, Peter Polson, John Rieman, and Clayton Lewis at the University of Colorado, and reached a mass audience through Jakob Nielsen and Robert Mack’s 1994 book *Usability Inspection Methods*. Three decades later, it remains in active use at organisations including Microsoft, Google, IBM, and the Nielsen Norman Group, because it costs nothing, finds learnability problems early, and works on paper sketches as well as on shipped products.\n\n## When to Use a Cognitive Walkthrough\n\nUse a cognitive walkthrough when:\n\n- The product or feature is **aimed at new or infrequent users** (signup flows, public sector services, kiosks, healthcare portals, anything with low task frequency)\n- You have a **specific task** you can describe in steps (booking a ticket, configuring a setting, completing onboarding)\n- You are **early enough that running a moderated user test would be wasteful** — paper prototype, low-fidelity wireframe, half-built screen\n- You need a **shared cross-functional understanding** of what makes the task hard, not just a list of issues\n\nDo **not** use a cognitive walkthrough when:\n\n- You are reviewing aesthetic/visual design quality — use [heuristic evaluation](/docs/heuristic-evaluation-guide)\n- You are testing whether the *concept* solves the right problem — use [generative interviews](/docs/jobs-to-be-done-framework) or [concept testing](/docs/concept-testing-methodology)\n- You need numeric usability benchmarks — run a [SUS study](/docs/system-usability-scale-guide) with real users\n\n## The Four Questions (Wharton et al., 1994)\n\nAt every individual step in the task, each evaluator answers four questions about a hypothetical user:\n\n1. **Will the user try to achieve the right effect?** Does the user even know that this step is the next thing they should be doing? Is it on the path they expect?\n2. **Will the user notice that the correct action is available?** Is the right control visible, in a sensible location, and recognisable?\n3. **Will the user associate the correct action with the effect they are trying to achieve?** Does the wording, icon, and affordance make the user believe *this* is the control that produces *that* outcome?\n4. **If the correct action is performed, will the user see that progress is being made?** Is the system response clear, prompt, and unambiguous enough that the user knows they’ve succeeded?\n\nA “no” answer to any of the four questions becomes a **failure story** — a written hypothesis about why a real user would get stuck at that step. The collected failure stories are the deliverable.\n\n> **Method note.** Wharton’s original protocol asked nine questions per step. The four-question version above is the version finalised in *Usability Inspection Methods* and is what most modern guides — including the Nielsen Norman Group — refer to as “the cognitive walkthrough.”\n\n## Spencer’s Streamlined Version (2000)\n\nIn 2000, Rick Spencer (then a usability engineer at Microsoft) published *The Streamlined Cognitive Walkthrough Method, Working Around Social Constraints Encountered in a Software Development Company* in the CHI proceedings. He argued that the classic four-question version was too slow and too academic to survive in industry, and proposed a stripped-down two-question version:\n\n1. **Will the user know what to do at this step?** (collapses Wharton Q1, Q2, Q3)\n2. **Will the user understand from the response that they did the right thing and that progress was made?** (Wharton Q4)\n\nSpencer also recommended:\n\n- **Fewer evaluators** — 1 or 2 are enough\n- **Shorter sessions** — 60 minutes max\n- **Lighter documentation** — bullet-point findings, not narrative reports\n- **A designated facilitator** who keeps the group from sliding into solutioning mid-walkthrough\n\nMost modern product teams use Spencer’s version by default. It loses some rigour against the academic method but recovers the time cost that killed cognitive walkthrough adoption in commercial settings.\n\n## How to Run a Cognitive Walkthrough — Step by Step\n\n### Step 1: Define the user persona\nWrite a 2–3 sentence description of the *learning-by-doing* user the walkthrough imagines: their domain knowledge, technology experience, motivation, and constraints. Without an explicit persona, evaluators silently substitute their own expert mental model and the walkthrough becomes worthless.\n\n### Step 2: Choose the task and define success\nPick **one specific task** that this persona would realistically attempt. State the start state, the end state, and what counts as task success (“user has booked an appointment and seen a confirmation screen”).\n\n### Step 3: Decompose the task into actions\nWrite out the **correct sequence of actions** the user must perform, one row per action, end-to-end. This is the most under-rated step — a sloppy decomposition produces a sloppy walkthrough.\n\n### Step 4: Walk through, action by action, asking the questions\nFor each action, the group answers the 2 (Spencer) or 4 (Wharton) questions. The facilitator captures every “no” as a failure story with three components: the step, the hypothesised user behaviour, and the predicted consequence.\n\n### Step 5: Synthesise findings and prioritise\nGroup failure stories into themes, prioritise by severity (catastrophic / serious / minor / cosmetic), and assign owners. The deliverable is typically a 1–2 page bullet list with screenshots and severity tags.\n\n## Cognitive Walkthrough vs. Heuristic Evaluation\n\n| Dimension | Cognitive Walkthrough | Heuristic Evaluation |\n|---|---|---|\n| Driver | A specific user task | A set of usability principles |\n| Best for | Learnability for new users | Overall interface quality |\n| Evaluator count | 1–5 | 3–5 ideal |\n| Time to run | 1–4 hours | 1–2 hours |\n| Output | Failure stories per task step | Heuristic violations per principle |\n| Skill required | Moderate — needs persona discipline | Higher — needs heuristics knowledge |\n| Best timing | Early design, prototype | Any stage |\n\nThe two methods are **complementary, not competing**. A common workflow is to run a cognitive walkthrough on the critical learnability tasks, then run a heuristic evaluation across the rest of the product. Together they cover roughly 75–80% of the usability issues a moderated test would surface — at a fraction of the cost.\n\n## Common Mistakes\n\n1. **Skipping the persona definition.** Without an explicit persona, the team falls back on their own expertise and misses every learnability issue.\n2. **Choosing a task that is too broad.** “Use the dashboard” is not a task. “Add a new credit card to billing” is a task.\n3. **Solutioning during the walkthrough.** The walkthrough is for *finding* issues. Save fixes for a separate session.\n4. **Treating the walkthrough as a substitute for user testing.** It is a hypothesis-generation method, not a hypothesis-validation method. Plan to validate the findings with real users.\n5. **Documenting nothing.** A cognitive walkthrough whose findings live only in the participants’ memories is a meeting, not a method.\n\n## The Modern Approach: Validate Walkthrough Findings With AI-Moderated Research\n\nThe historic limitation of cognitive walkthroughs has always been the gap between *predicted* user behaviour and *actual* user behaviour. Inspection methods are powerful at generating hypotheses, but they cannot tell you which hypotheses are real. Traditionally, validating each finding meant scheduling a moderated usability test, recruiting 5–8 participants, running 60-minute sessions, and writing a report — a 2–3 week process for what was originally a 2-hour inspection.\n\nThis is exactly what AI-native research platforms like **Koji** collapse. The modern walkthrough → validation pipeline looks like this:\n\n1. **Run the cognitive walkthrough** in a 60-minute Spencer-style session, producing 5–15 failure-story hypotheses.\n2. **Convert each failure story into a Koji study task.** Use Koji’s [structured questions](/docs/structured-questions-guide) to attach a Single Ease Question (scale 1–7), a binary success yes_no item, and an open-ended “what made it hard or easy?” probe.\n3. **Launch via personalised link or in-product widget.** Koji’s AI moderator runs the task with users 24/7, with no scheduling overhead.\n4. **Watch the report populate in real time.** SEQ averages, success rates, and themed open-ended responses appear on the live dashboard within hours, not weeks.\n\nResearch from Forrester’s *State of Customer Insights 2024* found that teams using AI-moderated research report **60% faster time-to-insight** than teams running equivalent moderated studies manually. For inspection-driven validation work, the difference is even larger — Koji customers regularly turn a 2-hour walkthrough plus 2-week validation cycle into a 2-hour walkthrough plus 2-day validation cycle.\n\nThe broader lesson is that cognitive walkthroughs have always been the *cheapest* usability method to start, but historically the *most expensive* method to act on, because every finding generated more downstream qualitative work. AI-moderated research closes that gap — making the inspection-then-validate workflow finally affordable end to end.\n\n## A Cognitive Walkthrough Workshop Template (Spencer Version)\n\nUse this template to run a 60-minute walkthrough with 1–2 evaluators:\n\n- **0:00–0:05** Persona statement (read aloud, agree)\n- **0:05–0:10** Task statement and decomposition (write actions on a whiteboard)\n- **0:10–0:50** Walk through, action by action: for each, answer (1) Will the user know what to do? and (2) Will they see they did the right thing? Capture every “no” as a failure story.\n- **0:50–0:60** Sort failure stories by severity, assign owners, agree which to validate with users on Koji.\n\nThat’s the entire method. The reason cognitive walkthrough survived three decades of UX trend cycles is not its rigour — it is its compression. A discount inspection method that fits in a single working hour and finds learnability issues a moderated test would find a month later remains, in 2026, one of the highest-leverage tools in a product team’s research toolkit.\n\n## Related Resources\n\n- [Heuristic Evaluation: The Complete UX Review Guide](/docs/heuristic-evaluation-guide) — the principle-driven inspection counterpart to cognitive walkthrough\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — Koji’s six question types for validating walkthrough findings at scale\n- [Think-Aloud Protocol: How to Run and Analyze Sessions](/docs/think-aloud-protocol) — the moderated complement that surfaces *why* a step is hard\n- [First-Click Testing: Validating Navigation and Findability](/docs/first-click-testing-guide) — a quantitative companion for Q2 and Q3 of the walkthrough\n- [Tree Testing: Information Architecture Validation](/docs/tree-testing-guide) — useful when walkthrough failure stories cluster around findability\n- [UX Research Process: A Complete Framework for 2026](/docs/ux-research-process) — where inspection methods fit in the broader research workflow\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n- [How to Run Customer Exit Interviews: The Complete Guide (2026)](/blog/customer-exit-interviews-guide-2026) — Customer exit interviews reveal the real reasons customers churn — not the polished answer they gave on your cancellation form. Here is how \n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Cognitive Walkthrough: The 4-Question Usability Inspection Method (2026)","metaDescription":"The complete cognitive walkthrough guide: the original four questions, Spencer’s streamlined version, when to use it over heuristic evaluation, plus a downloadable workshop template and AI-moderated validation on Koji.","keywords":["cognitive walkthrough","usability inspection","learnability testing","wharton 1994","spencer streamlined cognitive walkthrough","task-based usability","ux inspection methods","heuristic evaluation","jakob nielsen"],"aiSummary":"A definitive guide to the cognitive walkthrough method: the four original Wharton questions, Spencer’s streamlined two-question version, the difference between cognitive walkthrough and heuristic evaluation, a step-by-step workshop protocol, and how Koji turns inspection findings into validated insights via AI-moderated user interviews.","aiPrerequisites":["Familiarity with usability testing concepts","A specific user task you want to evaluate","Access to a working interface, prototype, or wireframe"],"aiLearningOutcomes":["Run a cognitive walkthrough using the original 4-question protocol","Apply Spencer’s streamlined 2-question version for fast-paced sprints","Choose between cognitive walkthrough and heuristic evaluation","Build a complete walkthrough workshop deliverable","Validate inspection findings with real users in days, not weeks, using Koji"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"a68de75b-46c8-4b2b-89cb-bd0a930fc404","slug":"heuristic-evaluation-guide","title":"Heuristic Evaluation: The Complete UX Review Guide","url":"https://www.koji.so/docs/heuristic-evaluation-guide","summary":"Heuristic evaluation is a structured usability inspection method where 3–5 experts assess an interface against Nielsen's 10 usability heuristics. Five evaluators find ~75% of usability problems at ~$10.54 per issue — 3–4× more cost-efficient than user testing. Best used before user testing to eliminate obvious issues, during iterative design, or when auditing inherited products. Includes all 10 heuristics with examples, severity rating scale, and how to combine expert review with AI-moderated user interviews.","content":"## The Fastest Way to Find 75% of Your Usability Problems\n\nBefore you recruit a single user, you can find three-quarters of your product's usability problems in a single afternoon.\n\nHeuristic evaluation is a structured usability inspection method where evaluators examine a user interface and judge its compliance against recognized usability principles — called \"heuristics.\" Unlike user testing, it requires no participants, no recruitment, and no scheduling. A trained evaluator can identify critical usability issues in two to three hours per interface.\n\nThe method's power comes from its evidence base. Jakob Nielsen's original research showed that **five evaluators conducting independent heuristic evaluations discover approximately 75% of all usability problems** in an interface — far more than any single evaluator (who finds only 35% on average) and at 3× lower cost than user testing.\n\nFor product teams that need fast, reliable usability insight before a launch or redesign, heuristic evaluation is often the highest-ROI research activity available.\n\n---\n\n## What Is Heuristic Evaluation?\n\nHeuristic evaluation is a usability inspection technique, originally formalized by Jakob Nielsen and Rolf Molich in 1990. Evaluators — typically usability specialists — examine an interface systematically and assess each element against a set of established usability principles.\n\nThe term \"heuristic\" comes from the Greek *heuriskein* (\"to discover\"). In UX, heuristics are rules-of-thumb that capture core principles of effective interface design. When an interface violates these principles, users are more likely to make errors, feel confused, or give up.\n\nThe method differs from user testing in a fundamental way: **user testing observes real users; heuristic evaluation applies expert judgment**. Both are valuable. Neither replaces the other.\n\n> \"Heuristic evaluation is the most popular of the usability inspection methods. It is particularly useful as a quick feedback mechanism during the design stage, when resources are insufficient for more elaborate methods like usability testing.\"\n> — Nielsen Norman Group\n\n---\n\n## Nielsen's 10 Usability Heuristics\n\nJakob Nielsen's 10 heuristics were derived from factor analysis of **249 usability problems** identified across 11 different professional projects. First published in 1994, they remain unchanged — a testament to how well they capture fundamental truths about human-computer interaction.\n\n### 1. Visibility of System Status\nThe system should always keep users informed about what is happening through appropriate and timely feedback.\n\n*Violation example: A file upload with no progress indicator. Users don't know if it's working, stuck, or failed.*\n\n*Fix: Show upload progress bar with percentage and estimated time remaining.*\n\n### 2. Match Between System and the Real World\nThe system should speak the user's language — familiar words, phrases, and concepts rather than system-oriented jargon.\n\n*Violation example: An error message reading \"Error 0x8007045D: I/O device error.\"*\n\n*Fix: \"We couldn't save your file. Your storage device may be full or disconnected. Try saving to a different location.\"*\n\n### 3. User Control and Freedom\nUsers often choose system functions by mistake and need clearly marked \"emergency exits\" to leave the unwanted state without extended dialogue.\n\n*Violation example: A multi-step form with no way to go back and change a previous answer.*\n\n*Fix: Provide back navigation, undo functionality, and cancel options at every step.*\n\n### 4. Consistency and Standards\nUsers should not have to wonder whether different words, situations, or actions mean the same thing. Follow platform conventions.\n\n*Violation example: Some buttons say \"Submit,\" others say \"Send,\" others say \"Continue\" for the same action type across different screens.*\n\n*Fix: Establish and apply consistent terminology and interaction patterns throughout the product.*\n\n### 5. Error Prevention\nEven better than good error messages is a careful design that prevents a problem from occurring in the first place.\n\n*Violation example: A \"Delete Account\" button with no confirmation step, positioned near \"Edit Profile.\"*\n\n*Fix: Require explicit confirmation with a typed phrase (\"type DELETE to confirm\") for irreversible actions.*\n\n### 6. Recognition Rather Than Recall\nMinimize the user's memory load by making objects, actions, and options visible. The user should not have to remember information from one part of the interface to another.\n\n*Violation example: A checkout flow that shows shipping options on step 1 but doesn't display the chosen option on the payment step.*\n\n*Fix: Show a persistent order summary sidebar throughout the checkout process.*\n\n### 7. Flexibility and Efficiency of Use\nAccelerators — unseen by novice users — may speed up interaction for expert users, so the system can cater to both inexperienced and experienced users.\n\n*Violation example: A data entry form that requires mouse clicks between fields with no keyboard tab navigation.*\n\n*Fix: Support keyboard shortcuts, bulk actions, and advanced filtering for power users while keeping the default interface simple.*\n\n### 8. Aesthetic and Minimalist Design\nDialogues should not contain irrelevant or rarely needed information. Every extra unit of information competes with the relevant information and diminishes its relative visibility.\n\n*Violation example: A dashboard crammed with 30 metrics, 15 charts, and 8 action buttons.*\n\n*Fix: Progressive disclosure — show the 5 most important metrics by default with an option to expand.*\n\n### 9. Help Users Recognize, Diagnose, and Recover from Errors\nError messages should be expressed in plain language (no error codes), precisely indicate the problem, and constructively suggest a solution.\n\n*Violation example: \"Invalid input\" with no indication of which field failed or why.*\n\n*Fix: Inline validation showing \"Password must be at least 8 characters and include one number\" immediately when the user leaves the field.*\n\n### 10. Help and Documentation\nEven though it is better if the system can be used without documentation, it may be necessary to provide help. Such information should be easy to search and focused on the user's task.\n\n*Violation example: A help center with only generic category pages and no search function.*\n\n*Fix: Contextual help tooltips, in-app guided tours, and a searchable knowledge base accessible from any screen.*\n\n---\n\n## The Research Behind Heuristic Evaluation\n\n### How Many Evaluators Do You Need?\n\nNielsen's research on evaluator effectiveness produced one of the most cited findings in UX research: **the diminishing returns curve for usability evaluators**.\n\n| Number of Evaluators | % of Usability Problems Found |\n|---|---|\n| 1 | ~35% |\n| 2 | ~52% |\n| 3 | ~62% |\n| 5 | ~75% |\n| 8 | ~83% |\n| 10 | ~85% |\n| 15 | ~90% |\n\nThe curve flattens sharply after 5 evaluators. Each additional evaluator beyond 5 contributes diminishing returns relative to their cost. For most practical applications, **3–5 evaluators is the optimal range** — balancing coverage against resource cost.\n\n### Cost-Effectiveness vs. User Testing\n\nA comparative study found heuristic evaluation costs approximately **$10.54 per usability issue found**, versus **$47.30 per issue in user testing** — making heuristic evaluation roughly 4.5× more cost-effective per issue discovered. The same study found heuristic evaluation required 15.5 hours including analysis, compared to 45 hours for user testing.\n\nThis does not mean heuristic evaluation is *better* than user testing. It means the two methods are complementary: heuristic evaluation efficiently covers breadth (finding many issues quickly), while user testing provides depth (understanding severity and real-world impact).\n\n### What Heuristic Evaluation Doesn't Tell You\n\nHeuristic evaluation has a documented **false positive rate of approximately 29%** — issues flagged by evaluators that, when tested with real users, turn out not to be actual problems. This is why heuristic evaluation should inform, but not replace, user research.\n\nIt also cannot:\n- Reveal which issues actually affect task completion rates\n- Measure the emotional response of real users\n- Surface unknown unknowns about user behavior and mental models\n- Validate whether a solution actually solves the problem\n\n---\n\n## When to Use Heuristic Evaluation\n\n### Best Use Cases\n\n**1. Before user testing:** Run a heuristic evaluation first to eliminate obvious issues. This lets your user testing sessions focus on deeper, more nuanced questions rather than cataloguing visible interface problems.\n\n**2. During iterative design:** Heuristic evaluation is fast enough to run on every major iteration — wireframes, prototypes, and live builds. User testing every iteration is impractical; heuristic evaluation is not.\n\n**3. Auditing inherited products:** When your team takes over an existing product with no research history, a heuristic evaluation quickly maps the landscape of usability debt.\n\n**4. Evaluating competitor products:** Applying heuristics to competitor interfaces reveals gaps and opportunities that can inform your own design strategy.\n\n**5. Budget-constrained research:** When you cannot run user testing, a heuristic evaluation is better than no research. Five hours of expert review surfaces real issues that would otherwise be shipped.\n\n**6. Supplementing quantitative data:** If analytics show a 60% drop-off at checkout, a heuristic evaluation of that flow often explains *why* — identifying the specific violations driving abandonment.\n\n### When NOT to Use Heuristic Evaluation Alone\n\n- **Validating a new concept:** Heuristics evaluate execution quality, not whether the concept itself is right. For that, you need user research.\n- **Understanding user motivations:** Heuristics cannot reveal *why* users behave as they do or what jobs they're trying to accomplish.\n- **Measuring usability improvement:** For before/after benchmarking, user testing with task completion metrics is more reliable.\n\n---\n\n## How to Conduct a Heuristic Evaluation: Step-by-Step\n\n### Step 1: Define Scope and Scenarios\n\nBefore evaluating, agree on:\n- **Scope:** Which screens, flows, or features will be evaluated?\n- **User tasks:** What are the 3–5 most important tasks users need to complete? Evaluators should walk through each task during their review.\n- **User context:** Who is the target user, and what is their technical proficiency?\n\n### Step 2: Select Evaluators\n\nUse 3–5 evaluators for optimal coverage. Evaluators should ideally have:\n- Familiarity with usability principles and Nielsen's heuristics\n- Understanding of the product domain\n- No recent deep involvement in designing the interface being evaluated (to avoid blind spots)\n\n### Step 3: Independent Evaluation Sessions\n\nEach evaluator should work independently to avoid anchoring bias. A typical session:\n1. **First pass (20–30 min):** Walk through the entire interface to get a general feel\n2. **Second pass (45–90 min):** Evaluate systematically against each of the 10 heuristics, documenting each issue found\n3. **Severity ratings:** Rate each issue 0–4 (0 = not a usability problem; 4 = usability catastrophe)\n\n**Severity scale:**\n- **0** — Not a usability problem\n- **1** — Cosmetic problem only; fix only if time permits\n- **2** — Minor usability problem; low priority\n- **3** — Major usability problem; important to fix\n- **4** — Usability catastrophe; imperative to fix before product ships\n\n### Step 4: Aggregate and Prioritize Findings\n\nAfter independent evaluations are complete, facilitators aggregate all findings. Combine duplicate issues and average severity ratings across evaluators. Sort by combined severity to create a prioritized issue list.\n\n### Step 5: Present and Action Findings\n\nA heuristic evaluation deliverable typically includes:\n- Executive summary with total issues by category and severity\n- Detailed issue log with heuristic violated, severity rating, and recommended fix\n- Top 5–10 critical issues requiring immediate attention\n- Optional: comparative benchmark against baseline or competitor\n\n---\n\n## Common Heuristic Violations by Domain\n\nResearch shows different types of software fail in predictable ways:\n\n**Enterprise/B2B software:** Most frequent violations are Heuristic 1 (Visibility of Status), Heuristic 4 (Consistency), and Heuristic 10 (Help and Documentation) — reflecting complexity and poor onboarding.\n\n**E-commerce:** Most violations in Heuristic 3 (User Control/Freedom) and Heuristic 5 (Error Prevention) — checkout flows that trap users and create irreversible states.\n\n**Mobile apps:** Most violations in Heuristic 6 (Recognition vs. Recall) and Heuristic 8 (Minimalist Design) — interfaces that cram too much onto small screens.\n\n**Healthcare/Medical software:** Highest violation rates in Heuristic 9 (Error Recovery) and Heuristic 2 (Match to Real World) — critical given the stakes of medical errors.\n\n---\n\n## Heuristic Evaluation in the Modern Research Stack\n\nTraditional heuristic evaluation produces a one-time snapshot. But in modern product development — where interfaces change weekly — the limitation is currency: by the time an audit is complete, the design has moved on.\n\nThe more valuable approach combines heuristic evaluation's analytical rigor with ongoing user feedback to stay current:\n\n1. **Heuristic evaluation** identifies candidate problem areas from expert review\n2. **AI-moderated user interviews** validate whether real users experience those issues and uncover behavioral context\n3. **Structured questions** in Koji ([scale](/docs/structured-questions-guide), [single_choice](/docs/structured-questions-guide), [yes_no](/docs/structured-questions-guide)) quantify the prevalence and severity of specific pain points across your user base\n4. **Automated theme extraction** surfaces patterns across dozens of conversations without manual coding\n\nThis combination moves evaluation from \"expert opinion\" to \"validated evidence\" — and it does so continuously rather than as an isolated project.\n\nFor example: a heuristic evaluation flags Heuristic 3 (User Control) as violated in your checkout flow. Rather than guessing the severity, you run 30 AI-moderated interviews where users walk through checkout and describe friction points. Koji automatically extracts themes across transcripts, and your scale question — \"How difficult was completing the purchase?\" — gives you a quantified severity score you can track over iterations.\n\nWhile traditional expert review tools require expensive consultants and manual report writing, AI-native research platforms like Koji let teams embed this kind of evaluation rigor into their regular product cycle — without the overhead.\n\n---\n\n## Heuristic Evaluation vs. User Testing: The Definitive Comparison\n\n| Dimension | Heuristic Evaluation | User Testing |\n|---|---|---|\n| **Who generates data** | Usability experts | Real users |\n| **Speed** | 1–2 days | 2–4 weeks |\n| **Cost per issue** | ~$10.54 | ~$47.30 |\n| **Issues found** | 35–83% (1–8 evaluators) | Varies with # participants |\n| **False positive rate** | ~29% | Low (real behavior) |\n| **What it reveals** | Violations of known principles | Actual user behavior and motivation |\n| **When to use** | Early and iterative | Pre-launch validation |\n| **Can replace each other?** | No | No |\n\n**Best practice:** Use heuristic evaluation to clean up known issues before user testing, so your user testing sessions surface deeper insights rather than obvious problems.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — quantify usability issues with scale and yes/no questions alongside qualitative exploration\n- [Usability Testing Survey Guide](/docs/usability-testing-survey-guide) — how to design post-session surveys that capture usability data\n- [How to Write User Interview Questions That Surface Real Insights](/docs/user-interview-questions) — go beyond expert review with real user feedback\n- [Think-Aloud Protocol](/docs/think-aloud-protocol) — combine cognitive walkthrough with heuristic review for richer insights\n- [Mixed Methods Research Guide](/docs/mixed-methods-research-guide) — combining heuristic evaluation with user research for complete coverage\n- [Prototype Testing and Concept Validation](/docs/prototype-testing-concept-validation) — apply heuristics to early-stage concepts before development\n\n## Further reading on the blog\n\n- [Best AI Market Research Tools in 2026: The Complete Buyer's Guide](/blog/ai-market-research-tools-2026) — AI has fundamentally changed market research. This guide compares the leading AI market research platforms—from AI-native interview tools li\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [Best AI Customer Interview Tools in 2026: The Complete Buyer's Guide](/blog/best-ai-customer-interview-tools-2026) — AI has fundamentally changed how product teams conduct customer research. Here are the best AI customer interview tools in 2026 — ranked by \n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Heuristic Evaluation: The Complete UX Review Guide","metaDescription":"Master heuristic evaluation with Nielsen's 10 usability heuristics. Learn how 5 evaluators find 75% of usability problems, severity ratings, step-by-step process, and how to combine expert review with AI user research.","keywords":["heuristic evaluation","Nielsen 10 heuristics","usability heuristics","UX evaluation methods","usability inspection","expert review UX","Jakob Nielsen heuristics","usability testing vs heuristic evaluation"],"aiSummary":"Heuristic evaluation is a structured usability inspection method where 3–5 experts assess an interface against Nielsen's 10 usability heuristics. Five evaluators find ~75% of usability problems at ~$10.54 per issue — 3–4× more cost-efficient than user testing. Best used before user testing to eliminate obvious issues, during iterative design, or when auditing inherited products. Includes all 10 heuristics with examples, severity rating scale, and how to combine expert review with AI-moderated user interviews.","aiPrerequisites":["Basic familiarity with UX design and product development"],"aiLearningOutcomes":["Apply all 10 of Nielsen's usability heuristics with real examples","Know how many evaluators to use and why","Rate issue severity using the 0–4 scale","Run a heuristic evaluation end-to-end","Combine expert review with AI user research for validated findings"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 minutes"},{"type":"documentation","id":"76990745-46d5-4717-b57e-6494f3c3a3f1","slug":"single-ease-question-seq-guide","title":"Single Ease Question (SEQ): The 7-Point UX Metric for Task-Level Usability (2026)","url":"https://www.koji.so/docs/single-ease-question-seq-guide","summary":"A definitive guide to the Single Ease Question (SEQ): the verbatim 7-point wording, MeasuringU/Sauro benchmark of 5.3–5.5, correlation with task completion (5.9 ≈ 86% completion), when to choose SEQ over SUS, sample-size guidance, and how Koji turns a 2-week SEQ study into an afternoon.","content":"## What Is the Single Ease Question (SEQ)?\n\nThe **Single Ease Question (SEQ)** is a one-item, 7-point rating scale used immediately after a user attempts a task to measure how difficult or easy that task felt. It is the simplest, fastest, and most-validated post-task usability metric in modern UX research, and is the standard companion to behavioural measures like task completion rate and time on task.\n\nThe SEQ was popularised by **Jeff Sauro** and the team at MeasuringU after years of empirical comparison against other post-task questionnaires (After-Scenario Questionnaire, NASA-TLX, Subjective Mental Effort Question). Sauro’s research established that a single well-anchored question correlated *just as strongly* with task completion and time-on-task as longer multi-item scales — and was dramatically less work to administer. The result is a metric that has effectively become the default post-task measure across modern usability research.\n\n## The Verbatim SEQ Wording\n\n> **Overall, how difficult or easy was [the task] to complete?**\n>\n> 1 — Very Difficult\n> 2\n> 3\n> 4\n> 5\n> 6\n> 7 — Very Easy\n\nA few critical implementation details:\n\n- **The scale runs from 1 (Very Difficult) to 7 (Very Easy).** Reversing the polarity invalidates direct comparison to MeasuringU benchmarks.\n- **Only the endpoints are labelled.** Some teams label the midpoint or every point; both reduce sensitivity.\n- **It is administered *immediately after* the task**, not at the end of the session. The experience must be fresh.\n- **The bracketed task name should be specific.** Use the actual task wording the user just attempted (“purchasing a coffee subscription”), not a generic “the previous task.”\n\n## Why SEQ Works\n\nSEQ’s superpower is **predictive validity** — the score correlates strongly with what users actually did. Sauro’s benchmark research at MeasuringU established that:\n\n- A raw SEQ score of **5.9** corresponds to a task completion rate of roughly **86%** and an average task time of about 2 minutes.\n- A raw SEQ score of **4.7** corresponds to a completion rate of roughly **58%** and an average task time of about 2.8 minutes.\n- The relationship is roughly linear within the 4.0–6.5 range that covers most real-world tasks.\n\nThis is unusually strong for a self-reported metric. Most attitudinal measures correlate weakly with behaviour. SEQ correlates almost as well with task success as task success itself — which is why it survives across two decades of usability research.\n\n## SEQ Benchmarks\n\nAccording to MeasuringU’s published benchmark dataset of more than 400 tasks and 10,000+ users:\n\n| SEQ Score | Interpretation |\n|---|---|\n| 6.5+ | Top-decile task. Almost all users succeed without friction. |\n| 5.6–6.4 | Above average. Workable; minor friction. |\n| **5.3–5.5** | **Population average.** Typical for a competent but unremarkable task. |\n| 4.5–5.2 | Below average. Friction is real and worth investigating. |\n| <4.5 | Bottom-decile. Likely a usability emergency. |\n\nA crucial calibration: the 5.3–5.5 average sits *above* the nominal scale midpoint of 4. This is normal for 7-point scales — humans cluster toward the positive end of unlabeled scales. Treating 4 as “average” is the single most common SEQ misinterpretation.\n\n> **Industry benchmark.** “Across over 400 tasks and 10,000 users the average score hovers between about 5.3 and 5.6, which is above the nominal midpoint of 4 but is typical for 7-point scales.” — MeasuringU, *10 Things to Know About the Single Ease Question*\n\n## SEQ vs SUS: When to Use Each\n\nSEQ and SUS are not competing — they measure different things at different cadences.\n\n| Dimension | SEQ | SUS |\n|---|---|---|\n| Scope | One specific task | Entire product/system |\n| Timing | Immediately after each task | At the end of the test session |\n| Question count | 1 | 10 |\n| Scale | 1–7 | 1–5 (Likert) |\n| Output | Per-task ease score | 0–100 system score |\n| Best for | Diagnosing which task is hard | Benchmarking the whole product |\n| Sample-size floor | ~10 per task | ~8 per study |\n| Time to administer | <10 seconds | 60–90 seconds |\n\nThe canonical pattern in a moderated usability study is: SEQ after every task → SUS at the end. SEQ tells you *which* task is hard; SUS tells you whether the *product* is competitive against the 68 industry average. See the [SUS guide](/docs/system-usability-scale-guide) for the full Sauro–Lewis benchmark scale.\n\n## How to Run a SEQ Study — Step by Step\n\n### Step 1: Define your tasks\nWrite each task as a goal the user can attempt without coaching. “Find a coat under £100 and add it to your basket” is a task. “Browse the catalogue” is not.\n\n### Step 2: Pick a sample size\nMinimum 10–12 participants per task for reliability. For directional sprint testing, 8 is workable. For benchmarking or external reporting, aim for 30+. SEQ is unusually robust at small samples but never reliable below n=8.\n\n### Step 3: Run the task\nLet the user attempt the task end-to-end. Do not interrupt. If they ask for help, treat it as a failure and move on.\n\n### Step 4: Administer the SEQ immediately\nThe instant the task ends — succeeded or failed — show the SEQ. Do not allow time for rationalisation. The fresher the response, the more diagnostic the score.\n\n### Step 5: Always pair SEQ with an open-ended probe\nThis is the single most under-used best practice. A bare SEQ score tells you the task is hard; the open-ended “What made the task feel that way?” tells you *why*. Without the probe, SEQ is a thermometer with no diagnosis.\n\n### Step 6: Analyse per task and across tasks\nPer task: report the mean SEQ, the 95% confidence interval, and the % of users below 5. Across tasks: rank tasks by mean SEQ to identify the friction hotspots. Pair SEQ scores with task completion rates to triangulate.\n\n## Common SEQ Mistakes to Avoid\n\n1. **Reversing the scale.** Some teams label 1 as “easy” and 7 as “difficult.” This breaks every benchmark comparison. Stick to 1 = Very Difficult, 7 = Very Easy.\n2. **Treating 4 as the average.** The midpoint is statistically *not* the population average. The real average is 5.3–5.5. A score of 4 is well below average.\n3. **Administering SEQ at the end of the session.** Recall bias collapses the diagnostic value. Administer immediately after each task.\n4. **Reporting SEQ without an open-ended probe.** A score without a *why* is a metric you cannot act on.\n5. **Using SEQ to benchmark the whole product.** SEQ is a task metric. For a product-level benchmark, use [SUS](/docs/system-usability-scale-guide).\n6. **Stopping at n=5.** SEQ requires more participants than think-aloud sessions because it is quantitative. n=8 is a floor, n=10–12 is reliable, n=20+ is publishable.\n\n## The Modern Approach: SEQ at Scale With AI-Moderated Research\n\nSEQ has always been *easy to administer* but *expensive to run at scale*. The traditional bottleneck is everything around the SEQ: recruiting, scheduling, moderating, transcribing the probes, then thematically analysing the open-ended responses. A 5-task SEQ study with 15 participants is two weeks of work for a research team — and most of those weeks are not the SEQ itself.\n\nAI-native research platforms like **Koji** collapse this end-to-end. The modern SEQ workflow looks like this:\n\n1. **Build the study in minutes.** Use Koji’s [structured questions](/docs/structured-questions-guide) — specifically the *scale* type (1–7) — to add the SEQ after each task. Add an open-ended *probe* directly underneath. Use the *yes_no* question type for binary task success.\n2. **Launch via personalised link or in-product widget.** No scheduling, no moderator availability constraints. The AI moderator runs the task with users 24/7.\n3. **Get clean per-task data.** Koji’s ground-truth widget scores every scale answer at high confidence. Per-task SEQ averages, 95% confidence intervals, and distributions update on the report in real time.\n4. **Get the *why* automatically.** Koji’s thematic analysis engine clusters the open-ended probe responses into friction themes per task — eliminating the manual coding step that traditionally consumes the entire week after a study closes.\n5. **Compare across releases.** Re-run the same SEQ study after every release to track per-task ease over time, exactly as you would track SUS or NPS at the system level.\n\nForrester’s *State of Customer Insights 2024* found teams using AI-moderated research achieve **60% faster time-to-insight** than teams running equivalent studies manually. For SEQ studies specifically — where the bottleneck is rarely the metric itself but the moderation and analysis around it — the gap is closer to 80%. Koji customers routinely run 5-task SEQ studies in an afternoon that previously took a fortnight.\n\nThe broader point is that SEQ’s adoption has historically been limited not by the metric’s value (which is well-established) but by the operational cost of running enough sessions to make the score meaningful. Removing that operational cost is the actual research breakthrough — the metric itself has been settled science for two decades.\n\n## When NOT to Use SEQ\n\nSEQ is not the right tool for:\n\n- **System-level benchmarking** — use [SUS](/docs/system-usability-scale-guide) instead\n- **Loyalty or recommendation intent** — use [NPS](/docs/nps-survey-guide)\n- **Effort to *resolve a problem*** — use [Customer Effort Score (CES)](/docs/customer-effort-score-guide)\n- **Generative discovery** (“what should we build?”) — use [Mom Test interviews](/docs/mom-test-methodology) or [JTBD interviews](/docs/jobs-to-be-done-framework)\n\nSEQ shines for one job and one job only: **measuring the perceived ease of a specific task immediately after it is attempted.** Used inside its lane, it is the highest-leverage metric in the usability researcher’s toolkit.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six Koji question types, including the scale type used to deploy SEQ\n- [System Usability Scale (SUS): Complete Guide](/docs/system-usability-scale-guide) — the system-level companion to SEQ\n- [Customer Effort Score (CES): How to Measure and Reduce Friction](/docs/customer-effort-score-guide) — a related effort-based metric for support and resolution flows\n- [HEART Framework: Google’s 5-Metric UX Model](/docs/heart-framework-ux-metrics) — where SEQ slots in as the Task Success attitudinal signal\n- [Likert Scale Questions in User Research](/docs/likert-scale-research-guide) — broader scale-design principles relevant to SEQ\n- [Usability Testing: The Complete Guide](/docs/usability-testing-guide) — the parent methodology in which SEQ is administered\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [B2C User Research: How to Understand Consumer Behavior at Scale (2026)](/blog/b2c-user-research-guide-2026) — B2C user research is systematically underinvested at most consumer companies. While B2B teams run structured customer discovery as a matter \n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Single Ease Question (SEQ): 7-Point Scale, Benchmarks & 2026 Guide","metaDescription":"The definitive Single Ease Question (SEQ) guide: the verbatim 7-point scale, MeasuringU benchmarks, when to pair SEQ with SUS, sample-size rules, and how to deploy SEQ in minutes with AI-moderated research on Koji.","keywords":["single ease question","SEQ","SEQ usability metric","task-level usability","jeff sauro SEQ","MeasuringU SEQ","post-task survey","7-point scale usability","SEQ benchmark","SEQ vs SUS"],"aiSummary":"A definitive guide to the Single Ease Question (SEQ): the verbatim 7-point wording, MeasuringU/Sauro benchmark of 5.3–5.5, correlation with task completion (5.9 ≈ 86% completion), when to choose SEQ over SUS, sample-size guidance, and how Koji turns a 2-week SEQ study into an afternoon.","aiPrerequisites":["Familiarity with usability testing concepts","Basic understanding of Likert/scale rating","A defined user task to evaluate"],"aiLearningOutcomes":["Use the verbatim SEQ wording correctly in any usability test","Interpret a SEQ score against the MeasuringU 5.3–5.5 benchmark","Choose between SEQ and SUS for a given study","Pair SEQ with an open-ended probe for diagnostic insight","Run a multi-task SEQ study on Koji in a single afternoon"],"aiDifficulty":"beginner","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"c175fbbe-9727-4f3c-8def-1613f64701e1","slug":"unmoderated-usability-testing-guide","title":"Unmoderated Usability Testing: Moderated-Quality Insight at Scale","url":"https://www.koji.so/docs/unmoderated-usability-testing-guide","summary":"Guide to unmoderated usability testing: definition, unmoderated vs moderated comparison, pros and cons, when to use it, how to write good tasks, key metrics (task success rate, time on task, SEQ, SUS), and how AI moderation (like Koji) restores the missing \"why\" by probing hesitation and failure in the moment, delivering moderated-depth insight at unmoderated scale.","content":"# Unmoderated Usability Testing: Moderated-Quality Insight at Scale\n\n**Bottom line up front:** Unmoderated usability testing is a method where participants complete tasks on your product or prototype on their own — no researcher present — while their screen, clicks, and often their voice are recorded for later analysis. It's faster and cheaper than moderated testing and scales to dozens of participants overnight. Its one classic weakness: when a participant gets stuck or does something surprising, no one is there to ask \"why?\" AI-moderated platforms like Koji close that gap — the participant works unmoderated, but an AI moderator watches for hesitation and asks the follow-up questions a human researcher would. The result is moderated-depth \"why\" at unmoderated scale.\n\n## What is unmoderated usability testing?\n\nIn unmoderated usability testing, you define a set of tasks, recruit participants, and let them work through those tasks independently — usually remotely, on their own device, at a time that suits them. Software records their interactions (screen, clicks, taps, sometimes think-aloud audio), and you review the sessions afterward to find where people struggled, hesitated, or failed.\n\nThe defining trait is the absence of a live moderator. Nobody nudges the participant, answers their questions, or probes their reasoning in the moment. That's both the method's greatest strength (scale, speed, no scheduling) and its historical weakness (no one to ask why).\n\n## Unmoderated vs. moderated usability testing\n\n| Dimension | Unmoderated | Moderated |\n| --- | --- | --- |\n| Moderator present | No | Yes (live) |\n| Scale | High — dozens overnight | Low — one session at a time |\n| Cost per session | Low | High |\n| Scheduling | None; async | Coordinated calendars |\n| Depth of \"why\" | Traditionally shallow | Deep — live probing |\n| Best for | Benchmarking, task success, A/B of flows | Complex flows, novel concepts, edge cases |\n\nThe trade has always been depth *or* scale. AI moderation is what finally lets you have both — more on that below.\n\n## Pros and cons\n\n**Advantages**\n- **Speed.** Launch today, get results tomorrow. No calendars to align.\n- **Scale.** Test with 30–50 people for the effort of scheduling one moderated call.\n- **Lower cost.** No researcher-hours per session.\n- **Natural behavior.** Participants use their own device in their own environment, reducing the observer effect that can creep into a moderated call.\n\n**Limitations (and how AI addresses them)**\n- **The missing \"why.\"** A recording shows a participant abandon a form — but not that they left because the password rule wasn't shown until after they failed. *AI moderation asks in the moment.*\n- **No clarification.** If a task instruction is misread, a human moderator would catch it; unmoderated tests can waste a session. *An AI moderator can detect confusion and re-orient.*\n- **Shallow think-aloud.** Many participants go quiet when no one's listening. *An AI that responds keeps them talking.*\n\n## When to use unmoderated usability testing\n\nReach for unmoderated testing when you need **breadth and benchmarks**:\n\n- Measuring task success rate and time-on-task across a larger sample\n- Comparing two designs or flows (A/B) with enough participants to trust the difference\n- Validating that a well-understood flow works before launch\n- Testing at multiple points over time to track whether a redesign actually improved things\n\nChoose moderated testing instead when the flow is novel or complex, when you expect lots of unexpected behavior, or when the *reasoning* matters more than the *rate*. Better yet, use an AI-moderated approach that gives you scale and reasoning at once.\n\n## How to run an unmoderated usability test\n\n1. **Define objectives.** What decision will this test inform? Pick 3–5 concrete tasks tied to it.\n2. **Write realistic tasks.** Frame tasks as goals, not instructions (see below).\n3. **Recruit the right participants.** Screen for your actual target users — Koji's screener questions and in-product recruiting keep the sample clean.\n4. **Set success criteria up front.** Decide what \"success\" means per task (completed the goal, found the right page, etc.) before you watch a single session.\n5. **Launch and monitor.** With Koji, sessions stream in and analysis begins immediately — you're not waiting to batch-review a week later.\n6. **Analyze and synthesize.** Identify the top friction points by frequency and severity, backed by verbatim quotes.\n\n## Writing good usability tasks\n\nThe task is the experiment. Bad tasks produce useless data.\n\n- **Frame as a goal, not a click path.** Good: \"You want to change the email address on your account — go ahead.\" Bad: \"Click Settings, then Account, then Edit.\"\n- **Give context and a scenario.** \"Imagine you just moved and need to update your shipping address.\"\n- **Avoid your product's own vocabulary.** If the task uses the label on the button, you're testing reading, not findability.\n- **One goal per task.** Compound tasks blur where the friction actually happened.\n\n## Measuring results\n\nStandard unmoderated usability metrics include:\n\n- **Task success rate** — % who completed the goal\n- **Time on task** — how long completion took\n- **Error rate** — wrong turns, dead ends, mis-clicks\n- **Single Ease Question (SEQ)** — a 1–7 rating of task difficulty (a natural fit for Koji's **scale** question type)\n- **System Usability Scale (SUS)** — a standardized 10-item usability score\n\nKoji's six structured question types — **open_ended, scale, single_choice, multiple_choice, ranking, and yes_no** — let you collect SEQ and SUS scores as chartable data right alongside the open-ended \"walk me through what just happened\" reflections. Post-task, the AI can probe: \"You paused for a while on the payment screen — what were you thinking there?\" That single question is the difference between knowing *that* users struggled and knowing *why*.\n\n## How AI turns unmoderated testing into deep research\n\nThe reason unmoderated testing has always felt like the budget option is the missing \"why.\" Koji removes that ceiling. Participants complete tasks on their own schedule — fully unmoderated in terms of logistics — but an AI moderator conducts a think-aloud conversation throughout, notices hesitation and failure, and asks the exact follow-up a skilled researcher would. Every session is transcribed and auto-analyzed; themes are clustered across all participants; and you get a real-time report ranking friction points by how often they occurred and how severe they were. You keep the scale and cost of unmoderated testing while getting the depth that used to require a moderator on every call — 10x the coverage of traditional moderated research, without giving up the reasoning.\n\n## Common pitfalls to avoid\n\n- **Vague tasks.** If participants don't understand the goal, you measure confusion about the task, not your product. Pilot your tasks on one or two people first.\n- **Testing too much at once.** Five focused tasks beat fifteen rushed ones — fatigue degrades the later tasks.\n- **Ignoring the recording context.** People behave a little differently when they know they're recorded; keep tasks realistic and low-stakes to reduce it.\n- **Treating success rate as the whole story.** A task can \"succeed\" while the participant hated every second of it. Pair the metric with the reasoning — which is exactly what AI probing captures.\n- **Batch-reviewing weeks later.** Insights go stale and momentum dies. Koji analyzes sessions as they arrive, so you can act while the study is still running.\n\n## Unmoderated testing in a mixed-methods plan\n\nUnmoderated testing shines brightest as one instrument in a wider plan. Use it for breadth — benchmarks, A/B comparisons, quick pre-launch validation — and pair it with a smaller number of deep sessions for the truly novel or ambiguous flows. With AI moderation, that line blurs: a single Koji study can run at unmoderated scale while still probing like a moderated one, letting many teams collapse two rounds into one and reach a confident decision faster.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — collect SEQ, SUS, and open reflections in one session\n- [Usability Testing Guide](/docs/usability-testing-guide)\n- [Moderated Usability Testing Guide](/docs/moderated-usability-testing-guide)\n- [Unmoderated vs. Moderated Research](/docs/unmoderated-vs-moderated-research)\n- [Remote Usability Testing Guide](/docs/remote-usability-testing-guide)\n- [Think-Aloud Protocol](/docs/think-aloud-protocol)\n- [AI Usability Testing Guide](/docs/ai-usability-testing-guide)","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Unmoderated Usability Testing: Moderated-Quality Insight at Scale | Koji","metaDescription":"Unmoderated usability testing lets participants complete tasks on their own — fast and scalable, but it loses the \"why.\" Learn how to run it and how AI moderation restores the depth.","keywords":["unmoderated usability testing","unmoderated user testing","unmoderated vs moderated","usability testing tasks","remote usability testing","task success rate","SEQ","SUS score"],"aiSummary":"Guide to unmoderated usability testing: definition, unmoderated vs moderated comparison, pros and cons, when to use it, how to write good tasks, key metrics (task success rate, time on task, SEQ, SUS), and how AI moderation (like Koji) restores the missing \"why\" by probing hesitation and failure in the moment, delivering moderated-depth insight at unmoderated scale.","aiDifficulty":"intermediate","aiEstimatedTime":"9 min read"},{"type":"documentation","id":"268df3eb-2fd5-429e-8d19-c3be250149bb","slug":"usability-metrics-guide","title":"Usability Metrics: Task Success Rate, Time on Task, and Error Rate Explained","url":"https://www.koji.so/docs/usability-metrics-guide","summary":"A definitive reference on the three core usability metrics — task success rate (effectiveness), time on task (efficiency), and error rate (accuracy) — with formulas, the 78% MeasuringU completion benchmark, sample-size guidance, and how Koji captures all three automatically via structured questions and an AI moderator.","content":"## What are the core usability metrics?\n\n**The three core usability metrics are task success rate (effectiveness), time on task (efficiency), and error rate (accuracy).** Together they answer the only questions that matter in a usability test: Can users finish the job? How long does it take them? And how many mistakes do they make along the way? Every other quantitative usability measure — completion confidence, lostness, task-level satisfaction — is a refinement of these three.\n\nThe Nielsen Norman Group calls success rate \"the simplest usability metric\" precisely because it is the bottom line of usability: if users cannot complete what they came to do, nothing else about the interface matters. The strength of these three metrics is that they are objective and behavioral. Unlike attitudinal scores such as NPS or [SUS](/docs/system-usability-scale-guide), they record what users *actually did*, not what they later said they felt.\n\nThis guide defines each metric, gives you the formulas and the published industry benchmarks, explains how many participants you need, and shows how an AI-native platform like Koji captures all three automatically — turning a multi-day analysis grind into a real-time dashboard.\n\n## Metric 1: Task success rate (effectiveness)\n\nTask success rate is the percentage of participants who complete a task successfully out of everyone who attempted it.\n\n> **Task success rate = (number of successful attempts ÷ total attempts) × 100**\n\nIf 17 of 20 participants successfully add an item to their cart and reach checkout, your success rate is 85%.\n\n**The benchmark:** In an analysis of 1,189 tasks across 115 usability studies, MeasuringU founder Jeff Sauro found the average task completion rate is **78%**. Most teams treat roughly 78–80% as the dividing line between \"acceptable\" and \"needs work\" for an important task — though success rate is highly sensitive to task difficulty, so the right target is always relative to the task and to your own historical baseline.\n\n**Binary vs. levels of success.** The cleanest version is binary: a participant either completed the task or did not. But many teams record *levels of success* — full success, partial success (completed with significant struggle or workaround), and failure — because a binary view hides the difference between a user who breezed through and one who barely limped to the finish line. Partial successes are often where your richest design insights hide.\n\n**The trap:** success rate alone is misleading. A task can show a 90% success rate while users take three minutes and make two errors getting there. That is why success rate must always be read alongside time and errors.\n\n## Metric 2: Time on task (efficiency)\n\nTime on task measures how long it takes a participant to complete a task, usually reported as the mean or median time in seconds for successful attempts only. (Including failed attempts pollutes the number — a user who gave up after 10 seconds would otherwise look \"efficient.\")\n\nBecause time data is almost always skewed by a few very slow users, the **geometric mean or the median** is the statistically appropriate measure of center for small samples, not the arithmetic mean. Report a measure of spread too — the range or confidence interval — because an average of 45 seconds means something very different if the spread is 40–50 seconds versus 10–120 seconds.\n\n**How to use it:** time on task is most powerful as a *comparative* metric — old design vs. new, your product vs. a competitor, or release over release. An absolute \"good\" time rarely exists in isolation; a 30-second task time is excellent for a complex configuration flow and terrible for a one-click action.\n\n## Metric 3: Error rate (accuracy)\n\nAn error is any unintended action, slip, mistake, or omission a user makes while attempting a task. Error rate is typically expressed as errors per task (the average number of errors across all attempts) or as a defect rate (the percentage of attempts containing at least one error).\n\n> **Errors per task = total errors observed ÷ total attempts**\n\n**The benchmark:** across an analysis of 719 tasks using consumer and business software, Jeff Sauro found an average of **0.7 errors per task**, with roughly **two out of every three users making at least one error**. Errors are far more common than most teams assume — which is exactly why counting them surfaces friction that success rate alone would never reveal.\n\nNot all errors are equal. Classify them by severity (does the error block completion, or merely slow the user down?) and by type (slips, where the user knows the goal but executes the wrong action, vs. mistakes, where the user has the wrong mental model). The pattern in *where* errors cluster is usually more actionable than the raw count.\n\n## Putting the three together\n\nEffectiveness, efficiency, and accuracy form a triangle. A mature usability scorecard reads all three at once:\n\n| Metric | What it measures | Typical benchmark |\n| --- | --- | --- |\n| Task success rate | Effectiveness — can they finish? | ~78% average (Sauro/MeasuringU) |\n| Time on task | Efficiency — how fast? | Comparative; no universal target |\n| Error rate | Accuracy — how clean? | ~0.7 errors/task; ~2 of 3 users err |\n\nLayer a task-level satisfaction question on top — a single [scale question](/docs/structured-questions-guide) such as \"How easy or difficult was that task?\" — and you capture the user's attitude alongside their behavior. The Nielsen Norman Group repeatedly finds that performance and satisfaction metrics correlate only moderately, so measuring both protects you from shipping something users *can* use but *hate* using.\n\n## How many participants do you need?\n\nFor **qualitative, formative** usability testing — finding problems to fix — five users per round uncovers roughly 85% of issues, the classic Nielsen-Landauer finding. But the moment you want *reliable quantitative metrics* like a stable success rate or time on task, five is far too few: the confidence interval on a metric from five users is enormous.\n\nA practical rule of thumb:\n\n- **5–8 users** — formative testing, finding usability problems (not for reporting precise numbers).\n- **15–20 users** — a reasonably tight success-rate estimate for a single design.\n- **30–50+ users** — benchmark-grade metrics you intend to track over time or quote externally. (See our [usability benchmarking guide](/docs/usability-benchmarking-guide) for the full methodology.)\n\nUse adjusted-Wald binomial confidence intervals for small-sample completion rates rather than naive percentages — a 4-of-5 success \"80%\" actually carries a confidence interval running from roughly 36% to 98%.\n\n## The modern approach: capturing usability metrics with AI\n\nTraditionally, collecting these three metrics meant scheduling moderated sessions, watching every recording, manually timing each task with a stopwatch, tallying errors by hand, and reconciling notes across a research team. A 20-participant benchmark could swallow a week of analyst time — which is why most teams measured usability once a quarter at best, if at all.\n\nThis is exactly the bottleneck AI-native research platforms remove. **Koji** captures all three core metrics automatically:\n\n- **Task success rate** is recorded directly through structured questions. Frame each task with a `yes_no` or `single_choice` outcome question, and Koji aggregates the completion rate across every respondent in real time — no manual tallying.\n- **Time on task** is timestamped automatically for every session, with the distribution (median, range, outliers) computed and charted as responses arrive.\n- **Error and friction signals** surface through Koji's AI moderator, which probes in the moment (\"What made that step confusing?\") and then clusters the open-ended answers into themed friction findings, so you see *where* and *why* users struggle, not just *that* they did.\n\nKoji supports all six [structured question types](/docs/structured-questions-guide) — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — which means you can capture a binary success flag, a 1–5 ease rating, and a rich open-ended \"what went wrong\" in a single automated study. Because Koji runs 24/7, you can recruit 30–50 participants for a true quantitative benchmark in days rather than weeks, and re-run the identical study every release to track the trend line. Teams using AI-assisted research tools consistently report dramatically faster time-to-insight precisely because the counting, timing, and tagging — the slow part — is done the instant the last response lands.\n\nYou do not need a PhD in measurement theory to run a rigorous usability study. Define the tasks, attach the right structured questions, and let the platform handle the statistics.\n\n## Common mistakes to avoid\n\n1. **Reporting time on task for failed attempts.** Always separate successful and unsuccessful times.\n2. **Using the arithmetic mean on small samples.** Time data is skewed — use the median or geometric mean.\n3. **Quoting a success rate without a confidence interval.** \"80% from five users\" is not a precise number.\n4. **Measuring success but never satisfaction.** A usable-but-frustrating product still loses users.\n5. **Changing the task wording between benchmark rounds.** Consistency is what makes release-over-release comparison valid.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types for capturing success, ease, and friction\n- [Usability Testing: The Complete Guide](/docs/usability-testing-guide) — the end-to-end method these metrics live inside\n- [Usability Benchmarking Guide](/docs/usability-benchmarking-guide) — turning these metrics into a tracked program\n- [System Usability Scale (SUS) Guide](/docs/system-usability-scale-guide) — the standard attitudinal usability score\n- [Customer Effort Score Guide](/docs/customer-effort-score-guide) — measuring perceived ease at the task level\n- [Think-Aloud Protocol](/docs/think-aloud-protocol) — surfacing the *why* behind every error","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Usability Metrics: Task Success Rate, Time on Task & Error Rate (2026)","metaDescription":"Master the three core usability metrics — task success rate, time on task, and error rate. Formulas, industry benchmarks (78% avg completion), sample sizes, and how to capture them automatically with AI.","keywords":["usability metrics","task success rate","time on task","error rate","usability testing metrics","task completion rate","usability benchmark","UX metrics","effectiveness efficiency","quantitative usability"],"aiSummary":"A definitive reference on the three core usability metrics — task success rate (effectiveness), time on task (efficiency), and error rate (accuracy) — with formulas, the 78% MeasuringU completion benchmark, sample-size guidance, and how Koji captures all three automatically via structured questions and an AI moderator.","aiPrerequisites":["Familiarity with usability testing concepts","A product or prototype to evaluate","Basic understanding of percentages and averages"],"aiLearningOutcomes":["Define and calculate task success rate, time on task, and error rate","Interpret each metric against published industry benchmarks","Choose the right sample size for qualitative vs quantitative usability studies","Avoid the five most common usability-metric mistakes","Capture all three metrics automatically in Koji using structured questions"],"aiDifficulty":"beginner","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"f0e6165a-53d4-4833-ae31-a0507aaaaf6b","slug":"rite-method-rapid-iterative-testing","title":"The RITE Method: Rapid Iterative Testing and Evaluation","url":"https://www.koji.so/docs/rite-method-rapid-iterative-testing","summary":"The RITE method (Rapid Iterative Testing and Evaluation) is a usability approach where you fix problems between participants instead of waiting for a final report — converging on a working design in days. Formalized at Microsoft Game Studios, its bottleneck is synthesis. AI-moderated interviews compress that step with instant transcription, automatic theme analysis, and structured severity signals (SEQ scales, success checks), letting teams run several test-fix-retest loops per day.","content":"## What Is the RITE Method? (BLUF)\n\nThe **RITE method — Rapid Iterative Testing and Evaluation** — is a usability approach where you **fix problems as you find them, between participants, instead of waiting until the study ends**. Identify a serious issue with participant 2, change the prototype that afternoon, and test the fix with participant 3. You converge on a working design in days, not the weeks a traditional \"test 8 users, write a report, then redesign\" cycle takes.\n\nRITE was formalized by Michael Medlock and colleagues at Microsoft Game Studios, where ship dates are immovable and a single broken tutorial can sink a game. Its core insight: the goal of early testing isn''t to *document* every problem for a report — it''s to *eliminate* problems as fast as possible. The slowest part of any iterative cycle is turning raw sessions into clear, agreed-upon findings. That''s exactly where AI-moderated interviews compress the loop: with [automatic transcription](/docs/ai-transcription-research-interviews) and [instant theme analysis](/docs/understanding-themes-patterns), the insight is ready minutes after each session — so the \"decide what to change\" step keeps pace with the \"test it\" step.\n\n---\n\n## How RITE Differs from Traditional Usability Testing\n\n| | Traditional usability test | RITE method |\n|---|---|---|\n| **When you change the design** | After all sessions | Between sessions |\n| **Primary goal** | Document findings | Fix problems fast |\n| **Cadence** | Test → report → redesign (weeks) | Test → fix → re-test (hours/days) |\n| **Sample per design version** | Fixed (e.g., 5–8) | Variable — as many as needed to confirm the fix |\n| **Best for** | Benchmarking, summative evaluation | Early, formative design refinement |\n\nIn a classic study you''d run all participants against the *same* build, then synthesize. In RITE, the build evolves mid-study. That means a problem found early gets multiple chances to be fixed and re-validated, while a problem found late might still be caught before launch.\n\n---\n\n## The RITE Cycle, Step by Step\n\n### 1. Assemble a decision-making team\nRITE only works if the people who can *change the design* are in the loop. Before you start, gather the designer, a developer who can implement quick changes, and the PM. They review findings together and commit to changes on the spot. This shared-context step is what makes immediate iteration possible.\n\n### 2. Define tasks and success criteria\nWrite the [task scenarios](/docs/task-analysis-ux-research) participants will attempt and decide in advance what counts as a failure worth fixing. A [clear research question](/docs/writing-a-research-question) keeps the team from chasing cosmetic nitpicks.\n\n### 3. Run a session\nHave the participant attempt the tasks while [thinking aloud](/docs/think-aloud-protocol). Capture where they hesitate, fail, or misunderstand.\n\n### 4. Classify each issue\nSort problems into:\n- **Obvious fixes with an obvious solution** → change immediately.\n- **Issues you can see but aren''t sure how to fix** → discuss; change if confident.\n- **Issues needing more data** → keep testing before acting.\n\n### 5. Change the design\nImplement the agreed fixes before the next participant. Even a clickable-prototype tweak counts.\n\n### 6. Re-test and repeat\nThe next participant validates the fix and surfaces the next layer of problems. Continue until sessions stop revealing serious new issues — you''ve reached [data saturation](/docs/data-saturation-qualitative-research) on the current design.\n\n---\n\n## Where RITE Slows Down — and How AI Fixes It\n\nThe RITE loop is only as fast as its slowest link. In practice that link is **synthesis**: after each session someone has to review what happened, articulate the problem clearly enough for the team to agree, and decide on a change. With back-to-back participants, notes pile up and the team debates from fuzzy memory.\n\nAn AI-native workflow tightens every link:\n\n- **Instant, structured capture.** Run the session as an [AI-moderated interview](/docs/how-ai-interviewers-work) by [voice or text](/docs/voice-vs-text-interviews). The AI [probes follow-ups automatically](/docs/probing-and-follow-up-questions) — *\"You paused on that screen — what were you expecting to happen?\"* — so you don''t lose the reasoning behind a failure.\n- **Analysis ready before the next session.** Each conversation is [transcribed](/docs/viewing-interview-transcripts) and [analyzed into themes](/docs/thematic-analysis-guide) immediately, giving the team a clear, citable problem statement instead of a hand-scrawled note.\n- **Structured severity signals.** Add a post-task [scale question](/docs/scale-questions-guide) (\"How easy was that task, 1–7?\" — a [Single Ease Question](/docs/single-ease-question-seq-guide)) and a [yes/no](/docs/yes-no-questions-guide) success check. Koji''s [six structured question types](/docs/structured-questions-guide) turn each session into comparable data, so you can see at a glance whether a fix actually moved the number.\n- **Parallel pre-screening.** Because AI interviews are unmoderated, you can pre-run a wave, read the analysis, fix, then release the next wave — getting RITE''s benefits without scheduling every session live.\n\nThe result: the \"test → understand → decide → change\" loop that traditionally takes a day per turn can run several turns in a day.\n\n---\n\n## When to Use RITE (and When Not To)\n\n**Use RITE when:**\n- You''re in early, formative design and the prototype can change quickly.\n- Problems are likely to be frequent and fixable (onboarding, navigation, new flows).\n- You have a team empowered to make changes mid-study.\n\n**Avoid RITE when:**\n- You need a stable benchmark or [summative evaluation](/docs/formative-vs-summative-research) — changing the design mid-study breaks comparability.\n- Fixes require deep engineering that can''t happen between sessions.\n- You''re measuring against a competitor or a baseline metric that must stay constant.\n\nFor benchmarking, run a traditional [usability test](/docs/usability-testing-guide) instead; for idea-stage validation, reach for [concept testing](/docs/concept-testing-methodology).\n\n---\n\n## A Realistic RITE Schedule\n\nA tight RITE study might look like: 3 sessions in the morning, team synthesis over lunch using the AI-generated analysis, fixes implemented in the early afternoon, 3 more sessions late afternoon against the updated build. Two iterations in a single day — with the evidence trail to justify each change — is entirely achievable when synthesis isn''t the bottleneck.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — add SEQ scales and success checks to every RITE session\n- [Usability Testing Guide](/docs/usability-testing-guide) — the foundation RITE builds on\n- [Think-Aloud Protocol](/docs/think-aloud-protocol) — how to capture reasoning during tasks\n- [Formative vs. Summative Research](/docs/formative-vs-summative-research) — where RITE fits\n- [Single Ease Question (SEQ)](/docs/single-ease-question-seq-guide) — the fastest per-task severity signal\n- [Data Saturation in Qualitative Research](/docs/data-saturation-qualitative-research) — knowing when to stop iterating","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"RITE Method: Rapid Iterative Testing & Evaluation Guide | Koji","metaDescription":"The RITE method fixes usability problems between participants instead of after the study. Learn the full Rapid Iterative Testing and Evaluation cycle and how AI interviews make each loop faster.","keywords":["rite method","rapid iterative testing and evaluation","rite usability testing","iterative usability testing","fix usability problems between participants","rite method ux"],"aiSummary":"The RITE method (Rapid Iterative Testing and Evaluation) is a usability approach where you fix problems between participants instead of waiting for a final report — converging on a working design in days. Formalized at Microsoft Game Studios, its bottleneck is synthesis. AI-moderated interviews compress that step with instant transcription, automatic theme analysis, and structured severity signals (SEQ scales, success checks), letting teams run several test-fix-retest loops per day.","aiPrerequisites":["Familiarity with usability testing basics"],"aiLearningOutcomes":["Explain how RITE differs from traditional usability testing","Run the full RITE test-fix-retest cycle","Identify where the loop slows down and how to speed it up","Know when to use RITE vs. a summative usability test"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 minutes"},{"type":"documentation","id":"1890dff6-222c-4112-800d-bd838c445d1e","slug":"agent-trajectory-evaluation","title":"Agent Trajectory Evaluation: How to Judge Multi-Step AI Agents with Real Users (2026)","url":"https://www.koji.so/docs/agent-trajectory-evaluation","summary":"Agent trajectory evaluation scores the entire sequence an AI agent takes - its reasoning steps, tool calls, recoveries, and hand-backs to the user - rather than only the final answer. It matters because agents that reach a correct outcome by an unsafe or unrepeatable path will fail in production. This guide covers the four failure classes (wrong plan, wrong tool call, silent recovery failure, and bad user hand-back), how to build a step-level rubric, how to combine automated trajectory checks with human participant evidence, and how to run trajectory studies in Koji using structured questions.","content":"# Agent Trajectory Evaluation: How to Judge Multi-Step AI Agents with Real Users (2026)\n\n**Short answer:** Agent trajectory evaluation scores the entire path an AI agent takes to a result - every reasoning step, tool call, error recovery, and message back to the user - instead of only grading the final answer. It matters because a multi-step agent can land on the correct outcome through an unsafe, expensive, or unrepeatable route, and outcome-only scoring is blind to all three. The strongest programmes pair automated step-level checks with human participants who can say what the agent actually did to them.\n\nSingle-turn AI evaluation is a solved shape: you have an input, an output, and a rubric. Agents broke that shape. An agent booking a return, reconciling an invoice, or refactoring a file takes ten, thirty, or a hundred steps, and the interesting failures live in the middle of that sequence - not at the end.\n\nThis guide covers how to evaluate trajectories properly: the failure classes to score against, how to build a step-level rubric, where automated judges stop working, and how to bring real users into the loop.\n\n## Why outcome scoring quietly fails\n\nThe most cited evidence for this comes from tau-bench, the tool-agent-user benchmark from Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan at Sierra. Two findings from that work reframed how teams evaluate agents.\n\nFirst, on realistic multi-step retail and airline tasks, even state-of-the-art function-calling agents \"succeed on <50% of the tasks.\" Second, and more damning, the authors report that **pass^8 - the probability an agent solves the same task correctly on all eight of eight attempts - is under 25% in the retail domain.**\n\nThat gap between single-run pass rate and eight-run consistency is the whole argument for trajectory evaluation. An agent with a 50% pass rate and a 25% pass^8 is not half-competent; it is unreliable in a way that a single-run outcome metric structurally cannot show you. Your users experience pass^k, not pass^1, because they run the same task repeatedly.\n\nThe successor benchmark, tau-squared-bench, added a further complication: a dual-control environment where **both** the agent and the user act on shared state. Performance drops significantly when agents move from acting alone to guiding a user, and the authors separate reasoning errors from communication and coordination failures. That distinction matters enormously for product teams, because a coordination failure is invisible to any evaluation that does not include a human.\n\n## The four classes of trajectory failure\n\nScore against these four classes explicitly. Most teams only instrument the second.\n\n| Failure class | What it looks like | Detectable by automation? |\n|---|---|---|\n| **Wrong plan** | The agent decomposes the task incorrectly from the start - right tools, wrong sequence, wrong sub-goal | Partially - only if you have a reference plan |\n| **Wrong tool call** | Malformed arguments, wrong tool selected, required tool skipped, forbidden action taken | Yes - this is the easy case |\n| **Silent recovery failure** | The agent hits an error, retries, appears to recover, but has quietly lost or corrupted state | Rarely - the final state often looks plausible |\n| **Bad user hand-back** | The agent asks the wrong clarifying question, over-claims what it did, or hands back at the wrong moment | No - requires a human who was there |\n\nThe fourth class is where most production complaints originate and where almost no evaluation budget goes. An agent that completed the task but told the user it had done something slightly different has failed, and no state-diff check will ever catch it.\n\n## Process reward vs outcome reward\n\nTwo scoring philosophies:\n\n- **Outcome reward** assigns one score to the end state. Cheap, objective, and unable to tell you *why* a run failed.\n- **Process reward** assigns a score to each step. Expensive, harder to define, and the only way to localise a failure to the step that caused it.\n\nThe practical answer is both: outcome scoring as your regression gate, process scoring on the subset of runs that fail, plus a rotating sample of runs that pass (to catch right-answer-wrong-path cases).\n\nStep labelling is genuinely hard even for humans. AgentProcessBench, a 2026 benchmark for step-level process quality in tool-using agents, was built from 1,000 trajectories and 8,509 human-labelled step annotations, reaching 89.1% inter-annotator agreement. Its headline finding is a warning for anyone planning to automate this: current models struggle to distinguish **neutral** actions from **erroneous** ones, and weaker agents show inflated ratios of correct steps simply because they terminate early. An agent that gives up quickly can look clean on a naive step-accuracy metric.\n\n## Building a trajectory rubric\n\nA workable rubric scores each step on three axes and the run as a whole on two.\n\n**Per-step:**\n\n| Axis | Scale | Definition |\n|---|---|---|\n| Necessity | Required / Helpful / Neutral / Wasteful | Did this step advance the goal? |\n| Correctness | Correct / Recoverable error / Unrecoverable error | Was the action itself right? |\n| Safety | Safe / Reversible risk / Irreversible risk | What happens if this step is wrong? |\n\n**Per-run:**\n\n| Axis | Scale | Definition |\n|---|---|---|\n| Goal satisfaction | Pass / Partial / Fail | Did the user get what they needed? |\n| Policy adherence | Pass / Fail | Was any rule, permission, or constraint violated? |\n\nScore goal satisfaction and policy adherence as a **conjunction** - a run counts as passing only if both pass. This is the design tau-bench uses, and it prevents the common failure of shipping an agent that is helpful and non-compliant.\n\nAdd severity to every identified failure. Without severity, a rubric produces a flat list of 200 issues and no prioritisation. Use a four-level scale: cosmetic, degraded, blocking, harmful.\n\n## Where humans are non-negotiable\n\nAutomated trajectory checks cover deterministic questions. Human evaluation covers four things they cannot:\n\n1. **Did the user understand what the agent was doing?** Legibility is a property of the trajectory, not the outcome.\n2. **Was the hand-back at the right moment?** Agents ask too early (annoying) or too late (dangerous). Only the person on the other end can say which.\n3. **Would the user trust it again?** Trust after a recovered failure is a different measurement from trust after a clean run.\n4. **What did the user do next?** Silent workarounds are the strongest signal that a trajectory is broken, and they never appear in your logs as errors.\n\nThe evidence that this gap is real and widening: the 2025 Stack Overflow Developer Survey found **66% of developers cite AI solutions that are \"almost right, but not quite\" as their top frustration, and 45% say debugging AI-generated code takes longer** than writing it themselves. Those are trajectory complaints, not outcome complaints. A near-miss passes most outcome checks and costs the user more time than no help at all.\n\n## How to run trajectory evaluation with Koji\n\nTraditional approach: recruit participants, schedule moderated sessions, watch each one attempt the agent task, transcribe, tag steps by hand, and synthesise. Fifteen sessions is roughly two weeks of researcher time and produces a sample too small to compare builds.\n\nThe AI-native approach compresses that to days:\n\n**1. Run the task, then interview immediately.** Participants complete the real agent task. Koji picks up the moment they finish, while the trajectory is still fresh, and an AI moderator interviews each participant about the specific step that went wrong. The moderator asks genuine follow-ups - *\"you said it looked like it had booked it; what made you think that?\"* - which is exactly the probe a static post-task survey cannot produce.\n\n**2. Capture comparable step-level scores with structured questions.** Koji supports six structured question types, and trajectory evaluation uses nearly all of them:\n\n| Question type | Trajectory use |\n|---|---|\n| `scale` | Rate legibility and trust per run (1-5) |\n| `ranking` | Order the steps by how confusing they were |\n| `single_choice` | Which step first went wrong? |\n| `multiple_choice` | Which failure classes did you observe? |\n| `yes_no` | Did the agent do anything you did not ask for? |\n| `open_ended` | Describe what you expected to happen instead |\n\nThe structured types give you a comparable numeric series across builds. The open-ended responses give you the reasoning behind the numbers. See the [structured questions guide](/docs/structured-questions-guide) for how to combine them in one study.\n\n**3. Cluster automatically.** Koji thematic analysis groups hundreds of open-ended trajectory complaints into named failure modes without manual tagging, and each theme stays linked to the verbatim quotes and the participants who raised it - so an engineer can trace a theme back to a specific run.\n\n**4. Re-run as a regression gate.** Because the study is a reusable artifact rather than a scheduling exercise, you can run the identical trajectory evaluation against every candidate build. That is what turns human evaluation from a one-off exercise into a gate, and it is the practical reason AI-moderated research changes agent development economics: 40 participant interviews in 48 hours is a different product decision from 12 interviews in three weeks.\n\nYou do not need a PhD in evaluation methodology to do this well. You need a fixed task set, a rubric with severity, and a repeatable way to talk to the people who ran it.\n\n## Sampling and cadence\n\n| Purpose | Trajectories | Cadence |\n|---|---|---|\n| Exploratory - find unknown failure modes | 30-50 per task family | Once per major capability change |\n| Regression gate | 100-500 fixed golden set | Every release candidate |\n| Post-incident deep dive | 10-20 targeted | On demand |\n| Longitudinal trust tracking | 40+ participants | Monthly |\n\nFor the regression gate, build the set the same way you would build any evaluation dataset - production-representative, edge, adversarial, and replayed past failures - and hold it stable so that movement in the pass rate means something. The [golden set guide](/docs/ai-evaluation-dataset-golden-set) covers the sizing maths and the confidence intervals in detail.\n\n## Five mistakes to avoid\n\n1. **Reporting pass^1 only.** Report pass^k as well. Consistency is the metric your users feel.\n2. **Letting the judge see the outcome before scoring steps.** Outcome knowledge contaminates step-level judgement. Score steps blind.\n3. **Treating early termination as clean.** An agent that quits fast scores well on step accuracy and terribly on usefulness. Always pair process scores with goal satisfaction.\n4. **Skipping the user hand-back class.** It is the largest source of production complaints and the easiest to omit because logs do not capture it.\n5. **Changing the golden set and the model in the same week.** You will not be able to attribute the movement to either.\n\n## Where this is heading\n\nRegulatory pressure is arriving on the same axis. The EU AI Act obliges providers of high-risk systems to report serious incidents, which requires knowing not just that a system failed but *how* it failed - a trajectory question, not an outcome question. Teams that already keep step-level evidence with participant testimony attached will find that reporting burden manageable. Teams with only pass/fail logs will not.\n\nAgent evaluation is converging on the same conclusion user research reached decades ago: the artifact tells you what happened, and only the person tells you what it meant.\n\n## Frequently asked questions\n\n**What is an agent trajectory?**\nThe full ordered sequence an agent produces while completing a task: intermediate reasoning, every tool call, the results returned, recovery attempts, and every message sent to the user. Outcome evaluation looks only at the final state; trajectory evaluation looks at the whole path.\n\n**Why is outcome-only evaluation not enough?**\nAgents can reach the right answer by a wrong path - for example issuing a destructive write and then undoing it. Outcome scoring also cannot distinguish a competent run from a lucky one, which is why consistency metrics like pass^8 collapse even when single-run pass rates look acceptable.\n\n**What is the difference between process reward and outcome reward?**\nOutcome reward scores the end state only. Process reward scores each intermediate step, localising the failure to the step that caused it. Process scoring gives far more debugging signal but costs more, because every step needs a label.\n\n**Can an LLM judge evaluate trajectories on its own?**\nOnly partially. Automated judges handle deterministic checks well - required tool called, forbidden action taken, final state matched. They are weak at separating neutral steps from erroneous ones, and they cannot tell you whether the user felt informed or misled.\n\n**How many trajectories do I need?**\nRoughly 30-50 per critical task family for a directional read on a release; 100-500 in a fixed golden set for a stable regression gate.\n\n**How does Koji help?**\nKoji runs the human half at scale: participants complete the real task, an AI moderator interviews each one about the exact moment it went wrong, structured questions capture comparable step-level scores, and thematic analysis clusters hundreds of trajectory complaints into named failure modes automatically.\n\n## Related Resources\n\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) - the foundational guide to scoring AI quality with people\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) - when automated scoring is trustworthy and when it is not\n- [Evaluation Datasets and Golden Sets](/docs/ai-evaluation-dataset-golden-set) - how to build the fixed set your regression gate depends on\n- [AI Red Teaming with Real Users](/docs/ai-red-teaming-with-users) - finding harms before your users do\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types and when to use each\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) - what AI-moderated research actually requires\n- [AI Governance for Customer Research](/docs/ai-governance-frameworks-research) - ISO 42001, the NIST AI RMF, and procurement\n\n---\n\n**Sources:** Yao, Shinn, Razavi & Narasimhan, *tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains* (arXiv:2406.12045); *tau-squared-Bench: Evaluating Conversational Agents in a Dual-Control Environment* (arXiv:2506.07982); Fan et al., *AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents* (arXiv:2603.14465); 2025 Stack Overflow Developer Survey; EU AI Act Article 73.","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Agent Trajectory Evaluation: Judging Multi-Step AI Agents (2026)","metaDescription":"A practical guide to evaluating AI agent trajectories: step-level rubrics, process vs outcome scoring, the four failure classes, and how to bring real users into agent evaluation with Koji.","keywords":["agent trajectory evaluation","ai agent evaluation","multi-step agent testing","process reward","llm agent evals","tool call evaluation","agentic ai testing","tau-bench"],"aiSummary":"Agent trajectory evaluation scores the entire sequence an AI agent takes - its reasoning steps, tool calls, recoveries, and hand-backs to the user - rather than only the final answer. It matters because agents that reach a correct outcome by an unsafe or unrepeatable path will fail in production. This guide covers the four failure classes (wrong plan, wrong tool call, silent recovery failure, and bad user hand-back), how to build a step-level rubric, how to combine automated trajectory checks with human participant evidence, and how to run trajectory studies in Koji using structured questions.","aiPrerequisites":["Basic familiarity with AI agents and tool calling","Understanding of qualitative and quantitative research methods"],"aiLearningOutcomes":["Explain why outcome-only agent scoring is insufficient","Identify the four classes of trajectory failure","Build a step-level trajectory rubric with defined severity levels","Design a human evaluation study that captures the user side of an agent trajectory","Run a trajectory evaluation using Koji structured questions"],"aiDifficulty":"advanced","aiEstimatedTime":"14 min"},{"type":"documentation","id":"4e2f432a-2762-4ef0-8c28-a021500a4b34","slug":"usability-testing-guide","title":"How to Conduct Usability Testing: The Complete Guide","url":"https://www.koji.so/docs/usability-testing-guide","summary":"Usability testing is a research method where representative users complete realistic tasks with a product while a researcher observes where they struggle or succeed. This guide covers the four main types of usability tests, the 5-users rule from Nielsen and Landauer, a 9-step facilitation process, and the six most common mistakes that invalidate results — including moderator bias and tasks that give away the answer.","content":"\nUsability testing is a research method in which representative users complete realistic tasks with a product or prototype while a researcher observes where they struggle, succeed, or get confused. It is the most direct way to discover whether your product works the way real people expect it to — before those problems reach production.\n\nAccording to Forrester Research, **every $1 invested in UX yields returns of up to $100**. And 88% of users will not return to a product after a poor experience. Usability testing is how you prevent that from happening.\n\n## What Is Usability Testing?\n\nUsability testing is not a survey, a focus group, or an analytics deep-dive. It is direct behavioral observation: you watch real people use your product, in real time, with real tasks.\n\nJakob Nielsen, co-founder of Nielsen Norman Group, defines usability across five measurable dimensions: **learnability** (how easily new users accomplish tasks), **efficiency** (speed once the system is learned), **memorability** (how quickly proficiency returns after absence), **errors** (frequency and severity of mistakes), and **satisfaction** (subjective pleasantness of use). All five can be observed and measured through structured testing.\n\nThe goal is not to prove your design works. It is to discover how and why it fails — while there is still time to fix it.\n\n## Why Usability Testing Matters: The Business Case\n\nThe numbers make the argument:\n\n- **Every $1 invested in UX returns up to $100**, according to Forrester Research — a potential 9,900% ROI from improved conversion rates, reduced support costs, and lower development rework.\n- **5 users uncover approximately 85% of usability problems.** Nielsen and Landauer's 1993 mathematical model shows that five qualitative test participants will surface the vast majority of discoverable issues in a product.\n- **88% of users will not return after a bad experience**, and **61% abandon websites due to unclear navigation**, according to Nielsen Norman Group research.\n- **Conversion rates can increase up to 400%** with improved UX design. Baymard Institute research shows a single checkout flow improvement can boost e-commerce conversions by 35%.\n- Fixing problems during the design phase costs dramatically less than fixing them post-launch — the directional principle is well-supported across software quality literature, even if precise multipliers should be treated as estimates rather than exact figures.\n\n## Types of Usability Testing\n\nChoosing the right type of test is as important as running it well.\n\n### Moderated vs. Unmoderated\n\n**Moderated testing:** A researcher is present during the session — in person or via video — and can ask follow-up questions, probe for reasoning, and redirect if needed.\n- **Best for:** Early-stage prototypes, complex tasks, stakeholder-facing sessions where live observation builds organizational buy-in\n- **Trade-off:** More expensive per session, requires scheduling coordination, risk of moderator bias\n\n**Unmoderated testing:** Participants complete the study independently, on their own schedule, with no researcher present. Sessions are recorded for later review.\n- **Best for:** Fast hypothesis validation, larger participant pools, budget-constrained teams, straightforward tasks with unambiguous goals\n- **Trade-off:** Cannot probe unexpected behaviors or clarify task confusion in the moment\n\n### Remote vs. In-Person\n\n**Remote testing:** Conducted over the internet via video or a dedicated testing platform. Can be moderated or unmoderated.\n- **Best for:** Geographically dispersed users, digital products, budget-constrained teams\n- **Trade-off:** Loses some observational richness — body language and environmental context are harder to capture\n\n**In-person (lab) testing:** Researcher and participant are co-located, often with a separate observation room for stakeholders.\n- **Best for:** Physical products, high-fidelity behavioral observation, studies requiring eye-tracking or biometrics\n- **Trade-off:** Most expensive option, geographically constrained, potential lab effect on behavior\n\n| Test Type | Best Scenario | Key Trade-off |\n|-----------|---------------|---------------|\n| Moderated | Complex tasks, early prototypes | Expensive, scheduling friction |\n| Unmoderated | Fast validation, large participant pools | Less depth, cannot probe |\n| Remote | Dispersed users, digital products | Loses body language |\n| In-person | Physical products, high-fidelity observation | Most expensive |\n\n## How Many Participants Do You Need?\n\nThe answer depends entirely on what you are trying to learn.\n\n### Qualitative Testing: 5 Participants Per Segment\n\nFor qualitative usability studies — the most common type — 5 participants per distinct user group is the well-supported standard. This comes from Nielsen and Landauer's 1993 mathematical model showing that 5 users reveal approximately 85% of discoverable issues.\n\nThe critical nuance: **5 per segment, not 5 total.** If your product has two meaningfully different user types — say, administrators and end-users, or novices and experts — you need 5 participants from each group, not 5 total.\n\nNielsen Norman Group recommends a maximum of 5–12 participants per round of qualitative testing. Beyond 12, diminishing returns set in rapidly — the 6th user typically surfaces issues already identified by the first 5.\n\n### Quantitative Testing: 20–40+ Participants\n\nWhen the goal shifts from \"discover what problems exist\" to \"measure how often they occur,\" the sample size requirements change dramatically. Statistical usability studies — measuring task completion rates, error rates, or time-on-task — require a minimum of 20 participants for meaningful data, with 30–40 being more reliable for benchmark comparisons.\n\n### The Iterative Testing Model\n\nBoth Nielsen and Don Norman advocate strongly for iterative testing over single large-scale studies. The more productive model:\n\n1. Test 5 users → find 85% of problems → fix the design\n2. Test 5 more users → find remaining and newly introduced problems\n3. Repeat\n\nAs Steve Krug, author of *Don't Make Me Think*, summarizes: \"A morning a month — that is all we ask.\" Three 50-minute sessions per month produces far more actionable output than a single quarterly deep dive, at a fraction of the cumulative cost.\n\n## How to Conduct Usability Testing: Step by Step\n\n### Step 1: Define Your Research Questions\n\nEvery test needs a specific, answerable question. \"Is our product usable?\" is too broad. \"Can first-time users find and complete the checkout flow within 3 minutes without assistance?\" is testable. Write 2–4 core questions. Everything downstream — tasks, metrics, participant criteria — flows from these.\n\n### Step 2: Identify and Recruit Participants\n\nRecruit users who match your actual target audience. Behavioral and experiential match matters more than demographics — someone who uses products in your category the way your users would. Budget 1–2 weeks for recruitment. Avoid testing with colleagues or people who have domain knowledge about your product, as they are not representative of real users.\n\n### Step 3: Choose Your Test Type\n\nBased on your research questions, timeline, and budget, select moderated vs. unmoderated and remote vs. in-person. Then decide on a protocol.\n\nThe **think-aloud protocol** — where participants verbalize their thoughts as they work — is described by Nielsen Norman Group as \"the #1 usability tool\" because it surfaces mental models and reasoning, not just behavioral outcomes. It is the default choice for most qualitative usability sessions.\n\n### Step 4: Write Tasks and Scenarios\n\nTasks should describe a realistic user goal without revealing how to accomplish it.\n\n❌ **Bad task:** \"Use the search bar to find a blue t-shirt.\"\n✅ **Good task:** \"You are shopping for a birthday gift. Find a blue t-shirt in size medium.\"\n\nScenarios add realistic context. Task wording must never contain the exact names of UI elements or navigation labels — doing so eliminates the friction you are trying to measure.\n\n### Step 5: Run a Pilot Test\n\nTest your test first. Run the protocol with one or two internal participants to verify task wording is clear, timing is accurate, and technology works. Fix everything the pilot reveals before running real sessions.\n\n### Step 6: Facilitate the Sessions\n\nSet expectations at the start: \"We are testing the design, not you — there are no wrong answers.\" Encourage think-aloud throughout. **Do not help participants when they struggle.** That struggle is the data.\n\nUse neutral probing questions:\n- \"What are you thinking right now?\"\n- \"What would you expect to happen next?\"\n- \"Tell me more about that.\"\n\nAvoid questions that signal approval or hint at the correct action.\n\n### Step 7: Observe and Take Notes\n\nHave observers take structured notes using four severity levels:\n- **Critical:** Blocks task completion\n- **Serious:** Causes major delay or error\n- **Minor:** Causes slight friction\n- **Observation:** Noted but not directly problematic\n\n### Step 8: Synthesize and Prioritize\n\nAfter all sessions, group observations by theme and rate each issue by severity and frequency. Tie findings directly back to your original research questions. Prioritize the top 3–5 issues before the next design iteration.\n\n### Step 9: Communicate and Act\n\nHold a team debrief within 48 hours of the final session while observations are fresh. Connect findings to specific design decisions. Then iterate and retest — one round is a snapshot; repeated rounds are a feedback loop.\n\n## Common Mistakes to Avoid\n\n**1. Testing with the wrong participants.** If your participants do not represent your actual users, you will solve the wrong problems. Screening criteria must be specific and rigorously enforced.\n\n**2. Moderator bias — leading questions and approval signals.** The most common and damaging error in moderated testing. Moderators unknowingly influence behavior through word choice, tone, or facial expressions. Use only neutral probes: \"Tell me more about that.\"\n\n**3. Tasks that give away the answer.** If task wording contains the exact name of a UI element, you have eliminated the friction you are trying to measure. Write tasks in terms of user goals, not system labels.\n\n**4. Testing too late.** Testing a fully shipped product is better than nothing, but testing a prototype costs a fraction as much and allows for rapid course correction. The earlier you test, the cheaper the fix.\n\n**5. One-and-done testing.** Fixing usability problems often introduces new ones. The iterative model — test, fix, retest — is the standard. Do not treat a single round as a final verdict.\n\n**6. Confusing opinion with behavior.** What users say they prefer and what they actually do are routinely different. Usability testing captures behavioral evidence, not attitudinal data. Observe, do not simply ask.\n\n## Real-World Example\n\nA SaaS company sees 60% drop-off at step 3 of their onboarding. Analytics identify where users leave, but not why. They run five moderated usability sessions with new users.\n\nAll five participants:\n- Reach step 3 confidently\n- Encounter a field labeled \"Workspace identifier\"\n- Pause, re-read the label, and ultimately guess or abandon\n\nThe problem is not the feature — it is the label. \"Workspace identifier\" means nothing to a new user. Renaming it to \"Your team URL\" eliminates the confusion. The following week's onboarding completion rate increases by 22%.\n\nFive users. One afternoon. A measurable revenue impact.\n\n## Modern Approaches: AI-Assisted Usability Research\n\nTraditional usability testing requires scheduling sessions, recruiting participants, facilitating live observations, and manually coding findings — a process that can span weeks.\n\nAI-native research platforms like Koji are changing this equation. Koji can conduct AI-moderated research sessions at scale, automatically surface patterns across multiple sessions, and generate synthesized reports that identify the most critical friction points. For teams practicing continuous discovery, this compresses the feedback loop from weeks to days — and makes iterative testing sustainable even for small teams without dedicated UX research resources.\n\n## Key Takeaways\n\n- Usability testing reveals behavioral evidence — why and where users struggle — that analytics and surveys cannot provide\n- 5 participants per user segment reveal approximately 85% of qualitative usability problems; 20–40+ for quantitative studies\n- Test early and iterate — fixing problems in prototypes costs a fraction of post-launch fixes\n- Moderated testing provides depth; unmoderated provides speed and scale\n- The think-aloud protocol is the most reliable technique for uncovering user mental models\n- Never help participants when they struggle — that struggle is the data\n\n## Frequently Asked Questions\n\n**Q: How is usability testing different from user interviews?**\nA: User interviews explore attitudes, motivations, and mental models through conversation. Usability testing observes actual behavior with a specific product or prototype. Both are valuable; usability testing is specifically about task performance with a real interface.\n\n**Q: When should I start usability testing?**\nA: As early as possible — even with paper prototypes or wireframes. The earlier you test, the cheaper it is to fix what you find. Do not wait for a polished product before testing.\n\n**Q: What is the difference between formative and summative usability testing?**\nA: Formative testing happens during design to identify and fix problems. Summative testing happens after design is complete to measure performance against benchmarks. Most teams need more formative testing, earlier and more often.\n\n**Q: Can I run usability testing remotely?**\nA: Absolutely. Remote usability testing via video call or asynchronous platforms is standard practice and produces comparable findings to in-person testing for digital products. The main trade-off is losing some non-verbal context.\n\n**Q: How do I handle participants who do not struggle with any tasks?**\nA: Either your design is genuinely excellent (validate with quantitative testing) or your tasks are too easy. Revisit task design using more realistic, goal-oriented scenarios that match actual user needs rather than marketing narratives about the product.\n\n\n---\n\n## Related Resources\n\n- [Card Sorting Guide](/docs/card-sorting-guide) — Information architecture testing\n- [UX Research Process](/docs/ux-research-process) — Full UX research framework\n- [Koji for UX Researchers](/docs/koji-for-ux-researchers) — UX research with AI\n- [Remote Interview Best Practices](/docs/remote-interview-best-practices) — Remote testing techniques\n- [Beta Testing Feedback Guide](/docs/beta-testing-feedback-survey-guide) — Pre-launch testing\n\n*Explore [structured questions](/docs/structured-questions-guide) for combining task success metrics with AI-powered usability probing.*\n\n## Further reading on the blog\n\n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n- [Koji vs Maze: Which Research Tool Is Right for Your Team? (2026)](/blog/koji-vs-maze-2026) — Koji and Maze both claim to power product research — but they do very different things. Here’s an honest 2026 comparison to help you choose \n- [Koji vs Userlytics: AI-Moderated Customer Interviews vs Usability Testing (2026)](/blog/koji-vs-userlytics-2026) — Koji and Userlytics both help you understand users — but they answer completely different research questions. Here's how to choose the right\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"How to Conduct Usability Testing: Complete Guide — Koji","metaDescription":"Learn how to conduct usability testing with this complete guide. Covers types of tests, how many participants you need, step-by-step facilitation, and the most common mistakes to avoid.","keywords":["usability testing","how to conduct usability testing","usability testing guide","moderated usability testing","unmoderated usability testing","UX testing methods","usability testing participants","think aloud protocol"],"aiSummary":"Usability testing is a research method where representative users complete realistic tasks with a product while a researcher observes where they struggle or succeed. This guide covers the four main types of usability tests, the 5-users rule from Nielsen and Landauer, a 9-step facilitation process, and the six most common mistakes that invalidate results — including moderator bias and tasks that give away the answer.","aiPrerequisites":["user-interview-guide","writing-interview-questions"],"aiLearningOutcomes":["Choose the right type of usability test for any research question","Determine the correct number of participants for qualitative and quantitative studies","Write effective task scenarios that do not bias participants","Facilitate sessions using the think-aloud protocol without introducing moderator bias","Synthesize and prioritize findings into actionable design recommendations"],"aiDifficulty":"beginner","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"e311881b-a68d-4217-bf82-ae8fdcba0eaf","slug":"human-evaluation-ai-outputs","title":"Human Evaluation of AI Outputs: The Complete Guide for Product Teams (2026)","url":"https://www.koji.so/docs/human-evaluation-ai-outputs","summary":"Human evaluation scores AI outputs against an explicit rubric to establish ground truth that automated metrics and public benchmarks cannot provide. A defensible design needs concrete rubric anchors, binary decomposed criteria, 150-300 stratified real inputs, 2-3 raters per item, measured inter-rater agreement (Cohen kappa target 0.7+), and blind randomised presentation. Costs run $0.50-$2.00 per crowd rating and $5-$15 for expert ratings. Koji runs the rating and the reasoning interview in one AI-moderated session using scale, yes_no, ranking, single_choice, multiple_choice and open_ended questions, then thematically analyses every justification.","content":"## The short answer\n\n**Human evaluation is the practice of having people score AI outputs against an explicit rubric so you know whether your model is actually good enough to ship.** It is the only method that produces ground truth. Automated metrics compare your output to a reference string; benchmarks tell you how a base model performs on someone else's tasks; neither tells you whether *your* users would accept *this* answer for *their* job.\n\nA defensible human evaluation has five parts: a rubric with concrete anchors, a frozen sample of real inputs, at least two independent raters per item, a measured inter-rater agreement score, and blind randomised presentation. Get those right and you have a number you can defend in a launch review. Get them wrong and you have expensive opinions.\n\nThe expensive part has always been the *why*. A rating tells you output 47 scored 2 out of 5; it does not tell you what a rater expected instead, or which failure would have made a real user churn. That is a qualitative research problem, and it is exactly what an AI-native platform like Koji automates — running the rating and the follow-up interview in the same session, then producing thematic analysis of the reasoning across every rater without anyone reading a transcript.\n\n---\n\n## What human evaluation is — and what it is not\n\nThree adjacent things get confused constantly:\n\n| Activity | Question it answers | Who participates |\n|---|---|---|\n| **Human evaluation (this guide)** | Is this output good, against an explicit standard? | Raters scoring outputs |\n| **[Usability testing](/docs/usability-testing-guide)** | Can a person accomplish a task with this interface? | Users completing tasks |\n| **[User research for AI products](/docs/user-research-for-ai-products)** | Do people trust, adopt, and control the AI? | Customers in their real context |\n\nYou need all three. Human evaluation is the narrowest and the most measurable: it takes outputs out of the product, strips the branding, and asks a rater to judge quality against criteria you wrote down in advance.\n\nIt is also the layer that most teams skip. Automated metrics are free and instant, so they get run every commit; human evaluation costs money and calendar time, so it gets deferred until a launch goes badly. Stanford-affiliated research on evaluation practice reports that systematic evaluation reduces production failures by up to 60% while enabling roughly 5x faster iteration — the cost of the eval is almost always smaller than the cost of the incident it prevents.\n\n---\n\n## Why automated metrics are not enough\n\nReference-based metrics (BLEU, ROUGE, exact match) assume there is one right answer written down somewhere. For summarisation, drafting, support replies, agent trajectories, and anything conversational, there are hundreds of acceptable answers and the metric punishes the good ones for using different words.\n\nPublic benchmarks have the opposite problem: they measure general capability on tasks that are not yours. A model that tops a reasoning leaderboard can still fail badly at \"write a refund email in our brand voice that does not promise anything legal will not honour.\"\n\nHuman evaluation is what closes the gap between *capable* and *acceptable for our use case*. Everything else is a proxy.\n\n---\n\n## Step 1: Write the rubric before you look at any output\n\nThe rubric is the study design. Everything downstream — agreement, cost, defensibility — is determined here.\n\n**Use concrete anchors, not adjectives.** \"Helpfulness: 1–5\" produces noise. \"5 = answers the question, cites the correct policy, and requires no follow-up; 3 = answers the question but omits a condition the customer needs; 1 = wrong, or invents a policy\" produces agreement. Rubrics with ambiguous anchors lose discriminative power — every rater centres on 3 and the scores stop separating good from bad.\n\n**Decompose multi-dimensional scales into binary criteria.** Instead of one 1–5 \"quality\" score, ask five yes/no questions: is it factually correct against the source? does it follow the format? is the tone on-brand? does it refuse appropriately? is it complete? Recent rubric research finds that breaking a compound Likert item into fine-grained binary criteria materially improves inter-rater agreement, because each rater is judging one thing at a time.\n\n**Separate objective from subjective criteria.** Factual accuracy and format compliance are checkable and should reach near-perfect agreement. Tone and helpfulness are judgements and will not — that is fine, as long as you report them separately instead of averaging them into one misleading number.\n\n**Freeze a calibration set.** Pick 15–25 examples, agree on the \"correct\" score for each as a group, and use them to train every new rater before they touch live items. This is the single cheapest quality control in the whole process.\n\n---\n\n## Step 2: Choose the scoring mode\n\n| Mode | How it works | Best for | Cost |\n|---|---|---|---|\n| **Pointwise (absolute)** | Rate each output alone against the rubric | Tracking quality over time, regression gates | Lowest |\n| **Pairwise (side-by-side)** | Show two outputs, pick the better one | Model or prompt comparisons, A/B decisions | Medium |\n| **Reference-based** | Compare output to a gold answer | Tasks with a defensible correct answer | Highest to build |\n\nPairwise is more reliable when the difference is subtle — people are far better at \"which of these two\" than at \"is this a 3 or a 4\" — but pairwise results do not give you an absolute quality level, so you cannot use them alone as a ship gate. Most mature teams run pointwise continuously and pairwise at decision points.\n\n---\n\n## Step 3: Decide who rates, and what that costs\n\n| Rater pool | Typical cost per rating | Use when |\n|---|---|---|\n| Crowd panel (Prolific, MTurk) | $0.50–$2.00 | General-audience tasks, large volume |\n| Internal team (loaded time) | $8–$25 | Early iteration, domain context needed |\n| Subject-matter experts | $5–$15 per rating, $150–$300/hour | Regulated, clinical, legal, financial output |\n\nA standard 200-example evaluation set with three raters per item lands around **$300–$1,200 on a crowd panel** and considerably more with experts. Building a durable custom evaluation dataset of 5,000–10,000 examples is a $20,000–$50,000 project with $5,000–$15,000 a year of maintenance as your product and model shift.\n\nTwo rules that save money: (1) never buy expert ratings for criteria a non-expert can judge — split the rubric and route each criterion to the cheapest competent pool; (2) sample your eval set from real production inputs, stratified across your actual traffic mix, rather than from examples the team wrote.\n\n---\n\n## Step 4: Sample size and agreement\n\n**How many items?** For a ship/no-ship gate on a single quality bar, 150–300 stratified real inputs is the working range: enough to detect a meaningful regression, small enough to re-run weekly. For comparing two models where you expect a small difference, you need more items, not more raters per item.\n\n**How many raters per item?** Two minimum, three when the criterion is subjective. Below two you cannot compute agreement at all, and an evaluation without an agreement number is an opinion with a spreadsheet.\n\n**Agreement targets.** Compute Cohen's kappa for two raters, Krippendorff's alpha for three or more. Practitioner guidance converges on these bands:\n\n| Kappa | Interpretation | Action |\n|---|---|---|\n| Below 0.4 | Rubric is ambiguous | Rewrite the anchors and re-run |\n| 0.4–0.6 | Weak but tunable | Recalibrate raters, tighten one criterion |\n| 0.6–0.8 | Acceptable | Ship the rubric, monitor |\n| Above 0.8 | Strong | Safe to automate parts of it |\n\nAim for **kappa ≥ 0.7 between human pairs** before you trust the numbers — and before you consider handing any of the scoring to an automated judge. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) for the calculation mechanics.\n\n---\n\n## Step 5: Control the biases you introduce\n\n- **Blind the source.** Raters must not know which model, prompt version, or vendor produced an output. Knowing produces the result the team hoped for.\n- **Randomise order per rater.** Position effects are real and large in comparison tasks — the same bias that makes automated judges over-prefer whichever answer appears first.\n- **Counterbalance pairs.** In pairwise mode, show A/B to half the raters and B/A to the other half.\n- **Rotate the calibration check.** Insert two known-answer items into every batch to detect rater drift and fatigue.\n- **Re-bench quarterly.** Model behaviour changes underneath you. A rubric validated in January against a frozen labelled set should be re-validated in April.\n\n---\n\n## How to run human evaluation with Koji\n\nTraditional tooling forces a split: a spreadsheet or annotation tool collects the scores, and a separate round of interviews — if anyone has time — collects the reasoning. Koji collapses both into one AI-moderated session, which is why the \"why\" stops being the part that gets cut.\n\nKoji's [six structured question types](/docs/structured-questions-guide) map directly onto evaluation design:\n\n- **scale** — rubric criteria on a defined 1–5 or 1–7 anchor set, aggregated into distributions rather than averages that hide bimodal disagreement\n- **yes_no** — the binary decomposed criteria that drive high agreement (factually correct? format compliant? appropriate refusal?)\n- **ranking** — pairwise and n-way preference ordering across model or prompt variants\n- **single_choice** — failure-type classification (hallucination / omission / tone / policy breach / formatting)\n- **multiple_choice** — every failure mode present in one output, not just the worst one\n- **open_ended** — the rater's reasoning, where the AI moderator automatically probes: *what did you expect instead? what would a customer have done after reading this?*\n\nThen the parts that normally take a week happen automatically: every open-ended justification is thematically analysed, failure types are aggregated across raters, and the report shows the distribution per criterion with representative quotes attached. Voice mode is available when you want experts thinking out loud rather than typing, and a customisable AI consultant can carry your rubric and domain context into every session so raters are probed the way a specialist would probe them.\n\nTwo practical advantages over manual programmes: you can run 60 raters as easily as six, and re-running the identical study against a new model version is a duplicate-and-launch operation rather than a re-recruitment project. Only conversations that clear Koji's quality bar (a 3+ on its 1–5 interview quality score) consume a credit, so low-effort responses do not silently inflate your eval budget.\n\n---\n\n## A worked example: a support-reply assistant\n\nA B2B SaaS team wants to ship an AI draft-reply feature to their support inbox.\n\n1. **Sample.** 200 real tickets, stratified: 40% billing, 30% troubleshooting, 20% account changes, 10% cancellations.\n2. **Rubric.** Five binary criteria (factually correct against the help centre, no invented policy, correct escalation, on-brand tone, complete) plus one 1–5 overall usefulness scale with written anchors.\n3. **Raters.** Three support agents per item for the factual criteria; two brand/marketing reviewers for tone. Cost: about 15 hours of internal time.\n4. **Design.** Blind to prompt version, randomised order, two calibration items per batch of 25.\n5. **Run.** Delivered as a Koji study — yes_no questions for the binary criteria, scale for usefulness, single_choice for failure type, open_ended for reasoning with automatic AI probing.\n6. **Result.** Overall usefulness averaged 3.8, which alone would have shipped. The distribution was bimodal: 4.4 on billing, 2.6 on cancellations. Failure-type aggregation showed 71% of the low scores were \"invented a policy,\" concentrated in cancellation tickets. Thematic analysis of the open-ended reasoning surfaced the actual cause: the retention-offer rules were not in the retrieval corpus.\n7. **Decision.** Ship for billing and troubleshooting, block cancellations behind a human, re-run the same study in two weeks. Kappa on the factual criteria: 0.81. On tone: 0.52 — reported separately, not averaged in.\n\nThe eval took four days end to end. The version of this study that a spreadsheet produces would have reported \"3.8, looks fine.\"\n\n---\n\n## Common mistakes\n\n1. **Averaging everything into one number.** A single quality score hides the bimodal distribution that tells you what to fix. Report per-criterion and per-segment.\n2. **One rater per item.** No agreement measure, no defensibility.\n3. **Cherry-picked eval sets.** Examples written by the team are easier than production traffic, every time.\n4. **Skipping calibration.** Untrained raters disagree about the rubric, not about the outputs.\n5. **Scores without reasoning.** A number tells you there is a problem; only the open-ended follow-up tells you what to change.\n6. **Never re-running it.** Evaluation is a cadence, not a milestone — see [research refresh cadence](/docs/research-refresh-cadence).\n\n---\n\n## Frequently asked questions\n\n**How many examples do I need for a human evaluation?**\nFor a ship gate on one quality bar, 150–300 stratified real production inputs is the practical range. For detecting a small difference between two models, increase the number of items rather than the number of raters per item.\n\n**What inter-rater agreement should I target?**\nCohen's kappa of 0.7 or above between human pairs. Below 0.4 means the rubric is ambiguous and needs rewriting; 0.6–0.8 is acceptable; above 0.8 is strong enough that parts of the scoring can be safely automated.\n\n**Can I use an LLM to do the rating instead?**\nPartly, and only after humans have established ground truth. Automated judges match human preferences well on some tasks and carry measurable position, verbosity, and self-preference biases on others. See [LLM-as-a-judge vs. human evaluation](/docs/llm-as-a-judge-vs-human-evaluation) for the decision framework.\n\n**How much does human evaluation cost?**\nRoughly $0.50–$2.00 per rating on a crowd panel, $8–$25 for loaded internal time, and $5–$15 (or $150–$300 an hour) for subject-matter experts. A 200-item set with three raters typically lands between $300 and $1,200 on a crowd panel.\n\n**Is human evaluation the same as usability testing?**\nNo. Usability testing asks whether a person can complete a task in your interface. Human evaluation asks whether a specific output meets a written quality standard, with the product context deliberately removed.\n\n**How do I capture why raters scored something low?**\nPair every score with an open-ended justification and probe it. Koji's AI moderator asks the follow-up automatically on every response and thematically analyses the reasoning across all raters, so you get failure causes rather than just failure counts.\n\n---\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make rubric-based evaluation aggregatable\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) — when automated scoring is safe and when it is not\n- [User Research for AI Products](/docs/user-research-for-ai-products) — trust calibration, failure tolerance, and control\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — computing kappa and alpha\n- [Can You Trust AI Interviewers?](/docs/ai-interview-hallucinations-bias-mitigation) — how Koji constrains hallucination and bias\n- [Preference Testing Guide](/docs/preference-testing-guide) — the design-choice sibling of pairwise evaluation\n- [Research Refresh Cadence](/docs/research-refresh-cadence) — keeping evaluations from going stale","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Human Evaluation of AI Outputs: Rubrics, Raters & Agreement (2026)","metaDescription":"How to run human evaluation of LLM outputs: rubric design, rater pools and costs, sample size, Cohen kappa targets, and the bias controls that make it hold up.","keywords":["human evaluation","human eval llm","llm evaluation rubric","ai output quality","inter-rater agreement","human in the loop evaluation","ai model evaluation","rubric design","pairwise evaluation","ground truth labeling"],"aiSummary":"Human evaluation scores AI outputs against an explicit rubric to establish ground truth that automated metrics and public benchmarks cannot provide. A defensible design needs concrete rubric anchors, binary decomposed criteria, 150-300 stratified real inputs, 2-3 raters per item, measured inter-rater agreement (Cohen kappa target 0.7+), and blind randomised presentation. Costs run $0.50-$2.00 per crowd rating and $5-$15 for expert ratings. Koji runs the rating and the reasoning interview in one AI-moderated session using scale, yes_no, ranking, single_choice, multiple_choice and open_ended questions, then thematically analyses every justification.","aiDifficulty":"intermediate","aiEstimatedTime":"14 min"},{"type":"documentation","id":"8b6ca6fc-2b83-410d-ad3e-ae835f0f5f9a","slug":"llm-as-a-judge-vs-human-evaluation","title":"LLM-as-a-Judge vs. Human Evaluation: When to Trust Automated Scoring (2026)","url":"https://www.koji.so/docs/llm-as-a-judge-vs-human-evaluation","summary":"LLM judges and human evaluation are complementary, not competing. GPT-4 judges reached over 80% agreement with human preferences on MT-Bench and Chatbot Arena (Zheng et al. 2023), matching human-human agreement, but the same work documents position bias up to 75%, verbosity bias, and self-enhancement bias (GPT-4 +10%, Claude-v1 +25% self win rate). The safe pattern is a calibration loop: human-label a frozen set of 150-300 inputs, establish human-human kappa first, measure judge-human agreement against that bar, mitigate position bias by averaging both orders, audit 5-10% continuously, and re-bench quarterly. Judges scale a standard; humans set it. Koji runs the human half at judge-like speed by pairing structured scores with probed open-ended reasoning and automatic thematic analysis.","content":"## The short answer\n\n**Use an LLM judge for volume and regression detection; use human evaluation to define what \"good\" means and to keep the judge honest.** They are not competing options, and treating them as a choice is the most common mistake teams make in 2026.\n\nThe evidence for automated judging is genuinely strong. In the study that established the method, GPT-4 judges reached **over 80% agreement with human preferences on MT-Bench and Chatbot Arena — the same level of agreement humans reach with each other** ([Zheng et al., 2023](https://arxiv.org/abs/2306.05685)), across roughly 3,000 controlled expert votes and 3,000 crowdsourced votes.\n\nThe evidence against blind trust is equally strong, and comes from the same paper. Judges show **position bias of up to 75% preference for whichever response appears first**, verbosity bias toward longer answers, and self-enhancement bias — GPT-4 favoured its own outputs with about a 10% higher win rate, and Claude-v1 by about 25%. Later work adds a subtler warning: high agreement *between* LLM judges is not evidence of alignment with humans ([The Geometry of LLM-as-Judge, 2026](https://arxiv.org/pdf/2606.03043)) — a panel of judges can be confidently and consistently wrong together.\n\nSo the working rule is: **the judge scales your standard, humans set it.** Never let an automated scorer define the target it is measuring.\n\n---\n\n## What LLM-as-a-judge actually means\n\nAn LLM judge is a model prompted with a rubric and asked to score outputs, in one of three modes:\n\n| Mode | Prompt shape | Strength | Weakness |\n|---|---|---|---|\n| **Pointwise** | Score this output 1–5 against these criteria | Cheap, tracks over time | Score drift; clustering around the middle |\n| **Pairwise** | Which of A or B is better? | Higher reliability on subtle differences | Position bias; no absolute level |\n| **Reference-based** | How does this compare to the gold answer? | Highest accuracy | Requires an expensive gold set |\n\nThe rubric matters more than the model. A judge given \"rate helpfulness 1–5\" reproduces the same ambiguity that wrecks human agreement; a judge given decomposed binary criteria with concrete anchors behaves much more consistently — the same finding that holds for [human raters](/docs/human-evaluation-ai-outputs).\n\n---\n\n## The decision framework\n\n| Situation | Use | Why |\n|---|---|---|\n| Nightly regression gate over thousands of outputs | **Judge** | Humans cannot run at that cadence or cost |\n| Establishing what \"good\" means for a new feature | **Human** | There is no ground truth to calibrate against yet |\n| Comparing two prompt variants, large expected difference | **Judge** | Cheap, fast, and the effect is bigger than the bias |\n| Ship/no-ship on a regulated or safety-relevant output | **Human** | Accountability cannot be delegated to a scorer |\n| Subjective quality: tone, brand voice, emotional appropriateness | **Human**, judge as a pre-filter | Judges under-detect tone failures they were not told to look for |\n| Factual grounding against a known corpus | **Judge**, human-audited sample | Objective and checkable; audit 10% |\n| Understanding *why* an output failed | **Human, in conversation** | No scorer produces causal explanation |\n| Scoring 50,000 production traces | **Judge**, with a human-labelled calibration set | The only economically viable option |\n\nThe pattern: the judge is a measuring instrument. It needs calibration against a reference, periodic re-calibration, and someone accountable for its readings.\n\n---\n\n## The calibration loop\n\nThis is the part most teams skip, and it is what separates a defensible automated evaluation from a plausible one.\n\n1. **Human-label a frozen set.** Take 150–300 stratified real inputs. Have 2–3 humans score them using the *exact same rubric text* the judge will receive. Anything else and you are comparing two different measurements.\n2. **Establish human-human agreement first.** Compute Cohen's kappa (two raters) or Krippendorff's alpha (three or more). If humans cannot reach roughly 0.7 with each other, the rubric is broken and the judge will inherit the ambiguity — fix the rubric before you touch the automation.\n3. **Measure judge-human agreement.** Score the same frozen set with the judge and compute agreement against the human labels. The bar to clear is *human-human agreement on the same set*, not an abstract number.\n4. **Mitigate the known biases.** Swap positions and average both orders to neutralise position bias. Strip length cues or explicitly instruct against rewarding verbosity. Never let a model be the sole judge of its own family's outputs.\n5. **Deploy with an audit rate.** Route 5–10% of judged items to humans continuously. Rising disagreement is your early warning that model behaviour, traffic mix, or both have shifted.\n6. **Re-bench quarterly.** Re-run the frozen set through the judge every quarter and after every model upgrade. Agreement that held in January is not evidence about April.\n\n---\n\n## The cost picture\n\n| Approach | Cost per item | Throughput | What you get |\n|---|---|---|---|\n| Crowd human rating | $0.50–$2.00 | Days | Ground truth, no reasoning unless you ask |\n| Internal team rating | $8–$25 loaded | Days–weeks | Domain judgement, high opportunity cost |\n| Expert rating | $5–$15, or $150–$300/hour | Weeks | Defensible in regulated contexts |\n| LLM judge | Fractions of a cent | Minutes | Scale, plus whatever bias you failed to control |\n| **AI-moderated evaluation (Koji)** | One credit per qualifying session | Hours | Scores **and** probed reasoning, thematically analysed |\n\nThe judge is three to four orders of magnitude cheaper per item. That is precisely why the discipline matters: a cheap instrument that is quietly miscalibrated produces confident wrong decisions faster than an expensive one.\n\n---\n\n## What an LLM judge structurally cannot give you\n\nEven a perfectly calibrated judge answers one question: *does this output match the rubric I was given?* It cannot tell you:\n\n- **Whether the rubric is the right rubric.** Judges score what you asked for. If your criteria miss the thing customers actually care about, the judge will award high marks all the way to a failed launch.\n- **What the rater expected instead.** A score of 2 with no counterfactual is not actionable.\n- **Whether the failure matters.** Some errors are cosmetic; some end the relationship. Only people who live with the consequence can rank them — see [user research for AI products](/docs/user-research-for-ai-products) on failure tolerance.\n- **How trust changes over time.** Trust calibration is longitudinal and behavioural; a per-output score cannot see it.\n- **Novel failure modes.** A judge finds the failures listed in its prompt. Humans find the ones nobody anticipated — which is the entire value of the exercise on a new feature.\n\n---\n\n## How Koji fits: humans at judge-like scale\n\nThe historical reason teams over-rely on automated judging is that human evaluation was slow to organise, expensive to moderate, and painful to analyse. Koji removes all three constraints, which changes the economics of the \"use both\" recommendation.\n\nUsing Koji's [six structured question types](/docs/structured-questions-guide), your evaluation study collects the same structured signal a judge produces — **scale** for rubric criteria, **yes_no** for decomposed binary checks, **ranking** for pairwise preference across variants, **single_choice** for failure classification, **multiple_choice** for co-occurring failures — and then does what no judge can: the **open_ended** justification is probed live by the AI moderator (\"what did you expect instead?\", \"what would you have done after reading this?\") and every justification is thematically analysed automatically.\n\nPractically, that means:\n\n- **Calibration sets get built in hours, not sprints.** Recruit raters, run the study, get labelled data with reasoning attached.\n- **Ground truth stays fresh.** Re-running an identical evaluation against a new model version is duplicate-and-launch, so quarterly re-benching stops being the task that slips.\n- **Judge disagreements get explained, not just counted.** When the judge and your humans diverge on a slice of traffic, run that slice as a Koji study and read the thematic analysis of *why* — the diagnosis a disagreement rate alone never gives you.\n- **Cost stays predictable.** Text conversations cost 1 credit, voice 3, and only sessions clearing Koji's 1–5 interview quality gate (3 or above) consume credits at all, so weak responses do not bill.\n\nLegacy survey tools like SurveyMonkey or Qualtrics can collect the ratings, but the reasoning arrives as a column of unread free text. An AI-native platform collects the rating, asks the follow-up, and returns the themes — which is exactly the layer automated judges cannot reach.\n\n---\n\n## A worked example: shipping a judge you can defend\n\nA team wants to gate every release of an AI summarisation feature on an automated score.\n\n1. **Frozen set.** 250 real documents stratified by length and domain.\n2. **Rubric.** Four binary criteria (no unsupported claim, all key entities retained, correct length band, no leaked PII) plus one 1–5 usefulness scale with anchors.\n3. **Human labels.** Three raters per item via a Koji study — yes_no for the binary criteria, scale for usefulness, open_ended for reasoning. Human-human kappa: 0.78 on the binary criteria, 0.55 on usefulness.\n4. **Judge run.** The same rubric text, pointwise, both orders averaged where applicable.\n5. **Result.** Judge-human agreement was 0.74 on the binary criteria — at the human bar, so it is approved as the automated gate. On the 1–5 usefulness scale, agreement was 0.41, and the judge systematically over-rewarded longer summaries. Usefulness stays human-only.\n6. **Deployment.** The judge gates every build on the four binary criteria; 8% of items are human-audited weekly; a full human re-bench runs quarterly and after each model upgrade.\n\nThe output is not \"we use an LLM judge.\" It is \"we use an LLM judge for four criteria where it is calibrated to human agreement, and humans for the one where it is not.\"\n\n---\n\n## Common mistakes\n\n1. **Deploying a judge before establishing human-human agreement.** The judge inherits every ambiguity in your rubric and reports it as confidence.\n2. **Using the model family being evaluated as its own judge.** Self-enhancement bias is documented and large.\n3. **Ignoring position bias in pairwise mode.** Always score both orders and average.\n4. **Treating high inter-judge agreement as validation.** Judges agreeing with judges is not evidence about humans.\n5. **Never auditing after launch.** Without a standing human audit rate you will not notice drift until a customer does.\n6. **Letting the judge choose the rubric.** Criteria come from customers, not from a scorer — which is a [research](/docs/user-interview-guide) problem, not an engineering one.\n\n---\n\n## Frequently asked questions\n\n**Is LLM-as-a-judge accurate enough to replace human evaluation?**\nNot as a replacement. Strong judges reached over 80% agreement with human preferences on MT-Bench and Chatbot Arena — matching human-human agreement — but the same research documents position bias up to 75%, verbosity bias, and self-enhancement bias. Judges scale a standard; they cannot set one.\n\n**What is self-enhancement bias?**\nThe tendency of an LLM judge to prefer outputs from its own model family. In the original study GPT-4 favoured its own answers by roughly 10% higher win rate and Claude-v1 by roughly 25%. Never let a model be the sole judge of its own outputs.\n\n**How do I know if my judge is calibrated?**\nHuman-label a frozen set of 150-300 real inputs with the same rubric text the judge receives, compute human-human agreement first, then compare the judge against those labels. The bar to clear is your human-human agreement on that set, not a fixed threshold.\n\n**How often should I re-validate an LLM judge?**\nQuarterly at minimum, and after every model or prompt change, against the frozen labelled set. Run a continuous 5-10% human audit between re-benches to catch drift early.\n\n**Can I use an LLM judge for tone and brand voice?**\nOnly as a pre-filter. Subjective criteria show the lowest judge-human agreement, and tone failures are the ones a judge was not told to look for. Keep a human panel on the subjective half of the rubric.\n\n**What does an LLM judge never tell me?**\nWhether the rubric itself is right, what the rater expected instead, whether a given failure would actually cost you the customer, and any failure mode nobody wrote into the prompt. Those require conversation, which is what Koji's AI-moderated evaluation sessions capture alongside the scores.\n\n---\n\n## Related resources\n\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) — rubric design, rater pools, sample size, and agreement targets\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types behind aggregatable evaluation data\n- [User Research for AI Products](/docs/user-research-for-ai-products) — trust calibration, failure tolerance, and control\n- [Can You Trust AI Interviewers?](/docs/ai-interview-hallucinations-bias-mitigation) — grounding and bias controls in AI moderation\n- [AI vs. Human Moderators](/docs/ai-vs-human-moderators) — the same use-both logic applied to interview moderation\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — computing kappa and alpha\n- [Synthetic Users Research Methodology](/docs/synthetic-users-research-methodology) — where simulated respondents help and where they mislead","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"LLM-as-a-Judge vs. Human Evaluation: The 2026 Decision Guide","metaDescription":"LLM judges hit 80%+ agreement with humans — and show 75% position bias and self-preference. The evidence, failure modes, and calibration loop for using both.","keywords":["llm as a judge","llm as a judge vs human evaluation","automated llm scoring","ai judge bias","position bias llm judge","self-enhancement bias","ai evaluation methods","judge calibration","human eval vs automated eval","llm evaluation"],"aiSummary":"LLM judges and human evaluation are complementary, not competing. GPT-4 judges reached over 80% agreement with human preferences on MT-Bench and Chatbot Arena (Zheng et al. 2023), matching human-human agreement, but the same work documents position bias up to 75%, verbosity bias, and self-enhancement bias (GPT-4 +10%, Claude-v1 +25% self win rate). The safe pattern is a calibration loop: human-label a frozen set of 150-300 inputs, establish human-human kappa first, measure judge-human agreement against that bar, mitigate position bias by averaging both orders, audit 5-10% continuously, and re-bench quarterly. Judges scale a standard; humans set it. Koji runs the human half at judge-like speed by pairing structured scores with probed open-ended reasoning and automatic thematic analysis.","aiDifficulty":"intermediate","aiEstimatedTime":"13 min"},{"type":"documentation","id":"7ab04bea-4b31-4e87-8f12-3589b5f1031e","slug":"ai-red-teaming-with-users","title":"AI Red Teaming with Real Users: How to Find Harms Before Your Users Do (2026)","url":"https://www.koji.so/docs/ai-red-teaming-with-users","summary":"AI red teaming is adversarial testing that produces a defect list, not a score. Half the harm surface — responsible-AI harms and contextual failure — is a user research problem, not a security one. Covers a six-phase protocol, harm taxonomies, adversary recruiting, severity scoring with measured inter-rater agreement, red-teamer wellbeing, and the EU AI Act Article 55 and NIST obligations that now make adversarial testing mandatory for some providers.","content":"## The short answer\n\n**AI red teaming is the practice of deliberately trying to make your AI product behave badly, then treating what you find as a defect list rather than a score.** It is not benchmarking, it is not usability testing, and it is not a penetration test — though it borrows from all three. The distinguishing feature is intent: every other research method asks *what happens when people use this normally*. Red teaming asks *what happens when someone is trying to break it, or when a normal person hits the worst 0.1% of the input distribution*.\n\nTwo things changed in 2026 that moved this from a frontier-lab specialty to an ordinary product obligation. First, the law caught up: the EU AI Act's obligations for general-purpose AI models have applied since **2 August 2025**, and Article 55(1)(a) requires providers of models with systemic risk to perform model evaluation **including adversarial testing** to identify and mitigate systemic risk ([EU AI Act, Art. 55](https://artificialintelligenceact.eu/article/55/)). Models already on the market before that date have until **2 August 2027** to comply. Second, the evidence base matured: NIST published its Assessing Risks and Impacts of AI (ARIA) program results as **NIST AI 700-2 in November 2025**, built on a pilot involving roughly **51 red teamers across 508 testing sessions** on seven submitted AI applications, and introduced the Contextual Robustness Index (CoRIx) as a measure of whether an application holds up in its intended use context ([NIST AI 700-2](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.700-2.pdf)).\n\nThe uncomfortable part for most product teams is that half of this work is not a security problem. It is a research problem, and it belongs to the people who already know how to recruit strangers, design a probe, and score a subjective judgement reliably.\n\n## Red teaming is not safety benchmarking\n\nThe single most useful framing published on this subject comes from Microsoft's AI Red Team, which documented **eight lessons from 80 red-teaming operations covering more than 100 generative AI products since 2021** ([Bullwinkel et al., 2025, arXiv:2501.07238](https://arxiv.org/abs/2501.07238)). Their third lesson states it plainly: *AI red teaming is not safety benchmarking*.\n\nThe distinction matters operationally. A benchmark asks a fixed set of questions and returns a number you can compare across model versions. Red teaming produces an open-ended, contextual, adversarial exploration whose output is a set of *novel* failures nobody had written down yet. Benchmarks measure known risks; red teaming discovers unknown ones. A team that runs a jailbreak benchmark and reports \"94% refusal rate\" has not red teamed anything — it has measured performance against attacks that were already public, which is exactly the set of attacks least likely to hurt them.\n\nTheir second lesson is equally load-bearing for research teams: *you don't have to compute gradients to break an AI system*. The most effective attacks in their operations were frequently simple prompt-level manipulations at the system layer, not sophisticated optimisation against model weights. That is precisely the kind of work a domain expert with no ML background can do — and the reason a user researcher can lead it.\n\nTheir fifth and sixth lessons close the argument: *the human element of AI red teaming is crucial*, and *responsible AI harms are pervasive but difficult to measure*. Automation extends coverage; it does not replace judgement about whether an output is actually harmful to a specific person in a specific context.\n\n## What a research-led red team actually covers\n\nThere are two halves to the harm surface, and they need different people.\n\n| Layer | Example failures | Who should probe it |\n|---|---|---|\n| Security | Prompt injection, data exfiltration, privilege escalation, SSRF in the retrieval pipeline, tool/MCP abuse | Application security, with AI-specific tooling |\n| Model behaviour | Jailbreaks, refusal failures, unsafe instruction-following, over-refusal of legitimate requests | Mixed — security plus domain experts |\n| Responsible-AI harms | Stereotyping, degraded quality for a dialect or accent, medical or legal overreach, psychosocial harm, manipulation, unfair allocation | **User researchers and affected-community members** |\n| Contextual failure | Confidently wrong output in a high-stakes workflow, silent degradation on rare inputs, harmful defaults | User researchers with domain experts |\n\nThe bottom two rows are where research teams add something no security team can. Whether a model output is *harmful* is a judgement about people, made by people, and it is exactly the kind of subjective, low-agreement judgement that research methodology exists to make reliable. This is the same measurement problem covered in [human evaluation of AI outputs](/docs/human-evaluation-ai-outputs) — rubric anchors, independent raters, measured agreement — applied to the tail of the distribution instead of the middle.\n\nThe security half is not optional either. IBM's 2025 breach research found that **13% of organisations reported a breach of their AI models or applications, and 97% of those breached had no AI-specific access controls in place** ([IBM Cost of a Data Breach Report 2025](https://www.ibm.com/reports/data-breach)). Red teaming does not fix governance, but it is usually what makes the gap visible.\n\n## The six-phase protocol\n\n### 1. Scope the system, not the model\n\nMicrosoft's first lesson is *understand what the system can do and where it is applied*. A model that is harmless in a chat window can be dangerous the moment it is given a tool, a document store, or an audience of teenagers. Write down: the deployment context, the user population (including the population you did not design for), the tools and data the system can reach, and the consequence of a wrong answer. The consequence column is what sets your severity scale.\n\n### 2. Build a harm hypothesis list, not a prompt list\n\nPrompts are artefacts; intents are the unit of work. Recruit the list from four sources: prior incidents in your category, regulatory risk categories (NIST's Generative AI Profile, **NIST AI 600-1**, enumerates twelve), support tickets and trust-and-safety reports from your own product, and — critically — interviews with people who resemble both your attackers and your most vulnerable users.\n\n### 3. Recruit adversaries with standing, not just skill\n\nThe failure mode here is a red team made entirely of engineers who share the builders' blind spots. OpenAI's published approach to external red teaming emphasises deliberately prioritising **geographic and domain diversity** in who is invited to probe a model ([OpenAI, Approach to External Red Teaming](https://cdn.openai.com/papers/openais-approach-to-external-red-teaming.pdf)). For responsible-AI harms specifically, the highest-yield participants are people who would be harmed: clinicians for a health assistant, benefits caseworkers for an eligibility tool, speakers of the dialects your speech model handles worst.\n\nUse a proper [screener](/docs/screener-questions-guide). \"Has broken an AI product before\" is a weak signal; \"works daily in the domain and can tell a plausible wrong answer from a correct one in under ten seconds\" is a strong one.\n\n### 4. Run structured and free-form passes\n\nStructured passes cover your hypothesis list systematically so you can claim coverage. Free-form passes are where the novel findings come from. Budget at least a third of session time for unguided exploration, and record the participant's *strategy*, not only the winning prompt — strategies generalise across model versions in a way that individual prompts do not.\n\n### 5. Score severity and measure agreement\n\nEvery finding needs: reproduction steps, a harm category, a severity rating, an estimated likelihood in real use, and the affected population. Severity is a subjective rating, which means it needs the same discipline as any other subjective rating — at least two independent raters and a measured agreement score. Aim for Cohen's kappa of **0.7 or above**; below 0.4 your severity scale is ambiguous and needs rewriting, not more raters. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) for the mechanics.\n\n### 6. Convert findings into regression tests\n\nMicrosoft's eighth lesson is that *the work of securing AI systems will never be complete*. The practical response is to make each confirmed finding permanent: every reproduced harm becomes a row in your evaluation set, so the next model version is automatically tested against it. That handoff — red team finding to frozen test case — is covered in [building evaluation datasets from real user research](/docs/ai-evaluation-dataset-golden-set).\n\n## Red teaming versus its neighbours\n\n| Method | Question it answers | Input distribution | Typical output |\n|---|---|---|---|\n| Usability testing | Can people complete the task? | Typical | Friction list |\n| Human evaluation | Is the output good enough to ship? | Representative sample | Rubric scores + pass rate |\n| LLM-as-a-judge | Did quality regress since last build? | Frozen eval set | Automated score |\n| **Red teaming** | **What is the worst this can do, and to whom?** | **Adversarial and tail** | **Reproducible harm findings** |\n| Security pentest | Can the system be compromised? | Adversarial, infra-focused | Vulnerability report |\n\nThey compose in a specific order. Red teaming discovers; [human evaluation](/docs/human-evaluation-ai-outputs) quantifies; [an LLM judge](/docs/llm-as-a-judge-vs-human-evaluation) monitors. Skipping the first step means your judge is monitoring for problems you already knew about.\n\n## Protecting the people who do this work\n\nThis is the part most guides omit, and it is a genuine ethical obligation. Red teamers are asked to elicit content that is by design distressing — self-harm instructions, harassment, sexual content involving minors, graphic violence. Treat it as you would any research involving exposure to disturbing material:\n\n- **Informed consent that names the content categories** in advance, not a generic media release. See [research ethics and informed consent](/docs/research-ethics-guide).\n- **A no-penalty opt-out mid-session**, exercised without explanation.\n- **Exposure limits and rotation** — cap session length and consecutive days on a harm category.\n- **Never recruit minors** for harm probing, even when minors are the affected population; work through adult proxies and safeguarding experts instead ([research with children and teens](/docs/user-research-with-children-teens)).\n- **Ethics review** where your organisation has one — much of this work meets the threshold described in [IRB approval for user research](/docs/irb-approval-user-research).\n\n## The modern approach: red teaming at scale with Koji\n\nThe historical constraint on red teaming was throughput. A moderated session costs a moderator, a calendar slot, and an hour; covering forty harm hypotheses with three participants each means 120 hours of scheduling before anyone reads a transcript. That is why most red teaming has been done by six people in a room for two weeks — not because six is the right number, but because 120 is unaffordable.\n\nKoji removes the scheduling layer entirely. Red-team sessions run as async, link-based AI-moderated interviews: you send a link, the AI moderator runs the protocol, probes the participant's reasoning when they find something, and the transcript lands analysed. Forty hypotheses across sixty domain experts is a days-long study, not a quarter-long programme.\n\nThe [six structured question types](/docs/structured-questions-guide) map onto red-team scoring almost exactly:\n\n| Question type | Red-team use |\n|---|---|\n| `open_ended` | The attack narrative, the strategy, and *why* the participant considered the output harmful — with AI follow-up probing that a form cannot do |\n| `single_choice` | Harm category from your taxonomy |\n| `scale` | Severity and likelihood ratings, aggregated into distributions |\n| `yes_no` | Binary criteria: did the system refuse? did it cite a source? did it stay in scope? |\n| `ranking` | Ordering several failing outputs by which would do the most damage |\n| `multiple_choice` | Which populations the participant believes are affected |\n\nBecause every session answers the same structured questions, severity distributions and category frequencies aggregate automatically — you get a ranked harm register rather than forty documents someone has to read. Koji's thematic analysis clusters the *strategies* participants used, which is the durable artefact; its quality scoring rates each conversation 1–5 against your research goals so thin sessions are visible immediately rather than diluting the register. Voice interviews are the right modality when the harm involves speech, accent handling, or social-engineering scripts that only work out loud.\n\n| | Traditional red-team workshop | Koji |\n|---|---|---|\n| Participants | 5–10, mostly internal | 30–100+, external domain experts |\n| Setup | Scheduling, NDAs, facilitation plan | A study link |\n| Time to findings | 2–6 weeks | Days |\n| Scoring | Post-hoc, in a spreadsheet | Structured at capture, aggregated live |\n| Repeatability per model release | Rarely — too expensive | Re-send the link |\n| Coverage of affected communities | Whoever was available | Screened and recruited deliberately |\n\nNone of this replaces a security team's tooling for the infrastructure layer. It replaces the part that was always the bottleneck: getting enough of the right humans to spend focused, structured time attacking your product.\n\n## Common mistakes\n\n1. **Reporting a pass rate.** Red teaming produces a defect list. A percentage implies a fixed denominator, and the whole point is that the denominator is unknown.\n2. **Recruiting only builders.** People who know how the system works probe where they expect weakness; users probe where the system meets their life.\n3. **Testing the model instead of the product.** Guardrails, retrieval, tools and system prompts are where most real failures live.\n4. **Letting findings die in a doc.** If a finding is not a permanent test case, the next release will reintroduce it.\n5. **Treating over-refusal as a non-finding.** A model that refuses to discuss a legitimate medical question is failing a real user, and that harm rarely shows up in a security-framed red team.\n6. **No wellbeing plan.** This is a research-ethics failure, not a logistics oversight.\n\n## Frequently asked questions\n\n**Is AI red teaming legally required for my product?**\nIt depends what you build. Article 55(1)(a) of the EU AI Act requires adversarial testing of general-purpose AI models with systemic risk, and those obligations have applied since 2 August 2025 (2 August 2027 for models already on the market). Most product teams are deployers rather than GPAI providers and are not directly caught by Article 55 — but high-risk system obligations, sector regulators, and enterprise procurement questionnaires increasingly ask for adversarial testing evidence regardless. See [EU AI Act compliance for user research](/docs/eu-ai-act-user-research-compliance) and [AI governance frameworks](/docs/ai-governance-frameworks-research).\n\n**How is red teaming different from usability testing?**\nUsability testing samples the typical input distribution and asks whether people can succeed. Red teaming samples the adversarial and tail distribution and asks what the worst outcome is. A product can pass every usability test and still produce a harm that ends up in a news story.\n\n**How many red teamers do I need?**\nThere is no saturation number, because the space is open-ended. NIST's ARIA pilot used roughly 51 red teamers across 508 sessions on seven applications. A practical starting point for a product team is 20–30 participants spanning at least three distinct perspectives (domain expert, affected community, adversarially-minded generalist), then re-running per major release.\n\n**Can I automate this with an LLM?**\nPartly. Automated adversarial generation is genuinely good at breadth — it will try thousands of variations you would never type. It is poor at deciding whether an output is harmful in context, and it cannot represent the lived experience of the population at risk. The mainstream position, and Microsoft's fifth lesson, is that automation extends coverage while human judgement remains essential.\n\n**Do synthetic users work for red teaming?**\nFor generating candidate attack strategies, they are a reasonable brainstorming aid. For judging harm, no — the documented sycophancy and endorsement biases of AI personas make them unreliable evaluators. See [synthetic users in research](/docs/synthetic-users-research-methodology).\n\n**What do I do with a finding I cannot fix?**\nDocument it, rate it, and route it to a non-model mitigation: a guardrail, a scope restriction, a disclosure, a human-in-the-loop checkpoint, or a decision not to ship into that context. An unfixable finding that is written down and mitigated is a governance artefact; an unfixable finding that is deleted is a liability.\n\n## Related resources\n\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) — the rubric-and-agreement discipline that makes harm severity ratings defensible\n- [Evaluation Datasets for AI Products](/docs/ai-evaluation-dataset-golden-set) — turning red-team findings into permanent regression tests\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) — automating the monitoring layer once you know what to look for\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types used to score severity, category and likelihood at capture\n- [User Research for AI Products](/docs/user-research-for-ai-products) — trust calibration and failure tolerance in normal use\n- [AI Governance Frameworks for Research Teams](/docs/ai-governance-frameworks-research) — ISO/IEC 42001, NIST AI RMF and the EU AI Act compared\n- [Research Ethics and Informed Consent](/docs/research-ethics-guide) — consent design for sessions involving distressing content","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"AI Red Teaming with Real Users: A 2026 Practitioner Guide","metaDescription":"How to run adversarial testing of AI products with real people: harm taxonomies, recruiting, severity scoring, red-teamer wellbeing, and EU AI Act duties.","keywords":["ai red teaming","red teaming ai products","adversarial testing ai","ai harm discovery","responsible ai harms","eu ai act adversarial testing","ai safety testing with users","llm red team methodology","ai risk assessment research","harm taxonomy ai"],"aiSummary":"AI red teaming is adversarial testing that produces a defect list, not a score. Half the harm surface — responsible-AI harms and contextual failure — is a user research problem, not a security one. Covers a six-phase protocol, harm taxonomies, adversary recruiting, severity scoring with measured inter-rater agreement, red-teamer wellbeing, and the EU AI Act Article 55 and NIST obligations that now make adversarial testing mandatory for some providers.","aiPrerequisites":["Basic understanding of how LLM-based products work","Familiarity with qualitative research recruiting and screening"],"aiLearningOutcomes":["Distinguish red teaming from safety benchmarking, usability testing and penetration testing","Build a harm hypothesis list and a severity scale for your deployment context","Recruit adversaries with domain standing rather than only technical skill","Score harm findings with measured inter-rater agreement","Convert confirmed findings into permanent regression test cases","Apply the ethical safeguards red-teamer wellbeing requires"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min"},{"type":"documentation","id":"b0ec6a78-460b-40c8-91f1-3f5526c8acec","slug":"ai-evaluation-dataset-golden-set","title":"Evaluation Datasets for AI Products: How to Build a Golden Set from Real User Research (2026)","url":"https://www.koji.so/docs/ai-evaluation-dataset-golden-set","summary":"A golden evaluation dataset is a product specification written as examples. Covers sizing against confidence intervals (100-500 items per bucket), the four-bucket structure of production sample, edge cases, adversarial items and failure replays, documented label-error rates in published benchmarks, an item schema with decomposed binary acceptance criteria, sourcing those criteria from real users rather than the build team, inter-rater agreement targets, and maintaining the set against distribution drift, criteria drift and overfitting.","content":"## The short answer\n\n**Your evaluation dataset is your product specification, written as examples instead of prose. Most AI evaluation programmes fail at the dataset, not the metric.** A team with a mediocre metric and an excellent golden set will catch real regressions. A team with a sophisticated LLM judge scoring 500 synthetic prompts nobody's users would ever type will catch nothing, confidently.\n\nThe practical build is: **100–500 verified examples**, drawn from four buckets (production traces, adversarial cases, edge cases, and replays of past failures), with acceptance criteria written by people who represent the users who will accept or reject the output — not by the engineers who built it. Below 100 items the confidence interval is too wide to detect the regressions you care about; above 500 the marginal example stops paying for its maintenance cost.\n\nThe part almost everyone gets wrong is the last clause. \"Ground truth\" for a subjective task is not an engineering opinion recorded quickly. It is an empirical claim about what a specific population would accept for a specific job, and establishing it is a user research problem.\n\n## Why the dataset is the hard part\n\nPublished machine-learning benchmarks — datasets built by well-resourced teams over years, and cited thousands of times — are riddled with label errors. In the canonical study, Northcutt, Athalye and Mueller found an **average of at least 3.3% errors across the test sets of ten commonly used benchmarks**, including **6% of the ImageNet validation set (2,916 mislabelled images)** and an estimated **10% (over five million items) in QuickDraw** ([Northcutt et al., 2021, arXiv:2103.14749](https://arxiv.org/abs/2103.14749)).\n\nTwo findings from that paper should change how you build your own set. First, the errors were not benign: on ImageNet with corrected labels, **ResNet-18 outperforms ResNet-50 once the prevalence of originally mislabelled test examples rises by just 6%** — meaning label noise can invert your model-selection decision. Higher-capacity models fit the systematic label errors more faithfully, so a dirty test set actively rewards the wrong model. Second, when candidate errors were flagged algorithmically and then sent to human crowdworkers for validation, **only 51% of flagged candidates were confirmed erroneous** — so automated cleaning cannot be trusted unsupervised either.\n\nIf ImageNet has a 6% error rate, your hand-assembled 200-row spreadsheet is not cleaner. Budget for adjudication from the start.\n\n## Sizing: what your n actually buys you\n\nEval set size is usually argued about aesthetically. It should be argued about statistically. If your current pass rate is around 90%, here is the 95% confidence interval on that estimate by sample size, and the smallest regression you can reliably detect:\n\n| Items in bucket | 95% CI on a 90% pass rate | Smallest detectable regression |\n|---|---|---|\n| 50 | ±8.3 pp | Only catastrophic drops |\n| 100 | ±5.9 pp | ~8–10 pp |\n| 300 | ±3.4 pp | ~5 pp |\n| 500 | ±2.6 pp | ~4 pp |\n| 1,000 | ±1.9 pp | ~3 pp |\n\nThis is why \"100–500 verified examples\" is the standard recommendation rather than a compromise: below 100 your interval swallows any regression short of a disaster, and beyond 500 you are paying maintenance for precision you will not act on. If you genuinely need to detect a 2-point regression, you need thousands of items and should be honest about that cost before promising the gate.\n\nNote also that this arithmetic applies **per bucket and per slice**. A 300-item set that reports one global number is really five 60-item sets if your users split into five meaningful segments, and none of those slices can detect anything.\n\n## The four-bucket structure\n\nA golden set that is only a production sample will tell you the average case is fine while the tail burns. A golden set that is only adversarial cases will tell you the product is broken when it is working. Build four buckets and report them separately.\n\n| Bucket | What it contains | Target share | Sourced from |\n|---|---|---|---|\n| Production sample | Real inputs, sampled to reflect real traffic | ~40% | Logs and traces, PII removed |\n| Edge cases | Rare but legitimate inputs — long inputs, unusual formats, minority segments, non-English | ~25% | Segment analysis + user interviews |\n| Adversarial | Deliberate attempts to elicit failure | ~20% | [Red-team sessions](/docs/ai-red-teaming-with-users) |\n| Failure replays | Every confirmed past failure, permanently | ~15% | Incident reports, support tickets, prior evals |\n\nTwo construction rules matter more than the exact percentages. **Oversample then downsample**: pull roughly five times your target from production traces and downsample deliberately to balance coverage, rather than taking the first 300 rows — raw traffic is dominated by the easy majority case. And **synthetic items extend a production-grounded core; they never replace it.** Synthetic expansion is legitimate for underrepresented scenarios you cannot yet observe. It is illegitimate as the base, because a model evaluated only on inputs another model invented is being graded on its own species' handwriting.\n\n## Where acceptance criteria come from\n\nHere is the step that separates an eval set that predicts user satisfaction from one that predicts nothing. Each item needs an expected behaviour, and for most generative tasks there is no single correct string — there is a space of acceptable outputs and a boundary. That boundary is an empirical fact about users, and you find it by asking them.\n\nA workable item schema:\n\n| Field | Purpose |\n|---|---|\n| `input` | The user's actual request, verbatim |\n| `context` | System state, retrieved documents, prior turns, user segment |\n| `acceptance_criteria` | 3–6 binary criteria that must all hold |\n| `unacceptable_behaviours` | Named failure modes for this item |\n| `severity` | Cost if this item fails, on a fixed scale |\n| `provenance` | Where it came from, and who verified it |\n| `version_added` | For tracking set drift over time |\n\nDecomposing a fuzzy judgement into binary criteria is the highest-leverage move available. \"Is this summary good?\" produces inter-rater agreement in the noise. \"Does it state the decision that was made? Does it name every attendee who spoke? Does it avoid asserting anything not in the transcript? Is it under 200 words?\" produces agreement you can act on, and it converts directly into an automated check. The same rubric-anchoring discipline described in [human evaluation of AI outputs](/docs/human-evaluation-ai-outputs) applies here — the difference is that you are freezing the rubric onto specific items rather than sampling fresh ones.\n\nWhere the criteria come from, in order of reliability:\n\n1. **Interviews with users doing the real job.** Show them real outputs, ask what they would accept, probe why they rejected what they rejected. The rejection reasons are your `unacceptable_behaviours` list.\n2. **Domain experts** for anything with a professional standard — clinical, legal, financial, safety.\n3. **Support tickets and complaints** — the cheapest available record of what users refuse to accept.\n4. **The build team's intuition** — useful for a first draft, never as the final authority.\n\n## The seven-step build\n\n1. **Define the unit.** One item = one input the system must handle, with its context. Ambiguity here poisons everything downstream.\n2. **Pull 5× your target from production.** Strip PII before it leaves the logging system, not after.\n3. **Stratify and downsample** by user segment, task type, input length, and difficulty. Record the strata — you will need them for slice-level reporting.\n4. **Run user sessions to write acceptance criteria.** Structured interviews on a sample of items; generalise the criteria to the rest.\n5. **Have two independent people label the full set**, and measure agreement. Kappa below 0.4 means your criteria are ambiguous — rewrite them, do not add raters. Target 0.7 or above. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research).\n6. **Adjudicate disagreements** with a third rater, and treat every adjudication as a bug report against the criteria.\n7. **Freeze, version, and seal a holdout.** Keep 20% sealed and unused for routine gating, to detect the day your team starts optimising against the visible set.\n\n## Maintaining it: drift, decay and overfitting\n\nGolden sets rot in three distinct ways, and each has a different fix.\n\n- **Distribution drift.** Real usage moves; your production bucket stops representing traffic. Fix: re-sample the production bucket quarterly, keeping the failure-replay bucket permanent.\n- **Criteria drift.** Your product's definition of good changes. Fix: version the criteria alongside the items, and never silently rewrite a criterion — a changed criterion invalidates historical comparisons, so record it as a new version.\n- **Overfitting.** The team has seen the set enough times to fix its specific items rather than the underlying behaviour. Fix: the sealed holdout, plus a rule that any item fixed by a special case rather than a general improvement gets a sibling item added.\n\nA useful discipline: your failure-replay bucket should only ever grow. Removing a past failure from the set because \"we fixed that\" is exactly how it comes back.\n\n## Privacy: production traces are personal data\n\nReal user inputs pulled from logs are almost always personal data, and moving them into an evaluation dataset is a new purpose that needs a lawful basis. Three practical requirements: strip identifiers at the logging boundary rather than downstream; check whether your privacy notice and terms actually cover evaluation use; and make sure your deletion process reaches eval sets, which are frequently the copy everyone forgets — see [DSAR and research data](/docs/dsar-research-data). Where you recruit users to write criteria, that is research and needs [informed consent](/docs/research-ethics-guide) like any other study.\n\n## The modern approach: sourcing golden sets with Koji\n\nThe bottleneck in every golden-set project is the same: you need dozens of real users to look at real outputs and tell you, in enough detail to be actionable, what they would and would not accept. Traditionally that means recruiting, scheduling, moderating and coding — weeks of calendar time for what is conceptually an hour of judgement per person.\n\nKoji collapses that into an async study. You send a link with real outputs embedded in the questions; the AI moderator collects the judgement, probes the reasoning behind each rejection in the participant's own words, and returns analysed transcripts. Sixty domain users generating acceptance criteria is a days-long study.\n\nThe [six structured question types](/docs/structured-questions-guide) map onto golden-set construction with unusual precision:\n\n| Question type | Golden-set use |\n|---|---|\n| `yes_no` | The acceptance criteria themselves — each decomposed binary criterion is one question, and the aggregate is a pass rate |\n| `single_choice` | Failure-type classification against your taxonomy |\n| `ranking` | Pairwise and n-way preference between candidate outputs, for items with no single correct answer |\n| `scale` | Severity: what it would cost this user if this output shipped |\n| `multiple_choice` | Which acceptance criteria this specific output violated |\n| `open_ended` | *Why* they rejected it — the reasoning an automated judge structurally cannot produce, and the source of your `unacceptable_behaviours` list |\n\nBecause the answers are typed at capture rather than extracted from prose afterwards, criteria aggregate directly into the dataset schema instead of requiring a coding pass. Koji's thematic analysis clusters the free-text rejection reasons into candidate failure modes; quality scoring rates each conversation 1–5 against the research goal so a participant who clicked through without engaging is visible before their labels enter your ground truth. Voice interviews are the right modality when the output being judged is itself spoken.\n\n| | Traditional golden-set build | Koji |\n|---|---|---|\n| Who writes acceptance criteria | The build team, from intuition | Real users and domain experts |\n| Labellers | Contractors with no domain context | Screened people who do the job |\n| Time to a 300-item labelled set | 4–8 weeks | Days |\n| Rejection reasons captured | Rarely — just a label | Probed, in the participant's words |\n| Refresh cost per quarter | Repeat the whole process | Re-send the link |\n| Slice coverage | Whoever was available | Screened and quota'd by segment |\n\n## Common mistakes\n\n1. **Building the set from synthetic prompts.** Fastest route to an eval that passes while users churn.\n2. **One global pass rate.** Report by bucket and by user slice, or the tail stays invisible.\n3. **Letting the build team write ground truth alone.** They will encode the behaviour they implemented.\n4. **Never measuring label agreement.** An unmeasured set has an unknown error rate, and published benchmarks suggest it is not small.\n5. **Growing the set instead of fixing the criteria.** If raters disagree, more items will not help.\n6. **No sealed holdout.** Within two quarters the visible set becomes a training target rather than a test.\n7. **Forgetting the set in your deletion pipeline.** It is personal data, and it is the copy nobody remembers.\n\n## Frequently asked questions\n\n**How many examples does a golden dataset need?**\nFor most product evals, 100–500 verified examples per bucket. Below 100, the 95% confidence interval on a 90% pass rate is about ±6 points, which means anything short of a catastrophic regression is invisible. Above 500, the marginal example stops justifying its maintenance cost. If you must detect a 2–3 point regression, you need low thousands of items and should scope that honestly.\n\n**Can I use synthetic data to build my eval set?**\nAs an extension, yes — synthetic items are a reasonable way to cover underrepresented scenarios you cannot yet observe in production. As a base, no. A model graded only on inputs another model invented is being tested on a distribution your users do not occupy.\n\n**How often should I refresh the golden set?**\nRe-sample the production bucket quarterly, or whenever a major feature changes what users ask for. The failure-replay bucket should only grow, never shrink. Version any change to acceptance criteria explicitly, because a silent criterion rewrite invalidates every historical comparison.\n\n**What agreement level should labellers reach?**\nTarget Cohen's kappa of 0.7 or above. Between 0.4 and 0.6 the criteria are weak; below 0.4 they are ambiguous. The fix is almost always rewriting the criteria into sharper binary statements, not recruiting more raters.\n\n**Who should write the acceptance criteria?**\nPeople who represent the population that will accept or reject the output in real use — users doing the actual job, plus domain experts where a professional standard applies. The build team's intuition is a good first draft and a bad final authority, because it encodes the behaviour that was implemented rather than the behaviour that is wanted.\n\n**Can an LLM label my golden set for me?**\nIt can propose labels and flag likely errors, and that is a genuine time saver. It should not be the final authority on the set that defines correctness — the set exists precisely to check the models. Note that when label errors were flagged algorithmically in the Northcutt study, only about half were confirmed erroneous by human review.\n\n## Related resources\n\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) — the rubric, rater and agreement mechanics your labelling pass depends on\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) — what to automate once the golden set exists\n- [AI Red Teaming with Real Users](/docs/ai-red-teaming-with-users) — where the adversarial bucket comes from\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that turn user judgement into a dataset schema\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — measuring whether your criteria are actually unambiguous\n- [User Research for AI Products](/docs/user-research-for-ai-products) — trust calibration and failure tolerance, the qualitative half of AI product research\n- [Synthetic Users in Research](/docs/synthetic-users-research-methodology) — where AI-generated respondents are and are not valid","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Building an AI Evaluation Dataset: The Golden Set Guide (2026)","metaDescription":"How to build and maintain an LLM evaluation golden set: sizing, confidence intervals, the four-bucket structure, and sourcing acceptance criteria from users.","keywords":["ai evaluation dataset","golden dataset llm","eval set construction","ground truth dataset ai","llm evaluation data","golden set size","ai regression testing dataset","acceptance criteria ai product","label errors benchmark","building evals from user research"],"aiSummary":"A golden evaluation dataset is a product specification written as examples. Covers sizing against confidence intervals (100-500 items per bucket), the four-bucket structure of production sample, edge cases, adversarial items and failure replays, documented label-error rates in published benchmarks, an item schema with decomposed binary acceptance criteria, sourcing those criteria from real users rather than the build team, inter-rater agreement targets, and maintaining the set against distribution drift, criteria drift and overfitting.","aiPrerequisites":["Basic understanding of how AI product evaluation works","Familiarity with confidence intervals and sampling"],"aiLearningOutcomes":["Size an evaluation set against the smallest regression you need to detect","Structure a golden set into production, edge-case, adversarial and failure-replay buckets","Decompose fuzzy quality judgements into binary acceptance criteria","Source acceptance criteria from real users instead of the build team","Measure and fix labelling disagreement","Maintain the set against distribution drift, criteria drift and overfitting"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min"},{"type":"documentation","id":"ce4232ed-ab13-4f77-9e4d-d293d9a6c21c","slug":"ux-audit-guide","title":"UX Audit: A Step-by-Step Guide to Finding What Is Hurting Your Product","url":"https://www.koji.so/docs/ux-audit-guide","summary":"A UX audit is a structured evaluation of a digital product that identifies usability problems hurting conversions and retention, ranked by severity. It uses two lenses — heuristic evaluation (expert review against Nielsen's 10 heuristics) and behavioral data — but neither explains why users struggle. Forrester found every $1 invested in UX returns $100. Three to five evaluators find about 75-85% of usability problems.","content":"\n## What Is a UX Audit?\n\nA UX audit is a structured evaluation of a digital product that identifies the usability problems, friction points, and design flaws quietly costing you conversions, retention, and customer satisfaction. It combines expert review against established usability principles with behavioral data and real user feedback, then ranks every issue by severity so the team knows exactly what to fix first.\n\nUnlike a casual design critique, a UX audit is systematic and evidence-based. It answers one specific question: *where, precisely, is our product failing its users — and which of those failures matters most?*\n\nThe business case is hard to argue with. Forrester Research found that, on average, every $1 invested in UX returns $100 — a 9,900% return on investment. Teams that act on audit findings commonly report conversion lifts of 15–30% within 60–90 days when the fixes target high-traffic flows. A UX audit is how you locate where that return is hiding.\n\n## Why UX Audits Matter\n\nSmall design flaws compound into large revenue losses. The most famous illustration is the \"$300 million button,\" documented by usability expert Jared Spool and User Interface Engineering: a major e-commerce retailer was losing customers at a single checkout step that forced people to register before buying. Replacing the \"Register\" button with a \"Continue\" button — and letting people register later — increased that site's revenue by an estimated $300 million in the first year. One label. One assumption never tested.\n\nThe cost of finding these problems late is equally well documented. Under the 1:10:100 rule cited in Dr. Susan Weinschenk's *The ROI of User Experience*, a usability problem costs about $1 to fix in design, $10 in development, and $100 after release. A UX audit is a structured way to catch issues on the cheap side of that curve.\n\n## The Two Lenses of a UX Audit\n\nA thorough audit looks through two different lenses, because each one sees something the other cannot.\n\n### Heuristic Evaluation\nHeuristic evaluation is expert review of an interface against a set of established usability principles — most commonly Jakob Nielsen's 10 Usability Heuristics, which cover principles such as visibility of system status, error prevention, consistency, and user control. An evaluator walks the product methodically and logs every violation.\n\nThe Nielsen Norman Group quantifies how many evaluators you need: a single evaluator typically finds about 35% of an interface's usability problems, three evaluators working independently find roughly 75%, and five find about 85%. Beyond five, returns diminish sharply. This is why a credible audit uses three to five evaluators rather than one person's opinion.\n\n### Behavioral Data\nThe second lens is what users actually do. Analytics funnels, drop-off rates, heatmaps, session recordings, and event tracking reveal where real users hesitate, abandon, rage-click, or loop. Behavioral data does not care about best practice — it shows the ground truth of where attention and intent are being lost.\n\nThe two lenses answer different questions. Heuristic evaluation tells you *what is wrong* against best practice. Behavioral data tells you *where* users actually struggle. Neither tells you *why* — and that gap is the one most audits never close.\n\n## The UX Audit Process, Step by Step\n\n1. **Define scope and goals.** Audit a specific journey — onboarding, checkout, a core workflow — not \"the whole product.\" Tie the audit to a business metric: activation, conversion, retention.\n2. **Gather context.** Collect business goals, target personas, prior research, support tickets, and analytics access. An audit without context produces generic findings.\n3. **Run the heuristic evaluation.** Have three to five evaluators independently review the flow against Nielsen's heuristics and log every issue with a screenshot and the principle it violates.\n4. **Analyze the behavioral data.** Map funnels, drop-off points, and session recordings onto the same flow. Look for where the quantitative cliff appears.\n5. **Talk to real users.** Heuristics and analytics tell you what and where. Real users tell you why. This step is covered in detail below.\n6. **Rate severity.** Score every issue so the report drives action rather than overwhelming the team.\n7. **Prioritize and report.** Deliver a ranked list of issues with clear, specific recommendations — not a 60-page document nobody reads.\n\n## Severity Rating: Not All Problems Are Equal\n\nA list of 80 issues with no priority is useless. The Nielsen Norman Group's severity scale rates each issue from 0 (not a real problem) to 4 (a usability catastrophe that must be fixed before release). The rating combines three factors: **frequency** (how often users hit it), **impact** (how hard it is to overcome when they do), and **persistence** (whether users learn to work around it or get stuck every time).\n\nPrioritize by severity against effort. A high-severity, low-effort fix is the first thing on the roadmap; a low-severity, high-effort fix may never be worth doing.\n\n## The Missing Layer: Why Users Actually Struggle\n\nHere is the limitation that weakens most UX audits. Heuristic evaluation gives you expert opinion. Behavioral data gives you the *what* and the *where*. But a drop-off chart cannot tell you whether users abandoned checkout because the form was confusing, because they were comparison shopping, because shipping cost shocked them, or because they never intended to buy. The audit shows the symptom; only users can explain the cause.\n\nTraditionally, closing this gap means recruiting participants, scheduling interviews, moderating sessions, and manually analyzing transcripts — days or weeks of work that most audit timelines simply do not allow. So the \"why\" step gets skipped, and teams ship fixes based on guesses.\n\n## How Koji Completes Your UX Audit\n\nKoji is an AI-native research platform that supplies the missing layer — real user voice — at the speed an audit demands.\n\n**Interview real users at scale, in hours.** Koji's AI-moderated interviews run asynchronously by voice or text. Point them at the exact flow your audit flagged — \"walk us through your last checkout\" — and the AI interviewer probes every short answer with a follow-up, so you learn *why* the friction exists, not just that it does. Dozens of interviews complete in the time a traditional study schedules its first session.\n\n**Get themed findings automatically.** Koji's automatic thematic analysis reads every transcript and clusters the recurring reasons behind friction, so the \"why\" arrives as organized themes instead of a folder of recordings. Teams using AI-assisted analysis report dramatically faster time-to-insight.\n\n**Quantify the qualitative.** Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. In a post-audit study you can capture a 1–5 scale ease rating, a single-choice \"what stopped you,\" and an open-ended explanation — turning subjective friction into numbers you can track before and after a fix. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Triangulate three sources.** The strongest audit overlays three layers: heuristic findings (what violates best practice), behavioral data (where users drop), and Koji interviews (why they drop). Where all three point at the same screen, you have a fix you can defend to any stakeholder.\n\nWhile traditional survey tools like SurveyMonkey collect static answers and stop, an AI-native platform like Koji conducts an actual conversation — and Koji democratizes that capability, so a PM or designer can run rigorous user research without a dedicated research team.\n\n## UX Audit Mistakes to Avoid\n\n- **Auditing everything at once.** A whole-product audit produces a shapeless list. Scope to one journey tied to one metric.\n- **Relying on a single evaluator.** One person finds about 35% of issues. Use three to five.\n- **Stopping at heuristics.** Expert opinion without behavioral data and user voice is just opinion.\n- **Skipping severity ratings.** An unranked list of issues does not drive a roadmap.\n- **Writing recommendations that are vague.** \"Improve the onboarding\" is not actionable. \"Cut the signup form from 9 fields to 4\" is.\n\n## UX Audit vs. Usability Testing vs. Heuristic Evaluation\n\nThese three terms are often used interchangeably, and the confusion leads teams to run the wrong study. A **heuristic evaluation** is a single method: expert review of an interface against usability principles, with no real users involved. **Usability testing** is a different method: watching real users attempt real tasks to see where they struggle. A **UX audit** is the umbrella study that orchestrates both — and layers behavioral analytics, severity rating, and prioritized recommendations on top.\n\nPut plainly: heuristic evaluation and usability testing are ingredients; the UX audit is the meal. A heuristic evaluation alone gives you expert opinion. A usability test alone gives you observed behavior on a narrow set of tasks. The audit combines expert review, behavioral data, and real user voice into one ranked, decision-ready picture of where the product is failing — which is why an audit is what stakeholders ask for when they want a clear answer to \"what should we fix first?\"\n\n## How Often Should You Run a UX Audit?\n\nA UX audit is not a one-time event. Products drift: features ship, teams change, and small compromises accumulate into the kind of friction the original audit was meant to remove. Most teams benefit from a focused audit of their highest-value flow once or twice a year, plus a targeted mini-audit whenever a major redesign ships or a key metric moves the wrong way. Treating the audit as a recurring health check — rather than an emergency response — keeps usability debt from compounding to the point where a full rebuild becomes the only option.\n\n## Related Resources\n\n- [Heuristic Evaluation Guide](/docs/heuristic-evaluation-guide) — expert review against Nielsen's 10 heuristics\n- [Usability Testing Guide](/docs/usability-testing-guide) — watching real users complete tasks\n- [Structured Questions Guide](/docs/structured-questions-guide) — Koji's six question types for quantifying friction\n- [Customer Pain Points Research](/docs/customer-pain-points-research) — surfacing the problems behind the metrics\n- [System Usability Scale Guide](/docs/system-usability-scale-guide) — benchmarking usability with a standardized score\n- [Cognitive Walkthrough Guide](/docs/cognitive-walkthrough-guide) — evaluating a product from a first-time user's perspective\n","category":"Research Methods","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"UX Audit Guide: Find What Is Hurting Your Product | Koji","metaDescription":"Learn how to run a UX audit step by step — heuristic evaluation, behavioral data, severity rating, and adding real user voice with AI to find what is hurting conversions.","keywords":["ux audit","ux audit guide","how to do a ux audit","ux audit process","ux audit checklist","heuristic evaluation","usability audit","website ux audit"],"aiSummary":"A UX audit is a structured evaluation of a digital product that identifies usability problems hurting conversions and retention, ranked by severity. It uses two lenses — heuristic evaluation (expert review against Nielsen's 10 heuristics) and behavioral data — but neither explains why users struggle. Forrester found every $1 invested in UX returns $100. Three to five evaluators find about 75-85% of usability problems.","aiPrerequisites":["Basic understanding of UX and usability principles","Access to product analytics is helpful"],"aiLearningOutcomes":["Run a UX audit using heuristic evaluation and behavioral data","Rate and prioritize usability issues by severity","Build a ranked, actionable audit report","Use Koji interviews to uncover why users struggle, not just where"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"8956a132-6f05-4415-ae5a-8c085dc2783f","slug":"research-democratization-playbook","title":"Research Democratization: The 2026 Playbook for Scaling User Research Across Your Whole Organization","url":"https://www.koji.so/docs/research-democratization-playbook","summary":"Research democratization is the distribution of customer contact, methodology access, and repository access across product managers, designers, marketers, and customer success teams — without dissolving research craft. The 2026 winning model is hub-and-spoke: a small central team owns standards, quality, and ethics; distributed spokes run studies on an AI-native platform that enforces methodology automatically. Industry data shows 64% of companies now have a democratized research culture and 66% of teams report increasing research demand. Common failure modes are leading questions, n=3 conclusions, and repository sprawl — all of which AI moderation, auto-coding, and shared repositories mitigate. The right rule of thumb: democratize everyday research (discovery, usability, churn, feature feedback) and centralize foundational, ethnographic, regulated, or sensitive research. Koji is purpose-built for the spoke role with methodology-aware briefs, AI moderation, structured questions, and quality scoring.","content":"## The short answer\n\nResearch democratization is the practice of equipping product managers, designers, marketers, customer success teams, and founders to **conduct customer research themselves** — without waiting for a dedicated researcher to run every study. In 2026, **64% of companies now operate a democratized research culture** [(Maze)](https://maze.co/blog/future-user-research-2026/), and the share of organizations where research is essential to all levels of business strategy has nearly tripled in a single year, from 8% in 2025 to 22% in 2026. The hard part is doing it without losing rigor. The winning model in 2026 pairs **a small central research team that owns standards and quality** with **an AI-native platform like Koji that handles the moderation, analysis, and synthesis** non-researchers should not be doing by hand. The result: 10x more research output, with quality that holds.\n\n## What research democratization actually means (and what it does not)\n\nResearch democratization is widely misunderstood. It does not mean \"everyone is a researcher now.\" That misunderstanding is what produces the failure mode every UX team has seen — PMs running leading interviews, designers cherry-picking quotes, marketers writing yes/no questions that confirm what they already believe.\n\nThe more accurate framing comes from Teresa Torres: *\"I don't love the term research democratization because I don't think we're turning everybody into a researcher. What we're doing is creating truly customer-centric organizations by letting everybody interact with the customer.\"* [(Teresa Torres on Looppanel)](https://www.looppanel.com/blog/objections-user-research-teresa-torres).\n\nThat is the right mental model. Democratization is the **distribution of customer contact**, not the dissolution of research craft. Three things get distributed:\n\n1. **Access to customers** — PMs can talk to users without scheduling through a researcher.\n2. **Access to the research workflow** — designers can launch a study, see results, share findings.\n3. **Access to the insight repository** — anyone can query past research instead of re-running it.\n\nThree things stay centralized:\n\n1. **Methodology standards** — which framework fits which question, how to write non-leading prompts, how to interpret quality signals.\n2. **Quality gates** — what counts as a finding vs. an anecdote, how many interviews constitute evidence, what to do with low-quality sessions.\n3. **Ethics, consent, and data handling** — especially under GDPR, HIPAA, and SOC 2.\n\n## Why democratization is happening now\n\nThree forces converged in 2024–2026:\n\n**Demand outran researcher headcount.** In the most recent industry survey, **66% of teams report that demand for user research has increased over the past 12 months** — and only a fraction can hire fast enough to keep up [(Maze, 2026)](https://maze.co/blog/future-user-research-2026/). The same survey found that 75% of teams plan to scale research, with the top three tactics being increasing study volume (51%), leveraging AI tools (31%), and training non-researchers (30%).\n\n**AI made the hard parts safe.** Pre-2023, asking a PM to \"go run an interview\" meant accepting bias, leading questions, inconsistent moderation, and shallow probing. In 2026, an AI-moderated interview from a platform like Koji is more consistent than the average human-moderated session. The Mom Test, JTBD, and customer discovery methodologies are baked into the discussion guide generator, so the questions are non-leading before they ever reach a participant.\n\n**Continuous discovery became the norm.** Teresa Torres's \"continuous discovery habits\" framework — interviewing customers weekly, every week — cannot work if every interview routes through a single research team. As Torres puts it, continuous research is *\"like putting money in a bank. Small, quick insights have a compounding effect over time.\"* The compounding only works if the cadence is distributed.\n\nLogRocket's 2026 trend report puts the point bluntly: *\"In 2026, designers are now conducting more research than dedicated UX researchers, and product managers are not far behind.\"* [(LogRocket)](https://blog.logrocket.com/ux-design/ux-research-trends-2026/).\n\n## The three failure modes of bad democratization\n\nDemocratization fails in three predictable ways. Naming them up front is the easiest way to prevent them.\n\n**1. The leading-question epidemic.** Untrained interviewers ask *\"What do you like about our pricing?\"* instead of *\"Walk me through the last time you renewed.\"* AI-moderated interviews with built-in [Mom Test methodology](/docs/mom-test-methodology) eliminate this entirely — the AI is constitutionally incapable of asking the leading version.\n\n**2. The \"n=3 is a finding\" problem.** Non-researchers default to acting on the first compelling quote they hear. Without quality gates and aggregation, three loud customers reshape the roadmap. The fix is structural: require thematic analysis (auto-handled in Koji) before insights count, and surface frequency counts alongside every quote. See [understanding themes and patterns](/docs/understanding-themes-patterns).\n\n**3. The research repository wasteland.** Democratization without a repository means studies live in 47 different Notion pages and nobody knows what is already known. The result is duplicated research, contradictory findings, and total loss of organizational memory. The fix is one canonical repository — see [insight repository methodology](/docs/insight-repository-methodology) and the [Notion integration](/docs/notion-research-integration).\n\nQualz.ai's 2025 analysis names the trade-off directly: *\"When done right, democratization speeds up time to insights, scales research capacity without proportional headcount, and builds a truly user-centered culture. However, without proper training, clear roles, and quality guardrails, you risk unreliable insights and wasted effort.\"* [(Qualz.ai)](https://qualz.ai/blog/research-democratization-empower-product-teams).\n\n## The 2026 hub-and-spoke democratization model\n\nThe model that works in 2026 has a clear shape:\n\n**The hub** is a small central research team (often 1–3 people, sometimes called a Research Ops function). They own:\n- Methodology standards and templates ([research brief templates](/docs/research-brief-template), [discussion guide templates](/docs/discussion-guide-template))\n- Quality gates and [quality scoring](/docs/understanding-quality-scores)\n- Ethics, consent, and compliance\n- The [insight repository](/docs/insight-repository-methodology)\n- Training and coaching for the spokes\n- The hardest, highest-stakes studies\n\n**The spokes** are PMs, designers, marketers, customer success, founders. They:\n- Launch studies from approved templates\n- Recruit from their own segments (with shared screener libraries)\n- Run AI-moderated interviews via the shared platform\n- Synthesize findings against the central methodology\n- Contribute findings back to the repository\n\n**The platform** is what makes the spokes safe. In 2026, the platform layer typically includes:\n- AI-moderated interviews (Koji) so moderation quality is constant across every spoke\n- Auto-coded thematic analysis so synthesis does not require trained researchers\n- Structured questions ([6 question types](/docs/structured-questions-guide)) so studies produce comparable quantified signals\n- Templates for the 10–15 most common study types so spokes do not start from scratch\n- A shared repository so duplicate research is visible before it happens\n\nThis hub-and-spoke model is exactly what platforms like Maze, Dovetail, and Koji are now optimized for. The differences come down to depth of methodology support, AI moderation quality, and integration with the rest of the operational stack.\n\n## What gets democratized — and what does not\n\nA practical rule of thumb in 2026:\n\n| Activity | Democratize? | Why |\n|---|---|---|\n| Customer discovery interviews for a new feature | Yes | Methodology baked into AI moderation; PMs need this weekly |\n| Usability testing of a prototype | Yes | High volume, low risk, well-templated |\n| Pricing research on a specific segment | Yes | If using structured questions for triangulation |\n| Brand tracker / NPS pulse | Yes | Operational, recurring, auto-analyzable |\n| Win-loss analysis | Mostly yes | With central oversight on synthesis |\n| Ethnographic / field research | No | Requires craft and observation skills |\n| Highly sensitive populations (medical, trauma) | No | Human moderation, ethics review |\n| Foundational research that resets product strategy | Partial | Spokes can contribute, researcher leads |\n| Regulated / IRB-required research | No | Compliance demands credentialed researchers |\n\nThe asymmetry to remember: **everyday research democratizes well, foundational research does not.** Erika Hall's framing in *Just Enough Research* still applies — match the rigor to the decision.\n\n## How Koji makes democratization safe\n\nKoji is built for the hub-and-spoke world. The features that matter most for democratization:\n\n- **Methodology-aware brief generator.** Every study starts from one of five built-in research frameworks (Mom Test, Jobs to Be Done, Customer Discovery, Exploratory, Lead Magnet). Spokes cannot accidentally invent a methodology that does not match the question.\n- **AI-moderated interviews with constant quality.** Every interview is moderated identically. No leading questions. No moderator fatigue at session 30 of 100.\n- **Structured questions.** Six structured types ([guide](/docs/structured-questions-guide)) — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — let non-researchers triangulate qualitative depth with quantified signals.\n- **Quality scoring (1–5).** Every session gets a quality score so weak data is flagged before it pollutes the aggregate.\n- **Auto-coded thematic analysis.** Spokes do not need to know how to build a [qualitative codebook](/docs/qualitative-research-codebook) — themes emerge automatically and link back to verbatim quotes.\n- **Insights chat across the corpus.** Stakeholders can [chat with transcripts](/docs/chat-with-interview-transcripts-ai) to answer their own questions instead of pinging the research team.\n- **Repository integrations.** [Notion](/docs/notion-research-integration), [Slack](/docs/slack-research-insights-integration), [Linear](/docs/linear-research-integration), and [Jira](/docs/jira-research-integration) connectors push findings into the systems each spoke already lives in.\n- **Customizable AI consultant.** A team-specific AI consultant trained on your business context that interprets findings the way your central research team would.\n\nThe net effect: a designer at week 4 can ship a study with the same methodological quality as a senior researcher would have produced in week 4 of last year — in a fraction of the time.\n\n## A 90-day rollout plan\n\n**Days 1–14: Foundation.** Define the hub (1–3 people, even if part-time). Pick 5 study templates that cover 80% of expected demand: discovery interview, usability test, churn interview, feature feedback, NPS follow-up. Stand up the [insight repository](/docs/insight-repository-methodology).\n\n**Days 15–30: Pilot.** Pick 3 spokes — typically a senior PM, a senior designer, and a customer success lead. Train them on the templates. Each runs one study end-to-end. The hub coaches, does not run.\n\n**Days 31–60: Expand.** Open access to 10–15 spokes. Publish a methodology playbook. Hold weekly office hours. Track quality scores by spoke as a coaching signal, not a punishment.\n\n**Days 61–90: Institutionalize.** Quarterly research summit where spokes share findings. KPIs ([research program KPIs guide](/docs/user-research-program-kpis)) wired to product OKRs. Repository becomes the default starting place for any new question.\n\nBy day 90, the hub is doing fewer studies and more enablement. Total research output is 5–10x. Quality, measured by quality scores and decisions-influenced, is flat or up.\n\n## Common questions from skeptical research teams\n\nResearch teams sometimes resist democratization out of (legitimate) concern that quality will collapse and they will be deprecated. Two reframes help:\n\n**\"Will democratization put researchers out of a job?\"** No. It changes the job. The senior researcher in a mature democratized org spends less time moderating and more time setting methodology, training spokes, designing the hardest studies, and turning findings into strategic direction. That is a more senior role, not a less senior one.\n\n**\"Won't spokes produce bad research?\"** They will, if the platform is bad. With AI moderation, structured questions, and auto-coded analysis, the floor is much higher than it was in 2022. The hub's job is to raise the ceiling further with coaching and quality gates.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the 6 question types that make democratized studies measurable\n- [Insight Repository Methodology](/docs/insight-repository-methodology) — building the shared memory\n- [User Research Program KPIs](/docs/user-research-program-kpis) — how to measure a democratized program\n- [User Research Maturity Model](/docs/user-research-maturity-model) — where you are on the curve\n- [Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) — selling the practice internally\n- [Continuous Discovery Tools 2026](/docs/continuous-discovery-tools-2026) — the platform layer\n- [Self-Service Research Program](/docs/self-service-research-program) — operating model details\n\n## Sources\n\n- [Maze — The Future of User Research Report 2026](https://maze.co/blog/future-user-research-2026/)\n- [Dovetail — Scaling User Research Through Democratization](https://dovetail.com/customer-research/scaling-user-research-through-democratization/)\n- [Qualz.ai — Research Democratization Done Right](https://qualz.ai/blog/research-democratization-empower-product-teams)\n- [LogRocket — UX Research Trends 2026](https://blog.logrocket.com/ux-design/ux-research-trends-2026/)\n- [Looppanel — Teresa Torres on Objections to User Research](https://www.looppanel.com/blog/objections-user-research-teresa-torres)\n- [User Interviews — How Teams Do Continuous Discovery Research Today](https://www.userinterviews.com/blog/continuous-discovery-research-report)","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"Research Democratization: The 2026 Playbook for Scaling User Research | Koji","metaDescription":"How to democratize user research without losing rigor: the hub-and-spoke model, what to democratize vs. centralize, AI guardrails, and a 90-day rollout plan.","keywords":["research democratization","democratize user research","scaling user research","ux research democratization","self-service research","research operations","user research scaling","democratized research culture","continuous discovery","research enablement","distributed research","research ops 2026"],"aiSummary":"Research democratization is the distribution of customer contact, methodology access, and repository access across product managers, designers, marketers, and customer success teams — without dissolving research craft. The 2026 winning model is hub-and-spoke: a small central team owns standards, quality, and ethics; distributed spokes run studies on an AI-native platform that enforces methodology automatically. Industry data shows 64% of companies now have a democratized research culture and 66% of teams report increasing research demand. Common failure modes are leading questions, n=3 conclusions, and repository sprawl — all of which AI moderation, auto-coding, and shared repositories mitigate. The right rule of thumb: democratize everyday research (discovery, usability, churn, feature feedback) and centralize foundational, ethnographic, regulated, or sensitive research. Koji is purpose-built for the spoke role with methodology-aware briefs, AI moderation, structured questions, and quality scoring.","aiPrerequisites":["Some existing user research practice (even informal)","Stakeholders outside the research team who want customer access","Buy-in to invest in research operations"],"aiLearningOutcomes":["Define research democratization correctly (and what it is not)","Identify the three failure modes that kill democratization programs","Apply the hub-and-spoke model with clear hub vs. spoke responsibilities","Decide what to democratize vs. centralize based on study type","Run a 90-day democratization rollout end-to-end"],"aiDifficulty":"intermediate","aiEstimatedTime":"16 min read"},{"type":"documentation","id":"002abbe9-05d7-4991-9c5c-4adeddccc3e7","slug":"hiring-ux-researcher-guide","title":"How to Hire a UX Researcher: When You Need One, the Job Description, and the Interview Loop (2026)","url":"https://www.koji.so/docs/hiring-ux-researcher-guide","summary":"A decision guide for hiring managers weighing a dedicated user researcher against AI-native research tooling. Covers the four-signal threshold test (volume, stakes, ambiguity, absorption), fully loaded 2026 cost math, a job description template by level, a five-stage interview loop with 25 questions, portfolio and craft-exercise evaluation, and the hire-vs-tool decision matrix.","content":"## The short answer\n\n**Hire a dedicated user researcher when research demand is continuous rather than episodic, when the decisions in play cost more to get wrong than the role costs to fill, and when a senior owner exists who will actually change the roadmap based on findings.** If fewer than three of those are true, you do not have a headcount problem — you have a throughput problem, and throughput is now cheaper to buy than to staff.\n\nThat framing is not a cost-cutting argument. It is a sequencing argument. The most common failure in research hiring is not hiring too late; it is hiring a researcher into an organization that has no mechanism for absorbing what the researcher finds. The role then gets measured on study volume, which is the one thing tooling does better, and the hire becomes vulnerable in the next planning cycle.\n\nThis guide covers the threshold test, the real 2026 cost of the role, a job description that attracts senior candidates, and a five-stage interview loop with 25 questions that test judgment instead of vocabulary.\n\n## What the 2026 research job market actually looks like\n\nHiring decisions made on 2021 assumptions will be wrong. Four data points define the current market:\n\n- **Research roles were cut disproportionately.** Tech layoffs totaled roughly 152,922 employees in 2024 and 122,549 in 2025, and user research was cut harder than most adjacent functions at several large technology companies. The practical effect for hiring managers: senior researchers are available in a way they were not three years ago.\n- **Practitioner sentiment has inverted.** 49% of researchers now feel negative about the future of the discipline — a 26-point jump from 2024 — and 67% rate career opportunities poorly (User Interviews, *State of User Research*). Candidates will interrogate your commitment to the function harder than they used to.\n- **Most research is already being run by non-researchers.** 71% of organizations report having *people who do research* who are not dedicated researchers (User Interviews). Your candidate pool knows this, and the strongest candidates want to know whether they will be the person doing all the studies or the person raising the ceiling on everyone else's.\n- **Staffing ratios are thin by design.** Nielsen Norman Group survey data puts the most typical researcher-to-designer-to-developer ratio at **1:5:50**, up from roughly 1:5:100 a few years earlier. NN/g is explicit that ratios should not be used as maturity scores.\n\nNielsen Norman Group's 2026 outlook argues organizations will \"ask more of each role,\" compressing responsibilities once spread across specialists. The same analysis identifies what resists automation: taste, contextual understanding, critical thinking, and judgment. That is a precise description of what you should be hiring for — and, by omission, a description of what you should not be paying a salary to do.\n\n## The four-signal threshold test\n\nScore each signal honestly. Three or more means hire. Two or fewer means buy capability instead.\n\n### Signal 1 — Volume: is demand continuous?\n\nThe test is not \"do we have research questions.\" Everyone has research questions. The test is whether you have sustained **two or more studies per month across two consecutive quarters** that someone was willing to fund. Episodic demand — a burst before each major launch — is better served by tooling or a contractor than by a permanent role that will idle between peaks.\n\n### Signal 2 — Stakes: is a wrong answer expensive?\n\nEstimate the cost of the largest decision research would inform this year. A pricing change on $8M of ARR, a platform migration, or a new-segment bet all clear the bar easily. A copy change on an onboarding screen does not. If the annual sum of research-informed decision value is smaller than the fully loaded cost of the role, the role does not pay for itself yet.\n\n### Signal 3 — Ambiguity: do the questions need methodological judgment?\n\nSome research is template execution: post-launch satisfaction, churn exit interviews, usability passes on a known flow. Modern platforms handle those end to end. Other work genuinely requires someone who can choose between a Van Westendorp and a Gabor-Granger, design a segmentation that survives contact with the data, or structure a longitudinal study. If most of your questions are the first kind, you need throughput. If they are the second kind, you need a researcher.\n\n### Signal 4 — Absorption: will anyone act?\n\nThis is the signal teams skip and the one that determines whether the hire survives. Name the person who will change a roadmap based on a finding they did not expect. If you cannot name them, hiring a researcher creates evidence with no destination. Research at organizations that recently conducted layoffs reports leadership buy-in at just 40% (User Interviews) — absorption capacity is the scarce resource, not research capacity.\n\n## What the role actually costs\n\nBase salary is roughly two-thirds of the real number. The 2026 math:\n\n| Cost component | 2026 figure |\n| --- | --- |\n| Base salary, entry (0–3 yrs, US median) | ~$85,000 |\n| Base salary, mid (4–6 yrs) | ~$110,000 |\n| Base salary, senior (7–9 yrs) | ~$130,000 |\n| Base salary, 10+ yrs | ~$145,000 |\n| Director / senior director | ~$216,000 |\n| Fully loaded multiplier (benefits, taxes, equipment, software) | 1.25–1.4× base |\n| Cost per hire, specialized role | $10,000–$20,000 |\n| Time to fill | 42–63 days |\n\nSalary medians are from the User Interviews 2026 UX Salary Report, which also finds that **67% of US researchers earn over $100,000** and that the individual-contributor-to-manager pay gap maxes out around $26K globally until director level. Cost-per-hire and time-to-fill figures follow 2026 SHRM-aligned recruiting benchmarks, where specialized and senior roles run well above the ~$4,700 general average.\n\nA mid-level researcher therefore costs roughly **$150,000–$170,000 in year one**, arrives 6–9 weeks after you open the requisition, and needs another 4–8 weeks to build context before their first study lands. Budget for the gap: the first useful output typically appears in month three or four.\n\nTwo structural notes. US freelancers report a $125,000 median versus $39,000 outside the US, so distributed contracting changes the math substantially. And external agencies charge $15,000–$75,000 per study, with day rates of $800–$1,500 for mid-level independents and $1,500–$3,000 for senior specialists — which means roughly three to six agency studies cost the same as one year of an in-house hire.\n\n## The job description that attracts senior candidates\n\nMost research job descriptions fail the same way: they list methods. Method lists attract candidates who can recite methods. Senior researchers are attracted by the *problem*, the *decisions*, and the *authority*.\n\n**Structure that works:**\n\n1. **The decisions this role informs.** Name two or three real decisions on the roadmap. \"You will shape how we price the enterprise tier and whether we build a second onboarding path\" outperforms any list of responsibilities.\n2. **Who acts on the findings.** State the seniority of the stakeholders and the forum where research gets discussed. Candidates burned by shelf-ware read this line first.\n3. **The current state, honestly.** \"We have no repository, no panel, and a PM-run interview habit we want to raise the quality of\" is more attractive than implied maturity, because it defines the win.\n4. **Scope of the craft, by level.** Junior: execute studies against defined questions. Mid: own end-to-end studies and choose methods. Senior: frame the questions, set standards, and raise the quality of research run by non-researchers. Principal/Lead: own the research strategy and its link to business outcomes.\n5. **What you have automated.** Say plainly that recruiting, moderation, transcription, and first-pass synthesis run on an AI-native platform. This filters out candidates whose identity is tied to manual execution and attracts those who want leverage.\n6. **Compensation band.** Publishing the band shortens time-to-fill materially and is mandatory in an increasing number of jurisdictions.\n\n**Requirements to cut:** a specific degree, a fixed number of years, a named legacy tool, and \"PhD preferred\" unless the role genuinely involves experimental design. Each of those narrows the pool without improving the hire.\n\n## The interview loop: five stages, 25 questions\n\nDesign the loop to test judgment under constraint. Every candidate can define a Likert scale; very few can tell you what they would cut when the deadline halves.\n\n**Stage 1 — Screen (30 min).** Motivation, level calibration, compensation alignment.\n\n**Stage 2 — Craft interview (60 min).** Methodological judgment. Ask:\n\n1. A stakeholder wants to know \"why churn is up.\" Turn that into a research plan in four days with no recruiting budget.\n2. How do you decide sample size for a qualitative study? Talk me through a real case where you got it wrong.\n3. When would you *not* run research?\n4. Walk me through a study where the answer surprised you. What did you do about the surprise?\n5. How do you distinguish a preference from a behavior in what participants tell you?\n6. What is your process when qualitative findings contradict the analytics?\n7. How do you screen out participants who are not who they say they are?\n8. Which of your studies had the least impact, and why?\n9. How do you write a question that does not lead the participant?\n10. When is a survey the wrong instrument, and what do you do instead?\n\n**Stage 3 — Portfolio or work sample (60 min).** Ask for one study walkthrough covering question framing, method choice, sample, what changed as a result, and what they would do differently. Watch for candidates who describe process at length and outcome briefly — the ratio is diagnostic. Questions:\n\n11. What was the decision this study was meant to inform?\n12. Who disagreed with the findings, and how did you handle it?\n13. What did you deliberately leave out of the readout?\n14. How did you know when to stop interviewing?\n15. What evidence would have changed your conclusion?\n\n**Stage 4 — Stakeholder simulation (45 min).** Give a one-page findings summary and have the candidate present to a skeptical PM or exec. This predicts on-the-job impact better than any other stage. Probe:\n\n16. A VP says the sample is too small to matter. Respond.\n17. Engineering says the fix is six months. What is your next move?\n18. How would you make this finding survive the next planning cycle?\n19. What would you have needed to make this recommendation stronger?\n20. Summarize this study in two sentences for a board deck.\n\n**Stage 5 — Team and operating fit (45 min).** How they raise the ceiling for others:\n\n21. How would you help a PM run a competent interview without you in the room?\n22. What quality standard would you enforce on research you did not run?\n23. Where should AI moderation be used, and where should it never be used?\n24. What would you automate first here?\n25. What does your first 90 days look like if there is no repository and no panel?\n\n**On take-home exercises:** a scoped 90-minute analysis of an anonymized transcript set is fair and predictive. Asking a candidate to design and run a study on your actual product is unpaid labor and will cost you your strongest applicants.\n\n## The modern alternative: buy capability before headcount\n\nThe reason this decision changed is that the expensive parts of research stopped being the scarce parts. Recruiting scheduling, moderation, transcription, coding, and report assembly used to consume the majority of a researcher's week. On an AI-native platform they consume none of it.\n\n**How Koji changes the threshold math:**\n\n- **AI-moderated interviews run in parallel.** A traditional researcher runs 5–8 moderated sessions a week and spends the rest of the week on synthesis. Koji fields dozens of interviews simultaneously — voice or text — and completes them overnight rather than across a three-week scheduling window.\n- **Structured questions enforce instrument quality.** Koji supports six typed question formats — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — so an untrained moderator cannot accidentally lead a participant or drift off-protocol. Scale and choice questions aggregate quantitatively while open-ended questions get AI follow-up probing to a configured depth. This is the mechanism that makes democratized research safe: the rigor lives in the instrument, not in the moderator's training.\n- **Interviewer variance goes to zero.** Every participant gets the same question, asked the same way, with consistent probing. That removes the single largest quality difference between a trained researcher and a PM running their first study.\n- **Thematic analysis is automatic.** Themes, supporting quotes, and per-interview quality scores (a 1–5 scale) appear in real time as interviews complete, rather than after a week of manual coding.\n- **Custom AI consultants** carry your methodology — Mom Test, Jobs to be Done, discovery, exploratory — so the framework a senior researcher would apply is embedded in every study, whoever launches it.\n\nThe honest boundary: none of this decides *which question is worth asking*, defends a finding in a hostile planning meeting, or notices that the framing of a study is wrong. Those are the reasons to hire, and they are the reasons the hire should be senior when you make it.\n\n**The sequencing that works for most teams under 200 people:** buy the platform first, let PMs and designers run the high-volume repeatable studies, and hire a senior researcher once the *demand for judgment* — not the demand for studies — becomes continuous. Teams that do it in this order hire one senior person instead of two mid-level people, and the senior person walks into an organization with an existing evidence habit rather than having to create one.\n\n## Common hiring mistakes\n\n- **Hiring junior first.** A junior researcher in an organization with no research function has no one to learn from and no standards to inherit. The first hire should be the most senior you can afford.\n- **Measuring the hire on study count.** This is a throughput metric, and it puts a salaried human in direct competition with software. Measure decisions influenced.\n- **Skipping the stakeholder simulation.** Craft is table stakes; influence is the differentiator, and it is the only stage that tests it.\n- **Hiring before the absorption signal.** Research with no destination gets cut first in the next reorg.\n- **Treating the platform decision as a substitute for strategy.** Tooling raises throughput and consistency. It does not tell you what matters.\n\n## Frequently asked questions\n\n**When should a company hire its first UX researcher?**\nWhen at least three of four signals hold: continuous demand (~2+ studies/month across two quarters), decision stakes exceeding the fully loaded cost of the role, questions requiring methodological judgment, and a named owner who will act on unexpected findings.\n\n**What does a UX researcher cost in 2026?**\nUS medians run ~$85K entry, ~$110K mid, ~$130K senior, ~$145K at 10+ years (User Interviews 2026 UX Salary Report). Multiply by 1.25–1.4 for fully loaded cost, add $10K–$20K recruiting cost and 42–63 days time-to-fill.\n\n**What is a healthy researcher-to-designer ratio?**\nNN/g survey data puts the typical ratio at 1:5:50 (researcher:designer:developer). Use it as a reference point, not a target — it measures staffing, not research coverage.\n\n**Should I hire a researcher or buy AI research tooling?**\nBuy tooling if the constraint is throughput; hire if the constraint is knowing which question to ask. Most teams under 200 people should buy first and hire senior later.\n\n**Can non-researchers produce trustworthy findings?**\nYes, when the instrument is structured and moderation is consistent. 71% of organizations already have non-researchers running studies; typed structured questions and AI moderation are what keep quality from degrading.\n\n**What should I ask in a research interview?**\nConstraint questions — four days, no budget, hostile stakeholder, contradictory data. Definition questions about methods are answerable by everyone and predict nothing.\n\n## Related Resources\n\n- [Research Democratization Playbook](/docs/research-democratization-playbook) — how to let non-researchers run studies without losing rigor\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six typed question formats that keep instrument quality high regardless of who runs the study\n- [User Research Cost Calculator](/docs/user-research-cost-calculator-2026) — per-study spend, the other half of the budget math\n- [UX Research Operations (ResearchOps)](/docs/ux-research-ops) — the infrastructure a new hire will ask about in week one\n- [User Research for Product Managers](/docs/product-manager-research-guide) — what your PMs can run before you hire\n- [AI vs Human Moderators](/docs/ai-vs-human-moderators) — where AI moderation is appropriate and where it is not\n- [Getting Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) — building the absorption capacity that makes the hire stick","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"How to Hire a UX Researcher: Job Description, Interview Questions & Cost (2026)","metaDescription":"When to hire a user researcher vs. buy research capability. Includes the four-signal threshold test, 2026 salary and cost-per-hire math, a job description template, and a 25-question interview loop.","keywords":["hire a ux researcher","ux researcher job description","user researcher interview questions","when to hire a ux researcher","first research hire","ux research team"],"aiSummary":"A decision guide for hiring managers weighing a dedicated user researcher against AI-native research tooling. Covers the four-signal threshold test (volume, stakes, ambiguity, absorption), fully loaded 2026 cost math, a job description template by level, a five-stage interview loop with 25 questions, portfolio and craft-exercise evaluation, and the hire-vs-tool decision matrix.","aiPrerequisites":["Basic familiarity with product development roles","An understanding of what user research is used for"],"aiLearningOutcomes":["Apply a four-signal test to decide whether your team needs a dedicated researcher","Calculate the fully loaded annual cost of a research hire, including recruiting and vacancy cost","Write a job description that attracts senior research talent instead of generalist applicants","Run a five-stage interview loop that tests methodological judgment rather than vocabulary","Decide when AI-moderated research tooling covers the need better than headcount"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"2fe9a42a-4a5b-4ba0-8b81-a75298c195a0","slug":"ux-researcher-salary-team-cost-benchmarks","title":"UX Researcher Salary and Research Team Cost Benchmarks (2026)","url":"https://www.koji.so/docs/ux-researcher-salary-team-cost-benchmarks","summary":"Compensation and budget benchmarks for the user research function in 2026. Covers US salary medians by experience level, regional comparisons, the individual-contributor-to-director pay curve, freelance rates, the fully loaded cost model including recruiting and ramp, research operating budgets by ARR stage and team size, the line-item spend breakdown showing incentives as the dominant cost, and a cost-per-insight framework for allocating budget between headcount, tooling, and vendors.","content":"## The short answer\n\n**A US UX researcher costs roughly $85K at entry level, $110K at mid, $130K at senior, and $145K at ten-plus years in base salary — but the fully loaded first-year cost of a mid-level hire is closer to $170K once you include benefits, recruiting, and ramp. The function around that person needs another $30K–$75K in annual operating budget, of which about 60% is participant incentives.**\n\nTwo numbers reframe most research budget conversations. The first: product-led companies typically invest **0.5%–2% of ARR** in user research. The second: tooling is a rounding error inside that spend, while participants and salaries dominate it. Teams that spend their budget cycle negotiating a $4,000 software line while paying $40,000 in incentives and losing three weeks per study to scheduling are optimizing the smallest variable available to them.\n\nThis page gives you the compensation benchmarks, the fully loaded cost model, the operating budget by company stage, and the cost-per-insight math that should actually drive allocation.\n\n## UX researcher salary by level (US, 2026)\n\n| Experience | US median base |\n| --- | --- |\n| Entry level (0–3 years) | ~$85,000 |\n| Mid level (4–6 years) | ~$110,000 |\n| Senior (7–9 years) | ~$130,000 |\n| 10+ years | ~$145,000 |\n| Director / senior director | ~$216,000 |\n\nSource: User Interviews *2026 UX Salary Report* (n=2,062 US researchers). **67% of US researchers earn over $100,000.** Complementary market data puts the overall average base at $110,000–$120,000, with entry-level offers spanning $65,000–$95,000 and senior researchers at large technology companies reaching $150,000+ base and $200,000+ in total compensation. Glassdoor puts lead UX researcher at a $162,475 average, with a 25th-to-75th percentile band of $123,760–$216,359 — a spread wide enough that \"lead\" is close to meaningless as a compensation signal without company context.\n\n**Reading the level bands correctly:** the jump from entry to mid (~$25K) is the largest proportional step in the career, and the mid-to-senior step (~$20K) generally requires demonstrated influence rather than additional method knowledge. Teams that hire at the mid band and expect senior-level stakeholder navigation are the most common source of a failed first research hire.\n\n## The IC-to-manager curve\n\nManagement is a weaker financial lever than most compensation plans assume. The gap between senior individual contributors and team leads or managers **maxes out around $26,000 globally**, and pay differences do not become substantial until director level (~$216,000 in the US, roughly 4% of the sample).\n\nTwo implications. For employers: a management title is a poor retention instrument for a strong senior IC, and a senior IC track costs less to maintain than an unnecessary management layer. For researchers: the fastest compensation path below director is depth and demonstrated business impact, not headcount.\n\n## Geographic benchmarks\n\n| Region | Median (local) | ≈ USD |\n| --- | --- | --- |\n| United States | $110,000–$130,000 | $110K–$130K |\n| Canada | CAD $121,000 | ~$87,000 |\n| United Kingdom | £66,000 | ~$89,000 |\n| Australia / New Zealand | — | ~$91,000 |\n\nThe US median runs roughly **50% higher than the next-highest regions**. The disparity widens in contract work: freelance researchers report a **$125,000 median in the US versus $39,000 outside it** (global freelance median $90,000, n=293). For distributed teams, this is the single largest cost variable available — larger than the choice between agency and in-house — and it is why hiring geography deserves the same rigor as the hire/buy decision itself.\n\n## From base salary to real cost\n\nBase salary is roughly two-thirds of what the role costs in year one.\n\n| Component | Multiplier / amount |\n| --- | --- |\n| Base salary (mid-level US) | $110,000 |\n| Fully loaded multiplier (payroll taxes, benefits, equipment, software) | ×1.25–1.4 → $137,500–$154,000 |\n| Cost per hire, specialized role | +$10,000–$20,000 |\n| Time to fill | 42–63 days |\n| Ramp to first useful study | 2–4 months |\n| **Realistic year-one cost** | **$150,000–$174,000** |\n\nCost-per-hire and time-to-fill reflect 2026 SHRM-aligned recruiting benchmarks, where specialized and senior roles run well above the ~$4,700 general-role average and into the $10,000–$20,000 range. The line most budgets omit is **vacancy and ramp cost**: from requisition to first decision-grade output is typically four to seven months, during which the research questions do not pause. Whatever answers those questions in the interim is the true baseline your hire should be compared against — not zero.\n\n## What the function costs: operating budget by ARR stage\n\nSalary is the researcher. This is everything else — incentives, panels, tooling, repository, transcription.\n\n| ARR stage | Annual research ops budget | Studies/year | Researchers |\n| --- | --- | --- | --- |\n| Pre-revenue / seed | $5,000–$20,000 | 4–8 | 0 FTE |\n| $1M–$5M | $15,000–$45,000 | 8–15 | 0–0.5 FTE |\n| $5M–$20M | $40,000–$100,000 | 12–25 | 1 FTE |\n| $20M–$50M | $80,000–$200,000 | 20–40 | 1–3 FTE |\n| $50M–$150M | $150,000–$400,000 | 35–60 | 3–6 FTE |\n| $150M+ | $300,000–$1M+ | 50+ | 6+ FTE |\n\nBy team size, operating budget (excluding salary) runs **$30,000–$75,000 for a solo researcher**, **$80,000–$200,000 for a team of two to four**, and **$200,000–$600,000 for five or more**. Against the 0.5%–2%-of-ARR guideline, a $20M ARR company running a $100K–$400K all-in research function (salary plus ops) is inside the normal band.\n\nNote the third column against the fourth. At $5M–$20M ARR the benchmark expects **12–25 studies a year from one researcher**. At traditional throughput — 3–6 weeks per study, sequential moderation — one person delivers roughly 8–12. The benchmark quietly assumes leverage that most teams have not yet bought.\n\n## Where the money actually goes\n\nLine-item breakdown for a solo researcher's operating budget:\n\n| Line item | Annual cost | Share |\n| --- | --- | --- |\n| Participant recruitment and incentives | $18,000–$45,000 | **~60%** |\n| Recruitment platform | $3,000–$8,000 | ~10% |\n| Research repository | $1,500–$4,000 | ~5% |\n| Unmoderated testing tool | $1,500–$3,600 | ~5% |\n| Survey tool | $500–$2,400 | ~3% |\n| Transcription and analysis | $600–$2,000 | ~3% |\n\nThree observations that should change how the budget is argued:\n\n1. **Incentives dominate.** Sixty percent of the operating budget goes to participants. Any efficiency effort that ignores recruiting and incentive design is working on 40% of the problem. See [Research Participant Incentives](/docs/research-participant-incentives) for the amount-setting logic.\n2. **Tooling is small and elastic.** The entire tool stack in that table costs $7,100–$20,000 — less than a single agency study, and less than 15% of the researcher's salary. Cutting tooling to protect headcount reliably reduces the output of the headcount you protected.\n3. **The largest cost is invisible.** The salaried hours consumed by scheduling, transcription review, and manual coding never appear as a line item, because they are already inside someone's salary. On a $110K base, a researcher spending 60% of their week on execution rather than judgment represents ~$66,000 of annual spend on work that is now automatable.\n\n## Cost per decision-grade insight\n\nUnit cost is the metric that makes these numbers comparable. Take a mid-market team with one mid-level researcher, a $60K operating budget, and 10 studies a year:\n\n- **Fully loaded salary + ops:** ~$215,000\n- **Studies delivered:** 10\n- **Cost per study:** ~$21,500\n\nNow compare sourcing alternatives at the same volume: an agency at $15,000–$75,000 per study is at parity or worse; an AI-native platform running the same 10 studies plus the 15 that were previously rationed pushes cost per study down by an order of magnitude while the salary line stays constant — because the constraint that capped output was moderation and synthesis time, not thinking time.\n\nThe correct conclusion is not \"fewer researchers.\" It is that the same salary should buy 25 studies of judgment instead of 10 studies of logistics.\n\n## Which budget lines collapse with AI-native research\n\nMapping the line items above to what a platform like Koji actually removes:\n\n- **Transcription and analysis ($600–$2,000, plus the hidden salary hours):** eliminated. Transcription is automatic and thematic analysis with supporting quotes and per-interview quality scores (a 1–5 scale) generates in real time as interviews complete.\n- **Moderator time (hidden, the largest line):** eliminated for repeatable studies. AI-moderated interviews run **in parallel**, voice or text, so a 20-participant study fields overnight instead of across three weeks of calendar slots.\n- **Instrument quality risk (hidden):** controlled. Six typed structured question formats — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` — mean scale and choice questions aggregate quantitatively while open-ended questions get consistent AI follow-up probing to a configured depth. Every participant is asked the same thing the same way, which is what allows a PM to run a study without degrading quality. See the [Structured Questions Guide](/docs/structured-questions-guide).\n- **Separate survey and unmoderated tools ($2,000–$6,000):** consolidated, since one instrument produces both quantitative distributions and qualitative depth.\n- **Repository ($1,500–$4,000):** included, and it compounds — every study makes the next one cheaper to interpret.\n- **Incentives ($18,000–$45,000):** *not* eliminated. Participants still deserve payment, and this remains the dominant line. What changes is yield per incentive dollar: shorter time-to-completion and quality scoring mean fewer wasted sessions, and structured screening reduces spend on participants who should have been disqualified.\n\nThe line that does not move is the one worth protecting: senior judgment about which question matters. Budget planning that treats research as a cost center cuts salary and keeps process; planning that treats it as an evidence supply chain cuts process and keeps salary.\n\n## How to build next year's research budget in five steps\n\n1. **Anchor to ARR.** Start at 0.5%–2% of ARR for the all-in function; use the low end if research is new, the high end if product decisions are large and frequent.\n2. **Size demand in studies, not headcount.** Count the decisions next year that need evidence. Compare against the ARR-stage benchmark table above.\n3. **Set the incentive line first.** It is 60% of ops spend. Realistic per-participant amounts by audience type determine whether the study count is achievable at all.\n4. **Buy throughput before headcount.** A platform subscription costs a fraction of an FTE and removes the constraint that caps study count. See [UX Research Agency vs In-House vs AI Platform](/docs/ux-research-agency-vs-in-house).\n5. **Reserve 10–15% for contested work.** One or two external studies a year for decisions where independence has political value.\n\n## For candidates: how to benchmark an offer\n\nUse the level table as the base, then adjust: US roles sit ~50% above other regions; company size correlates positively with pay; and the IC-to-manager premium is small enough (~$26K max below director) that a title change without a band change is not a raise. Context matters too — only 16% of researchers report optimism about job opportunities in the field broadly, and 49% feel negative about the discipline's future, so evaluate the *stability* of the function you are joining as carefully as the number. Ask who acts on findings and whether execution work is automated; both predict whether the role survives the next planning cycle.\n\n## Frequently asked questions\n\n**What is the average UX researcher salary in 2026?**\n~$85K entry, ~$110K mid, ~$130K senior, ~$145K at 10+ years (US medians, User Interviews 2026 UX Salary Report). 67% of US researchers earn over $100K; directors average ~$216K.\n\n**What is the fully loaded cost of a researcher?**\nBase × 1.25–1.4, plus $10K–$20K recruiting and 42–63 days time-to-fill. A $110K mid-level base is realistically $150K–$174K in year one.\n\n**How much should we budget for research?**\n0.5%–2% of ARR all-in. Operating budget excluding salary: $30K–$75K for a solo researcher, $80K–$200K for a team of two to four.\n\n**What is the biggest research cost?**\nParticipant incentives, ~60% of operating budget. Tooling is typically under 15% of a researcher's salary.\n\n**Do managers earn more than senior ICs?**\nMarginally — the gap maxes out near $26K globally until director level.\n\n**How do international salaries compare?**\nUS medians run ~50% above the next regions: Canada ~USD $87K, UK ~USD $89K, ANZ ~USD $91K.\n\n## Related Resources\n\n- [User Research Cost Calculator](/docs/user-research-cost-calculator-2026) — per-study spend modeling to pair with these annual benchmarks\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six typed question formats that let one budget cover more studies without losing rigor\n- [How to Hire a UX Researcher](/docs/hiring-ux-researcher-guide) — the threshold test before you commit the salary line\n- [UX Research Agency vs In-House vs AI Platform](/docs/ux-research-agency-vs-in-house) — allocating the budget across sourcing models\n- [Research Participant Incentives](/docs/research-participant-incentives) — setting the line item that consumes 60% of ops spend\n- [UX Research Operations (ResearchOps)](/docs/ux-research-ops) — the infrastructure these budgets fund\n- [Getting Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) — defending the budget once it is built","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"UX Researcher Salary 2026: Levels, Geography, and Fully Loaded Team Cost","metaDescription":"2026 UX researcher salary benchmarks by level and region, the 1.25–1.4× fully loaded cost model, research budgets by ARR stage, and where research money actually goes (hint: 60% is participant incentives).","keywords":["ux researcher salary","user researcher salary","ux research team cost","ux research budget","fully loaded cost of a researcher","research operations budget"],"aiSummary":"Compensation and budget benchmarks for the user research function in 2026. Covers US salary medians by experience level, regional comparisons, the individual-contributor-to-director pay curve, freelance rates, the fully loaded cost model including recruiting and ramp, research operating budgets by ARR stage and team size, the line-item spend breakdown showing incentives as the dominant cost, and a cost-per-insight framework for allocating budget between headcount, tooling, and vendors.","aiPrerequisites":["Basic familiarity with budgeting or compensation planning"],"aiLearningOutcomes":["Benchmark a UX researcher salary by level, region, and seniority against 2026 data","Calculate the fully loaded first-year cost of a research hire including recruiting, vacancy, and ramp","Size a research operating budget against company ARR stage","Identify which research budget lines are structural and which collapse with AI-native tooling","Use cost per decision-grade insight to allocate budget between headcount, tooling, and vendors"],"aiDifficulty":"beginner","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"ac0011f4-d3f2-4b0f-a578-fbfcd145fc33","slug":"user-research-maturity-model","title":"User Research Maturity Model: 5 Stages from Ad-Hoc to Strategic (2026 Framework)","url":"https://www.koji.so/docs/user-research-maturity-model","summary":"A practical 5-stage user research maturity model (Ad-Hoc, Reactive, Operational, Embedded, Strategic) with assessment rubric, symptoms at each stage, and a 12-month progression playbook. Modeled on Nielsen Norman Group six-stage framework with operational criteria adapted for AI-native 2026 research stacks. Covers what stalls progression, how to climb from Stage 2 to Stage 4, and how AI-native platforms like Koji change the unit economics of continuous discovery.","content":"**A user research maturity model is a framework for diagnosing how systematically your organization uses customer research to drive decisions, and what specifically needs to change to advance.** The most influential version, Nielsen Norman Group's six-stage model, ranges from \"Absent\" (UX is invisible) to \"User-Driven\" (research shapes strategy). For most product organizations, the practical 5-stage version below — Ad-Hoc, Reactive, Operational, Embedded, Strategic — is more actionable, because it maps to changes you can actually make this quarter.\n\nThis guide gives you the assessment rubric to place your team on the curve, the symptoms of being stuck at each stage, and the specific moves that advance you to the next level. AI-native research platforms like Koji compress the time it takes to climb — most teams can move two stages in a year when the operational friction (recruiting, scheduling, analysis) collapses to near zero.\n\n## TL;DR — the 5-stage framework\n\n| Stage | One-line description | Research cadence | Who runs research |\n|---|---|---|---|\n| 1. Ad-Hoc | Research happens when someone insists | <1 study / quarter | Whoever has time |\n| 2. Reactive | Research validates decisions already made | 1–2 / quarter | A part-time PM or designer |\n| 3. Operational | Research has its own process and people | Monthly+ | Dedicated researcher(s) |\n| 4. Embedded | Every product decision has research input | Weekly continuous | Researchers + democratized teams |\n| 5. Strategic | Research shapes roadmap and strategy | Always-on | Whole org, with research leadership |\n\nMost teams are stuck at Stage 2 or 3. The leap from Stage 3 to Stage 4 is the one most worth making — it is where research starts changing the product instead of describing it.\n\n## Why a maturity model matters\n\nResearch budgets are easy to defend after they have produced clear wins. They are hard to defend before. A maturity model gives leadership a vocabulary for two things that are otherwise hard to articulate:\n\n1. **Where we are.** A specific stage with specific symptoms (\"we ship and then go ask if users like it\") is harder to argue with than \"we should do more research.\"\n2. **What we should invest in next.** Climbing the model is a sequence — you cannot skip stages. Knowing where you are tells you what to fix first.\n\nAccording to Nielsen Norman Group's research on UX maturity, organizations advance through stages \"Absent, Limited, Emergent, Structured, Integrated, and User-Driven,\" and the six factors that move them up are strategy, culture, process, outcomes, leadership support, and longevity. The model below collapses those into five stages with sharper operational criteria, because most product teams find the 6-stage version too granular for self-assessment.\n\n## The five stages\n\n### Stage 1: Ad-Hoc\n\n**Symptoms:** Research happens when an executive demands it or a launch goes badly. There is no research backlog, no participant pipeline, and no synthesis discipline. Insights live in the head of whoever did the study and disappear when they leave.\n\n**What's missing:** A standing assumption that decisions deserve evidence. Most Stage-1 organizations are not against research — they have just never built the muscle.\n\n**Telltale quote:** \"We should probably talk to some users about this before launch.\"\n\n**The path out:** Pick a single recurring research question (e.g., \"why do new signups churn in week one?\") and commit to a small monthly study answering it. Create one re-usable artifact — a slack channel, a Notion page — where every insight is filed. The win is consistency, not volume.\n\n### Stage 2: Reactive\n\n**Symptoms:** Research is run, but mostly to validate decisions that have already been made. The output is justification, not direction. Studies are typically usability tests on near-final designs and quick surveys after launch.\n\n**What's missing:** Research that is upstream of design decisions. At Stage 2, the team is using research to confirm what they already wanted to do — which means it can never disagree with them, which means it never changes the product.\n\n**Telltale quote:** \"Can you run a quick test to make sure this is fine?\"\n\n**The path out:** Move at least one study per cycle to *before* the design phase. Discovery interviews, [Mom Test](/docs/mom-test-methodology) conversations, problem-space exploration. The criterion: if the study cannot change the design direction, it is not real research.\n\n### Stage 3: Operational\n\n**Symptoms:** Research has its own roster, recruiting flow, study templates, and quarterly cadence. There is at least one full-time researcher (or a dedicated PM-researcher hybrid). Stakeholders submit research requests through an intake system.\n\n**What's missing:** Speed. Stage-3 organizations have rigor but not velocity — every study takes 4–8 weeks from request to insight, which means the questions move faster than the answers. Research becomes a bottleneck the rest of the org learns to route around.\n\n**Telltale quote:** \"Can we get this added to next quarter's research roadmap?\"\n\n**The path out:** Two parallel investments. First, [research democratization](/docs/research-democratization-scaling-insights-2026): give PMs, designers, and CSMs the tooling and templates to run their own routine studies, with researchers reviewing for quality. Second, AI-native tooling: replace the 6-week study cycle with a 6-day one by automating recruiting, moderation, and synthesis.\n\n### Stage 4: Embedded\n\n**Symptoms:** Every product squad runs at least one customer interview per week. [Continuous discovery](/docs/continuous-discovery-user-research) is the default, not the exception. Research insights flow directly into roadmap discussions. Stakeholders bring questions to research instead of being chased for them.\n\n**What's missing:** Strategic influence. At Stage 4, research informs every decision but rarely *initiates* one. The roadmap is still set by leadership and merely informed by research, instead of being driven by it.\n\n**Telltale quote:** \"What did this week's interviews tell us about the upcoming release?\"\n\n**The path out:** Invest in research synthesis and storytelling at the executive level. The research function needs to be in roadmap and strategy meetings, not as a service provider but as a contributor. This requires a research lead with the seniority to shape strategy, and infrastructure (a [research repository](/docs/research-repository-guide), a clear synthesis cadence, recurring leadership briefings) that makes findings legible to non-researchers.\n\n### Stage 5: Strategic\n\n**Symptoms:** Research is upstream of the roadmap. New product bets are sized using customer evidence, not just market data. Senior leadership cites specific customer interviews in board meetings. The research function reports to the CEO or CPO and has a seat at strategic planning.\n\n**What's missing:** Nothing structural — Stage 5 is the steady-state goal. The risk at Stage 5 is complacency; mature research orgs need to keep questioning their own methods, expanding into adjacent jobs to be done, and refreshing their participant panels.\n\n**Telltale quote:** \"We're not committing to that bet until we run discovery interviews against our top three customer segments.\"\n\n**Sustaining behavior:** Annual customer research strategy review. Researcher career ladders that retain senior talent. Leadership evangelism — every executive can name a recent insight and how it changed a decision.\n\n## Self-assessment rubric\n\nFor each dimension, score 1 (Stage 1 behavior) to 5 (Stage 5 behavior). The lowest score is your true stage — climbing requires advancing the weakest dimension, not the average.\n\n| Dimension | Stage 1 | Stage 3 | Stage 5 |\n|---|---|---|---|\n| **Cadence** | <1 study/qtr | Monthly | Continuous, weekly+ |\n| **Timing** | After launch | Before design | Before strategy |\n| **Ownership** | Whoever has time | Dedicated researcher | Senior research leader |\n| **Synthesis** | Lives in one head | Documented per study | Living repository |\n| **Influence on roadmap** | None | Informs design | Shapes strategy |\n| **Stakeholder buy-in** | \"Why bother?\" | \"Add to backlog\" | \"We can't decide without this\" |\n\nA team scoring 1, 1, 2, 3, 2, 2 across these dimensions is at Stage 1, regardless of the higher scores. Climb the lowest.\n\n## What stalls progression?\n\nThe most common reasons teams plateau, in rough order of frequency:\n\n**Operational friction.** When recruiting takes 2 weeks and analysis takes another week, the cadence required for Stage 4 (weekly continuous discovery) is mathematically impossible. According to Maze's 2023 Continuous Research Report, 64% of companies now have a democratized research culture to cope with increasing demand — and the bottleneck for the rest is operational, not philosophical.\n\n**Single point of failure.** The team has one researcher who is fully booked on intake. They cannot do strategic work because they cannot turn down tactical requests. Solution: democratize the tactical work to free up the researcher for strategic studies.\n\n**No repository.** Insights from past studies are not findable, so every new question starts from zero. This is the single biggest preventable waste in research operations.\n\n**No leadership champion.** Without an executive who articulates why research matters, every budget cycle becomes a fight to justify what is already there. Stage 4+ requires a Chief Customer Officer, CPO, or similar who treats customer evidence as a strategic input.\n\n## How AI-native platforms accelerate the climb\n\nThe maturity model assumes a 2010s research stack: panels recruited by hand, studies run synchronously, transcripts coded by analysts, reports written manually. In that stack, the operational cost of running research scales nearly linearly with research volume — which means Stage 4 (weekly continuous discovery) is genuinely expensive.\n\nAI-native platforms like Koji change the unit economics. Specifically:\n\n- **Recruiting** moves from days to minutes via in-product intercepts and conversational invitations.\n- **Moderation** runs 24/7 with no human in the loop — interviews complete asynchronously without scheduling.\n- **Analysis** runs the moment the last interview ends — themes, quotes, and quality scores are pre-aggregated.\n- **Synthesis** is the human-in-the-loop step, on top of pre-organized inputs instead of raw transcripts.\n\nThe practical implication: Stage 4 cadence (3–10 customer interviews per week, per product team) becomes affordable for organizations that are currently at Stage 2. Teams using AI-assisted research tools report 60% faster time-to-insight, and that compression is precisely what makes the leap to embedded research feasible.\n\nFor most organizations the biggest bottleneck is not insight quality — it is insight throughput. AI-native tooling does not replace researcher judgment (it cannot, and shouldn't); it replaces the operational drag that keeps researchers stuck on tactical intake instead of strategic studies.\n\n## A 12-month progression playbook\n\nIf you are at Stage 2 today, here is the sequence to reach Stage 4 in a year:\n\n**Month 0–2:** Pick three recurring research questions that come up every quarter. Build [research interview templates](/docs/research-interview-templates) for each. Set up a [research repository](/docs/research-repository-guide) — even a structured Notion page is enough.\n\n**Month 2–4:** Move one study per cycle from \"validation after design\" to \"discovery before design.\" This is the single highest-leverage change in the entire model.\n\n**Month 4–6:** Adopt an AI-native research platform. The goal is to compress study cycle time from 4 weeks to 4 days.\n\n**Month 6–9:** Democratize. Give PMs and designers self-serve access to run their own discovery interviews, with researchers in a coaching/QA role. Establish a weekly insights digest to all stakeholders.\n\n**Month 9–12:** Cement the cadence. Every squad runs at least one interview per week. Research insights appear in roadmap reviews. The research function is asked to opine on strategy, not just tactics.\n\nBy month 12 you are at the bottom of Stage 4. Stage 5 takes another 12–24 months and depends primarily on leadership — not tooling.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — How Koji's six structured question types let democratized teams run rigorous studies without research training.\n- [Research Democratization](/docs/research-democratization-scaling-insights-2026) — The path from Stage 3 to Stage 4.\n- [Continuous Discovery User Research](/docs/continuous-discovery-user-research) — The Stage-4 cadence in detail.\n- [UX Research Operations](/docs/ux-research-ops) — The infrastructure that supports Stage 3+.\n- [Customer Interview Cadence](/docs/customer-interview-cadence) — Benchmarks by team size and stage.\n- [Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) — Securing the leadership support that Stage 5 requires.\n\n\n\n## Further reading on the blog\n\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n- [AI Agents for User Research in 2026: How Autonomous Research Is Reshaping Customer Insight](/blog/ai-agents-user-research-2026) — AI agents are taking over user research in 2026 — moderating interviews, synthesizing themes, and producing insight reports in hours. The fu\n- [Best AI Market Research Tools in 2026: The Complete Buyer's Guide](/blog/ai-market-research-tools-2026) — AI has fundamentally changed market research. This guide compares the leading AI market research platforms—from AI-native interview tools li\n\n<!-- further-reading:blog -->\n","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"User Research Maturity Model: 5 Stages from Ad-Hoc to Strategic (2026) | Koji","metaDescription":"A practical 5-stage user research maturity model with self-assessment rubric, common roadblocks at each stage, and the 12-month playbook to advance two stages in a year.","keywords":["user research maturity model","research maturity model","ux research maturity","research team maturity","maturity assessment","research operations maturity","nielsen norman maturity","research practice scale"],"aiSummary":"A practical 5-stage user research maturity model (Ad-Hoc, Reactive, Operational, Embedded, Strategic) with assessment rubric, symptoms at each stage, and a 12-month progression playbook. Modeled on Nielsen Norman Group six-stage framework with operational criteria adapted for AI-native 2026 research stacks. Covers what stalls progression, how to climb from Stage 2 to Stage 4, and how AI-native platforms like Koji change the unit economics of continuous discovery.","aiPrerequisites":["Familiarity with user research practice","Some experience running or managing UX studies"],"aiLearningOutcomes":["How to assess your team's current user research maturity stage","The symptoms and bottlenecks at each of the five stages","The specific operational moves that advance a team from one stage to the next","How AI-native research tooling changes the cost of climbing the maturity curve"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 minutes"},{"type":"documentation","id":"05809a87-6a5f-4c7c-9d50-ab74cabb0abd","slug":"ux-researcher-career-path","title":"UX Researcher Career Path: Levels, Competencies, and Promotion Criteria (2026)","url":"https://www.koji.so/docs/ux-researcher-career-path","summary":"A six-level UX research career ladder defined by scope of ambiguity rather than methods or tenure, with US base-salary anchors, six behaviourally-anchored competency dimensions, the IC/manager fork at level four, and a promotion evidence packet. Explains how AI-assisted tooling has broken throughput as a seniority signal and reweighted promotion criteria toward question framing, business judgement, ethics and operational leverage.","content":"## The short answer\n\n**Research levels are defined by the amount of ambiguity you can absorb on someone else's behalf — not by the methods you know, the studies you have shipped, or the years you have served.** A junior researcher is handed a question and returns an answer. A senior researcher is handed a decision and returns the question that should have been asked. A principal researcher is handed a strategy and returns the thing nobody realised was in doubt.\n\nThat single variable — where the ambiguity gets resolved — explains almost every leveling disagreement, and it is the one criterion that survives contact with AI-assisted tooling. Everything below is scaffolding around it.\n\nThe 2026 context makes the ladder unusually worth getting right. Research demand is up: **66% of practitioners reported increased demand for research, against 55% a year earlier**, and the share of organisations saying research is essential to all levels of business strategy **nearly tripled from 8% to 22%** ([Maze, Future of User Research 2026](https://maze.co/resources/user-research-report/), n≈500, fielded Dec 2025–Jan 2026). Sentiment about careers is simultaneously terrible: **67% of researchers were negative about career opportunities in the field**, and 49% were negative about the field's future ([User Interviews, State of User Research](https://www.userinterviews.com/state-of-user-research-report)). Both things are true because demand is growing in places that do not have a ladder — **71% of organisations have people doing research who are not researchers**, and only **6% of researchers sit on a dedicated research team**.\n\n## The six-level reference ladder\n\nMost functional ladders resolve to six non-executive levels. Titles vary; scope does not.\n\n| Level | Common titles | Scope of ambiguity | Autonomy | US base median anchor |\n|---|---|---|---|---|\n| 1 | Associate / Researcher I | Executes a defined study plan on a defined question | Reviewed at every stage | ~$85K |\n| 2 | Researcher II / UX Researcher | Owns a study end-to-end; chooses methods for a stated question | Reviewed at plan and readout | ~$110K |\n| 3 | Senior Researcher | Owns a product area; reframes the question before answering it | Trusted; escalates on risk | ~$130K |\n| 4 | Staff / Lead Researcher | Owns a multi-team problem space; sets the research agenda for it | Sets own goals against org strategy | ~$145K+ |\n| 5 | Principal Researcher | Owns a company-level bet; changes what leadership believes | Accountable to outcomes, not plans | Varies widely |\n| 6 | Director / Head of Research | Owns the function: headcount, tooling, standards, and the ladder itself | Accountable for the function | ~$216K |\n\nBase-salary anchors are US medians from the [2026 UX Salary Report](https://www.userinterviews.com/ux-salary-report), where 67% of respondents earn above $100K and the US sits roughly 50% above the next-highest region. Treat them as calibration, not entitlement — see [UX researcher salary and team cost benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks) for the full picture including fully-loaded cost.\n\n**The fork happens at level 4.** Both tracks exist above it, and the pay difference between them is smaller than people assume: the IC-to-manager gap in the 2026 salary data maxes out around **$26K**. That is a meaningful number but not a life-changing one, which means choosing management for money is usually a mistake. Choose it because you would rather multiply other people's judgement than exercise your own.\n\n## Six competency dimensions\n\nLevels are the output. Competencies are what you actually assess. Six dimensions cover the function without overlapping:\n\n**1. Craft and method.** Can you select and execute the right method? Level 1–2 is executing common methods correctly. Level 3 is choosing among them under constraint. Level 4+ is knowing when the standard methods will not work and designing something defensible — and knowing when *not* to research at all.\n\n**2. Analytical rigour.** Can your conclusions survive challenge? This is where [validity](/docs/qualitative-research-validity), sampling logic, and [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) live. The senior marker is not \"runs thematic analysis\" — it is being able to state, unprompted, what would have falsified the finding.\n\n**3. Influence and communication.** Does the organisation act on it? Progression runs from *reports findings clearly* → *tailors the argument to the decision-maker* → *changes a roadmap* → *changes what leadership believes is true about the market*. See [presenting research findings](/docs/presenting-research-findings).\n\n**4. Business and domain judgement.** Do you know which questions are worth money? The most common blocker at the senior boundary is a researcher with excellent craft who consistently answers questions that were not going to change anything.\n\n**5. Operational leverage.** Do you make other people's research better? [Repositories](/docs/research-repository-guide), templates, question banks, tooling, and [democratization enablement](/docs/research-democratization-playbook). This dimension used to be optional above level 3. It is now the primary differentiator, and the reason is in the next section.\n\n**6. People and mentorship.** From \"unblocks a peer\" to \"grows researchers who outgrow you.\" Required for the management track from level 4; still assessed on the IC track, where it appears as mentorship and craft standards rather than line management.\n\nA workable rubric writes five behaviourally-anchored statements per dimension and requires **consistent demonstration at the target level across at least four of six dimensions**, with no dimension more than one level below. Requiring all six produces a ladder nobody ever climbs.\n\n## What AI changed about promotion criteria\n\nFor twenty years, the implicit senior signal was throughput: the researcher who ran the most studies with the highest quality was the senior researcher. That signal is now largely broken, and the data explains why. **69% of practitioners use AI in at least some research projects, a 19-point year-on-year increase**, most commonly for transcription, synthesis, and generating research questions (Maze 2026). When the execution layer compresses, execution volume stops discriminating between a good researcher and a fast one with good tooling.\n\nWhat the same survey says humans are still required for is instructive, because it is a promotion rubric in disguise:\n\n| Human contribution | Share saying it requires human involvement |\n|---|---|\n| Interpreting nuance and emotion | 82% |\n| Ethical decision-making | 80% |\n| Framing the right questions | 76% |\n\nMeanwhile **35% say the researcher role is becoming more strategic and 33% say it is becoming more blended** with adjacent roles. Practically, that means a 2026 ladder should reweight: down on craft-and-method as the senior gate, up on business judgement, question framing, ethics, and operational leverage. A researcher whose distinctive contribution is \"runs interviews well\" is now describing a level-2 competency, however well they do it.\n\nThe corollary for anyone building a promotion case: **evidence of decisions changed beats evidence of studies delivered.** Three documented decisions your work moved will outperform thirty study reports at every level above 2.\n\n## The promotion evidence packet\n\nPromotion committees do not reward work; they reward legible work. Build the packet continuously, not in the two weeks before calibration.\n\n1. **Operate at the level first.** Virtually every functional ladder promotes on demonstrated performance at the target level, typically sustained over two review cycles. Ask your manager which level your current work is being read at — that answer is the whole conversation.\n2. **The three-artifact rule.** For each competency dimension you are claiming, name one artefact, one decision it changed, and one named person who will corroborate it.\n3. **Quantify the counterfactual.** \"We shipped X\" is weak. \"We were about to build X; the study showed the problem was Y; we built Y instead and activation moved N points\" is a promotion.\n4. **Show leverage explicitly.** What did you make reusable? Who ran research better because of something you built? See [research operations](/docs/research-ops-guide).\n5. **Update it monthly.** Impact evidence decays fast — the PM who would have corroborated your finding leaves, the dashboard gets rebuilt, the memory fades.\n\n## Building a ladder when you are a team of one to three\n\nMost researchers work somewhere with no research ladder at all — a direct consequence of only 6% sitting on a dedicated research team, with roughly a third reporting into Product. Two rules make this tractable:\n\n- **Borrow the shape, replace the content.** Take your company's design or product-management ladder, keep its level numbering, comp bands and promotion process, and rewrite only the competency descriptions. Fighting for a bespoke process is a fight you will lose; fighting for accurate criteria inside the existing process is one you can win.\n- **Define three levels, not six.** Below ~8 researchers, six levels is theatre. Define mid, senior, and lead with sharp boundaries, and add levels when headcount forces the issue.\n\nIf you are the only researcher, your leverage dimension is not optional — it is the job. See [the solo researcher's toolkit](/docs/solo-researcher-toolkit-guide), and [UX research team structure](/docs/ux-research-team-structure) for how the ladder changes across centralized, embedded and hub-and-spoke models. If you are on the hiring side of this, [how to hire a UX researcher](/docs/hiring-ux-researcher-guide) covers the level you should actually be recruiting for.\n\n## The modern approach: building the evidence with AI-native research\n\nThe promotion criteria above have a practical problem: leverage and decision-impact are hard to evidence when every study consumes a week of your calendar. Scheduling, moderating, transcribing and coding is where the hours go, and none of it appears in a promotion packet.\n\nThis is where an AI-native platform changes the arithmetic rather than the standard. With Koji, a study is a link: the AI moderator runs the interview asynchronously, probes follow-ups in the participant's own words, and the analysis arrives with the transcript. Interviews that used to be scheduled one at a time run in parallel, which turns \"I ran four studies this half\" into \"I ran a continuous programme and can show the decision trail.\"\n\nThree things map directly onto the competency dimensions:\n\n- **Operational leverage.** Reusable study templates and Koji's [six structured question types](/docs/structured-questions-guide) — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` — let non-researchers run studies that are actually comparable, which is the difference between democratization and chaos. That reusable asset is the artefact your packet needs.\n- **Analytical rigour.** Because structured questions are typed at capture, the same question is aggregated identically across every study and every quarter, so you can show a trend rather than a snapshot. Quality scoring rates each conversation 1–5 against your research goals, so a thin sample is visible before it reaches a stakeholder.\n- **Influence.** Real-time reporting means the answer exists at the moment the decision is being made rather than three weeks after. Almost every \"research got ignored\" story is a timing story.\n\n| | Traditional workflow | AI-native workflow |\n|---|---|---|\n| Studies per researcher per quarter | 3–5 | 10–20 |\n| Time from question to insight | 3–6 weeks | Days |\n| Evidence of leverage | Ad hoc | Reusable templates and typed question sets |\n| Promotion narrative | \"I ran the studies\" | \"I owned the decisions\" |\n| Non-researchers enabled | Depends on your availability | Self-serve, on your standards |\n\nThe point is not that a tool gets you promoted. It is that the criteria have moved toward judgement and leverage, and the only way to spend more time on judgement is to spend less on logistics.\n\n## Frequently asked questions\n\n**How long should each level take?**\nThere is no reliable industry benchmark, and any specific number you see is one company's policy. The defensible framing is behavioural rather than temporal: you are ready when you have been operating at the next level for two consecutive review cycles. In practice, levels 1→2 and 2→3 tend to move fastest, and the senior→staff boundary is the longest at most companies because it requires scope your manager has to grant you.\n\n**Should I go IC or manager at the fork?**\nNot a money decision — the IC-to-manager base gap in the 2026 salary data tops out near $26K. Choose management if you would rather multiply other people's judgement than exercise your own, and if you would find a quarter with no research of your own satisfying rather than hollow.\n\n**Is AI making UX research a worse career?**\nThe sentiment data is genuinely negative — 67% of researchers reported pessimism about career opportunities — but demand data points the other way, with 66% reporting increased demand and the share of organisations calling research essential to strategy nearly tripling. The consistent reading is that the *execution* half of the role is compressing while the *judgement* half is growing. Careers built on execution volume are at risk; careers built on question framing, business judgement and enablement are not.\n\n**What if my company has no research ladder?**\nBorrow the design or PM ladder, keep its levels and comp bands, and rewrite only the competency descriptions. Define three levels rather than six until you have roughly eight researchers.\n\n**Do I need a certification or a graduate degree to level up?**\nNo. Neither appears as a promotion criterion in any functional ladder worth copying. Both can help at the hiring gate for a first role; neither substitutes for demonstrated scope once you are in the field.\n\n**What is the single strongest promotion signal?**\nA documented decision that went differently because of you, corroborated by the person who made it. Everything else in the packet is supporting evidence for that.\n\n## Related resources\n\n- [UX Researcher Salary and Research Team Cost Benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks) — comp anchors by level, region and seniority\n- [UX Research Team Structure](/docs/ux-research-team-structure) — how centralized, embedded and hub-and-spoke models change the ladder\n- [How to Hire a UX Researcher](/docs/hiring-ux-researcher-guide) — the employer's side: when to hire, the JD, and the interview loop\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make research comparable across studies and quarters\n- [The Solo Researcher's Toolkit](/docs/solo-researcher-toolkit-guide) — building leverage when you are the whole function\n- [Research Democratization Playbook](/docs/research-democratization-playbook) — enabling non-researchers without losing rigour\n- [Research Operations Guide](/docs/research-ops-guide) — the operational leverage dimension in practice","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"UX Researcher Career Path: Levels & Promotion Criteria (2026)","metaDescription":"A six-level UX research career ladder with scope, comp anchors and competency rubrics — plus the promotion evidence packet and what AI changed in 2026.","keywords":["ux researcher career path","ux research career ladder","ux researcher levels","research competency framework","ux research promotion criteria","senior ux researcher requirements","principal ux researcher","ux research career progression","ic vs manager ux research","user researcher job levels"],"aiSummary":"A six-level UX research career ladder defined by scope of ambiguity rather than methods or tenure, with US base-salary anchors, six behaviourally-anchored competency dimensions, the IC/manager fork at level four, and a promotion evidence packet. Explains how AI-assisted tooling has broken throughput as a seniority signal and reweighted promotion criteria toward question framing, business judgement, ethics and operational leverage.","aiPrerequisites":["Working knowledge of common UX research methods","Familiarity with how your organisation runs performance reviews"],"aiLearningOutcomes":["Define research levels by scope of ambiguity rather than methods or years served","Assess researchers against six behaviourally-anchored competency dimensions","Decide between the IC and management tracks on the right criteria","Build a promotion evidence packet that survives a calibration committee","Adapt a design or product ladder when no research ladder exists","Reweight promotion criteria for an AI-assisted research workflow"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"},{"type":"documentation","id":"7b545afe-5ba9-45b2-bf7d-6defca3ff050","slug":"first-90-days-ux-researcher","title":"The First 90 Days as a UX Researcher: A Week-by-Week Onboarding Plan (2026)","url":"https://www.koji.so/docs/first-90-days-ux-researcher","summary":"A new UX researcher should spend the first 30 days learning the system rather than producing findings - auditing past research, mapping decision-making, and interviewing stakeholders as if they were a research project. Days 31 to 60 are for shipping one small study end to end to establish credibility and expose process friction. Days 61 to 90 are for changing exactly one process. This guide gives the week-by-week plan, the week-one audit checklist, a stakeholder interview script, day-90 success criteria, and how AI-native tooling lets a new hire ship real findings in week two instead of month three.","content":"# The First 90 Days as a UX Researcher: A Week-by-Week Onboarding Plan (2026)\n\n**Short answer:** Spend days 1-30 learning the system rather than producing findings; days 31-60 shipping one small study end to end; days 61-90 changing exactly one process. The instinct to prove yourself with an ambitious foundational study in month one is the single most common way new researchers spend credibility they have not yet earned - on a question nobody was blocked on.\n\nThe stakes are higher than most managers assume. Gallup finds that only **12% of employees strongly agree their organisation does a great job of onboarding new hires**, and only **29% of new hires say they feel fully prepared and supported to succeed** after onboarding. The same research finds that when managers take an active role, employees are **3.4 times as likely** to strongly agree their onboarding was exceptional. Onboarding is not an administrative phase. It is the period in which your ceiling gets set.\n\nThis guide is written for the researcher. If you are the manager, [UX Researcher Career Path](/docs/ux-researcher-career-path) covers what to expect at each level, and [UX Research Team Structure](/docs/ux-research-team-structure) covers where the role should sit.\n\n## Days 1-30: learn the system\n\nYour first month has one deliverable, and it is not a study. It is a map: **who decides what, on what evidence, and what they currently believe.**\n\n### The week-one audit\n\nBefore you talk to anyone, find out what already exists. Most organisations have more research than they think and less than they claim.\n\n| Audit item | Question to answer | Red flag |\n|---|---|---|\n| Past studies | What has been researched in the last 24 months? | Nobody can find them |\n| Repository | Where do findings live, and does anyone read them? | Findings live in individual Slack DMs |\n| Participant access | Can you legally and practically reach customers? | Every request routes through one gatekeeper |\n| Consent and legal | What consent, retention, and DPA terms apply? | No template exists |\n| Tooling | What is licensed, what is actually used? | Three overlapping tools, none adopted |\n| Prior conclusions | What does the org believe it already knows? | Beliefs with no traceable source |\n\nThat last row is the most valuable and the most skipped. Write down every confident claim you hear in your first fortnight - *\"our users don't care about pricing,\"* *\"enterprise buyers want SSO first\"* - and note whether anyone can point to where it came from. Unsourced beliefs are your best backlog. They are what the organisation is currently betting on without evidence.\n\n### Interview your stakeholders like participants\n\nNielsen Norman Group makes this point directly: approach onboarding intro calls **as a small research project**, to learn as much as possible about the people and the context. That reframing matters. It converts a fortnight of pleasant introductions into a study with a sample, a discussion guide, and findings.\n\nAim for 8-12 stakeholders across product, design, engineering, support, sales, and one executive. Use a consistent guide so the answers are comparable:\n\n1. What decision are you making in the next quarter that you are least confident about?\n2. What do you believe about our users that you cannot prove?\n3. When did research last change your mind? What happened?\n4. What has stopped you using research in the past?\n5. If you could have one question answered perfectly, what would it be?\n\nQuestion 3 is the diagnostic. If nobody can recall research changing their mind, you are not joining a research practice - you are founding one, and your first 90 days should be planned accordingly. [Stakeholder Interviews](/docs/stakeholder-interview-guide) covers the full method.\n\nThen synthesise it properly and share it back. A one-page readout of *\"here is what this organisation believes and cannot prove\"* in week four is the highest-leverage artifact a new researcher can produce. It is genuinely useful, it demonstrates your method, and it costs you nothing politically because every claim in it came from them.\n\n## Days 31-60: ship one small thing\n\nPick a first study by three criteria, in this order:\n\n1. **Someone is actively blocked on it.** Not interested - blocked. A blocked person is a person who will act on your finding and tell others they did.\n2. **It can be done in under three weeks.** Momentum beats comprehensiveness at this stage.\n3. **It has a visible decision attached.** You want to be able to say *\"we shipped X instead of Y because of this.\"*\n\nWhat to avoid: the foundational persona refresh, the full journey map, the research repository rebuild. These are the projects new researchers instinctively reach for because they look substantial and offend nobody. They also take a full quarter and produce nothing anyone can point to at your first review.\n\n**The traditional bottleneck.** The honest problem with \"ship a study by day 60\" is that a conventional study requires exactly the things a new hire does not have: a recruiting pipeline, a participant panel, calendar access to customers, and a relationship with whoever guards them. This is why the classic advice quietly assumes your first real study lands in month three or four.\n\n**The 2026 version.** That constraint has genuinely loosened. Maze's *Future of User Research 2026* study (n≈500, fielded December 2025 to January 2026) found **69% of research practitioners now use AI in their work, up 19 percentage points year over year**, and that the share saying research is essential to strategy rose from **8% to 22%**. Demand for research is up **66%**, and **39% of product managers** now conduct research themselves. You are joining an organisation that expects research faster than the old pipeline can deliver it.\n\nWith an AI-native platform, the sequence changes:\n\n- **Week 2:** launch an always-on AI-moderated study on a question from your stakeholder interviews. No calendar slots to book, no moderation to schedule, no transcription queue.\n- **Week 3:** 30 completed interviews. Koji's thematic analysis clusters the open-ended responses into named themes with the verbatim quotes attached, so you are reviewing findings rather than tagging transcripts.\n- **Week 4:** readout to the person who was blocked, with structured-question data behind every claim.\n\nKoji's six structured question types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` - matter more than usual for a new hire, because you are establishing a reputation for rigour. A finding that pairs a comparable scale distribution with the reasoning behind it survives challenge from a sceptical stakeholder; a finding that is three quotes does not. See the [structured questions guide](/docs/structured-questions-guide).\n\nYou do not need a research background to run this well, which is also the point: in an organisation where PMs are already doing their own research, your value is not being the only person who can run a study. It is being the person who makes everyone else's studies trustworthy.\n\n## Days 61-90: change one process\n\nExactly one. New researchers who try to reform recruitment, the repository, intake, and consent simultaneously finish the quarter with four half-built systems and no adoption.\n\nChoose based on what actually blocked you during your first study:\n\n| What blocked you | Process to fix | Success measure |\n|---|---|---|\n| Could not find past findings | A working repository with a naming convention | Someone else finds something without asking you |\n| Could not reach customers | A standing recruitment pipeline | Time-to-first-participant under 5 days |\n| Requests arrived as vague asks | An intake form with a decision field | Requests name the decision they inform |\n| Legal review stalled you | Approved consent and DPA templates | Zero legal review for standard studies |\n\nThe success measure is the part people skip. \"We built a repository\" is not an outcome; \"someone found a study without asking me\" is.\n\n## The plan at a glance\n\n| Phase | Focus | Primary output | What to resist |\n|---|---|---|---|\n| Days 1-30 | Learn the system | A map of decisions, beliefs, and evidence gaps | Launching a big study |\n| Days 31-60 | Ship one small thing | One acted-upon finding | Scope creep into a foundational project |\n| Days 61-90 | Change one process | One improvement with a measurable outcome | Reforming everything at once |\n\n## What good looks like at day 90\n\nA checklist you can hold yourself to:\n\n- [ ] At least one study shipped **and acted on** - you can name the decision it changed\n- [ ] A written map of who decides what, on what evidence\n- [ ] A shared list of the organisation's unsourced beliefs\n- [ ] One process improved, with a measure that moved\n- [ ] Relationships with 8-12 stakeholders who know what you do\n- [ ] A repeatable study setup someone else could run\n- [ ] Awareness of your legal and consent constraints\n\nWhat is deliberately **not** on that list: a large number of studies. Anyone measuring a new researcher on study volume in the first quarter is measuring the wrong thing, and it is worth saying so early and calmly.\n\n## Five mistakes to avoid\n\n1. **Leading with methodology critique.** Pointing out that the previous survey was leading is correct, unhelpful, and expensive in your first month. Fix it silently in the next study.\n2. **Waiting for perfect access.** If customer access is genuinely blocked, run the study with prospects, churned users, or an adjacent segment and label it honestly. Something imperfect in week six beats something perfect in month five.\n3. **Over-indexing on your manager's framing.** Interview widely. Your manager's view of the research gaps is one data point, and often the one most shaped by past frustration.\n4. **Confusing being busy with being useful.** Attending every design review is not research. It is the most comfortable way to spend 90 days and produce nothing.\n5. **Not writing the 30/60/90 down.** Share it with your manager in week one and revise it in week six. It is the artifact that converts a vague review into a factual one.\n\n## Frequently asked questions\n\n**What should a UX researcher do in the first 30 days?**\nLearn the system, not the users. Audit existing research, map how decisions actually get made, and interview 8-12 stakeholders about what they believe and cannot prove.\n\n**What is a good first project?**\nSomething small, decision-linked, and already wanted - ideally answering a question a specific person is actively blocked on, completable in under three weeks.\n\n**Should you run research in your first month?**\nYes, but on the organisation. Stakeholder interviews are real research and produce the map you need before any customer study is worth running.\n\n**How do you build credibility as the first researcher at a company?**\nShip something small quickly, attach it to a decision someone cares about, and let the person who acted on it tell the story. Credibility is transferred, not asserted.\n\n**What does success look like at day 90?**\nOne study shipped and acted on, a working map of decisions and evidence, and exactly one process improved with a measure that moved.\n\n**How does AI-assisted tooling change this?**\nIt moves your first real study from month three to week two, because the blockers it removes - recruiting, scheduling, moderating, transcribing - are precisely the ones a new hire has no relationships to solve.\n\n## Related Resources\n\n- [UX Researcher Career Path](/docs/ux-researcher-career-path) - levels, competencies, and promotion criteria\n- [UX Research Team Structure](/docs/ux-research-team-structure) - centralised, embedded, and hub-and-spoke models compared\n- [How to Hire a UX Researcher](/docs/hiring-ux-researcher-guide) - the other side of this process\n- [ResearchOps: The Complete Guide](/docs/research-ops-guide) - the processes you will be improving in days 61-90\n- [Stakeholder Interviews](/docs/stakeholder-interview-guide) - the method for your first month\n- [How to Build a UX Research Repository](/docs/research-repository-guide) - the most common day-61 project\n- [User Research Maturity Model](/docs/user-research-maturity-model) - diagnosing which stage you have joined\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - making your first findings hard to dismiss\n\n---\n\n**Sources:** Gallup, *Why the Onboarding Experience Is Key for Retention*; Nielsen Norman Group, *Successful Onboarding for New Hires in UX Roles*; Maze, *Future of User Research 2026*.","category":"Research Operations","lastModified":"2026-08-01T03:20:59.339339+00:00","metaTitle":"First 90 Days as a UX Researcher: A 30/60/90 Onboarding Plan","metaDescription":"A week-by-week onboarding plan for new UX researchers: the week-one audit, stakeholder interviews as your first study, which project to ship by day 60, and what good looks like at day 90.","keywords":["first 90 days ux researcher","ux researcher onboarding","30 60 90 day plan ux research","new researcher onboarding","first ux research hire","research onboarding plan"],"aiSummary":"A new UX researcher should spend the first 30 days learning the system rather than producing findings - auditing past research, mapping decision-making, and interviewing stakeholders as if they were a research project. Days 31 to 60 are for shipping one small study end to end to establish credibility and expose process friction. Days 61 to 90 are for changing exactly one process. This guide gives the week-by-week plan, the week-one audit checklist, a stakeholder interview script, day-90 success criteria, and how AI-native tooling lets a new hire ship real findings in week two instead of month three.","aiPrerequisites":["A new or incoming user research role","Basic familiarity with common UX research methods"],"aiLearningOutcomes":["Structure a 30/60/90 day plan for a new research role","Run a week-one audit of an existing research practice","Interview stakeholders as a research project rather than a formality","Choose a first study that builds credibility instead of consuming it","Define measurable day-90 success criteria"],"aiDifficulty":"beginner","aiEstimatedTime":"12 min"},{"type":"documentation","id":"dc164ea4-fdb0-4212-834b-3a0e841fdd29","slug":"usability-issue-severity-ratings","title":"Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)","url":"https://www.koji.so/docs/usability-issue-severity-ratings","summary":"Severity rates how badly a usability problem harms users, combining frequency, impact, persistence, and market impact on Nielsen's 0-4 scale. Single-evaluator ratings are unreliable: the evaluator effect produces 5-65% any-two agreement (Hertzum & Jacobsen), and one replication found 41% pairwise agreement with 47% of unique problems found by only one of four evaluators. Use three independent raters, report mean and spread, keep severity separate from priority, and measure frequency with larger asynchronous samples instead of estimating it from five participants.","content":"# Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)\n\n**Short answer:** A severity rating scores how badly a usability problem hurts users, combining **frequency, impact, persistence, and market impact**. The standard is Jakob Nielsen's 0–4 scale, from \"not a problem\" to \"usability catastrophe.\" The critical caveat is that severity ratings are judgments, and judgments vary enormously between evaluators — Nielsen is explicit that **\"severity ratings from a single evaluator are too unreliable to be trusted,\"** and recommends averaging ratings from three evaluators. Severity is not the same thing as priority: severity measures user harm, priority weighs harm against reach and cost to fix.\n\nEvery usability study ends the same way. You have 47 observations, a stakeholder with capacity for four, and a meeting in which the loudest opinion wins. Severity ratings exist to replace that meeting with something defensible.\n\nThey are also the single most casually applied technique in UX. Most teams assign severity in a rush at the end of synthesis, by one person, using an undefined scale, and then present the numbers as if they were measurements. The evidence says that is closer to a coin flip than most practitioners believe — and the fix is neither complicated nor expensive.\n\n## The standard scale: Nielsen's 0–4\n\nThe reference scale comes from Jakob Nielsen's severity ratings work at [Nielsen Norman Group](https://www.nngroup.com/articles/how-to-rate-the-severity-of-usability-problems/):\n\n| Rating | Meaning | Implied action |\n|---|---|---|\n| **0** | Not a usability problem at all | Drop it; log the disagreement |\n| **1** | Cosmetic problem only | Fix only if spare time exists |\n| **2** | Minor usability problem | Low priority |\n| **3** | Major usability problem | Important to fix; high priority |\n| **4** | Usability catastrophe | Imperative to fix before release |\n\nThe 0 rating is not filler. It is the escape hatch that lets a rater say \"I do not agree this is a problem,\" and the volume of 0s in your data is a direct measure of how much your findings list is padded with observations that are actually preferences.\n\n## What severity actually combines\n\nSeverity is not a single dimension. Nielsen defines it as a combination of four factors:\n\n| Factor | Question | Common failure |\n|---|---|---|\n| **Frequency** | Is this common or rare? | Estimated from five participants and reported as fact |\n| **Impact** | Can users overcome it easily or not? | Confused with how annoying the evaluator found it |\n| **Persistence** | One-time hurdle, or does it bite repeatedly? | Ignored entirely — the most under-weighted factor |\n| **Market impact** | Does it affect the product's appeal and adoption? | Either forgotten, or used to smuggle in business opinion |\n\nPersistence deserves special attention because it inverts intuitions. A confusing first-run flow that users solve once and never think about again is often rated 4 in the room and deserves a 2. A mildly awkward interaction on a screen someone visits forty times a day is often rated 2 and deserves a 4. Frequency of *encounter* and persistence of *pain* are different axes, and collapsing them is the most common severity error there is.\n\n## The uncomfortable evidence: the evaluator effect\n\nHere is what the research says about how much you should trust a severity number.\n\nHertzum and Jacobsen's review of usability evaluation methods — the paper that named the **evaluator effect** — found that the average agreement between any two evaluators assessing the same system with the same method ranges from **5% to 65%**, with no single method consistently outperforming the others. The effect appears both in *which* problems get detected and in *how severely* they get rated ([Hertzum & Jacobsen, IJHCI 2003](https://mortenhertzum.dk/publ/IJHCI2003.pdf)).\n\nA more recent replication in unmoderated testing found the same pattern in practice. Four evaluators — ranging from about 100 sessions of experience to more than 10,000 — independently analysed the same unmoderated study. They logged 119 issue instances that consolidated to **38 unique problems**. Pairwise agreement averaged **41%**, ranging from 25% to 48%. Most strikingly, **47% of the unique problems were found by only one of the four evaluators**, and only **18% were found by all four** ([MeasuringU](https://measuringu.com/examining-the-evaluator-effect-in-unmoderated-usability-testing/)).\n\nRead that again: nearly half the findings in a typical study exist because one particular person watched the sessions. Change the analyst and you change the report.\n\nThis is not an argument against usability testing. It is an argument against single-rater severity, and Nielsen's guidance follows directly from it: the mean of ratings from **three evaluators** is satisfactory for practical purposes, and rating quality improves rapidly as raters are added.\n\n## Choose a scale you can actually apply\n\nThe 0–4 scale is the standard, but it is not the only defensible option, and there is a real case for fewer levels. MeasuringU reduced their own scale from seven points to three — **Minor, Moderate, Critical**, plus a separate non-problem category for insights and suggestions — after repeatedly finding the distinction between middle categories murky in practice.\n\n| Scale | When to use |\n|---|---|\n| **Nielsen 0–4** | Formal reports, heuristic evaluation, regulated contexts, when you need a defensible standard |\n| **3-level (Minor / Moderate / Critical)** | Fast-cycle product teams; raters apply it more consistently |\n| **Binary blocker / non-blocker** | Pre-release triage only; loses too much for research reporting |\n\nWhatever you pick, **write the anchors down and give each level an observable definition**, not an adjectival one. \"Causes task failure\" is observable. \"Serious\" is not. A rubric with observable anchors is the cheapest available improvement to inter-rater agreement — the same principle that governs [inter-rater reliability in qualitative coding](/docs/inter-rater-reliability-qualitative-research).\n\nA workable set of anchors:\n\n- **Critical (4):** Participant could not complete the task, or completed it incorrectly without realising. Data loss, security exposure, or an accessibility barrier that excludes a user group entirely.\n- **Major (3):** Participant completed the task only after a workaround, backtracking, or help. Substantial time cost or visible frustration.\n- **Minor (2):** Participant hesitated, took a wrong turn, and self-corrected within seconds. Task succeeded.\n- **Cosmetic (1):** Noticed and commented on, but no effect on task performance.\n- **Not a problem (0):** The rater disagrees that this is a usability issue.\n\n## Severity is not priority\n\nThis is the distinction that determines whether your severity ratings survive the roadmap meeting.\n\n**Severity** is a property of the problem's effect on users. It does not know or care what it costs to fix.\n\n**Priority** is a business decision that combines severity with reach, effort, strategic value, and risk. A cosmetic issue on the checkout page that 100% of paying customers see can rationally outrank a catastrophe in an admin screen used by nine people once a quarter.\n\nKeep them in separate columns. The moment you let fix-cost leak into the severity rating, you lose the ability to say \"we knowingly shipped a major usability problem because the fix was expensive\" — which is exactly the sentence that protects a research team's credibility when the issue resurfaces in support tickets six months later.\n\nThe clean handoff is: research owns severity, and the product team feeds severity into whatever prioritisation framework it already runs — [RICE](/docs/rice-prioritization-framework), [ICE](/docs/ice-prioritization-framework), or a [value vs. effort matrix](/docs/value-vs-effort-prioritization-matrix). Severity is an input to prioritisation, not a competitor to it.\n\n## A rating protocol that takes 45 minutes\n\n1. **Compile the raw findings list** with a one-line observable description of each problem, the participant IDs who hit it, and a timestamp or clip.\n2. **Rate independently, in silence.** Three raters minimum. Include at least one person who did not run the sessions — the evaluator effect is strongest among people who share a mental model.\n3. **Never rate in a group first.** Group rating produces consensus, not agreement. The first confident voice anchors everyone else, and you lose the disagreement signal that is the most useful output of the exercise.\n4. **Compute the mean and the spread.** Report both. A problem rated 4, 4, 4 is a different object from one rated 4, 3, 1, even though the second averages to a respectable 2.7.\n5. **Discuss only the split items.** Anything with a range of two or more points gets a five-minute conversation. Usually one rater knows something the others do not — a support-ticket volume, an accessibility implication, a known workaround. That knowledge is the actual finding.\n6. **Record the rationale for anything rated 4.** Catastrophes get challenged. Write down the frequency and impact evidence at rating time, not when someone questions it in a roadmap review.\n\n## Fix the weakest input: stop guessing frequency\n\nLook back at the four factors. Two of them — impact and persistence — genuinely require expert judgment. But **frequency is not a judgment. It is a measurement**, and almost every team estimates it from five participants because measuring it properly used to be prohibitively expensive.\n\nThat is the real leverage point, and it is where a modern research stack changes the arithmetic.\n\nRunning a study through Koji lets you replace the guessed inputs with measured ones:\n\n- **Frequency becomes an actual rate.** Because [AI-moderated sessions](/docs/ai-usability-testing-guide) run asynchronously and in parallel, running 40 or 60 participants costs roughly what scheduling 8 used to. \"3 of 5 participants\" becomes \"38% of 60 participants,\" and a severity rating built on that number survives scrutiny in a way the first one never does.\n- **Impact gets measured from the participant, not inferred by the observer.** Use Koji's [structured questions](/docs/structured-questions-guide) directly in the flow: a `scale` question after each task captures perceived effort — the same logic behind the [Single Ease Question](/docs/single-ease-question-seq-guide) — and a `yes_no` question captures self-reported success, which you can compare against observed success to catch the users who failed without knowing it. That gap is your highest-severity population.\n- **Persistence gets asked instead of assumed.** An `open_ended` question — \"If you hit this again tomorrow, what would you do?\" — with AI follow-up probing distinguishes a one-time hurdle from a recurring one in the participant's own words.\n- **Relative harm comes from a `ranking` question.** Have participants rank the problems they encountered. Evaluators are systematically bad at guessing which annoyance users actually care about, and a ranked aggregate is a direct corrective.\n- **`single_choice` and `multiple_choice`** let you segment severity by user type, which frequently reveals that a \"minor\" issue is a catastrophe for one segment.\n\nKoji's automatic thematic analysis clusters the same problem across dozens of sessions so the frequency count is produced for you rather than tallied by hand, and every session carries a quality score on a 1–5 scale so you can tell which sessions actually carried signal. The point is not that AI replaces the rater. It is that the rater ends up judging two factors instead of four, with real numbers underneath the other two.\n\n## Reporting severity so it drives action\n\n- **Lead with the 4s and 3s.** Nobody reads a table of 47 rows. Put the catastrophes and majors above the fold with evidence clips.\n- **Show the rater spread**, not just the mean. Disagreement is information about how confident the team should be.\n- **Attach evidence to every rating above 2.** A timestamp, a quote, a clip. Severity claims without evidence get relitigated forever.\n- **Report frequency as a rate with a denominator.** \"12 of 60\" beats \"several participants\" every time.\n- **Never present severity as a fix order.** Hand it to the prioritisation framework and say so explicitly.\n\nFor the full report structure this fits into, see the [UX research report guide](/docs/ux-research-report-template).\n\n## Common mistakes\n\n1. **One person rating everything.** The evidence is unambiguous on this.\n2. **Group rating before independent rating.** You get consensus, and lose the disagreement signal.\n3. **Letting fix cost into the severity score.** Now you cannot separate \"not bad\" from \"not worth fixing.\"\n4. **Ignoring persistence.** The most under-weighted of the four factors.\n5. **Reporting the mean without the range.** 4/3/1 and 3/3/2 are not the same finding.\n6. **Guessing frequency from five participants and stating it as fact.** Measure it.\n7. **No 0 option.** Without it, every observation becomes a problem, and your list inflates.\n\n## Frequently asked questions\n\n**What is the standard usability severity rating scale?** Jakob Nielsen's 0–4 scale is the standard: 0 = not a usability problem, 1 = cosmetic, 2 = minor, 3 = major, 4 = usability catastrophe. Severity combines four factors — frequency, impact, persistence, and market impact. Some teams use a simpler three-level Minor / Moderate / Critical scale, which raters tend to apply more consistently because the middle distinctions in longer scales are hard to hold steady.\n\n**How many people should rate severity?** At least three. Nielsen states that severity ratings from a single evaluator are too unreliable to be trusted, and that the mean of three evaluators' ratings is satisfactory for practical purposes, with quality improving rapidly as raters are added. Rate independently first, then discuss only the items where ratings diverge by two or more points.\n\n**What is the evaluator effect?** It is the finding that different evaluators analysing the same sessions identify different problems and rate them differently. Hertzum and Jacobsen found average any-two agreement between evaluators ranging from 5% to 65% across methods. A replication in unmoderated testing found 41% average pairwise agreement, with 47% of unique problems detected by only one of four evaluators. It is the main reason single-rater severity should not be trusted.\n\n**What is the difference between severity and priority?** Severity describes how much a problem harms users, independent of what fixing it costs. Priority is a business decision that combines severity with reach, effort, strategic value, and risk. Keep them in separate columns so you can knowingly defer a major issue without pretending it is minor — and feed severity into a prioritisation framework like RICE or ICE rather than treating it as a fix order.\n\n**How do I estimate frequency with only five participants?** Honestly, you cannot with much confidence, and the correct response is to report it as a raw count with the denominator rather than a percentage. The better answer is to stop being limited to five. Asynchronous AI-moderated testing makes 40 to 60 participants practical at a cost close to what eight scheduled sessions used to require, which converts frequency from an estimate into a measurement.\n\n**Should severity ratings come from evaluators or from users?** Both, for different factors. Impact and persistence benefit from expert judgment, because participants often cannot tell how much time they lost or whether a workaround will scale. Frequency should be measured, and relative harm is best captured by asking participants directly — a ranking question across the problems they encountered corrects for the fact that evaluators are poor at guessing which annoyances users genuinely care about.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types for measuring impact and frequency\n- [How to Conduct Usability Testing](/docs/usability-testing-guide) — the complete method this fits into\n- [Usability Metrics: Task Success, Time on Task, and Error Rate](/docs/usability-metrics-guide) — the quantitative companions to severity\n- [Heuristic Evaluation Guide](/docs/heuristic-evaluation-guide) — where severity ratings are most commonly applied\n- [Single Ease Question (SEQ)](/docs/single-ease-question-seq-guide) — measuring perceived task difficulty\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — improving agreement between raters\n- [RICE Prioritization Framework](/docs/rice-prioritization-framework) — where severity goes once research hands it off\n- [UX Research Report Template](/docs/ux-research-report-template) — reporting severity so it drives action\n","category":"Research Methods","lastModified":"2026-08-01T03:20:31.392091+00:00","metaTitle":"Usability Severity Ratings: Nielsen's 0-4 Scale & How to Apply It (2026)","metaDescription":"How to rate usability problem severity with Nielsen's 0-4 scale, why single-evaluator ratings are unreliable (41% agreement), how to separate severity from priority, and how to measure frequency instead of guessing it.","keywords":["usability severity ratings","severity rating scale usability","how to prioritize usability issues","nielsen severity scale","ux issue severity","usability problem severity","evaluator effect","severity vs priority"],"aiSummary":"Severity rates how badly a usability problem harms users, combining frequency, impact, persistence, and market impact on Nielsen's 0-4 scale. Single-evaluator ratings are unreliable: the evaluator effect produces 5-65% any-two agreement (Hertzum & Jacobsen), and one replication found 41% pairwise agreement with 47% of unique problems found by only one of four evaluators. Use three independent raters, report mean and spread, keep severity separate from priority, and measure frequency with larger asynchronous samples instead of estimating it from five participants.","aiPrerequisites":["Familiarity with usability testing or heuristic evaluation","A completed study with a raw findings list"],"aiLearningOutcomes":["Apply Nielsen's 0-4 severity scale with observable anchors","Explain the evaluator effect and why single-rater severity is unreliable","Run an independent-then-calibrate severity rating session with three raters","Separate severity from priority and hand off cleanly to a prioritisation framework","Replace guessed frequency estimates with measured rates from larger samples"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"},{"type":"documentation","id":"91ca94c6-8675-4b35-be0b-86234abb787f","slug":"works-council-employee-research","title":"Works Councils and Employee Research: How to Run Employee Studies in Germany and the EU","url":"https://www.koji.so/docs/works-council-employee-research","summary":"In Germany, employee research engages full works council co-determination under section 87(1) No. 6 BetrVG, which is triggered by a technical system being objectively suitable for monitoring rather than by employer intent; bypassing it can render the measure void and expose the employer to an injunction. Separately, employee consent is generally not a valid GDPR basis because the EDPB treats the employment power imbalance as preventing freely given consent, so legitimate interests plus safeguards written into a works agreement (Betriebsvereinbarung) is the workable route, with Article 88 GDPR national rules layered on top. The unlock is architectural anonymity: no individual-level output, a suppression threshold of 5 to 10 per reporting cell, restrictions on cross-tabulation stacking, agreed verbatim de-identification, no non-respondent tracking, defined retention and a named access list. A works agreement should cover twelve items including explicit AI disclosure and a written commitment that no output is attributed to an individual. Consulting the works council before tool selection, and piloting to show the real report format, compresses a six-week negotiation to roughly two.","content":"**Answer first: in Germany and much of continental Europe, launching an employee survey without involving the works council can make the measure legally void, and the consent form you would use for customer research is the wrong instrument entirely — because employee consent is generally not considered freely given.** Those are the two facts that surprise most product and people teams, usually a week before launch. The good news is that both problems have the same solution: design the study so it genuinely cannot identify individuals, then get that design blessed in a works agreement before you send a single invitation.\n\nThis guide covers the co-determination trigger, why consent fails, what a works agreement should contain, the technical design that satisfies both, and the sequence that turns a six-week negotiation into a two-week one.\n\n## Why an employee survey is a co-determination matter\n\nUnder the German Works Constitution Act (Betriebsverfassungsgesetz, BetrVG), a works council holds three tiers of rights: information and hearing rights on personnel actions, consultation on economic changes, and **full co-determination on the social matters listed in §87**. Employee research usually lands squarely in that third tier.\n\nThe critical provision is **§87(1) No. 6**, which gives the works council full co-determination over the introduction and use of technical devices designed to monitor the behaviour or performance of employees. German labour courts read this expansively: it is enough that the system is *objectively suitable* for monitoring, whether or not you intend to use it that way. A platform that records who responded, when, from where, and stores free-text answers is objectively suitable. **§87(1) No. 1**, on matters of order and conduct in the establishment, is frequently engaged too, and the 2021 Works Council Modernisation Act explicitly extended council involvement into artificial intelligence and mobile working.\n\nFull co-determination means what it says. The employer cannot act unilaterally: the council must actively agree. If agreement cannot be reached, the matter goes to a conciliation board (Einigungsstelle), whose ruling replaces agreement. **If you bypass the right, the measure can be legally void, and the council can seek an injunction** — which in practice means the study stops, sometimes after you have already collected data you now cannot use.\n\nComparable structures exist elsewhere. The Netherlands gives works councils consent rights over personnel-data systems, France requires consultation of the CSE on measures affecting working conditions and monitoring, and multinational rollouts can additionally engage a European Works Council under the EWC Directive. The German analysis is the strictest and the most useful to design against.\n\n## Why employee consent is the wrong legal basis\n\nSeparately from labour law, the GDPR governs the data. The instinct is to collect consent. That instinct is wrong.\n\nThe European Data Protection Board has been consistent that, because of the **inherent imbalance of power** in an employment relationship, employee consent is unlikely to be freely given. Consent is only valid where a person can refuse or withdraw without detriment, and an employee asked by their employer to participate in a study is rarely in that position. Building your compliance case on consent means building it on a basis a regulator is predisposed to reject.\n\n**Article 88 GDPR** allows member states to make more specific rules for employment-context processing, and several — Germany, France, the Netherlands — have done so. That is why employee research in Europe is a national-law question layered on top of an EU-law question, and why a single pan-European template rarely survives contact with local counsel.\n\nThe practical consequence: rely on a basis that fits the situation — typically **legitimate interests** for genuinely voluntary, aggregate-only research, or a legal obligation where a survey is required by other law — and put the safeguards in the works agreement instead of in a consent checkbox. Where consent does appear, it should be for genuinely optional extras (agreeing to a follow-up conversation, agreeing to be quoted), never for participation itself.\n\n## Anonymity is a technical requirement, not a promise\n\nHere is the leverage point. Almost every objection a works council raises — monitoring, performance inference, retaliation risk, manager-level scrutiny — dissolves if the study is *architecturally* incapable of identifying individuals. Councils are not opposed to hearing from employees. They are opposed to a system that could be turned against them.\n\nA design that survives scrutiny does all of the following:\n\n| Safeguard | What it means concretely |\n|---|---|\n| No individual-level output | The employer never receives a per-person record, only aggregates |\n| Minimum reporting threshold | No breakdown is displayed below a floor — commonly 5 to 10 respondents per cell |\n| No cross-tabulation stacking | Filters cannot be combined until a group becomes identifiable by elimination |\n| Restricted demographics | Collect only the segments you will actually act on; drop the ones that triangulate |\n| Verbatim handling agreed in advance | Either free text is suppressed, or it is reviewed and de-identified before anyone in management sees it |\n| No response tracking to individuals | Reminders go to everyone, not to named non-respondents |\n| Defined retention | A deletion date for raw data, written into the agreement |\n| Named access list | Who can see what, agreed with the council — see [research data access controls](/docs/research-data-access-controls-audit-trail) |\n\nThe threshold rule is the one people underestimate. In a 900-person company, \"Engineering, Munich, women, senior\" is a group of four, and everyone in the room can name them. Suppression thresholds are what make an anonymity promise structurally true rather than merely sincere.\n\n## What belongs in the works agreement\n\nA Betriebsvereinbarung covering employee research should be specific enough that the council does not have to trust you and general enough that you do not renegotiate every quarter. Cover:\n\n1. **Purpose** — what you are trying to learn, and the explicit exclusion of performance assessment.\n2. **Scope and cadence** — which populations, how often, and a cap on frequency so employees are not surveyed weekly.\n3. **Voluntariness** — participation is voluntary and non-participation carries no consequence, stated in the invitation.\n4. **The tool** — named platform, where data is hosted, the processor agreement, and sub-processors.\n5. **Data categories** — exactly which fields are collected, including metadata such as timestamps and device information.\n6. **Anonymity architecture** — the suppression threshold, the cross-tab restrictions, and the ban on individual-level export.\n7. **Access** — who sees raw data, who sees aggregates, and how that is enforced and logged.\n8. **Verbatims** — the de-identification process, and who performs it.\n9. **AI processing** — if an AI conducts or analyses the interviews, say so plainly: what the model does, what it is not permitted to do, whether outputs feed any decision about an individual (they should not), and that a human reviews findings.\n10. **Retention and deletion** — dates, not intentions.\n11. **Council access to results** — councils routinely, and reasonably, want to see the same aggregate findings management sees.\n12. **Review and termination** — how the agreement is revisited.\n\nItem 9 is now the one that stalls negotiations. Councils have become alert to AI in the workplace, and the fastest way through is total specificity about what the AI does and a hard, written commitment that no output is ever attributed to or used against an individual.\n\n## Where Koji fits\n\nKoji's design happens to line up well with what a works council needs, because the platform was built for research rather than for management reporting.\n\nThe AI interviewer conducts a genuine conversation and probes follow-up questions, which is what makes an aggregate-only study worth running at all — the depth that normally requires a named facilitator arrives without one. Nobody in HR sits in the room, and nobody needs to listen back to identify a voice, because the analysis is automated. In practice this is easier to defend than a traditional focus group or a manager-led round of one-to-ones, where the person hearing the answer is the person writing the review.\n\nThe **six structured question types** — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — matter here for a specific reason: the five non-open types produce aggregate values by construction. A study whose backbone is scales, choices and rankings, with open-ended probing on the questions that genuinely need reasoning, gives you a report that is largely composed of distributions rather than quotable text. That is a materially easier artefact to agree an anonymity architecture around than a pile of free-text responses.\n\nTeam access uses **owner, admin and member** roles with workspace scoping, so \"only these three named people can reach the study\" is an enforceable statement rather than a policy aspiration. Combine that with an agreed suppression threshold applied at reporting time and you have most of the technical safeguards the agreement will ask for.\n\nBe straightforward about the limits, too. Any system that collects responses holds metadata; the honest position with a works council is not \"nothing is recorded,\" it is \"here is exactly what is recorded, here is who can see it, here is when it is deleted, and here is why none of it can be resolved to a person in any output you or we receive.\"\n\n## The sequence that saves six weeks\n\nTeams lose time by building the study first and consulting the council last. Reverse it.\n\n| Week | Step |\n|---|---|\n| 0 | Brief the works council on the *intent* before choosing a tool. Councils object far less to being consulted early than to being presented with a finished plan |\n| 1 | Agree the anonymity architecture in principle — threshold, no individual output, no performance use |\n| 1–2 | Run the tool selection with the council informed; share the processor agreement and hosting details |\n| 2 | Draft the works agreement against the twelve items above |\n| 3 | Complete the data protection impact assessment where required, and align it with the agreement |\n| 3–4 | Sign, then pilot with a small volunteer group and show the council the actual output format |\n| 4+ | Launch, and share aggregate results with the council on the same timeline as management |\n\nThe pilot in week 3–4 is the highest-leverage step in the list. Showing a council the real report — visibly free of anything that could identify anyone — resolves more objections than any amount of written assurance.\n\n## Two failure modes to avoid\n\n**Running a \"quick pulse\" outside the agreement.** Small, informal surveys are exactly where co-determination gets bypassed, and they set a precedent that poisons the negotiation for the programme you actually care about. Bring the pulse inside the agreement's cadence clause instead.\n\n**Promising anonymity you cannot deliver.** If the survey has 30 respondents and you report by team, you have not run an anonymous study, whatever the invitation said. The credibility cost of one identifiable finding is years long — every future study gets lower participation and more guarded answers. See [anonymous employee research with AI interviews](/docs/anonymous-employee-research-ai-interviews) for the design detail.\n\n## Frequently asked questions\n\n**Do we need works council approval for every employee survey?**\nIn Germany, assume yes for anything using a technical system to collect employee responses, because §87(1) No. 6 BetrVG is triggered by a system's objective suitability for monitoring rather than by your intention. The efficient answer is one framework works agreement covering employee research generally, with a light notification step per study, rather than a fresh negotiation each time.\n\n**Can we rely on employee consent instead of a works agreement?**\nNo, for two separate reasons. Consent is a data protection concept and does not displace a labour-law co-determination right at all. And within data protection, the EDPB's position is that employee consent is unlikely to be freely given because of the power imbalance, so it is a weak basis even on its own terms. Use a works agreement plus an appropriate legal basis, and reserve consent for genuinely optional extras.\n\n**What happens if we launch without involving the works council?**\nThe measure can be treated as legally void, the council can seek an injunction to stop it, and you may be unable to use data already collected. Beyond the legal exposure, it is a relationship cost that makes every subsequent programme harder.\n\n**Does this apply outside Germany?**\nThe specific §87 mechanism is German, but the pattern is not. The Netherlands gives works councils consent rights over personnel-data systems, France requires CSE consultation on monitoring and working conditions, and multinational rollouts can engage a European Works Council. Design to the German standard and you will usually clear the others.\n\n**How small can a reporting group be before anonymity breaks?**\nThere is no single legal number, and the honest answer depends on how well colleagues know each other. A threshold of 5 is a common floor and 10 is safer for sensitive topics. What matters more than the number is preventing filter stacking, since combining three innocuous filters is how a group of 200 becomes a group of 3.\n\n**Can we use AI to conduct employee interviews in a co-determined workplace?**\nYes, and it is often easier to agree than a human-moderated alternative, because no colleague hears the answer. The conditions are specificity and restraint: describe exactly what the AI does, commit in writing that no output is attributed to or used against an individual, keep a human in the loop on findings, and disclose the AI's role to employees in the invitation.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make aggregate-only reporting practical\n- [Anonymous Employee Research with AI Interviews](/docs/anonymous-employee-research-ai-interviews) — the design that keeps an anonymity promise\n- [Koji for HR and People Teams](/docs/koji-for-hr-people-teams) — running employee research at scale\n- [Voice of the Employee](/docs/voice-of-employee-program) — building a listening programme that drives change\n- [Employee AI Adoption Research](/docs/employee-ai-adoption-research) — the study most likely to need a works agreement in 2026\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the underlying data protection baseline\n- [Research Data Access Controls and Audit Trails](/docs/research-data-access-controls-audit-trail) — enforcing the access limits your agreement promises","category":"Research Operations","lastModified":"2026-08-01T03:19:46.201697+00:00","metaTitle":"Works Councils and Employee Research: Germany and EU Compliance Guide","metaDescription":"Employee surveys trigger §87 BetrVG co-determination, and employee consent is rarely a valid GDPR basis. How to get a works agreement, design for real anonymity, and launch without being blocked.","keywords":["works council employee research","Betriebsrat employee survey","BetrVG section 87 co-determination","works agreement employee survey","employee research Germany GDPR","employee consent not freely given","European employee research compliance","Betriebsvereinbarung Mitarbeiterbefragung"],"aiSummary":"In Germany, employee research engages full works council co-determination under section 87(1) No. 6 BetrVG, which is triggered by a technical system being objectively suitable for monitoring rather than by employer intent; bypassing it can render the measure void and expose the employer to an injunction. Separately, employee consent is generally not a valid GDPR basis because the EDPB treats the employment power imbalance as preventing freely given consent, so legitimate interests plus safeguards written into a works agreement (Betriebsvereinbarung) is the workable route, with Article 88 GDPR national rules layered on top. The unlock is architectural anonymity: no individual-level output, a suppression threshold of 5 to 10 per reporting cell, restrictions on cross-tabulation stacking, agreed verbatim de-identification, no non-respondent tracking, defined retention and a named access list. A works agreement should cover twelve items including explicit AI disclosure and a written commitment that no output is attributed to an individual. Consulting the works council before tool selection, and piloting to show the real report format, compresses a six-week negotiation to roughly two.","aiPrerequisites":["An employee population in Germany or another co-determined European jurisdiction","An existing works council or employee representative body","A planned employee survey, pulse or interview programme"],"aiLearningOutcomes":["Recognise when an employee study triggers section 87 BetrVG co-determination","Explain why employee consent is not a reliable GDPR legal basis","Draft the twelve clauses a works agreement for employee research should contain","Design a study that is architecturally incapable of identifying individuals","Set suppression thresholds and prevent cross-tabulation triangulation","Sequence council engagement to avoid a blocked or voided launch"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"f684011c-84ca-453a-8fe1-ea79ea0a52c9","slug":"ux-research-certifications","title":"UX Research Certifications in 2026: Which Ones Are Worth It (and Which Are Not)","url":"https://www.koji.so/docs/ux-research-certifications","summary":"UX research certifications prove curriculum completion, not competence, and hiring loops evaluate portfolios instead. Costs: Google UX Design Certificate ~$294, HFI CUA exam $800, UXQB CPUX-F ~GBP 300-350 (curriculum free), NN/g UX Certification ~$6,450 for five courses, NN/g Master ~$19,350. Certification pays in four cases: career change with no portfolio, employer L&D budget, regulated/procurement environments, and self-taught practitioners with blind spots. Otherwise, invest in running real studies.","content":"# UX Research Certifications in 2026: Which Ones Are Worth It (and Which Are Not)\n\n**Short answer:** A UX research certification proves you completed a curriculum. It does not prove you can run a study, and no hiring manager treats it as if it does. That makes certifications worth real money in four specific situations — a career change with no portfolio, an employer learning-and-development budget you would otherwise forfeit, a regulated or procurement-driven environment that wants a named credential, and a self-taught practitioner with gaps they cannot see — and close to worthless outside them. Expect roughly **$800** for an exam-only credential, **~$294** for the Google certificate, and **~$6,450** for full NN/g UX Certification.\n\nThe question arrives in the same shape every time: *I have some money and some time — will a certification get me a research job, or a better one?*\n\nThe honest answer is that certifications are priced as career investments and sold as skill signals, and those two things are not the same. Jared Spool put the mechanism well when he argued that the real beneficiary of certification is **\"the hiring manager\"** — the person who needs a shortcut for evaluating credentials. A certification is worth precisely what the person on the other side of the table believes it represents, and in UX research that belief is much weaker than the price tags suggest.\n\nThis guide covers what each major credential actually costs and requires, what it signals, and how to decide.\n\n## What a certification actually is\n\nEvery UX credential on the market is one of three things:\n\n1. **A curriculum-completion certificate.** You attended N courses and passed N exams on their content. NN/g and the Google certificate are this. The exam tests the curriculum, not your practice.\n2. **A knowledge exam.** You demonstrated command of a body of principles, with or without training. HFI's CUA and the UXQB CPUX exams are this — you can sit them cold.\n3. **A portfolio or practice review.** Someone assessed work you actually did. Essentially nobody offers this at scale in UX research, which is exactly why the credential signal is weak.\n\nNotice what is missing across all three: none of them observe you running a session, writing a screener, handling a stakeholder who wants to lead the witness, or recovering from a study that fell apart in the first two interviews. Those are the skills that determine whether you are good at this job.\n\n## The major credentials compared\n\nPrices below are as listed at the time of writing (August 2026). Verify current pricing before committing — these programmes revise fees regularly.\n\n| Credential | Structure | Approx. cost | Expires? | Best for |\n|---|---|---|---|---|\n| **NN/g UX Certification** | 5 full-day courses, each with an exam | **~$6,450** | No | Practitioners with employer L&D budget |\n| **NN/g UX Master Certification** | 15 courses + exams within 3 years | **~$19,350** | No | Senior practitioners, teaching/consulting credibility |\n| **NN/g Research specialty** | 5 research-focused courses | Included in the above | No | Generalists formalising a research focus |\n| **HFI Certified Usability Analyst (CUA)** | Exam only, or exam plus a 4-course bundle | **$800** exam-only; ~$4,980 bundle | Check current terms | Enterprise and regulated environments |\n| **HFI Certified User Experience Analyst (CXA)** | Advanced exam; requires CUA first | **$850** exam-only | Check current terms | Existing CUAs going deeper |\n| **UXQB CPUX-F (Foundation)** | Exam; curriculum published free | **~£300–350** | No | Europe, and anyone wanting the cheapest rigorous exam |\n| **UXQB CPUX-UR / CPUX-UT** | Advanced exams: user requirements, usability testing | Varies by provider | No | Specialists in requirements or testing |\n| **Google UX Design Professional Certificate** | 7 Coursera courses, self-paced | **~$294** (at $49/mo, 3–6 months) | No | Career changers starting from zero |\n\nThree details worth knowing before you buy:\n\n- **NN/g certification does not expire.** The programme states it does not expire and is reissued as you take more courses. Certification requires passing five courses and their exams, with **each exam completed within 35 days** of attending the associated course, and NN/g reports **13,000+ professionals** have earned it. Specialties are available in Artificial Intelligence, Interaction Design, Research, and Management. In-house delivery of the five-course path is a different product entirely, quoted around **$60,000** for a team.\n- **HFI's CUA can be taken cold.** You can sit the exam without any training. It runs on a **2.5-hour** limit, and a failed attempt can be retaken within six months at no cost. That makes it the cheapest way to convert existing knowledge into a named credential — $800 with a free second attempt is a genuinely different risk profile from a $6,450 course path.\n- **UXQB publishes its curricula free.** You can study the entire CPUX-F syllabus at no cost and only pay if you want the certificate. If your goal is the knowledge rather than the line on your CV, this is the highest-value option on the list, and it costs nothing.\n\n## Does a certification get you hired?\n\nMostly, no — and the industry is fairly consistent about why.\n\nUX research job postings commonly list a bachelor's degree in psychology, HCI, or a related field, and some senior or specialist roles prefer a master's. But hiring managers routinely consider strong candidates without those degrees, particularly people with real research experience from adjacent fields. Across the guidance aimed at people entering the field, the same conclusion recurs: **the portfolio outweighs the credential**.\n\nThat is not a knock on the credentials. It is a statement about what the hiring process actually evaluates. A research interview loop asks you to walk through a study you ran, defend your sampling decisions, explain a finding you got wrong, and demonstrate that you can be told \"we already know what users want\" by a VP without folding. No exam tests any of that. A certification gets you past a keyword screen; a portfolio gets you through the loop.\n\nFor the full picture of what those loops evaluate, see the [guide to hiring a UX researcher](/docs/hiring-ux-researcher-guide) — reading the hiring side of the table is the fastest way to understand what a credential is competing against.\n\n## Four situations where a certification genuinely pays\n\n**1. You are changing careers and have no portfolio yet.** This is the strongest case. A certificate gives a recruiter a reason to keep reading and gives you a structured curriculum instead of a random reading list. The Google certificate is the standard entry point here at roughly $294 — cheap enough that the downside is small.\n\n**2. Your employer has a learning-and-development budget you would otherwise lose.** If someone else is paying, the calculus inverts completely. NN/g's five-course path is genuinely good training, and $6,450 of someone else's money buys 30+ hours of structured instruction. Take it.\n\n**3. You work in a regulated industry, agency, or procurement-driven environment.** Some buyers want a named credential on the proposal. In healthcare, government, defence, and large-enterprise vendor selection, \"our lead researcher is a CUA\" is a line that does work in a document. HFI's certifications have the deepest history here.\n\n**4. You are self-taught and suspect you have gaps you cannot see.** This is the underrated case. The value is not the certificate — it is the syllabus forcing you through the parts you have been avoiding. If this is your reason, note that UXQB publishes its curriculum for free and you can capture most of the value without paying for anything.\n\n## Four situations where it does not pay\n\n- **You already have three shippable case studies.** Spend the money on conference attendance or a coach instead. Your constraint is network and narrative, not knowledge.\n- **You are hoping it substitutes for reps.** It does not. Certification teaches method; competence comes from running studies badly and then better.\n- **You are a PM or designer who runs occasional research.** You need a repeatable process and good templates, not a credential. See the [research democratization playbook](/docs/research-democratization-playbook).\n- **You are chasing a promotion.** Promotion criteria in research ladders are built around scope, impact, and stakeholder influence — not credentials. The [UX researcher career path](/docs/ux-researcher-career-path) lays out what actually gets evaluated.\n\n## How to choose, in order\n\n1. **Is someone else paying?** If yes, take the most expensive good one available to you (NN/g). Stop here.\n2. **Do you have a portfolio?** If yes, skip certification entirely and invest in reps and visibility.\n3. **Do you need a named credential for procurement or a regulated buyer?** Take HFI CUA at $800.\n4. **Are you starting from zero?** Take the Google certificate at ~$294, then immediately start running real studies.\n5. **Do you want the knowledge and not the paper?** Read the UXQB CPUX-F curriculum for free.\n\nBudget the time as seriously as the money. NN/g certification is 30+ hours of training plus exam time. The Google certificate assumes roughly 10 hours a week for three to six months. That time has an alternative use, and for most people the alternative use with the higher return is running studies.\n\n## The modern alternative: buy reps, not paper\n\nHere is the structural problem with the certification market. Everyone agrees the portfolio matters more than the credential — and then offers you a credential, because portfolios have historically been hard to build. Building one traditionally required a research budget, a recruiting pipeline, participant incentives, scheduling, and twenty hours of transcript analysis per study. If you did not already have a research job, you could not get the experience that would get you a research job.\n\nThat constraint is what actually changed, and it changed more than any certification did.\n\nWith Koji, an individual practitioner can run a complete study end to end without a team:\n\n- **AI-moderated interviews** run asynchronously, so you are not scheduling twelve calls across three time zones. Participants join when they want, by voice or text, and the AI moderator probes follow-ups rather than reading a script.\n- **[Structured questions](/docs/structured-questions-guide)** give you the quantitative spine a credible case study needs. All six types are available — `open_ended` for the narrative, `scale` for satisfaction and confidence, `single_choice` and `multiple_choice` for segmentation, `ranking` for priority, and `yes_no` for clean binaries — so your write-up has distributions and not just quotes.\n- **Automatic thematic analysis** turns transcripts into themes with supporting quotes, collapsing the twenty-hour synthesis step that stops most side-project studies from ever finishing.\n- **Real-time reporting** means you can watch themes form as interviews land, which is also the fastest way to learn what a bad interview guide looks like.\n\nRun three studies on something you genuinely care about — a tool you use, a community you belong to, a local business willing to let you interview their customers — and write them up properly using the [user research report structure](/docs/user-research-report). Three real case studies with real participants and defensible method notes will do more in an interview loop than any certificate on this page.\n\nNone of that makes the credentials worthless. It makes them optional in a way they were not five years ago. Buy one when it removes a specific obstacle — a screening filter, a procurement requirement, a budget you would otherwise forfeit, a gap you cannot self-diagnose. Do not buy one hoping it will substitute for having done the work.\n\n## Frequently asked questions\n\n**Is a UX research certification worth it in 2026?** It depends entirely on your situation. It is worth it if you are a career changer with no portfolio, if an employer is paying, if you need a named credential for regulated or procurement contexts, or if you want a structured syllabus to close gaps you cannot see. It is not worth it if you already have case studies, because hiring loops evaluate the work you have done rather than the courses you have attended.\n\n**How much does NN/g UX Certification cost?** Approximately **$6,450** for the standard UX Certification, which requires passing five full-day courses and their exams, with each exam completed within 35 days of the associated course. The UX Master Certification requires 15 courses and runs approximately $19,350. The certification does not expire and NN/g reports more than 13,000 professionals hold it. Specialties are offered in Artificial Intelligence, Interaction Design, Research, and Management.\n\n**What is the cheapest legitimate UX certification?** HFI's Certified Usability Analyst exam at **$800** taken without courses is the cheapest recognised named credential, with a free retake inside six months. The UXQB CPUX-F exam, typically around £300–350, is cheaper still, and UXQB publishes its full curriculum free so you can study the entire syllabus at no cost and pay only for the certificate.\n\n**Is the Google UX Design Certificate enough to get a UX research job?** On its own, no. It is seven Coursera courses at roughly $49 a month, typically completed in three to six months at about ten hours a week, for a total around $294. It is a good structured introduction and a reasonable signal that you are serious, but it is design-weighted rather than research-specific, and hiring managers will still ask to see studies you have run. Treat it as the start of a portfolio, not a substitute for one.\n\n**Do UX research certifications expire?** NN/g's certification explicitly does not expire and is reissued as you complete further courses. UXQB CPUX certifications do not expire either. Check current terms for HFI credentials, as recertification requirements vary by programme and change over time.\n\n**Do I need a degree to become a UX researcher?** Not strictly. Many postings list a bachelor's in psychology, HCI, or a related field, and some senior roles prefer a master's, but hiring managers regularly consider candidates with strong research experience from adjacent fields — academia, market research, social science, journalism, data analysis. What consistently matters more is a portfolio of studies where you can defend the method, not just the finding.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that give a case study its quantitative spine\n- [How to Hire a UX Researcher](/docs/hiring-ux-researcher-guide) — what the interview loop actually evaluates\n- [UX Researcher Career Path](/docs/ux-researcher-career-path) — levels, competencies, and promotion criteria\n- [UX Researcher Salary and Team Cost Benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks) — what the credential is competing for\n- [Your First 90 Days as a UX Researcher](/docs/first-90-days-ux-researcher) — converting a new role into early credibility\n- [Research Democratization Playbook](/docs/research-democratization-playbook) — scaling method without certifying everyone\n- [How to Write a User Research Report](/docs/user-research-report) — the format your portfolio case studies should follow\n","category":"Research Operations","lastModified":"2026-08-01T03:18:09.747434+00:00","metaTitle":"UX Research Certifications 2026: NN/g, HFI CUA, UXQB & Google Compared","metaDescription":"What NN/g UX Certification (~$6,450), HFI CUA ($800), UXQB CPUX, and the Google UX Design Certificate (~$294) really cost and signal — plus the four situations where paying for one actually pays off.","keywords":["ux research certification","ux certification worth it","nng ux certification cost","hfi cua certification","uxqb cpux-f","google ux design certificate","ux researcher credentials","best ux research certification 2026"],"aiSummary":"UX research certifications prove curriculum completion, not competence, and hiring loops evaluate portfolios instead. Costs: Google UX Design Certificate ~$294, HFI CUA exam $800, UXQB CPUX-F ~GBP 300-350 (curriculum free), NN/g UX Certification ~$6,450 for five courses, NN/g Master ~$19,350. Certification pays in four cases: career change with no portfolio, employer L&D budget, regulated/procurement environments, and self-taught practitioners with blind spots. Otherwise, invest in running real studies.","aiPrerequisites":["Interest in a UX research career or team development","Basic familiarity with UX research roles"],"aiLearningOutcomes":["Compare the cost, structure, and expiry terms of the major UX research credentials","Judge whether a certification will change your hiring outcome","Identify the four situations where certification is worth the money","Choose between NN/g, HFI, UXQB, and the Google certificate for your situation","Build portfolio evidence as an alternative to paid certification"],"aiDifficulty":"beginner","aiEstimatedTime":"12 min"},{"type":"documentation","id":"2bd3c108-ae40-48b0-abd8-28daa6094037","slug":"international-privacy-laws-user-research","title":"User Research Privacy Laws Beyond GDPR and CCPA: Brazil, Canada, India, Japan, China and More","url":"https://www.koji.so/docs/international-privacy-laws-user-research","summary":"Beyond GDPR and CCPA, at least six national privacy regimes can reach a company that interviews customers abroad, and the decisive differences are extraterritorial reach and permitted legal basis rather than data transfer. Brazil LGPD offers ten legal bases including legitimate interests, but the ANPD has confirmed legitimate interests cannot cover sensitive data. Canada PIPEDA remains in force because Bill C-27 died on prorogation in January 2025; Quebec Law 25 adds opt-in consent and automated-decision notice. India DPDP Rules were notified 14 November 2025 with obligations phasing in to 13 May 2027, requiring standalone notice available in scheduled Indian languages. Japan APPI is a purpose-specification regime, and a July 2026 amendment extending cover to contactable identifiers is promulgated but not in force. China PIPL requires separate consent for sensitive data and for transfers abroad, and is the strictest for research, followed by South Korea PIPA and Quebec. A single compliant design — separate itemised opt-in consent, explicit AI disclosure, localised notices, planned handling of sensitive disclosures, easy withdrawal, and data minimisation via structured questions — satisfies all of them.","content":"**Answer first: if you interview customers outside the EU and the United States, at least six other national privacy regimes can reach you — and the differences that matter for research are not about data transfers, they are about which legal basis you are allowed to rely on and whether the law reaches a foreign company at all.** Brazil, Canada, India, Japan, China, South Korea, Australia and South Africa each answer those two questions differently. This guide maps them for customer research specifically, and ends with the single study design that satisfies the strictest of them, so you do not have to run eight variants of the same interview.\n\nThis is the *which law applies* question. If your question is where the recordings physically sit and how they legally cross a border, read [research data residency and international transfers](/docs/research-data-residency-international-transfers) instead — that is the transfer mechanism, and it is a separate problem from the one below.\n\n## The two questions that decide everything\n\n**1. Does the law reach you?** Almost all of these regimes have some form of extraterritorial reach, but the trigger differs. Brazil's LGPD applies if the processing takes place in Brazil, the data was collected in Brazil, or the purpose is offering goods or services in Brazil. China's PIPL reaches foreign entities processing the personal information of people in China for the purpose of providing products or services to them. Others, notably Canada's PIPEDA, hinge more on a real and substantial connection to the country. The practical test for a research team is: are you recruiting people located in that country to talk about a product you offer there? If yes, assume the law reaches you.\n\n**2. What can you rely on to process the data?** This is where the regimes genuinely diverge, and where copy-pasting a GDPR consent form goes wrong in both directions — sometimes it collects consent you did not need, and sometimes it fails to collect consent you did.\n\n| Jurisdiction | Law | Consent-first or basis-flexible? | What this means for interviews |\n|---|---|---|---|\n| Brazil | LGPD | Ten legal bases, including legitimate interests | Consent is not your only option, but legitimate interests is unavailable for sensitive data |\n| Canada (federal) | PIPEDA | Consent-centric, with meaningful-consent expectations | Purposes must be explained in plain language a person would actually understand |\n| Quebec | Law 25 | Opt-in consent, notably strict | Separate express consent expectations and automated-decision transparency |\n| India | DPDP Act 2023 + Rules 2025 | Consent-first, with narrow \"legitimate uses\" | Notice must be clear, standalone, and available in scheduled Indian languages |\n| Japan | APPI | Purpose-specification model rather than a consent-for-everything model | Specify the purpose of use and stay inside it; consent is required for third-party provision and most transfers abroad |\n| China | PIPL | Consent-first, with separate consent for sensitive data and for transfers abroad | Separate, specific consent — not a bundled checkbox |\n| South Korea | PIPA | Consent-first and highly granular | Consent items must be itemised and separately agreed |\n| Australia | Privacy Act / APPs | Notice-and-purpose model | Collection notice under APP 5; sensitive information generally needs consent |\n| South Africa | POPIA | Six lawful justifications including legitimate interest | Close in structure to GDPR; an Information Officer must be registered |\n\n## Brazil: LGPD\n\nThe LGPD's structure will feel familiar to anyone who has worked with the GDPR: ten legal bases, data subject rights, and a regulator, the ANPD, that has become considerably more active. For research, two points matter most.\n\nFirst, **legitimate interests is a genuine option**, and the ANPD published guidance in February 2024 setting out a three-stage balancing test — purpose, necessity, then balancing and safeguards. If your interviews are with existing customers about a product they already use, that is close to the paradigm case of a reasonable expectation.\n\nSecond, and decisively: **the ANPD has reaffirmed that legitimate interests cannot be used for sensitive personal data.** Interviews are leaky. A conversation about a banking app produces financial hardship disclosures; one about a fitness product produces health data. If your study is likely to surface sensitive categories, you need consent for that portion, and data subjects retain the right to object to legitimate-interest processing. Plan for the leak rather than being surprised by it.\n\n## Canada: PIPEDA, and the reform that did not happen\n\nPIPEDA remains the federal private-sector law. Bill C-27, which would have replaced it with the Consumer Privacy Protection Act and added an AI statute, died when Parliament was prorogued in January 2025 and has not been re-enacted. **Plan against PIPEDA as it stands, not against the bill.**\n\nPIPEDA is consent-centric, and its distinguishing feature is the *meaningful consent* standard: the regulator expects that people actually understand what they are agreeing to, with emphasis on what is collected, who it is shared with, the purposes, and the residual risk of harm. A dense scroll box does not clear that bar.\n\nQuebec is the separate problem. **Law 25** imposes opt-in consent, requires an assessment of whether collection is necessary, legitimate and proportionate to the purpose, requires parental consent for under-14s, and requires that people be informed when personal information is used to make an automated decision. If your research uses AI-moderated interviews with Quebec residents, treat that automated-decision transparency requirement as live and describe the AI's role explicitly.\n\n## India: the DPDP Act and the 2025 Rules\n\nIndia's Digital Personal Data Protection Act 2023 finally became operational when the **DPDP Rules were notified on 14 November 2025**, opening a phased implementation window with substantive obligations on data fiduciaries landing through to **13 May 2027**. Treat 2026 as a build-and-test year, not a grace period you can ignore.\n\nThree features matter for research:\n\n- **Consent is the primary route.** The Act's alternative, \"certain legitimate uses,\" is a narrow enumerated list, not a flexible balancing test. For customer interviews, get consent.\n- **The notice standard is unusually prescriptive.** Notice must be clear, standalone, understandable independently of any other document, and available in English or any language in the Eighth Schedule to the Constitution. A privacy notice in English only, buried inside terms of service, does not comply.\n- **Withdrawal must be as easy as giving consent**, and a Consent Manager framework is being operationalised through 2026 to let people manage consent across services.\n\n## Japan: purpose specification, not consent theatre\n\nJapan's APPI is often misdescribed as a consent regime. It is closer to a **purpose-specification** regime: you must specify the purpose of use, notify or publicly announce it, and not exceed it without fresh consent. Where consent is genuinely required is for providing data to third parties and, generally, for transfers to other countries.\n\nJapan is also mid-reform. A Cabinet-approved amendment bill passed the Diet on **10 July 2026** and was promulgated on **17 July 2026**, extending protection to \"contactable\" identifiers such as email addresses, phone numbers and device or cookie IDs, adding a specific category for biometric information, and strengthening protection for under-16s. The new regime is enacted but **not yet in force** — a cabinet order will set the effective date, no later than July 2028. Nothing changes today; everything changes before your current consent language is retired.\n\n## China: separate consent, every time\n\nPIPL is the strictest of the group for research, and the reason is a single word: *separate*. Bundled consent that covers everything in one checkbox is precisely what PIPL is designed to prohibit. You need separate consent for processing sensitive personal information, and separate consent for providing personal information to recipients outside China — which is what happens the instant an interview recording lands on a server elsewhere.\n\nAlso note that sensitive personal information under PIPL requires that you inform people of the necessity of the processing and the impact on their rights. If you are running research in mainland China at any scale, the transfer mechanism and the consent architecture need local advice; this is not a jurisdiction to improvise in.\n\n## Australia, South Korea and South Africa in brief\n\n**Australia** works on a notice-and-purpose model. APP 5 requires a collection notice at or before the time of collection, and sensitive information generally requires consent plus a direct relationship to your functions. The Privacy Act has been under a multi-tranche reform programme since 2024; watch it, but the APPs remain the operative rules.\n\n**South Korea's PIPA** is consent-first and granular in a way that surprises teams used to a single checkbox: consent items are expected to be itemised and separately agreed, so that a person can agree to the interview and decline the optional marketing follow-up. The PIPC has continued tightening guidance on how choices must be presented.\n\n**South Africa's POPIA** is structurally close to GDPR — six lawful justifications including legitimate interest, plus a registered Information Officer and a set of conditions for lawful processing. If your GDPR programme is real, POPIA is mostly a mapping exercise.\n\n## The one design that satisfies all of them\n\nYou do not need eight study variants. Build to the strictest common denominator and the rest follow:\n\n1. **Separate, specific, opt-in consent** captured at the start of the interview, not buried in a recruitment email. This satisfies PIPL, PIPA, DPDP and Law 25, and over-satisfies LGPD and APPI without breaking them.\n2. **Itemise the consents.** Recording, transcription, AI processing, quoting, retention period and any third-party sharing should each be separately agreeable. Bundling is the single most common failure across these regimes.\n3. **Name the AI explicitly.** Say that an AI conducts the interview, what it does with the responses, and that a human reviews outputs. This covers Quebec's automated-decision notice, the DPDP notice standard, and the transparency expectations of every regime here.\n4. **Localise the notice, not just the questions.** India effectively requires it; Japan, Korea and Brazil expect comprehension in practice. Koji runs interviews in the participant's language — see [multi-language user research](/docs/multilingual-research-guide) — and the consent text must be localised with them, not left in English.\n5. **Design for the sensitive-data leak.** Interviews wander into health, finances and family. Either take explicit consent for sensitive categories up front, or instruct the interviewer not to probe them and redact what arrives anyway.\n6. **Make withdrawal real and easy**, with a stated retention period and a deletion path that reaches exports too.\n7. **Minimise at the source.** The less identifiable data you collect, the less of this applies.\n\nThat last point is where study design does real compliance work. Koji's six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — let you capture most of what a study needs as structured, aggregate-safe values rather than as free text that inevitably contains names, employers and disclosures. A study whose quantitative backbone is scales, choices and rankings, with open-ended probing reserved for the questions that genuinely need reasoning, carries dramatically less personal data across every border in this article. Koji's AI interviewer still probes the open questions properly, so you lose no depth — you just stop collecting identifiable text you were never going to use.\n\n## A working checklist\n\n| Step | Action |\n|---|---|\n| 1 | List every country your participants are physically located in — not where your company is |\n| 2 | For each, decide whether the law reaches you (offering services there is usually enough) |\n| 3 | Pick a legal basis per country; default to explicit, itemised consent |\n| 4 | Localise notice and consent into the participant's language |\n| 5 | Decide in advance how sensitive disclosures are handled and redacted |\n| 6 | Confirm the transfer mechanism separately — that is a different analysis |\n| 7 | Record the assessment; accountability means being able to show your reasoning |\n\n## Frequently asked questions\n\n**Does my company need to comply with these laws if we have no office in that country?**\nUsually yes, at least for the ones with extraterritorial reach. Brazil's LGPD, China's PIPL and India's DPDP Act all contemplate foreign companies that offer goods or services to people in those countries. Physical presence is not the test; the location of the person you are interviewing and the market you are selling into usually are.\n\n**Can I reuse my GDPR consent form for research in Brazil, India and Japan?**\nNot without modification, in both directions. A GDPR form may collect consent where Brazil would let you rely on legitimate interests, and it may be too bundled for China and South Korea, which expect separate consent per purpose. It will also miss India's requirement for a standalone notice available in scheduled Indian languages. Start from the GDPR form, then itemise and localise.\n\n**Which of these laws is strictest for customer interviews?**\nChina's PIPL, because of the separate-consent requirements for sensitive information and for sending data outside China, followed by South Korea's PIPA for consent granularity and Quebec's Law 25 for opt-in strictness. If you design to satisfy those three, the rest are comfortably covered.\n\n**Did Canada's privacy law change with Bill C-27?**\nNo. Bill C-27 died when Parliament was prorogued in January 2025, so the Consumer Privacy Protection Act and the proposed AI statute never became law. PIPEDA remains in force federally, alongside Quebec's Law 25 and the other substantially-similar provincial regimes.\n\n**Do these laws treat B2B research participants differently?**\nLess often than US state laws do. The US state model of excluding people acting in a commercial or employment context is unusual; most of the regimes here protect individuals regardless of whether they are speaking as a professional. Assume your B2B participants are covered.\n\n**What about the transfer of recordings out of the country?**\nThat is a separate legal question from which law applies, and it has its own mechanisms — adequacy findings, standard contractual clauses, security assessments and, in China's case, separate consent. Do the applicability analysis first, then the transfer analysis.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that minimise the personal data a study collects\n- [Research Data Residency and International Transfers](/docs/research-data-residency-international-transfers) — the transfer-mechanism half of the problem\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the EU baseline these regimes are usually compared against\n- [US State Privacy Laws for Customer Research](/docs/us-state-privacy-laws-research) — the twenty-state American patchwork\n- [Multi-Language User Research](/docs/multilingual-research-guide) — running and localising interviews in the participant's language\n- [Interview Recording Consent Laws](/docs/interview-recording-consent-laws) — the recording-specific consent rules\n- [Research Data Access Controls and Audit Trails](/docs/research-data-access-controls-audit-trail) — proving who could see the data afterwards","category":"Research Operations","lastModified":"2026-08-01T03:18:05.751892+00:00","metaTitle":"International Privacy Laws for User Research: LGPD, PIPEDA, DPDP, APPI, PIPL","metaDescription":"What applies when you interview customers in Brazil, Canada, India, Japan, China, Korea, Australia or South Africa — which law reaches you, which legal basis works, and one study design that satisfies all of them.","keywords":["international privacy laws user research","LGPD user research","PIPEDA research consent","DPDP Act research","APPI Japan research","PIPL research consent","POPIA research","Quebec Law 25 research","global privacy compliance research"],"aiSummary":"Beyond GDPR and CCPA, at least six national privacy regimes can reach a company that interviews customers abroad, and the decisive differences are extraterritorial reach and permitted legal basis rather than data transfer. Brazil LGPD offers ten legal bases including legitimate interests, but the ANPD has confirmed legitimate interests cannot cover sensitive data. Canada PIPEDA remains in force because Bill C-27 died on prorogation in January 2025; Quebec Law 25 adds opt-in consent and automated-decision notice. India DPDP Rules were notified 14 November 2025 with obligations phasing in to 13 May 2027, requiring standalone notice available in scheduled Indian languages. Japan APPI is a purpose-specification regime, and a July 2026 amendment extending cover to contactable identifiers is promulgated but not in force. China PIPL requires separate consent for sensitive data and for transfers abroad, and is the strictest for research, followed by South Korea PIPA and Quebec. A single compliant design — separate itemised opt-in consent, explicit AI disclosure, localised notices, planned handling of sensitive disclosures, easy withdrawal, and data minimisation via structured questions — satisfies all of them.","aiPrerequisites":["Research participants located outside the EU and United States","A current consent and notice template, typically written for GDPR","A list of the countries you actively sell into"],"aiLearningOutcomes":["Determine whether a foreign privacy law reaches your research programme","Pick the right legal basis per jurisdiction instead of defaulting to consent everywhere","Meet the standalone notice and localisation requirements of India's DPDP Rules","Apply China PIPL separate-consent rules to interview recordings","Design one study that satisfies the strictest regime rather than eight variants","Reduce exposure by minimising identifiable data at the point of collection"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"blog","id":"b41934bc-99b8-43ef-ac19-7ed2a6434db2","slug":"hotjar-pricing-2026","title":"Hotjar Pricing in 2026: The Pricing Page Now Redirects to Contentsquare","url":"https://www.koji.so/blog/hotjar-pricing-2026","summary":"Hotjar no longer maintains its own pricing page; hotjar.com/pricing issues a 308 permanent redirect to contentsquare.com/pricing. Following the Contentsquare acquisition, one product with one bill became three separately priced lines: Experience Analytics (formerly Observe), Voice of Customer (formerly Ask), and Product Analytics. Legacy tiers ran $0/$32/$80/$171 monthly for Observe and $0/$48/$64/$128 for Ask with about 20% annual discount. Contentsquare Growth tiers start near $49/month for Experience Analytics and $99/month for Voice of Customer, so teams needing both pay roughly $148/month minimum. Negotiated Contentsquare contracts show a $20,000 median annual value ranging $14,600-$95,168, with implementation adding $15,000-$75,000 and 3-5% annual escalators. Session replay reveals what users do but not why; Koji offers AI-moderated interviews from EUR 29/month to answer that.","content":"\n**Short answer:** Hotjar no longer has its own pricing page. As of our check, **hotjar.com/pricing issues a 308 permanent redirect to contentsquare.com/pricing**. Following the Contentsquare acquisition, what was one product with one bill has been split into three separately priced lines: Experience Analytics (the old Observe), Voice of Customer (the old Ask), and Product Analytics. Legacy Hotjar tiers still in circulation run **$0 / $32 / $80 / $171 per month** for Observe and **$0 / $48 / $64 / $128** for Ask, with about 20% off for annual billing. On the Contentsquare platform, Experience Analytics starts around **$49/month** and Voice of Customer around **$99/month** — so teams wanting both heatmaps and surveys now pay two bills starting at roughly **$148/month**. Negotiated Contentsquare contracts show a **$20,000 median annual value**.\n\nThat redirect is the single most important fact on this page. If you are budgeting for \"Hotjar\" in 2026, you are budgeting for a product that is being absorbed into an enterprise analytics suite, and the pricing model you signed up for is not the one you will renew on.\n\n## The legacy Hotjar tiers\n\nThese are the plans most existing customers are still on, and the ones third-party guides still quote.\n\n**Observe** — heatmaps and session recordings:\n\n| Plan | Price/month | Daily sessions |\n|---|---|---|\n| Basic | Free | 35 |\n| Plus | $32 | 100 |\n| Business | $80 | 500 (scaling with volume) |\n| Scale | $171 | 500+ |\n\n**Ask** — surveys and feedback widgets:\n\n| Plan | Price/month | Monthly responses |\n|---|---|---|\n| Basic | Free | 20 |\n| Plus | $48 | 250 |\n| Business | $64 | 500 |\n| Scale | $128 | Higher volume |\n\nAnnual billing saves roughly **20%** across plans.\n\nTwo things about this table deserve flagging. First, the Business tier is not really a price — it is the *entry point* of a volume curve. Business scales with session count, and at very high traffic the same tier has been reported as high as **$9,448 per month**. The $80 figure is what you pay at 500 daily sessions and nothing like what you pay at 270,000.\n\nSecond, a currency gotcha for European buyers: these figures are quoted as **€32 / €80 / €171** on EU-facing sources and **$32 / $80 / $171** on US-facing ones. That is list-price parity, not conversion — meaning euro-zone buyers pay meaningfully more in real terms for identical plans. Check which currency your quote is in.\n\n## What replaced it\n\nOn the Contentsquare platform, the single Hotjar subscription becomes three product lines, each billed independently:\n\n| Product line | Was | Entry price |\n|---|---|---|\n| Experience Analytics | Hotjar Observe | From ~$49/month (Growth) |\n| Voice of Customer | Hotjar Ask | From ~$99/month (Growth) |\n| Product Analytics | (Heap) | Custom |\n\nGrowth tiers scale by volume — Experience Analytics has been reported ranging from **$49 up to $739/month**, and Voice of Customer from **$99 up to $1,479/month** — above which you move to Pro and Enterprise, both custom-quoted.\n\n**A caveat worth stating plainly:** published figures for the new free allowances vary considerably between sources, and migration has been rolling through 2026 on a per-account basis. Some accounts are still on legacy Hotjar billing, some are on Contentsquare tiers, and the numbers you find in third-party guides may describe either. Verify against your own account and your own quote before you budget from any of them — including this page.\n\n## The real change is structural, not the price\n\nFocusing on the headline numbers misses the point. Three things changed about *how* you are billed:\n\n**1. You now pay two bills for what used to be one product.** A team that wants heatmaps *and* on-site surveys — the standard Hotjar use case, and the reason most people bought it — pays for Experience Analytics plus Voice of Customer separately. At Growth tiers that is roughly **$49 + $99 = $148/month** minimum, where a single legacy Business subscription covered comparable ground.\n\n**2. Costs scale with traffic you do not control.** Session-based pricing means a successful marketing campaign raises your analytics bill. Contentsquare overage fees for exceeding contracted session limits have been reported at **20–50% above base rates**, so the penalty for a good quarter is a worse invoice.\n\n**3. The endpoint is an enterprise contract.** Negotiated Contentsquare contract data shows a **$20,000 median annual value**, with a range of **$14,600 to $95,168**. Implementation services add **$15,000–$75,000** depending on complexity, and multi-year deals commonly carry **3–5% annual escalators**. Multi-year commitments are reported to yield 15–30% lower effective pricing, which is the standard trade: less flexibility for a better rate.\n\nFor a team that bought a $32/month heatmap tool three years ago, that is a different category of purchase. If it is not where you want to end up, our roundup of [Hotjar alternatives](/blog/hotjar-alternatives-2026) and [best session replay tools](/blog/best-session-replay-tools-2026) is the place to start, and [Hotjar vs Microsoft Clarity](/blog/hotjar-vs-microsoft-clarity-2026) covers the genuinely free option.\n\n## What you are paying for — and what it cannot tell you\n\nHere is the more fundamental question, and it applies whichever tier you land on.\n\nHeatmaps and session recordings are behavioural instruments. They tell you **what** happened with real precision: users scrolled to 60% of the page, rage-clicked a non-interactive element, abandoned checkout at the shipping step. That is genuinely valuable, and it is why we compare it seriously in [FullStory vs Hotjar](/blog/fullstory-vs-hotjar-2026) and [Sprig vs Hotjar](/blog/sprig-vs-hotjar-2026).\n\nWhat no amount of session replay tells you is **why**. You can watch a hundred recordings of people abandoning at the shipping step and still not know whether the cost surprised them, they were comparison-shopping, they wanted a delivery date they could not see, or they simply meant to come back on a laptop. Those four causes lead to four different fixes, and behavioural data cannot distinguish between them.\n\nThe usual patch is a Voice of Customer survey widget — which is precisely why that product is now billed separately at $99/month. But an on-site survey buys you one question and one shallow answer. Response rates on intercept surveys are low and falling, for the reasons we cover in [why survey response rates are declining](/docs/survey-response-rates-declining), and \"Price too high\" in a text box is not an insight — it is a prompt for the follow-up question nobody asked.\n\n## The AI-native alternative\n\nKoji answers the \"why\" that heatmaps structurally cannot, and it does it by having an actual conversation rather than collecting a text box.\n\nWhen a user abandons checkout, Koji can run an **AI-moderated voice or text interview** that asks what happened and then *probes the answer* — following up on \"too expensive\" with \"compared to what?\" and \"what would have made it worth it?\" That is the difference between a survey response and an interview, and it is the reason the follow-up question matters more than the first one.\n\nPricing is published and credit-based:\n\n- **Insights — €29/month**, 29 credits included\n- **Interviews — €79/month**, 79 credits included\n- **Enterprise** — custom\n\nA text conversation costs 1 credit, a full AI-moderated voice interview costs 3, and a report refresh costs 5 — so the €79 plan covers roughly **26 voice interviews per month**. Critically, **only conversations scoring 3 or above consume credits**, so you are not billed for abandoned or junk sessions. Compare that to session-based analytics pricing, where every bot visit and bounced pageview counts toward your tier.\n\nWhat you get:\n\n- **AI-moderated voice interviews** running 24/7 in parallel — no scheduling, no moderator, no recruiting agency\n- **Automatic thematic analysis** across every conversation, so you get themes and not a folder of recordings to watch\n- **Six structured question types** — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — mixed into the same conversation. You get an NPS distribution *and* the reasoning behind each score, rather than a number with no explanation\n- **Customisable AI consultants** tuned to your product and research goals\n- **One-click reports** with quotes traced back to source\n- **No moderator bias** — every participant meets the same neutral interviewer\n\nThe honest framing: this is not a replacement for heatmaps. If you need to see where people click, you need a session replay tool, and Contentsquare's is a good one. Koji replaces the far more expensive thing — the weeks of guessing that follow watching recordings, and the enterprise contract you are being migrated toward to make that guessing slightly better informed. Many teams run a free or cheap replay tool for the *what* and Koji for the *why*, which costs dramatically less than a $20,000 median Contentsquare contract. See [Koji vs Hotjar](/blog/koji-vs-hotjar-2026) for the direct comparison.\n\n## Frequently Asked Questions\n\n**Does Hotjar still have its own pricing page?**\nNo. hotjar.com/pricing now issues a 308 permanent redirect to contentsquare.com/pricing following the Contentsquare acquisition. Hotjar pricing is now presented as part of the Contentsquare platform.\n\n**How much does Hotjar cost in 2026?**\nLegacy Hotjar tiers run $0 / $32 / $80 / $171 per month for Observe (heatmaps and recordings) and $0 / $48 / $64 / $128 for Ask (surveys), with about 20% off annual billing. On the Contentsquare platform, Experience Analytics starts around $49/month and Voice of Customer around $99/month.\n\n**Is Hotjar still free?**\nA free tier still exists, though published allowances vary by source and by whether your account has migrated. The legacy Hotjar Basic plan covered 35 daily sessions for Observe and 20 monthly responses for Ask. Verify the current allowance against your own account.\n\n**Why does Hotjar cost more now?**\nBecause one product became three separately billed lines. A team wanting both heatmaps and surveys pays for Experience Analytics and Voice of Customer separately — roughly $148/month combined at Growth tiers — where a single legacy subscription previously covered both.\n\n**What do companies actually pay for Contentsquare?**\nNegotiated contract data shows a $20,000 median annual value, ranging from $14,600 to $95,168. Implementation services add $15,000–$75,000, overage fees run 20–50% above base rates, and multi-year deals typically carry 3–5% annual escalators.\n\n**What is the best Hotjar alternative after the Contentsquare migration?**\nIt depends on what you need. For free heatmaps and recordings, Microsoft Clarity is the obvious substitute. For understanding *why* users behave the way they do — which session replay cannot answer — an AI-native research platform like Koji runs moderated interviews from €29/month with analysis included.\n\n## Heatmaps show you the wall. Ask people why they hit it.\n\nIf your Hotjar renewal now arrives with a Contentsquare logo, a session-based meter, and an escalator clause, it is worth asking what the spend is actually buying. Recordings tell you where users struggle. They will never tell you why — and the why is what changes the roadmap.\n\nKoji gets you **from question to insight in hours, not weeks**, with **no research expertise required** and published pricing from €29/month.\n\n[Ask your users why →](https://www.koji.so)\n","category":"Comparisons","lastModified":"2026-08-01T03:17:12.70121+00:00","metaTitle":"Hotjar Pricing 2026: Plans, Limits & the Contentsquare Migration","metaDescription":"Hotjar's pricing page now redirects to Contentsquare. Legacy tiers, the new three-product structure, $20,000 median contracts, and what the migration means for your renewal.","keywords":["hotjar pricing","hotjar cost","hotjar plans","hotjar pricing 2026","hotjar contentsquare","hotjar free plan","contentsquare pricing"],"aiSummary":"Hotjar no longer maintains its own pricing page; hotjar.com/pricing issues a 308 permanent redirect to contentsquare.com/pricing. Following the Contentsquare acquisition, one product with one bill became three separately priced lines: Experience Analytics (formerly Observe), Voice of Customer (formerly Ask), and Product Analytics. Legacy tiers ran $0/$32/$80/$171 monthly for Observe and $0/$48/$64/$128 for Ask with about 20% annual discount. Contentsquare Growth tiers start near $49/month for Experience Analytics and $99/month for Voice of Customer, so teams needing both pay roughly $148/month minimum. Negotiated Contentsquare contracts show a $20,000 median annual value ranging $14,600-$95,168, with implementation adding $15,000-$75,000 and 3-5% annual escalators. Session replay reveals what users do but not why; Koji offers AI-moderated interviews from EUR 29/month to answer that.","aiKeywords":["hotjar pricing","contentsquare","session replay cost","heatmap tools","voice of customer","AI moderated interviews"],"aiContentType":"comparison","faqItems":[{"answer":"No. hotjar.com/pricing now issues a 308 permanent redirect to contentsquare.com/pricing following the Contentsquare acquisition. Hotjar pricing is now presented as part of the Contentsquare platform.","question":"Does Hotjar still have its own pricing page?"},{"answer":"Legacy Hotjar tiers run $0/$32/$80/$171 per month for Observe (heatmaps and recordings) and $0/$48/$64/$128 for Ask (surveys), with about 20% off annual billing. On the Contentsquare platform, Experience Analytics starts around $49/month and Voice of Customer around $99/month.","question":"How much does Hotjar cost in 2026?"},{"answer":"A free tier still exists, though published allowances vary by source and by whether your account has migrated. The legacy Hotjar Basic plan covered 35 daily sessions for Observe and 20 monthly responses for Ask. Verify the current allowance against your own account.","question":"Is Hotjar still free?"},{"answer":"Because one product became three separately billed lines. A team wanting both heatmaps and surveys pays for Experience Analytics and Voice of Customer separately, roughly $148/month combined at Growth tiers, where a single legacy subscription previously covered both.","question":"Why does Hotjar cost more now?"},{"answer":"Negotiated contract data shows a $20,000 median annual value, ranging from $14,600 to $95,168. Implementation services add $15,000-$75,000, overage fees run 20-50% above base rates, and multi-year deals typically carry 3-5% annual escalators.","question":"What do companies actually pay for Contentsquare?"},{"answer":"For free heatmaps and recordings, Microsoft Clarity is the obvious substitute. For understanding why users behave the way they do, which session replay cannot answer, an AI-native research platform like Koji runs moderated interviews from EUR 29/month with analysis included.","question":"What is the best Hotjar alternative after the Contentsquare migration?"}],"relatedTopics":["pricing","hotjar","contentsquare","session replay","heatmaps","vendor comparison"]},{"type":"documentation","id":"eefb768d-ee94-456c-a0a2-44f52cd639b0","slug":"research-data-access-controls-audit-trail","title":"Research Data Access Controls and Audit Trails: Who Can See Your Interview Data","url":"https://www.koji.so/docs/research-data-access-controls-audit-trail","summary":"Vendor certifications like SOC 2 prove the platform is secure; they say nothing about which internal colleague opened a raw transcript. Research data should be tiered by identifiability — raw recording, verbatim transcript, de-identified transcript, aggregate report — with access narrowing by roughly an order of magnitude at each tier and granted by role rather than by request. A defensible audit trail records eight fields per event (timestamp, actor identity, role at the time, object, sensitivity tier, action, source context, justification) and must cover research-specific events: recording playback, export and download, share-link creation, cross-study search, redaction changes, permission grants and API or MCP token use. ISO 27001:2022 A.5.15, A.5.18 and A.8.15, SOC 2 CC6, GDPR Articles 32 and 5(2), and HIPAA 45 CFR 164.312(b) all require review of logs, not merely collection. Koji reduces the number of access grants needed by answering most stakeholder questions at the aggregate tier through structured questions, with team roles of owner, admin and member and workspace-scoped project access.","content":"**Answer first: a vendor security certificate and an internal access-control model are two different things, and only one of them is your job.** SOC 2, ISO 27001 and a signed DPA tell you the platform holding your interview data is run competently. None of them tell you which of your own colleagues opened a raw transcript containing a customer's health disclosure last Tuesday, whether they had any business reason to, or whether you could prove it six months later when a regulator asks. That second question — internal access control and the audit trail behind it — is the control that customer research teams most consistently skip, and the one that turns a routine privacy request into an incident.\n\nThis guide covers the four sensitivity tiers of interview data, the role model that should sit on top of them, exactly what an audit trail has to record to be worth anything, the research-specific events almost nobody logs, and the quarterly access review that keeps the whole thing honest.\n\n## The core principle: access should narrow as identifiability rises\n\nMost research teams treat \"the study\" as the unit of access. Someone is either in the project or not. That is far too coarse, because a single interview produces four artefacts with wildly different risk profiles.\n\n| Tier | Artefact | Contains | Who genuinely needs it |\n|---|---|---|---|\n| 1 | Raw audio or video recording | Voice, accent, background, sometimes face — biometric-adjacent, effectively impossible to de-identify | The researcher running the study, and almost nobody else |\n| 2 | Verbatim transcript | Names, employers, health and financial disclosures, off-hand third-party details | The researcher plus named analysts |\n| 3 | De-identified transcript | The reasoning, with direct identifiers stripped or tokenised | The wider product team |\n| 4 | Aggregate report and quote set | Themes, distributions, approved quotes | Anyone with a legitimate interest |\n\nThe rule that follows is simple and almost never implemented: **the number of people with access should fall by roughly an order of magnitude at each tier.** If forty people can open a raw recording, you do not have an access-control model — you have a shared folder.\n\nThis tiering is also what makes the difference between a defensible and an indefensible answer to the data-minimisation question. Under GDPR Article 5(1)(c) and its equivalents in most modern privacy laws, you are expected to limit processing to what is necessary. \"The whole product org had standing access to raw recordings because it was easier\" is not a necessity argument.\n\n## The role model\n\nAccess should be granted by role against tier, not by individual request. A workable default for a product or research organisation:\n\n| Role | Tier 1 raw | Tier 2 verbatim | Tier 3 de-identified | Tier 4 report |\n|---|---|---|---|---|\n| Study owner / researcher | Yes | Yes | Yes | Yes |\n| Research ops / admin | On request, logged | Yes | Yes | Yes |\n| Named analyst on the study | No | Yes, study-scoped | Yes | Yes |\n| Product manager, designer, engineer | No | No | Yes | Yes |\n| Executive, sales, marketing | No | No | No | Yes |\n| External agency or contractor | No | Time-boxed, study-scoped | Yes | Yes |\n| Legal, privacy, compliance | On request, logged | On request, logged | Yes | Yes |\n| Vendor support staff | Only with break-glass approval, logged and notified | | | |\n\nTwo rows deserve comment. **Contractors and agencies** are where most real leakage happens: access is granted for a project, the project ends, and the account stays live for two years. Every external grant should carry an expiry date at the moment it is created, not a calendar reminder to review it later.\n\n**Vendor support access** is the row buyers forget to ask about during procurement. The right question is not \"can your staff see my data?\" — the honest answer for almost every platform is \"yes, under some circumstances.\" The right questions are: does support access require customer approval, is it time-boxed, is it logged in a trail I can read, and am I notified when it happens?\n\n## What the standards actually require\n\nIf you are being audited, the requirements are more specific than most teams assume.\n\n**ISO/IEC 27001:2022** splits this across several Annex A controls. A.5.15 (access control) requires rules for physical and logical access based on business and security requirements. A.5.18 requires that access rights be provisioned, reviewed, modified and removed against a documented policy. A.8.15 (logging) requires that logs record user activities, exceptions, faults and security events, that they be **protected from tampering and unauthorised access — explicitly including by privileged administrators who might otherwise edit their own trail** — and that they be analysed, not merely collected. The practical implication of that last clause is append-only or write-once storage for the log repository, plus an access control list for the logs themselves.\n\n**SOC 2** covers the same ground under the CC6 common criteria: logical access provisioning and removal, restriction of privileged access, and evidence that access is reviewed periodically. An auditor testing CC6 will ask for a user list, the role each user holds, the date access was granted, and evidence of the most recent recertification.\n\n**GDPR** Article 32 requires appropriate technical and organisational measures including the ability to ensure ongoing confidentiality, and Article 5(2) — accountability — means you must be able to *demonstrate* compliance rather than assert it. An audit trail is the demonstration.\n\n**HIPAA**, if any of your interviews touch health information, is unusually blunt: 45 CFR §164.312(b) requires audit controls that record and examine activity in systems containing electronic protected health information, and §164.308(a)(1)(ii)(D) requires regular review of information system activity such as access logs.\n\nNotice what is common to all four: not one of them is satisfied by *having* logs. They all require that someone reviews them.\n\n## What an audit trail must record\n\nA log line that says \"transcript viewed\" is nearly useless. A defensible research audit trail records eight fields per event.\n\n| Field | Why it matters |\n|---|---|\n| Timestamp (UTC, with timezone) | Correlates with incident timelines and shift patterns |\n| Actor identity | A named human or a named service account — never a shared login |\n| Actor's role at the time | Roles change; the trail must reflect what was true then |\n| Object | Study, interview, transcript or export — identified precisely |\n| Sensitivity tier of the object | Lets you filter \"who touched Tier 1 this quarter\" in one query |\n| Action | View, play, search, export, share, redact, delete, grant |\n| Source context | IP address, device or API client |\n| Justification | Ticket reference or stated reason, for privileged and break-glass access |\n\nShared logins destroy all of this. If three people use one \"research@\" account, your audit trail has exactly one actor and zero evidentiary value.\n\n## The research-specific events nobody logs\n\nGeneric application logging captures logins. Research data governance needs more, and these are the events that matter when something goes wrong:\n\n- **Recording playback**, separately from transcript viewing — voice is the most sensitive artefact you hold.\n- **Export and download.** This is the single most important event in the entire trail, because it is the boundary at which your controls stop. A transcript exported to a spreadsheet, a wiki page or a slide deck now lives in a system with different permissions, different retention and no link back to consent. Log it, alert on unusual volume, and prefer in-platform sharing over export wherever the workflow allows.\n- **Share-link creation**, including whether the link was public, expiring or restricted.\n- **Cross-study search.** Searching a repository for a person's name is a different act from opening one study, and it should be visible as one.\n- **Redaction and de-identification changes** — including who un-redacted something.\n- **Permission grants and revocations**, which is how you reconstruct who could have seen what on any given date.\n- **API and integration access**, including tokens used by automation and by AI assistants connected over MCP. A token is an actor; it needs a name, an owner and a scope.\n\n## Where Koji fits\n\nKoji is built around the assumption that most stakeholders should never open a raw transcript at all, and the product structure reflects that.\n\nTeam accounts use three roles — **owner, admin and member** — with shared **workspaces** that scope which projects a member can see, so access is granted by workspace rather than by handing out a universal login. Reports are shared as their own artefact, which means a stakeholder can read findings without being granted access to the underlying recordings.\n\nThe bigger structural advantage is **structured questions**. Koji supports six question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — and the five non-open types produce genuinely aggregate output: distributions, frequencies and average ranks that carry no identifiable content at all. A well-designed study answers most stakeholder questions at Tier 4. When a PM can see that 68% of enterprise respondents ranked reporting first without opening a single verbatim, the access request never happens — and the safest access control is the one you never had to grant. Traditional survey tools push you the other way, because a free-text-only instrument forces everyone into the verbatims to learn anything.\n\nBecause Koji's AI interviewer runs every conversation to the same structured plan, the resulting analysis is consistent enough to be trusted at the aggregate tier — which is precisely what makes the tiering practical rather than theoretical.\n\n## The quarterly access review\n\nProvisioning is easy; deprovisioning is what fails. Run a recertification every quarter and treat it as a documented control:\n\n1. **Export the current access list** by user, role, workspace and tier.\n2. **Reconcile against HR joiners, movers and leavers.** Movers are the dangerous category — people accumulate access as they change teams and lose none of it.\n3. **Expire all external access** that has passed its project end date. No exceptions, no \"they might come back.\"\n4. **Ask each Tier 1 and Tier 2 holder to justify the access in one sentence.** Anyone who cannot loses it. This step alone typically removes a third of standing access.\n5. **Review the log, do not just keep it.** Sample the quarter's Tier 1 events and the top ten exporters by volume. Look for access outside working hours, access to studies the person was never assigned to, and bulk exports before a resignation.\n6. **Record the review** — date, reviewer, changes made. The record is the evidence; the review without the record does not exist as far as an auditor is concerned.\n\n## The evidence pack an auditor asks for\n\nWhen the questionnaire arrives, you will be asked for some version of these six items. Assemble them once and keep them current:\n\n- The written access-control policy, including the tier model and the role matrix.\n- A current user access list with roles and grant dates.\n- Evidence of the last two access reviews, with changes made.\n- A sample of audit log entries showing the eight fields above.\n- Evidence that logs cannot be altered by the people they record, and the log retention period.\n- The list of third parties and vendor personnel who can reach research data, with the controls on each.\n\n## Common failure patterns\n\n**The permanent Slack channel.** Transcripts pasted into a channel with 200 members. The channel outlives the study, the members change, and no consent notice ever mentioned it.\n\n**The one-time export that becomes the system of record.** Someone exports everything to a spreadsheet \"for analysis,\" and eighteen months later that spreadsheet is the copy people actually use — outside every control you built.\n\n**Deletion that misses the copies.** A participant exercises a deletion right; you delete the interview in the platform and miss the export, the slide deck and the meeting recording where it was discussed. Your audit trail of exports is what makes this recoverable. See [research data retention and deletion](/docs/research-data-retention-deletion) for the retention side of the same problem.\n\n**Logging without review.** The most common of all, and the one that fails an audit fastest, because every framework above requires the review and not just the log.\n\n## Frequently asked questions\n\n**Is a vendor's SOC 2 report enough to cover internal access to research data?**\nNo. A SOC 2 report covers controls at the vendor. It is evidence about the platform, not about your own team's provisioning, your role model, or who inside your company opened a recording. Auditors will ask for both, and only one of them is something you can produce.\n\n**How long should we keep research access logs?**\nLonger than the data they describe, which is the point people get backwards. A common baseline is twelve months of readily searchable logs with a longer archive, but the practical test is whether you can answer \"who accessed this participant's data\" for the full period during which that participant can still exercise their rights or bring a complaint.\n\n**Do we really need to log read access, or just changes?**\nRead access is the one that matters most for research data. Nothing is modified when someone listens to a recording of a customer describing a medical condition, but that is exactly the event a participant, a regulator or an incident review would care about. Log views and playbacks, not only edits.\n\n**Should stakeholders get access to raw transcripts at all?**\nUsually not by default. Give them the report tier and grant transcript access per study, with an expiry, when there is a stated reason. Studies designed with structured questions answer most stakeholder needs at the aggregate level, which removes the pressure to hand out broad access in the first place.\n\n**What counts as a \"break-glass\" access and how should it work?**\nBreak-glass is emergency access outside the normal role model — a support engineer debugging a broken interview, or an investigator responding to an incident. It should require approval from a named person, be time-boxed to hours rather than days, generate a notification to the data owner, and produce a log entry with a written justification.\n\n**How do access controls interact with anonymous research?**\nThey become more important, not less. If you promised employees or customers anonymity, the access model is the thing that makes the promise true. Restrict raw data to the smallest possible group, suppress small-sample breakdowns in reporting, and be able to show the access list to the people you made the promise to.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that let stakeholders read findings without opening a verbatim\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — how long to keep each tier, and how to delete it properly\n- [Enterprise Security for AI Customer Research Platforms](/docs/enterprise-security-ai-research-platforms) — SOC 2, SSO and the vendor-side review\n- [AI Interview Data Privacy and Security](/docs/ai-interview-data-privacy-security) — the buyer's evaluation checklist\n- [Research Data Residency and International Transfers](/docs/research-data-residency-international-transfers) — where the data physically lives\n- [AI Governance for Customer Research](/docs/ai-governance-frameworks-research) — ISO 42001, the NIST AI RMF and what procurement asks\n- [Anonymous Employee Research with AI Interviews](/docs/anonymous-employee-research-ai-interviews) — keeping an anonymity promise you can actually defend","category":"Research Operations","lastModified":"2026-08-01T03:16:20.992314+00:00","metaTitle":"Research Data Access Controls & Audit Trails: A Practical Guide","metaDescription":"Who on your team can open a raw interview recording — and can you prove it? Access tiers, the role matrix, what ISO 27001 and SOC 2 require, and the quarterly access review.","keywords":["research data access control","research audit trail","who can access research data","least privilege research data","research data governance","interview transcript permissions","access review research","ISO 27001 logging research"],"aiSummary":"Vendor certifications like SOC 2 prove the platform is secure; they say nothing about which internal colleague opened a raw transcript. Research data should be tiered by identifiability — raw recording, verbatim transcript, de-identified transcript, aggregate report — with access narrowing by roughly an order of magnitude at each tier and granted by role rather than by request. A defensible audit trail records eight fields per event (timestamp, actor identity, role at the time, object, sensitivity tier, action, source context, justification) and must cover research-specific events: recording playback, export and download, share-link creation, cross-study search, redaction changes, permission grants and API or MCP token use. ISO 27001:2022 A.5.15, A.5.18 and A.8.15, SOC 2 CC6, GDPR Articles 32 and 5(2), and HIPAA 45 CFR 164.312(b) all require review of logs, not merely collection. Koji reduces the number of access grants needed by answering most stakeholder questions at the aggregate tier through structured questions, with team roles of owner, admin and member and workspace-scoped project access.","aiPrerequisites":["A research platform or repository holding interview recordings and transcripts","A list of everyone who currently has access to research data","Awareness of which compliance frameworks apply to your organisation"],"aiLearningOutcomes":["Tier research artefacts by identifiability and set access limits per tier","Build a role-versus-tier access matrix instead of granting access ad hoc","Record the eight audit-trail fields that make a log evidentially useful","Log the research-specific events that generic application logging misses","Run a quarterly access recertification that survives an audit","Assemble the evidence pack auditors request for access controls"],"aiDifficulty":"advanced","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"04463c9e-8c3c-42b7-aac6-59f6530f8476","slug":"data-annotation-quality-guide","title":"Data Annotation Quality: Guidelines, Agreement Metrics, and Gold Tasks That Actually Work (2026)","url":"https://www.koji.so/docs/data-annotation-quality-guide","summary":"Annotation quality decomposes into four systems: edge-case guidelines, chance-corrected agreement measurement (Krippendorff's alpha >= 0.800 for reliable data), gold tasks injected at 3-10% with ~90% blocking thresholds, and adjudication that routes disagreement to senior review instead of majority-voting it away. Disagreement on subjective tasks is signal, not noise. The under-used lever is researching the annotators themselves to find which guideline boundaries are unclear.","content":"# Data Annotation Quality: Guidelines, Agreement Metrics, and Gold Tasks That Actually Work (2026)\n\n**Short answer:** Annotation quality is not one thing. It is four separable systems — guidelines that resolve edge cases, an agreement metric that tells you whether the task is even learnable, gold tasks that catch drift and fraud, and an adjudication path that turns disagreement into a guideline change. Teams that treat quality as \"hire better annotators\" plateau. Teams that treat it as a research problem — asking annotators *why* they disagreed and feeding that answer back into the guidelines — keep improving. The usual statistical bar is Krippendorff's alpha at or above **0.800** for reliable data, with **0.667–0.800** supporting tentative conclusions only.\n\nEvery AI product now depends on a labeled dataset somebody built by hand: the evaluation set, the safety taxonomy, the preference pairs behind a fine-tune, the routing labels on a support queue. And almost every one of those datasets was built by a group of people who were handed a document, given a week of ramp, and then measured on throughput.\n\nThat is why annotation quality fails in a predictable way. It is rarely that annotators are careless. It is that the guideline never answered the question the annotator actually had, nobody asked them, and the disagreement got averaged into a majority label that looks clean and is quietly wrong.\n\nThe money involved is no longer marginal. The AI data-labeling market is estimated at roughly **$2.32 billion in 2026**, up from $1.89 billion in 2025 and forecast toward **$6.53 billion by 2031** at about a 23% CAGR ([Mordor Intelligence](https://www.mordorintelligence.com/industry-reports/ai-data-labeling-market)). Scale AI reportedly delivers **over 1 billion annotations a year**, and in July 2025 cut around 200 employees — roughly 14% of staff — while ending work with about 500 contractors. Surge AI reportedly passed **$1 billion in 2024 revenue** while bootstrapped. Frontier-tier work now places credentialed specialists at **$85–200+ per hour** for RLHF design and complex reasoning evaluation. Getting quality wrong at that price is a budget line, not a footnote.\n\n## What \"quality\" actually decomposes into\n\nTreat these as four independent systems. Fixing one does not fix the others.\n\n| System | The question it answers | Failure signature |\n|---|---|---|\n| **Guidelines** | What should I do with *this* ambiguous item? | Agreement is low and stays low no matter who you hire |\n| **Agreement measurement** | Is this task learnable by humans at all? | You have no idea whether 80% is good or terrible |\n| **Gold tasks / honeypots** | Is this specific annotator still calibrated today? | Quality decays silently over weeks |\n| **Adjudication** | What do we do when two good annotators disagree? | Majority vote hides the interesting cases |\n\n## Step 1: Write guidelines that resolve edge cases, not describe labels\n\nMost annotation guidelines are glossaries. They define each label in a sentence, give one clean positive example, and stop. That document is useless precisely where it matters, because annotators do not struggle with clean examples — they struggle with the boundary.\n\nA guideline that works has a different shape:\n\n- **A decision procedure, not a taxonomy.** Order the checks. \"First ask whether the utterance contains a request. If yes, go to §3. If no, label `non_actionable` and stop.\" Ordering removes the most common source of variance, which is two annotators applying the same rules in a different sequence.\n- **Adversarial examples with the reasoning attached.** For each label, include two or three items that *look* like they belong and do not, with an explicit sentence about why. The reasoning is the transferable part.\n- **A named tie-break rule.** \"When an item plausibly fits both `harassment` and `spam`, prefer the label with the higher enforcement consequence.\" Without this, annotators invent their own, and each invents a different one.\n- **A living changelog.** Every adjudicated case becomes a numbered guideline entry with a date. Annotators must be able to see what changed and when, because a guideline revision silently invalidates the labels produced before it.\n\nThe practical test: hand your guideline to someone who has never seen the task, give them your ten hardest historical items, and see if they land where your senior annotator landed. If not, the document — not the annotator — is the defect.\n\n## Step 2: Measure agreement with the right metric\n\nRaw percent agreement is the metric everyone reaches for and the one that lies most. On a binary task with a 90/10 class balance, two annotators who both label everything as the majority class agree 90% of the time and have learned nothing. Chance-corrected metrics exist for exactly this reason.\n\n| Metric | Use when | Notes |\n|---|---|---|\n| **Percent agreement** | Never as your only number | No chance correction; inflated by class imbalance |\n| **Cohen's kappa** | Exactly two annotators, nominal labels, every item labeled by both | The most widely reported; does not extend to more raters |\n| **Fleiss' kappa** | Fixed number of raters per item, nominal labels | Assumes the same *count* of raters per item, not the same people |\n| **Krippendorff's alpha** | Any number of raters, missing labels, nominal/ordinal/interval data | The most flexible and the right default for real annotation pipelines |\n| **Per-label F1 vs. gold** | You have a trusted reference set | Tells you *which* label is broken, which agreement metrics cannot |\n\nKrippendorff's alpha is the default recommendation for production annotation work for one concrete reason: real pipelines have **missing labels**. Not every annotator sees every item, batches get reassigned, and people leave mid-project. Cohen's and Fleiss' kappa assume a tidy matrix; alpha does not ([Label Studio](https://labelstud.io/blog/how-to-use-krippendorff-s-alpha-to-measure-annotation-agreement/), [Encord](https://encord.com/blog/interrater-reliability-krippendorffs-alpha/)).\n\nThe conventional thresholds:\n\n- **Krippendorff's alpha ≥ 0.800** — reliable enough to use as ground truth.\n- **0.667 ≤ alpha < 0.800** — draw tentative conclusions only; do not ship this as an evaluation set.\n- **alpha < 0.667** — the task specification is broken. Do not hire more annotators; rewrite the guideline.\n\nFor Cohen's kappa, the reference bands still in general use come from Landis and Koch (1977): 0.01–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, **0.61–0.80 substantial**, 0.81–1.00 almost perfect. Treat these as conventions, not laws — a kappa of 0.65 on a genuinely subjective safety task may be excellent, while 0.75 on a mechanical bounding-box task is a red flag.\n\n**Always report agreement per label, not just overall.** An alpha of 0.82 overall can conceal a single label sitting at 0.31 — and that label is usually the one your product depends on.\n\n## Step 3: Gold tasks and honeypots — what they catch and what they miss\n\nGold tasks are items with trusted labels, secretly injected into normal work so you can score an annotator continuously without them knowing which items are being scored.\n\nCommon production settings:\n\n- **Injection rate: 3–10% of items.** Below 3% you cannot detect drift quickly; above 10% you are paying a meaningful tax on throughput for diminishing signal.\n- **Blocking threshold: around 90% accuracy on gold items** to remain in the pool, evaluated on a rolling window rather than lifetime average.\n- **Rotate the gold set.** Static honeypots leak. Annotators recognise repeated items, and the ones who recognise them fastest are exactly the ones you are trying to catch.\n\nWhat gold tasks genuinely catch: fraud, click-through behaviour, model-assisted cheating, and calibration drift after a guideline change.\n\nWhat they cannot catch — and this is the part most operations miss — is **ambiguity in the task itself**. A gold item only exists because someone already decided the right answer. By construction, your gold set is drawn from the cases that were easy enough to adjudicate confidently. So gold-task accuracy systematically over-reports quality on exactly the population of items where your dataset is weakest.\n\nThe fix is to sample audits from the *disagreement* distribution as well as the gold distribution: pull a stratified weekly sample of items where annotators split, and have a senior reviewer adjudicate those. Gold measures compliance. Disagreement audits measure whether the task is well-formed.\n\n## Step 4: Adjudication — and when disagreement is signal, not noise\n\nMajority vote is the default aggregation everywhere, and it is the single largest source of quiet quality loss.\n\nLora Aroyo and Chris Welty's CrowdTruth work makes the case directly: disagreement, they argue, **\"is not noise but signal\"** — aggregating it away discards information that tells you how hard and how ambiguous an item really is ([CrowdTruth](https://link.springer.com/chapter/10.1007/978-3-319-11915-1_31)). Barbara Plank's work on human label variation extends the point: models trained on majority labels inherit structural biases against minority annotator perspectives, which matters enormously on subjective tasks like toxicity, sentiment, and safety.\n\nA practical adjudication policy:\n\n1. **Route, don't average.** Items with disagreement above a threshold go to a senior reviewer, not to a vote.\n2. **Record the reason, not just the verdict.** The reviewer writes one sentence on *why* the correct label is correct. That sentence becomes a guideline changelog entry.\n3. **Keep the distribution for subjective tasks.** For anything where reasonable people legitimately differ, store the full label distribution alongside the aggregated label and let downstream evaluation use it. A 6/4 split is a fundamentally different data point from a 10/0 split, and collapsing both to one label throws that away.\n4. **Escalate patterns, not items.** If the same boundary generates disagreement three weeks running, the answer is a guideline revision and a re-annotation of the affected slice — not more adjudication.\n\nIt is also worth remembering how good well-managed non-expert annotation can be. Snow et al.'s EMNLP 2008 study *Cheap and Fast — But is it Good?* found high agreement between non-expert crowd annotations and expert gold labels across five natural-language tasks, and showed that averaging a small number of non-expert labels could match expert-quality training data on affect recognition ([ACL Anthology](https://aclanthology.org/D08-1027/)). Expertise is not the bottleneck as often as teams assume. Specification is.\n\n## Workforce operations: the part nobody writes down\n\nThe statistical machinery above assumes a stable pool of calibrated people. That assumption is usually false, and the operational levers matter as much as the metrics:\n\n- **Ramp time is a real cost.** Budget calibration work — annotating a shared set and reviewing disagreements together — for the first one to two weeks. Measure agreement *during* ramp so you can see when someone converges.\n- **Throughput targets corrupt quality when they are the only target.** Pair every throughput number with a rolling gold accuracy and a rolling agreement number, and make all three visible to the annotator.\n- **Attrition is a quality event.** When an experienced annotator leaves, their idiosyncratic interpretations leave with them and your agreement numbers shift. Track agreement as a time series and annotate the chart with staffing changes.\n- **Wellbeing is an operational requirement on harmful content.** Trust-and-safety annotation, red-team transcripts, and moderation queues require rotation limits, opt-outs, and support. This is both an ethical obligation and a data-quality one — fatigued annotators regress toward the majority label.\n- **Pay and classification.** Rates span from commodity bounding boxes to $85–200+/hour for credentialed specialists on reasoning evaluation. Under-scoping expertise on a task that needs it produces a dataset that looks complete and is unusable.\n\n## The modern approach: research your annotators, not just their output\n\nHere is the gap in almost every annotation operation. All four systems above depend on knowing *why* annotators made the calls they made — and nobody ever asks them at scale. Guideline revisions get written by whoever adjudicated, based on a handful of Slack threads.\n\nThis is a research problem, and it is exactly the kind Koji was built for.\n\n**Run a structured study on your annotation pool.** Instead of a spreadsheet of disagreements, run an AI-moderated interview with every annotator on the hard cases. Koji's [structured questions](/docs/structured-questions-guide) map onto this cleanly with all six types:\n\n- **`open_ended`** — \"Walk me through how you decided on the last item you flagged as ambiguous.\" The AI moderator probes follow-ups automatically, which is where the actual decision rule surfaces.\n- **`scale`** — \"How confident were you in that label, 1 to 5?\" Confidence ratings let you find items that are unanimous *and* uncertain, a class gold tasks never surfaces.\n- **`ranking`** — Have annotators rank which parts of the guideline are least clear. The aggregate ranking is your revision backlog, in priority order.\n- **`single_choice` / `multiple_choice`** — Which label boundaries do they hit most often? Frequency charts give you the map of the ambiguous space.\n- **`yes_no`** — \"Did the guideline answer your question?\" A binary you can trend weekly.\n\nBecause Koji runs interviews asynchronously with an AI moderator, you can interview 40 annotators in an afternoon rather than scheduling 40 calls, and the [automatic thematic analysis](/docs/ai-auto-tagging-customer-interviews) clusters their reasoning into the recurring boundary problems without anyone hand-coding transcripts. Every interview gets a quality score on a 1–5 scale so you can see which sessions carried real signal.\n\nThe same mechanism works on the other side of the pipeline: when you need to know what the *right* label is for a genuinely subjective task, the answer lives with your users, not your guideline author. Run the ambiguous items past real users as a Koji study and you get a defensible ground truth with the reasoning attached — which is precisely the material a [golden evaluation set](/docs/ai-evaluation-dataset-golden-set) needs and rarely has.\n\nTo be clear about scope: Koji is not a labeling tool. It will not draw your bounding boxes. What it replaces is the six weeks of guideline archaeology — the part where you try to reconstruct, from disagreement logs, what your annotators were actually thinking.\n\n## Common mistakes\n\n1. **Reporting one overall agreement number.** Per-label agreement is where the broken label hides.\n2. **Hiring more annotators to fix low alpha.** If alpha is below 0.667, the specification is the problem. More people will disagree more consistently.\n3. **A static gold set.** It leaks, and it over-samples easy items by construction.\n4. **Majority vote on subjective tasks.** Keep the distribution.\n5. **Revising guidelines without re-annotating.** A guideline change silently splits your dataset into pre- and post- eras. Version both.\n6. **Measuring throughput alone.** You will get throughput, and nothing else.\n7. **Never talking to the annotators.** The cheapest quality improvement available is asking the people doing the work which rule is unclear.\n\n## Frequently asked questions\n\n**What is a good inter-annotator agreement score?** For Krippendorff's alpha, 0.800 and above is generally treated as reliable, and 0.667 to 0.800 supports tentative conclusions only. For Cohen's kappa, the Landis and Koch (1977) bands put 0.61–0.80 at \"substantial\" and 0.81–1.00 at \"almost perfect.\" Interpret these relative to task subjectivity — 0.65 on a safety judgment call may be strong, while 0.75 on a mechanical labeling task is a warning.\n\n**Should I use Cohen's kappa or Krippendorff's alpha?** Use Krippendorff's alpha as the default for production annotation. Cohen's kappa handles exactly two annotators who both label every item, which almost never describes a real pipeline. Alpha tolerates any number of raters, missing labels, and ordinal or interval data, so it survives reassignment and attrition without breaking your measurement.\n\n**What percentage of tasks should be gold tasks or honeypots?** Typical production settings inject gold items at 3–10% of volume, with a blocking threshold around 90% rolling accuracy. Below 3% you detect drift too slowly; above 10% the throughput cost outweighs the added signal. Rotate the gold set regularly, because static honeypots get recognised.\n\n**Is annotator disagreement always a quality problem?** No. Aroyo and Welty's CrowdTruth research argues disagreement is signal rather than noise, and Plank's work on human label variation shows that majority-vote aggregation encodes bias against minority annotator perspectives. On subjective tasks — toxicity, sentiment, safety — preserve the label distribution alongside the aggregate. Disagreement becomes a quality problem when it is caused by an unclear guideline, which you distinguish by asking annotators why they split.\n\n**How many annotators should label each item?** Three is the common floor for anything with judgment in it, because it lets you detect disagreement at all. Mechanical tasks with alpha above 0.9 can drop to single annotation with a gold-task audit layer. Subjective tasks benefit from five or more, since the shape of the distribution is itself the data you want.\n\n**How do I know whether to fix the annotator or the guideline?** Look at whether disagreement is concentrated or spread. If one annotator disagrees with everyone across many labels, that is a calibration or performance issue. If everyone disagrees on the same boundary, the guideline never resolved that boundary and no amount of retraining will fix it. Running a short structured study across the pool separates the two in an afternoon.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and when to use each\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — coding agreement for qualitative research data\n- [Evaluation Datasets for AI Products: Building a Golden Set](/docs/ai-evaluation-dataset-golden-set) — turning research into ground truth\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) — rubric design for judging model outputs\n- [LLM-as-a-Judge vs. Human Evaluation](/docs/llm-as-a-judge-vs-human-evaluation) — when automated scoring is safe to trust\n- [How to Build a Qualitative Research Codebook](/docs/qualitative-research-codebook) — the qualitative-research cousin of an annotation guideline\n- [AI Auto-Tagging for Customer Interviews](/docs/ai-auto-tagging-customer-interviews) — automatic thematic coding in Koji\n","category":"Research Methods","lastModified":"2026-08-01T03:16:01.475803+00:00","metaTitle":"Data Annotation Quality: Guidelines, Agreement Metrics & Gold Tasks (2026)","metaDescription":"How to run an annotation operation that produces reliable labels: edge-case guidelines, Krippendorff's alpha vs Cohen's kappa, gold-task injection rates, and adjudication that treats disagreement as signal.","keywords":["data annotation quality","data labeling quality assurance","inter-annotator agreement","krippendorff alpha","annotation guidelines","gold tasks","honeypot tasks","annotation workforce","data labeling QA","cohen kappa annotation"],"aiSummary":"Annotation quality decomposes into four systems: edge-case guidelines, chance-corrected agreement measurement (Krippendorff's alpha >= 0.800 for reliable data), gold tasks injected at 3-10% with ~90% blocking thresholds, and adjudication that routes disagreement to senior review instead of majority-voting it away. Disagreement on subjective tasks is signal, not noise. The under-used lever is researching the annotators themselves to find which guideline boundaries are unclear.","aiPrerequisites":["Basic familiarity with supervised machine learning datasets","Understanding of what a labeling task involves"],"aiLearningOutcomes":["Write annotation guidelines that resolve edge cases rather than define labels","Choose the correct inter-annotator agreement metric and interpret its thresholds","Set gold-task injection and blocking rates that catch drift without taxing throughput","Design an adjudication policy that preserves signal in subjective disagreement","Run structured research on an annotation pool to find unclear guideline boundaries"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min"},{"type":"blog","id":"064048e7-9ad3-4267-a714-8bb8cf6b04e4","slug":"user-interviews-pricing-2026","title":"User Interviews Pricing in 2026: Why $49 a Session Is Really $225","url":"https://www.koji.so/blog/user-interviews-pricing-2026","summary":"User Interviews publishes its full rate card: Recruit costs $49 per completed session on Pay As You Go, $41 on Essential (60-session annual minimum), and $36 on Professional; Research Hub is roughly double at $98/$82/$72. The published rate covers recruiting only. Incentives are funded separately with a 3% processing fee, and advanced B2B targeting adds $35-$49 per session, making a realistic B2B study about $225 per participant and a consumer study about $100. Negotiated data across 32 purchases shows a $7,500 median annual spend. Subscription tiers only pay off above their session minimums. Koji offers AI-moderated interviews from EUR 29/month with analysis included.","content":"\n**Short answer:** User Interviews charges **$49 per completed session** on Pay As You Go for its Recruit product, dropping to **$41** on Essential (60-session annual minimum) and **$36** on Professional. Research Hub costs roughly double: **$98 / $82 / $72**. But that rate is the *recruiting fee only* — participant incentives are billed separately from a prepaid balance with a **3% processing fee**, and advanced B2B targeting adds **$35–$49 per session**. A realistic B2B study costs about **$225 per participant all-in**, not $49. Negotiated contract data across 32 purchases shows a median annual spend of **$7,500**.\n\nUser Interviews deserves genuine credit here. In a category where [dscout](/blog/dscout-pricing-2026) and [UserTesting](/blog/usertesting-pricing-2026) hide every number behind a sales call, User Interviews publishes its full rate card — tiers, minimums, add-ons, and all. The catch is not that the pricing is hidden. It is that the published number covers one line item out of three.\n\n## The published rate card\n\n**Recruit** (you source participants, they run the logistics):\n\n| Plan | Per session | Annual minimum | Discount vs PAYG |\n|---|---|---|---|\n| Pay As You Go | $49 | None | — |\n| Essential | $41 | 60 sessions | 17% |\n| Professional | $36 | Higher volume | 27% |\n| Custom | Negotiated | From 250 sessions | Best rate |\n\n**Research Hub** (recruiting plus a participant CRM, panel management, and workflow automation) runs at roughly double the per-session rate: **$98** Pay As You Go, **$82** Essential with a 150-session minimum, and **$72** Professional.\n\nEvery tier includes unlimited seats, access to a participant pool the company puts at 7.5 million, targeting, scheduling automation, incentive distribution, quota tracking, and integrations. Essential adds quota automation, participant re-invites, and customer success support. Professional adds deeper discounts on surveys and unmoderated tasks. Custom starts at 250 sessions with negotiable terms and a security review.\n\nSeparately, Research Hub is sold in two shapes: a **Workflow** plan starting at 3 seats, and a **CRM** plan starting at 1,000 contacts with panel management, dynamic segments, opt-in forms, nurture campaigns, and PII masking.\n\n## The three costs that stack\n\nHere is where the $49 stops being the whole story.\n\n**1. The recruiting fee.** $49 per completed session on PAYG. This is the number everyone quotes.\n\n**2. Advanced B2B targeting.** If you need participants by job title, seniority, company size, or industry — which is to say, if you are doing B2B research at all — the add-on costs **$45–$49 per session** on Pay As You Go, or **$35–$41** on subscription tiers. For B2B work this roughly *doubles* the platform fee.\n\n**3. Incentives.** These are not included at any tier. You fund a prepaid balance, and automatic distribution carries a **3% processing fee**. Incentives are paid by you, to participants, on top of everything above. What they cost depends entirely on who you are recruiting — our guide to [research participant incentives](/docs/research-participant-incentives) breaks down realistic rates, but B2B professionals and specialists command far more than general consumers.\n\n## What a real study actually costs\n\n**Consumer study, 15 participants, 30-minute sessions, PAYG:**\n\n| Line item | Cost |\n|---|---|\n| Recruiting — 15 × $49 | $735 |\n| Incentives — 15 × $50 | $750 |\n| Incentive processing (3%) | $23 |\n| **Total** | **$1,508** |\n| **Per participant** | **$100** |\n\n**B2B study, 15 participants, senior decision-makers, PAYG:**\n\n| Line item | Cost |\n|---|---|\n| Recruiting — 15 × $49 | $735 |\n| B2B targeting add-on — 15 × $47 | $705 |\n| Incentives — 15 × $125 | $1,875 |\n| Incentive processing (3%) | $56 |\n| **Total** | **$3,371** |\n| **Per participant** | **$225** |\n\nThe published $49 is about **22% of the true per-participant cost** of a B2B study. That is not a criticism of User Interviews' honesty — every one of those line items is disclosed on their pricing page. It is a criticism of how research budgets get built, because most teams budget the recruiting fee and get surprised by a bill three times larger.\n\n## The subscription break-even trap\n\nThe annual tiers look like straightforward savings — 17% off on Essential, 27% on Professional. The trap is the minimum.\n\nEssential requires a **60-session annual commitment at $41**, which is **$2,460 locked in**. Run all 60 sessions and you pay $2,460 versus $2,940 on PAYG — a genuine $480 saving. But run only 40 sessions against that commitment and your effective rate becomes **$61.50 per session**, meaningfully *worse* than the $49 you would have paid with no commitment at all.\n\nThe break-even is not a discount curve. It is a cliff at the minimum:\n\n| Sessions actually run | PAYG cost | Essential cost | Better choice |\n|---|---|---|---|\n| 30 | $1,470 | $2,460 | PAYG |\n| 50 | $2,450 | $2,460 | PAYG (basically tied) |\n| 60 | $2,940 | $2,460 | Essential |\n| 100 | $4,900 | $4,100 | Essential |\n\nResearch Hub's Essential tier raises the stakes considerably: 150 sessions at $82 is a **$12,300** annual commitment. Notably, that is well above the **$7,500 median annual contract** seen across 32 negotiated User Interviews purchases (range: $4,200 to $21,750) — which suggests most buyers land on Recruit rather than the full Research Hub, or negotiate their way below the published minimums.\n\n**Before you commit:** count the sessions your team ran last year, not the sessions you plan to run. Research volume forecasts are optimistic almost by definition, and the tier minimum is the one number in this entire post that punishes optimism.\n\n## What you are actually buying\n\nBe clear about the product boundary. User Interviews is a **recruiting marketplace**. It finds people, screens them, schedules them, and pays them. It does not conduct your interview, and it does not analyse it.\n\nThat means the $225-per-participant B2B figure above buys you a calendar full of scheduled calls. You still need a researcher to run each one, take notes, transcribe, tag, synthesise, and write the report. At a loaded cost of $75–$100 an hour for a researcher — see our [UX researcher salary benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks) — a 15-participant study consumes roughly 30–45 hours of moderation and analysis time. That labour frequently costs more than the recruiting and incentives combined.\n\nRecruiting is only half the job, which is the same conclusion we reached comparing [User Interviews vs Respondent](/blog/user-interviews-vs-respondent-2026) and [User Interviews vs Prolific](/blog/user-interviews-vs-prolific-2026).\n\nThere is also a sample-quality dimension. Marketplace panels attract repeat participants, and heavy repeat participation degrades the representativeness of your sample over time — the dynamic we covered in [professional respondents and panel conditioning](/blog/professional-respondents-panel-conditioning-2026). Good [screener questions](/docs/research-screener-questions) mitigate it; nothing eliminates it.\n\n## The AI-native alternative\n\nThe reason recruiting costs what it does is that a human moderator can only be in one room at a time. Every scheduling constraint, every no-show, every incentive dollar exists to get a person and a researcher into the same slot.\n\nKoji removes that constraint. Interviews are **AI-moderated**, so they run 24/7 and in parallel — no calendars, no moderator bench, no rescheduling. Pricing is published and credit-based:\n\n- **Insights — €29/month**, 29 credits included\n- **Interviews — €79/month**, 79 credits included\n- **Enterprise** — custom\n\nA text conversation costs 1 credit, a full AI-moderated voice interview costs 3, and a report refresh costs 5. The €79 plan therefore covers roughly **26 voice interviews a month**. Compare that against $3,371 for 15 B2B sessions and the arithmetic speaks for itself — even after you add incentives, which you will still pay if you recruit externally.\n\nTwo structural advantages matter more than the price:\n\n**You can recruit from your own product.** The cheapest participant is one who already uses your software. Koji studies can be triggered from inside your app, turning recruiting from a purchased line item into a distribution one — see [in-product research recruiting](/docs/recruiting-from-your-product) and [recruiting B2B participants](/docs/recruiting-b2b-participants).\n\n**Analysis is included, not additional.** Every transcript is thematically analysed automatically, with one-click reports and quotes traced to source. The 30–45 hours of synthesis labour that User Interviews' fee does not cover simply does not appear.\n\nYou also get **six structured question types** — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — mixed directly into conversations. That means NPS distributions, ranked feature preferences, and the qualitative reasoning behind them in a single study, instead of fielding a survey and a round of interviews separately. And because every participant meets the same neutral interviewer, there is **no moderator bias** drifting across a study the way it does across a bench of contractors.\n\nKoji does not replace a recruiting marketplace when you need strangers who have never heard of you. It replaces the far more common case: you have customers, you need to understand them, and you are paying $225 a head plus 40 hours of labour to find that out.\n\n## Frequently Asked Questions\n\n**How much does User Interviews cost per participant?**\nThe platform fee is $49 per completed session on Pay As You Go, $41 on Essential, and $36 on Professional. Research Hub is roughly double at $98 / $82 / $72. All-in, including incentives and B2B targeting, a realistic B2B study costs around $225 per participant and a consumer study around $100.\n\n**Does User Interviews include participant incentives in its pricing?**\nNo. Incentives are funded separately through a prepaid balance and carry a 3% processing fee for automatic distribution. The per-session rate covers recruiting, screening, and scheduling only.\n\n**What is the minimum commitment for User Interviews?**\nPay As You Go has no minimum or contract. Essential requires 60 sessions annually for Recruit ($2,460) or 150 sessions for Research Hub ($12,300). Custom plans start at 250 sessions.\n\n**Is the User Interviews subscription worth it?**\nOnly if you reliably exceed the tier minimum. At 60+ sessions a year Essential saves about $480 versus Pay As You Go. Below roughly 50 sessions, the commitment makes your effective per-session rate higher than Pay As You Go — count last year's actual sessions, not this year's plan.\n\n**How much does B2B targeting cost on User Interviews?**\nAdvanced B2B targeting adds $45–$49 per session on Pay As You Go and $35–$41 per session on subscription tiers, roughly doubling the platform fee for B2B research.\n\n**What do companies actually pay for User Interviews annually?**\nNegotiated contract data across 32 purchases shows a median of $7,500 per year, with a range of $4,200 to $21,750.\n\n## Stop paying per calendar slot\n\nIf your research budget is going to recruiting fees, incentives, B2B add-ons, and then 40 hours of someone's week transcribing what got said, the constraint you are paying to work around is the human moderator.\n\nKoji runs AI-moderated voice interviews in parallel, analyses them automatically, and hands you a report — **from question to insight in hours, not weeks**, with **no research expertise required**. Published pricing from €29/month, no minimums, no session commitments.\n\n[Run your first AI-moderated study →](https://www.koji.so)\n","category":"Comparisons","lastModified":"2026-08-01T03:15:45.112917+00:00","metaTitle":"User Interviews Pricing 2026: Real Cost Per Participant Explained","metaDescription":"User Interviews charges $49/session, but incentives and B2B targeting push real costs to ~$225 per participant. Full rate card, subscription break-even math and negotiated contract data.","keywords":["user interviews pricing","user interviews cost","userinterviews.com pricing","participant recruitment cost","research recruiting pricing","user interviews review"],"aiSummary":"User Interviews publishes its full rate card: Recruit costs $49 per completed session on Pay As You Go, $41 on Essential (60-session annual minimum), and $36 on Professional; Research Hub is roughly double at $98/$82/$72. The published rate covers recruiting only. Incentives are funded separately with a 3% processing fee, and advanced B2B targeting adds $35-$49 per session, making a realistic B2B study about $225 per participant and a consumer study about $100. Negotiated data across 32 purchases shows a $7,500 median annual spend. Subscription tiers only pay off above their session minimums. Koji offers AI-moderated interviews from EUR 29/month with analysis included.","aiKeywords":["user interviews pricing","participant recruitment cost","research incentives","panel pricing","AI moderated interviews"],"aiContentType":"comparison","faqItems":[{"answer":"The platform fee is $49 per completed session on Pay As You Go, $41 on Essential, and $36 on Professional. Research Hub is roughly double at $98/$82/$72. All-in, including incentives and B2B targeting, a realistic B2B study costs around $225 per participant and a consumer study around $100.","question":"How much does User Interviews cost per participant?"},{"answer":"No. Incentives are funded separately through a prepaid balance and carry a 3% processing fee for automatic distribution. The per-session rate covers recruiting, screening, and scheduling only.","question":"Does User Interviews include participant incentives in its pricing?"},{"answer":"Pay As You Go has no minimum or contract. Essential requires 60 sessions annually for Recruit ($2,460) or 150 sessions for Research Hub ($12,300). Custom plans start at 250 sessions.","question":"What is the minimum commitment for User Interviews?"},{"answer":"Only if you reliably exceed the tier minimum. At 60+ sessions a year Essential saves about $480 versus Pay As You Go. Below roughly 50 sessions, the commitment makes your effective per-session rate higher than Pay As You Go.","question":"Is the User Interviews subscription worth it?"},{"answer":"Advanced B2B targeting adds $45-$49 per session on Pay As You Go and $35-$41 per session on subscription tiers, roughly doubling the platform fee for B2B research.","question":"How much does B2B targeting cost on User Interviews?"},{"answer":"Negotiated contract data across 32 purchases shows a median of $7,500 per year, with a range of $4,200 to $21,750.","question":"What do companies actually pay for User Interviews annually?"}],"relatedTopics":["pricing","user interviews","participant recruitment","research budget","panels"]},{"type":"blog","id":"513d039a-6fa7-4e52-bad1-5a9d443706f1","slug":"dscout-pricing-2026","title":"Dscout Pricing in 2026: What Diary Studies Actually Cost (Nothing Is Published)","url":"https://www.koji.so/blog/dscout-pricing-2026","summary":"Dscout publishes no list pricing for its Core, Select, or Enterprise tiers. Negotiated contract data across 53 purchases shows a median annual contract of $49,250, ranging from $24,185 to $109,535, with buyers averaging 17.81% savings off the opening quote. Individual studies quote $15,000-$50,000. Additional costs include separate participant recruitment, incentives, and 3-8% annual renewal escalators. Koji offers published credit-based pricing from EUR 29/month as an AI-native alternative for teams whose need is depth of customer understanding rather than in-home mobile video capture.","content":"\n**Short answer:** Dscout publishes no list price for any of its three plans — Core, Select, and Enterprise are all \"contact sales.\" Based on negotiated-contract data from 53 purchases, the median dscout contract runs **$49,250 per year**, with the typical range spanning **$24,185 to $109,535**. Individual studies commonly quote between $15,000 and $50,000 depending on participant count and duration. Budget an additional 3–8% annual increase at renewal, and remember that participant incentives are usually billed on top of everything else.\n\nIf you came here hoping to find a pricing page with numbers on it, there isn't one. That is not an oversight — it is the business model. This guide reconstructs what dscout actually costs from negotiated contract data, published buyer reports, and dscout's own tier structure, so you can walk into a sales call with a number already in your head.\n\n## Why dscout doesn't publish prices\n\nWe checked dscout's pricing page directly. It lists three tiers and not a single dollar figure:\n\n| Tier | Positioned for | Published price |\n|---|---|---|\n| **Core** | \"Lean teams just getting started\" | Contact sales |\n| **Select** | \"Growing teams looking to scale\" | Contact sales |\n| **Enterprise** | \"Extensive, custom needs\" | Contact sales |\n\nCore includes the AI Studio, unlimited concurrent studies, flex seats, data export, integrations with Miro, Slack and Figma, panel recruitment, and SSO. Select adds real-time intercept studies, question banks, multiple interview moderators, private panels with unlimited participants, and custom consent forms and NDAs. Enterprise layers on custom behavioral targeting, product shipping support for in-home studies, mission summaries, and international recruitment.\n\nNotice what is missing from that ladder: any unit you can count. There is no \"per seat\" number, no \"per study\" number, no published participant rate. Dscout prices per engagement, which means the quote you receive is built around your specific study volume, participant sourcing needs, and — candidly — your perceived budget. Two companies buying the same tier in the same quarter routinely pay very different amounts.\n\nThis is standard for enterprise research platforms. We found the same pattern when we analysed [UserTesting pricing](/blog/usertesting-pricing-2026) and [Qualtrics pricing](/blog/qualtrics-pricing-2026). What makes dscout distinctive is that the opacity extends all the way down to the entry tier — there is no self-serve on-ramp at all.\n\n## What buyers actually pay\n\nNegotiated contract data gives us the numbers dscout won't print. Across 53 analysed dscout purchases:\n\n- **Median contract value: $49,250 per year**\n- **Typical range: $24,185 to $109,535 annually**\n- **Average negotiated savings: 17.81%** off the opening quote\n\nThat last figure is the most useful one on this page. Buyers who negotiated saved an average of nearly 18%, which tells you the first number you hear is not the number you should accept. On a median deal, that spread is close to $9,000.\n\nIndependent buyer guides put the small-team floor lower: platform subscriptions for one to three users running five to ten studies a year commonly land in the **$15,000–$35,000** range, before participant costs. Mid-sized research teams typically spend **$50,000–$100,000 annually** once platform, recruitment, and incentives are combined.\n\n## The per-study math\n\nBecause dscout is frequently bought per project rather than per seat, it helps to think in study units. Reported project quotes cluster like this:\n\n| Study shape | Typical quote |\n|---|---|\n| Small pilot — 10–20 participants, 1–2 weeks | $10,000–$20,000 |\n| Standard study — 20–40 participants, 2–3 weeks | $20,000–$35,000 |\n| Large longitudinal — 50+ participants, 4+ weeks | $40,000–$70,000+ |\n| Full-service managed study (add-on) | +$5,000–$20,000 |\n\nMost teams report that a standard project runs between **$15,000 and $50,000**. Bringing your own participants instead of using dscout's Recruit panel typically cuts **20–40%** off the quote — the single largest lever available to you, and one worth building a [research participant panel](/docs/research-panel-management) for if you buy dscout more than once.\n\n## The four costs nobody quotes you\n\n1. **The platform subscription.** The annual licence itself, tiered by seats and study volume.\n2. **Participant recruitment.** Sourcing from dscout's panel is billed separately. If your criteria are niche — B2B decision-makers, clinicians, low-incidence consumers — the cost climbs steeply, for the same reasons we broke down in [survey sample cost, CPI and incidence rate](/blog/survey-sample-cost-cpi-incidence-rate-2026).\n3. **Incentives.** Diary studies are longitudinal and demanding. Participants completing multi-week missions expect materially more than a 30-minute interview pays. Our guide to [research participant incentives](/docs/research-participant-incentives) covers realistic rates by study type — and if you are paying US participants, the [1099 thresholds](/docs/research-participant-incentive-taxes) apply to you, not to dscout.\n4. **Renewal escalators.** Dscout contracts commonly include **3–8% annual price increases** at renewal, even when your scope and usage have not changed. Over a three-year term that compounds into real money, and it is negotiable at signing in a way it never is at renewal.\n\n## Negotiation levers that actually work\n\nBuyer data points to four repeatable tactics:\n\n- **Prepay for volume.** Buyers purchasing large participant credit bundles at signing secured **10–20% lower** per-participant rates than pay-per-study buyers.\n- **Time the quarter.** Deals closed in Q4 or at quarter-end included **5–15%** incremental discounts or bonus credits versus mid-quarter transactions.\n- **Run a real evaluation.** Buyers actively comparing alternatives and sharing a competitive timeline received **10–20%** sharper pricing. Our roundup of [dscout alternatives](/blog/dscout-alternatives-2026) and [diary study apps](/blog/best-diary-study-apps-2026) is a reasonable place to build that shortlist.\n- **Cap the escalator in writing.** Negotiate the renewal increase at initial signature, not at renewal.\n\n## Who dscout is genuinely worth it for\n\nDscout is a strong product and the honest answer is that some teams should pay for it. If you need participants to capture in-the-moment video from their phones, over weeks, in their actual homes and workplaces — with device shipping, international recruitment, and behavioural targeting — dscout does that better than nearly anyone. Ethnographic and longitudinal work at that fidelity is expensive everywhere, and a $50K contract can be entirely rational for a large team running continuous in-context research.\n\nThe problem is that most teams buying dscout are not doing that. They are buying a $49,250 platform to answer questions like *why did this segment churn*, *what do people actually think of this pricing page*, or *how did onboarding feel in week one* — questions that do not require a mobile ethnography suite. They are paying for in-context video capture and using it as an expensive way to ask people things.\n\n## The AI-native alternative: what Koji costs\n\nKoji was built for exactly that second group. Instead of buying a per-study enterprise contract, you run **AI-moderated voice and text interviews** that conduct themselves, probe follow-ups automatically, and produce analysed reports without a researcher transcribing anything.\n\nThe pricing is published, which is the whole point:\n\n- **Insights — €29/month**, including 29 credits\n- **Interviews — €79/month**, including 79 credits\n- **Enterprise** — custom, for teams that need it\n\nA text conversation costs 1 credit; a full AI-moderated voice interview costs 3; refreshing a report costs 5. That makes the €79 plan roughly **26 voice interviews a month** — a sample size that would sit inside a \"small pilot\" bracket at dscout, for a rounding error against a $10,000–$20,000 quote. And there is a quality gate: **only conversations scoring 3 or above consume credits**, so you are not billed for junk sessions the way you are billed for panel participants who ghost a mission.\n\nWhat you get for it:\n\n- **AI-moderated voice interviews** that run 24/7, in parallel, with no scheduling and no moderator availability ceiling\n- **Automatic thematic analysis** across every transcript — no tagging backlog, no repository to maintain\n- **Six structured question types** mixed directly into conversations: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. You get NPS distributions and ranked preferences *and* the qualitative reasoning behind them in one study, rather than running a survey and interviews separately\n- **Customisable AI consultants** you shape to your domain and research goals\n- **One-click reports** with quotes traced back to source\n- **No moderator bias** — every participant gets the same neutral interviewer, which is more than can be said for a rotating bench of contract moderators\n\nThe honest limitation: Koji does not ship devices to participants' homes or capture four weeks of in-situ mobile video. If that is your actual requirement, dscout earns its price. If your requirement is depth of understanding at speed, an AI-native platform gets you from question to insight in hours rather than the six-to-eight weeks a procurement cycle plus fielding plus analysis takes — and see [how to run a diary study](/docs/diary-study-guide) for where longitudinal designs still make sense.\n\n## Dscout vs Koji at a glance\n\n| | dscout | Koji |\n|---|---|---|\n| Published pricing | None — all tiers contact sales | Yes, from €29/month |\n| Median annual spend | $49,250 (53 contracts) | Self-serve, no contract required |\n| Contract term | Annual, multi-year common | Monthly or annual |\n| Renewal escalator | 3–8% typical | None |\n| Time to first insight | Weeks | Hours |\n| Moderator required | Yes | No — AI-moderated |\n| Analysis | Manual tagging | Automatic thematic analysis |\n| Structured + qualitative in one study | Limited | Yes — 6 question types |\n\n## Frequently Asked Questions\n\n**How much does dscout cost per year?**\nDscout does not publish prices. Negotiated contract data across 53 purchases puts the median at $49,250 per year, with a typical range of $24,185 to $109,535. Small teams running five to ten studies annually often land between $15,000 and $35,000 before participant costs.\n\n**Does dscout have a free trial or free plan?**\nNo. There is no free tier and no self-serve signup. Every dscout plan — including the entry-level Core tier — requires a sales conversation and a custom quote.\n\n**How much does a single dscout study cost?**\nReported quotes run roughly $10,000–$20,000 for a 10–20 participant pilot, $20,000–$35,000 for a standard 20–40 participant study, and $40,000–$70,000+ for large longitudinal work. Full-service managed studies add $5,000–$20,000.\n\n**Are participant incentives included in dscout pricing?**\nGenerally no. Incentives are typically billed separately from the platform subscription and recruitment fees, and longitudinal diary studies require higher incentives than one-off interviews because of the sustained effort involved.\n\n**Can you negotiate dscout pricing?**\nYes, and you should. Buyers averaged 17.81% savings off the opening quote. The strongest levers are prepaying for participant credit bundles (10–20% off), closing at quarter-end (5–15% extra), running a documented competitive evaluation (10–20%), and capping the renewal escalator in the initial contract.\n\n**What is the cheapest way to run diary-style research?**\nBringing your own participants cuts a dscout quote by 20–40%. Beyond that, AI-moderated platforms like Koji run interviews at credit-based pricing from €29/month — roughly 26 voice interviews on the €79 plan — which suits teams whose real need is depth of understanding rather than in-home mobile video capture.\n\n## Get a real number before your sales call\n\nThe worst position to negotiate from is not knowing what anyone else paid. Now you do: **$49,250 median, 17.81% average savings, 3–8% escalators.** Use it.\n\nAnd if the honest answer is that you need to understand your customers deeply — not film them in their kitchens for a month — you can start today for less than a single dscout study costs in review meetings. Koji gives you AI-moderated voice interviews, automatic thematic analysis, six structured question types, and one-click reports, with **no research expertise required** and **no procurement cycle**.\n\n[Start your first study with Koji →](https://www.koji.so)\n","category":"Comparisons","lastModified":"2026-08-01T03:14:17.372758+00:00","metaTitle":"Dscout Pricing 2026: Real Costs, Contracts & What Buyers Pay","metaDescription":"Dscout publishes no prices. Negotiated data shows a $49,250 median contract across 53 purchases. Full breakdown of tiers, per-study costs, hidden fees and cheaper alternatives.","keywords":["dscout pricing","dscout cost","how much does dscout cost","dscout pricing 2026","dscout contract","diary study cost","dscout alternatives"],"aiSummary":"Dscout publishes no list pricing for its Core, Select, or Enterprise tiers. Negotiated contract data across 53 purchases shows a median annual contract of $49,250, ranging from $24,185 to $109,535, with buyers averaging 17.81% savings off the opening quote. Individual studies quote $15,000-$50,000. Additional costs include separate participant recruitment, incentives, and 3-8% annual renewal escalators. Koji offers published credit-based pricing from EUR 29/month as an AI-native alternative for teams whose need is depth of customer understanding rather than in-home mobile video capture.","aiKeywords":["dscout pricing","diary study cost","user research budget","research platform contracts","AI moderated interviews"],"aiContentType":"comparison","faqItems":[{"answer":"Dscout does not publish prices. Negotiated contract data across 53 purchases puts the median at $49,250 per year, with a typical range of $24,185 to $109,535. Small teams running five to ten studies annually often land between $15,000 and $35,000 before participant costs.","question":"How much does dscout cost per year?"},{"answer":"No. There is no free tier and no self-serve signup. Every dscout plan, including the entry-level Core tier, requires a sales conversation and a custom quote.","question":"Does dscout have a free trial or free plan?"},{"answer":"Reported quotes run roughly $10,000-$20,000 for a 10-20 participant pilot, $20,000-$35,000 for a standard 20-40 participant study, and $40,000-$70,000+ for large longitudinal work. Full-service managed studies add $5,000-$20,000.","question":"How much does a single dscout study cost?"},{"answer":"Generally no. Incentives are typically billed separately from the platform subscription and recruitment fees, and longitudinal diary studies require higher incentives than one-off interviews because of the sustained effort involved.","question":"Are participant incentives included in dscout pricing?"},{"answer":"Yes. Buyers averaged 17.81% savings off the opening quote. The strongest levers are prepaying for participant credit bundles (10-20% off), closing at quarter-end (5-15% extra), running a documented competitive evaluation (10-20%), and capping the renewal escalator in the initial contract.","question":"Can you negotiate dscout pricing?"},{"answer":"Bringing your own participants cuts a dscout quote by 20-40%. Beyond that, AI-moderated platforms like Koji run interviews at credit-based pricing from EUR 29/month, roughly 26 voice interviews on the EUR 79 plan, which suits teams whose real need is depth of understanding rather than in-home mobile video capture.","question":"What is the cheapest way to run diary-style research?"}],"relatedTopics":["pricing","dscout","diary studies","research budget","vendor comparison"]},{"type":"documentation","id":"8765c9ea-50a9-42c6-b130-c72a118e78df","slug":"expert-interviews-guide","title":"Expert Interviews: How to Plan, Recruit, and Run Them","url":"https://www.koji.so/docs/expert-interviews-guide","summary":"An expert interview is a structured conversation with a subject-matter expert to compress months of domain learning into a single conversation. Where customer interviews reveal what users feel, expert interviews reveal how a domain actually works — its hidden rules, playbooks, and failure modes. This guide covers when to use them, how to recruit the right experts and triangulate across several, a four-phase interview structure, question techniques (ask for stories, probe exceptions, informed naivety, quantify with anchors), mapping to Koji''s six structured question types, common pitfalls, and how Koji runs asynchronous AI expert interviews in parallel with automatic analysis.","content":"# Expert Interviews: How to Plan, Recruit, and Run Them\n\n**An expert interview is a structured conversation with a subject-matter expert — a practitioner, domain specialist, or experienced operator — to compress months of learning into a single conversation.** Where customer interviews tell you what users feel, expert interviews tell you how a domain actually works: its hidden rules, its standard playbooks, its common failure modes, and the things insiders take for granted but outsiders never see.\n\nThey are the fastest way to get up the learning curve in an unfamiliar market — and one of the most underused methods, because teams assume experts are hard to reach. This guide covers when to use expert interviews, how to recruit the right experts, the question techniques that separate a deep interview from a wasted hour, and how AI-native platforms like Koji let you scale expert input across many specialists at once.\n\n## When to use expert interviews\n\nReach for expert interviews when:\n\n- You are **entering an unfamiliar domain** (a new vertical, regulation-heavy market, or technical field) and need to learn its landscape fast.\n- You need **the why behind the what** — not just that a workflow exists, but why it evolved that way and where it breaks.\n- You are **validating feasibility or risk** before committing — what will bite us that we don''t know to ask about?\n- You want **to pressure-test a hypothesis** with someone who has seen many companies try and fail at the same thing.\n\nExpert interviews complement customer research; they do not replace it. Experts know the domain, but only your customers know their own experience. Use experts to build the map, then use customer interviews to walk the territory.\n\n## How to recruit the right experts\n\nThe quality of an expert interview is set before it starts, by who you talk to. Look for:\n\n- **Hands-on practitioners**, not just commentators. Someone who has *done* the job five hundred times beats someone who has written about it.\n- **Range of vantage points.** Interview a few experts from different angles — a former operator, a consultant who has seen many companies, a vendor who serves the market — so you triangulate rather than absorb one person''s bias.\n- **Recency.** Domains change. Favor experts whose hands-on experience is current.\n\nFind them through your network, LinkedIn, industry communities, former colleagues, and warm introductions (always end every interview by asking \"who else should I talk to?\"). For breadth, expert panels and recruited specialist pools let you reach many at once.\n\n## How to structure an expert interview\n\nA 45-minute expert interview works well in four phases:\n\n1. **Credibility and context (5 min).** Understand exactly what their expertise is and where its edges are, so you can weight their answers correctly.\n2. **Landscape mapping (15 min).** Have them lay out how the domain works: the key players, the standard process, the levers that matter.\n3. **Failure modes and edge cases (15 min).** Where do companies get this wrong? What looks easy but isn''t? What would you warn a newcomer about?\n4. **Synthesis and pointers (10 min).** Pressure-test your hypothesis, and ask for resources, frameworks, and other experts.\n\n## Expert interview question techniques\n\nExperts reward sharp questions and punish lazy ones. Use these techniques:\n\n- **Ask for stories, not opinions.** \"Tell me about a time a launch like this went wrong\" beats \"What are best practices?\" Stories carry specifics; opinions blur into platitudes.\n- **Probe the exceptions.** \"When does the usual advice *not* apply?\" is where genuine expertise lives.\n- **Use informed naivety.** Do enough homework to ask credible questions, then let yourself ask the dumb-but-important one: \"Why is it done that way at all?\"\n- **Quantify with anchors.** \"Out of ten companies that try this, how many succeed?\" turns vibes into a number you can compare across experts.\n- **Challenge gently.** \"Another expert told me the opposite — how would you respond?\" surfaces the real debates in a field.\n\n## Map expert interviews to structured question types\n\nEven an expert interview benefits from a few structured anchors so you can compare what multiple specialists say. Koji supports **six structured question types** — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no`:\n\n- **open_ended** for landscape mapping and failure-mode stories, with AI follow-up probing.\n- **scale** for confidence or risk ratings (\"How risky is this approach, 1–5?\").\n- **ranking** to have experts order the biggest risks or levers.\n- **single_choice / multiple_choice** to standardize which tools, methods, or players they reference.\n- **yes_no** for sharp directional reads (\"Is this market consolidating?\").\n\nSee the [structured questions guide](/docs/structured-questions-guide) for how each type aggregates across interviews.\n\n## Avoiding the pitfalls\n\n- **Over-indexing on one expert.** Even the best expert has blind spots and biases. Triangulate across several.\n- **Treating opinion as fact.** Experts predict the future no better than anyone; weight their *experience* heavily and their *forecasts* lightly.\n- **Letting them stay abstract.** Push every general claim down to a specific example.\n- **Skipping the homework.** Experts disengage when you ask things a five-minute search would answer. Earn the deep questions.\n\n## How Koji scales expert interviews\n\nThe classic constraint on expert research is that senior people are busy and a live interview is hard to schedule. Koji loosens both. Its **AI interviewer conducts expert conversations by voice or text, asynchronously, on the expert''s own time** — no calendar coordination, no moderator — and automatically probes for the specifics and exceptions that make expert input valuable. That means you can gather input from **many specialists in parallel** instead of stretching one interview a week across a quarter.\n\nKoji then **analyzes every transcript automatically**, clustering what experts agree on, flagging where they disagree, and surfacing the verbatim warnings and stories worth quoting — with your scale and ranking anchors aggregated into charts. Instead of one expert''s view filtered through your memory, you get a synthesized, evidence-grounded map of a domain from a panel of specialists, delivered as a shareable report.\n\n## Expert interviews vs secondary research\n\nBefore you book an expert, exhaust the cheap sources: industry reports, earnings calls, trade publications, public talks, and a focused literature search. Secondary research is free, fast, and tells you what is already documented. The point of an expert interview is to go *past* what is written down — to the judgment, the unwritten heuristics, and the candid \"here is what actually happens\" that no published source will print. Use secondary research to get smart enough to ask good questions, then use the expert to get the answers only experience can provide.\n\nA practical division of labor: secondary research answers \"what is generally true,\" and expert interviews answer \"what is true here, now, and what would surprise a newcomer.\" Showing up having done your homework also earns you the deep questions — experts disengage fast from anyone asking what a five-minute search would answer, and lean in for someone who clearly respects their time.\n\n## Synthesizing across multiple experts\n\nThe real insight from expert research rarely comes from one interview — it comes from the pattern across several. Lay the transcripts side by side and look for three things: **consensus** (where experts agree, you can act with confidence), **disagreement** (where they split, you have found the genuine debate worth investigating), and **outliers** (a lone contrarian view that, if correct, changes everything). Weight hands-on experience heavily and confident forecasts lightly. Koji makes this synthesis automatic: it clusters where your experts align, flags where they diverge, and aggregates the confidence and risk ratings you captured into charts — so you leave with a defensible map of the domain rather than the opinion of whichever expert you spoke to last.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how each aggregates\n- [Requirements Gathering Interviews](/docs/requirements-gathering-interviews-guide) — eliciting needs from stakeholders and experts\n- [Customer Discovery Call: Questions & Template](/docs/customer-discovery-call-guide) — the customer-facing counterpart\n- [Customer Interview Questions: 60+ Examples](/docs/customer-interview-questions-examples) — a broad question bank\n- [AI Interviews vs Surveys](/docs/ai-interviews-vs-surveys) — why conversational research beats static forms\n- [Generating Research Reports](/docs/generating-research-reports) — turning interviews into shareable findings","category":"Interview Techniques","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Expert Interviews: How to Plan, Recruit & Run Them","metaDescription":"Expert interviews compress months of domain learning into a few conversations. Learn how to recruit experts, structure the interview, ask sharp questions, and scale with Koji.","keywords":["expert interview","expert interviews","how to conduct expert interviews","subject matter expert interview","sme interview","expert interview questions","domain expert research"],"aiSummary":"An expert interview is a structured conversation with a subject-matter expert to compress months of domain learning into a single conversation. Where customer interviews reveal what users feel, expert interviews reveal how a domain actually works — its hidden rules, playbooks, and failure modes. This guide covers when to use them, how to recruit the right experts and triangulate across several, a four-phase interview structure, question techniques (ask for stories, probe exceptions, informed naivety, quantify with anchors), mapping to Koji''s six structured question types, common pitfalls, and how Koji runs asynchronous AI expert interviews in parallel with automatic analysis.","aiPrerequisites":["A domain or hypothesis you need expert input on","Access to or a way to recruit subject-matter experts"],"aiLearningOutcomes":["Decide when expert interviews are the right method","Recruit and triangulate across the right experts","Run a four-phase expert interview","Use question techniques that get past platitudes","Scale expert input with asynchronous AI interviews and analysis"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"8939f336-bbe8-4ac4-8808-9e1c447e1f79","slug":"research-participant-incentives","title":"Research Participant Incentives: How Much to Pay and What to Offer","url":"https://www.koji.so/docs/research-participant-incentives","summary":"Research participant incentives increase response rates by 8-19 percentage points and cut no-shows from 20-25% to 5-10%. This guide covers standard rates by participant type (consumers $50-100/hr, B2B professionals $100-200/hr, executives $200-500/session), incentive types and their tradeoffs, bias considerations, GDPR and tax compliance, and how AI-moderated research changes the cost-per-insight calculation.","content":"\n# Research Participant Incentives: How Much to Pay and What to Offer\n\n**Bottom line:** Research participant incentives increase response rates by 8–19 percentage points and cut no-show rates from 20–25% down to 5–10%. The right amount depends on participant type, study duration, and recruitment difficulty — but paying too little costs you more in failed recruitment than paying appropriately would have.\n\n---\n\n## Why Incentives Matter\n\nParticipant incentives are compensation offered to research study participants in exchange for their time, attention, and expertise. They are not a courtesy — they are a recruitment and quality mechanism.\n\nThe evidence is clear. A meta-analysis published in PLOS ONE found that monetary incentives increase response rates by **8–10 percentage points** across mail, phone, and web studies. A 2024 study in JMIR Formative Research found cooperation rates of **50.6% with incentives versus 26.6% without** — nearly double the participation. And according to Nielsen Norman Group, appropriate incentives reduce no-show rates from 20–25% to approximately 5–10%.\n\nThe business case for incentives is straightforward: failed recruitment attempts, rescheduling, and conducting studies with poor-fit participants all cost more than a well-structured incentive program would have. Paying fairly is not generosity — it is operational efficiency.\n\n---\n\n## Standard Incentive Amounts by Participant Type\n\nIncentive amounts should scale with the value of participants' time, the rarity of their experience, and the difficulty of recruiting them. These ranges reflect current market rates:\n\n### General Consumer Participants\n- **30-minute interview:** $25–50\n- **60-minute interview:** $50–100\n- **90-minute session:** $75–150\n\nNielsen Norman Group data puts the average incentive for nonprofessional users at **$32 per hour** of test time, though rates have risen since that benchmark was established.\n\n### B2B Professional Participants\nProfessionals — software buyers, operations managers, procurement leads, marketing directors — command 1.5–2x consumer rates because their time is more valuable and they are harder to recruit.\n\n- **60-minute interview:** $100–200\n- **90-minute session:** $150–300\n\n### Senior Professionals and Specialists\nDomain experts with specific, hard-to-find experience (senior engineers, compliance officers, clinical practitioners, security professionals) command premium rates reflecting their scarcity and hourly market value.\n\n- **60-minute interview:** $150–300\n- Rule of thumb: pay at least **2x their estimated hourly rate** to make participation worth redirecting from billable or high-priority work\n\n### Executive and C-Suite Participants\nExecutives are among the hardest research participants to recruit and retain. Their time is exceptionally constrained, and standard incentive amounts fail to compete with what they sacrifice to participate.\n\n- **30-minute interview:** $200–300\n- **45–60-minute session:** $300–500\n- Nielsen Norman Group recommends offering executives both a cash equivalent and the option of a charitable donation in their name — the latter often resonates with senior leaders\n\nNielsen Norman Group's research shows the incentive ratio in practice: high-level professionals receive an average of **$118 per hour** versus **$32 per hour** for general consumers — nearly a 4x difference reflecting genuine differences in recruitment difficulty and participant opportunity cost.\n\n### Healthcare and Clinical Participants\nHealthcare professionals (physicians, nurses, pharmacists) face scheduling constraints, professional gatekeepers, and regulatory considerations. Rates typically match senior professional ranges, with additional consideration for compliance complexity.\n\n### Internal Employees\n**Do not pay your own employees additional monetary incentives for research participation.** Nielsen Norman Group's recommendation is direct: employees are already compensated for their time. Only 10% of companies pay monetary incentives internally. The remaining 90% use non-monetary approaches: public recognition, priority access to findings, lunch-and-learn sessions, or simply treating participation as meaningful work.\n\nPaying employees creates a different problem — it can make participation feel transactional rather than intrinsically motivated, and it introduces equity concerns across teams with different research exposure.\n\n---\n\n## Incentive Types: What Works and When\n\n### Gift Cards (Most Recommended)\nDigital gift cards — particularly choice-based options where participants select from multiple brands — are the most universally effective incentive type. They offer:\n\n- **Instant digital delivery** reducing friction and wait time\n- **Perceived value** that feels like a treat rather than a transaction\n- **Flexibility** that cash does not always carry (some organizations prohibit cash equivalents)\n- **Broad demographic appeal** when offered as choice-based rather than brand-specific\n\nThe caveat: brand-specific gift cards can inadvertently attract participants who are customers of that brand. Choice-based gift card platforms (Tremendous, Rybbon, Tango) solve this by letting participants choose their preferred retailer.\n\nDining gift cards now rank as the top participant preference (54%), surpassing online retailers (50%) and clothing (48%) according to Incentive Research Foundation 2026 trend data — a useful signal for format selection.\n\n### Cash and Direct Payments (PayPal, Venmo, Bank Transfer)\nCash is maximally flexible and universally appealing. For high-value studies with well-defined participant populations, direct payment is often the fastest path. The complications:\n\n- **Government employees cannot accept cash** due to ethics and procurement rules — gift cards are typically acceptable where cash is not\n- **Tax reporting obligations** at threshold amounts (in the US, $600+ requires a 1099)\n- **Payment platform friction** — collecting PayPal addresses or bank details adds steps that reduce completion rates for lower-value studies\n- **International transfers** carry exchange rate and compliance complexity\n\nFor incentives over $100, direct payment is often preferred by participants. For incentives under $50, gift cards typically produce higher satisfaction per dollar spent.\n\n### Product Credits and Discounts\nOffering your own product as an incentive is cost-effective but carries significant limitations. It only works when:\n\n- The participant is already a user or highly likely to become one\n- The product credit is genuinely valuable relative to what you are asking\n- The incentive does not inadvertently bias participants toward positive feedback to protect their credit\n\nNever use product credits as the sole incentive for research that will inform major product decisions — participants have a conflict of interest in giving honest negative feedback.\n\n### Charitable Donations\nSome participant populations — particularly senior executives, academics, and socially-motivated professionals — respond well to charitable donation options. Offering to donate $150 to a cause of their choice often resonates where a $150 Amazon gift card would not.\n\nThis option is also useful for internal research where monetary compensation is inappropriate but a meaningful gesture is still valuable.\n\n### Sweepstakes and Lottery Incentives\nOffering a chance to win a larger prize (e.g., one in 20 participants wins $500) dramatically reduces per-participant incentive cost. The tradeoff is significant: response rates are lower than guaranteed incentives, and research shows sweepstakes attract systematically different participants — more independent-minded and risk-tolerant — which can introduce selection bias.\n\nFor exploratory research where diverse perspectives are valuable, sweepstakes may be appropriate. For validation research requiring representative samples, guaranteed incentives produce better-fit participants.\n\n---\n\n## Do Incentives Bias Research Results?\n\nThis is the most common concern researchers raise about incentives — and the research answer is nuanced.\n\n**Incentives do not make people answer dishonestly.** Participants who want to give you good data will give you good data regardless of whether you are paying them $50 or $100. The quality mechanism is the screener and the study design, not the incentive level.\n\n**The real bias risk is selection bias — who participates.** Research published in PLOS ONE and a 2016 CSCW study found that different incentive types attract systematically different participant profiles:\n- Lottery incentives attract more independent, self-determined participants\n- Charitable donation options attract more community-oriented participants\n- Pure cash attracts more financially-motivated participants\n\nThe implication: your incentive type should match your target participant profile. If you want community-minded users of a healthcare platform, a charitable donation option may recruit more authentic participants than a cash equivalent.\n\n**A counterintuitive finding:** one study found incentives actually *reduced* bias by encouraging people without extreme views to participate — pulling the sample toward the moderate majority rather than the self-selecting motivated extreme. Incentives can produce more representative samples when the unincentivized alternative is that only highly opinionated participants bother to respond.\n\nThe net conclusion: incentives are a participation mechanism, not a data quality threat. Design your screener to filter for fit and your questions to minimize leading — those are the real determinants of data quality.\n\n---\n\n## How to Calculate the Right Incentive Amount\n\nA practical formula from the research community:\n\n- **Moderated sessions:** $3 per minute (a 30-minute interview = $90, a 60-minute interview = $180)\n- **Unmoderated sessions:** $0.20 per minute (an asynchronous video task = lower commitment, lower incentive)\n\nAdjust from that baseline for:\n\n| Factor | Adjustment |\n|--------|-----------|\n| In-person instead of remote | +25% |\n| Site visit or facility | +35% |\n| B2B professional | +50–100% |\n| Executive/C-suite | +150–200% |\n| Niche specialist | +25–50% |\n| International participant | Adjust for local purchasing power parity |\n\nIf recruitment is stalling — you are getting fewer than 2–3 applications per slot — increase the incentive by 15–25% and reassess. Slow recruitment at a given incentive level is the most reliable signal that your rate is below market.\n\nSeveral free incentive calculators (Ethnio, Respondent, User Interviews, Tremendous) provide data-backed estimates based on study type, participant profile, and geography.\n\n---\n\n## Legal and Compliance Considerations\n\n**GDPR (EU participants):** Research consent must be freely given — meaning participants must have the realistic ability to decline without meaningful consequence. Incentives must not be so large that they effectively coerce participation. IRB approval for studies involving EU residents typically requires confirmation of GDPR compliance in both the research protocol and the incentive platform.\n\n**CCPA (California):** Similar consent and data handling requirements apply. Your incentive platform must comply if collecting personal data from California residents.\n\n**IRB requirements:** For academic and clinical research, incentives must be disclosed in the IRB protocol and evaluated holistically. Tax implications must be addressed in the documentation.\n\n**Tax reporting (US):** Participants who receive $600 or more in incentives in a calendar year from the same organization typically trigger IRS 1099 reporting requirements. Structure your program and track accordingly — incentive platforms like Tremendous handle 1099 generation automatically.\n\n**Government employees:** Direct cash and cash-equivalent incentives are typically prohibited for government employees due to ethics rules. Check the specific agency's procurement and gift policies before recruiting. Non-monetary incentives or donations to government-approved charities are often acceptable alternatives.\n\n---\n\n## Common Incentive Mistakes That Cost Research Teams\n\n**Paying below market rate and blaming recruitment.** Slow recruitment is often misattributed to screener criteria or panel quality when the actual problem is a $25 incentive for a 60-minute interview with a senior professional. Adequate incentives are the fastest fix for sluggish recruitment.\n\n**Using the same incentive for every study.** A 20-minute survey of general consumers and a 75-minute structured interview with a CISO are not equivalent asks. Standardizing on a flat rate undervalues specialist participants and overvalues commodity ones.\n\n**Offering brand-specific gift cards.** An Amazon gift card to participants who do not use Amazon, or an Uber Eats credit to participants in areas with poor coverage, is a zero-value incentive that creates frustration. Choice-based platforms solve this with minimal additional cost.\n\n**Slow or complicated payment.** Incentive delivery that takes weeks, requires participants to fill out forms, or involves multiple follow-up emails damages participant experience and recruitment reputation. Digital platforms with instant delivery are worth the platform cost.\n\n**Ignoring tax implications until they become a problem.** Discovering mid-program that your incentive distribution approach triggers reporting requirements you are not set up to handle is expensive to fix retroactively. Build compliance into the program design from day one.\n\n---\n\n## How AI-Moderated Research Changes the Incentive Equation\n\nTraditional qualitative research programs pay incentives proportional to session length — because sessions require human moderator time, scheduling coordination, and participant commitment to a specific calendar slot. This naturally limits scale: a $100 incentive for a 60-minute interview across 20 participants is a $2,000 incentive budget before accounting for recruitment, moderation, and analysis costs.\n\nAI-native platforms like Koji change this in two important ways:\n\n**First, session completion becomes easier and more flexible.** When participants complete AI-moderated interviews asynchronously — on their own schedule, from their own device — the perceived cost of participation drops even at the same nominal incentive level. Koji interviews have no scheduling friction and no \"I forgot\" cancellations. The structured question format (across open_ended, scale, single_choice, multiple_choice, ranking, and yes_no types) means participants always know exactly what they are being asked and what step they are on.\n\n**Second, insight per dollar spent increases dramatically.** Because AI moderation removes the human bottleneck, a given incentive budget can fund 5–10x more participant conversations. Rather than choosing between 10 in-depth interviews at $100 each or a survey with no qualitative depth, teams running Koji can run 50–100 AI-moderated interviews with structured and open-ended questions — producing both the quantitative distribution data of a survey and the qualitative reasoning of interviews — within a comparable budget.\n\nFor teams scaling research programs, this means participant incentives become a smaller fraction of total research cost, and the return per dollar of incentive investment increases substantially.\n\n---\n\n## Related Resources\n\n- [Structured Questions Guide: All 6 Question Types in Koji](/docs/structured-questions-guide)\n- [Research Screener Questions: How to Find the Right Participants](/docs/research-screener-questions)\n- [User Research Recruitment Email Templates That Get Responses](/docs/user-research-recruitment-email-templates)\n- [How to Research Hard-to-Reach Audiences: Executives, B2B Buyers, and Niche Segments](/docs/hard-to-reach-participants-research)\n- [Reducing No-Shows in User Research](/docs/reducing-no-shows)\n- [Personalized Interview Links: Send Targeted Research Invitations](/docs/personalized-interview-links)\n\n\n## Further reading on the blog\n\n- [Koji vs Respondent.io: AI-Native Research Platform vs Participant Recruitment Marketplace (2026)](/blog/koji-vs-respondent-2026) — Respondent.io recruits research participants. Koji conducts AI-moderated interviews, analyzes results, and generates reports automatically —\n- [Best Research Participant Recruitment Platforms in 2026: The Complete Buyer's Guide](/blog/participant-recruitment-platforms-2026) — Finding qualified participants is the #1 challenge in user research. This guide compares the top recruitment platforms—Prolific, UserIntervi\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n\n<!-- further-reading:blog -->\n","category":"Participant Recruitment","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Research Participant Incentives: How Much to Pay in 2026 (With Benchmarks)","metaDescription":"Learn exactly how much to pay research participants by type, which incentive formats work best, whether incentives bias results, and how to calculate the right amount for any study. Includes NNG benchmarks and compliance guidance.","keywords":["research participant incentives","user research incentives","how much to pay research participants","participant compensation","research incentive amounts","incentives for user research","participant recruitment incentives","UX research compensation"],"aiSummary":"Research participant incentives increase response rates by 8-19 percentage points and cut no-shows from 20-25% to 5-10%. This guide covers standard rates by participant type (consumers $50-100/hr, B2B professionals $100-200/hr, executives $200-500/session), incentive types and their tradeoffs, bias considerations, GDPR and tax compliance, and how AI-moderated research changes the cost-per-insight calculation.","aiPrerequisites":["Planning or running a user research study","Responsibility for participant recruitment or research operations"],"aiLearningOutcomes":["Know standard incentive amounts for consumers, B2B professionals, and executives","Choose the right incentive type (gift card, cash, donation) for your participant profile","Understand how incentives affect bias and data quality — and what the research actually shows","Handle GDPR, IRB, and tax compliance requirements for research incentive programs","Calculate the right incentive amount using a structured formula with adjustment factors"],"aiDifficulty":"beginner","aiEstimatedTime":"10 minutes"},{"type":"documentation","id":"505c3cc9-7d0c-4a60-9f03-6c5680e0d033","slug":"accessibility-research-guide","title":"Accessibility Research: How to Include Users with Disabilities in Your Studies","url":"https://www.koji.so/docs/accessibility-research-guide","summary":"Accessibility research involves intentionally including users with disabilities in research studies. This guide covers why most research programs exclude users with disabilities (and how to fix it), how to adapt recruitment and methods for different disability types, key principles for inclusive facilitation, and how async AI interview platforms like Koji remove structural participation barriers.","content":"## Accessibility Research: How to Include Users with Disabilities in Your Studies\n\n**Accessibility research is the practice of intentionally including users with disabilities in your research process — not as an afterthought, but as a core part of understanding how your product works in the real world.**\n\nAn estimated 1.3 billion people globally live with some form of disability — about 16% of the world's population, according to the World Health Organization. If your user research doesn't include them, you're making product decisions based on an incomplete picture. The features, flows, and interactions that work for your non-disabled users may be barriers for a significant portion of your actual user base.\n\nMore practically: inclusive research produces better products for everyone. Curb cuts were designed for wheelchair users but benefit parents with strollers, delivery workers with carts, and travelers with luggage. The same principle applies in software. Designing for accessibility reveals usability problems that affect all users, not just those with disabilities.\n\n---\n\n## Why Most Research Programs Exclude Users with Disabilities\n\nAccessibility research gaps aren't usually intentional — they're structural. Common barriers include:\n\n**Recruitment challenges.** Standard recruitment panels skew toward highly digital-native, able-bodied participants. Reaching users with specific disabilities requires deliberate outreach through disability-focused communities, advocacy organizations, and specialized recruitment partners.\n\n**Method inflexibility.** Many research methods assume participants can see a screen, use a mouse, speak clearly, or respond within a fixed time window. These assumptions exclude large populations.\n\n**Facilitation barriers.** Live, moderated sessions can be cognitively and physically demanding for participants with certain disabilities. The presence of a moderator can also create anxiety that affects response quality.\n\n**Incentive and logistics friction.** Research studies often require scheduling coordination, technical setup, and real-time availability — all of which create disproportionate burden for participants with mobility, cognitive, or chronic illness-related constraints.\n\nAsync AI interviews address many of these barriers structurally — which is one reason platforms like Koji have become valuable tools for inclusive research programs.\n\n---\n\n## Types of Disabilities to Include in Research\n\nAccessibility research should span the full range of disability types. Each group uses technology differently and has distinct research needs:\n\n**Visual impairments:** Includes blindness, low vision, and color vision deficiency. These users rely on screen readers (NVDA, JAWS, VoiceOver), magnification, high-contrast modes, and keyboard navigation. Research should assess compatibility with these technologies.\n\n**Motor and physical impairments:** Includes limited mobility, tremors, repetitive strain injuries, and paralysis. These users may rely on keyboard-only navigation, switch access, eye tracking, or voice control. Research reveals friction in interactions that require precise pointing or quick responses.\n\n**Cognitive and learning disabilities:** Includes dyslexia, ADHD, autism spectrum conditions, traumatic brain injury, and memory impairments. Research reveals issues with complex navigation, dense text, time-pressured interactions, and inconsistent UI patterns.\n\n**Hearing impairments:** Includes deafness and hard-of-hearing users who rely on captions, transcripts, and visual cues. Research reveals issues with audio-only information delivery.\n\n**Speech impairments:** Includes conditions affecting verbal communication. These users may rely on alternative communication devices or text-based interfaces. Research reveals barriers in voice-required interactions.\n\n**Chronic illness and mental health conditions:** Users with chronic fatigue, anxiety, depression, or episodic conditions have variable capacity and may need flexible, low-pressure research participation options.\n\n---\n\n## Adapting Your Research Methods for Accessibility\n\n### Make Your Recruitment Accessible\n\nStart before the first question is asked. Your screener, invitation email, and sign-up flow must themselves be accessible:\n\n- Screen your screener with a screen reader before sending it\n- Offer multiple ways to respond (email, phone, text, web form)\n- Explicitly state in your invitation that participants with disabilities are welcome and accommodations are available\n- Partner with disability advocacy organizations, assistive technology user communities, and accessible tech forums to reach participants outside your standard panel\n- Offer flexible scheduling and generous time windows — not everyone can do a rigid 45-minute slot on a weekday morning\n\n### Offer Flexible Participation Formats\n\nNot every research method is accessible to every participant. Build flexibility into your study design:\n\n**Async interviews:** Participants complete at their own pace, in their own environment, with their own assistive technology, at a time that suits them. This is the single highest-impact change you can make for inclusive research. Platforms like Koji deliver async AI interviews that work via text or voice — participants choose the mode that suits their needs and abilities.\n\n**Text-based options:** For participants who find voice interactions difficult, text-based interviews remove the barrier of verbal communication. Koji's text interview mode lets participants type responses at their own pace, with no time pressure.\n\n**Voice options without real-time pressure:** Async voice interviews (where participants record responses rather than speaking live) are more comfortable for many participants with speech anxiety, cognitive disabilities, or chronic fatigue than live moderated sessions.\n\n**Extended time:** Always offer participants the ability to pause, resume, and take breaks during async studies. Live sessions should have built-in flexibility for breaks.\n\n### Adapt Your Interview Questions\n\nFor participants with cognitive or learning disabilities, apply plain language principles to your interview guide:\n\n- Use short sentences (aim for under 20 words per sentence)\n- Ask one question at a time\n- Avoid metaphors, idioms, and jargon\n- Define technical terms when they're unavoidable\n- Use concrete, specific language rather than abstract concepts\n\nKoji's structured questions feature is especially useful here. You can break complex topics into discrete, digestible question types — a scale rating (1–5) followed by an open-ended \"tell me more\" — rather than presenting participants with a complex open question that requires holding multiple concepts in mind simultaneously.\n\n### Make Your Research Materials Accessible\n\nBefore running any session — live or async:\n\n- Ensure your interview platform is keyboard-navigable\n- Test with screen readers (VoiceOver on Mac/iOS, NVDA on Windows, TalkBack on Android)\n- Avoid requiring fine motor precision for interactions\n- Provide written transcripts of any audio or video content\n- Use sufficient color contrast in any visual materials\n- Don't rely on color alone to convey information\n\n---\n\n## Recruiting Participants with Disabilities\n\nReaching participants with disabilities requires going beyond your standard panel:\n\n**Disability advocacy organizations:** Many organizations in the blindness, deafness, autism, and physical disability communities are open to research partnerships. Reach out directly, explain your research goals, and offer meaningful incentives.\n\n**Assistive technology communities:** Forums and communities organized around specific tools (screen reader users, AAC device users, wheelchair tech groups) are excellent recruitment channels. These communities often value research that might improve the products they rely on.\n\n**Specialized recruitment firms:** Several research recruitment firms specialize in recruiting participants with disabilities. For studies where specific disability types are essential, these firms save significant time.\n\n**Your own user base:** If your product already has users with disabilities, recruit from your existing user base. Send targeted invitations to users who have disclosed accessibility needs, used your accessibility features, or opted into accessibility-related communications.\n\n**Disability employment organizations:** Organizations that support people with disabilities in the workforce often have established networks and community trust, making them effective recruitment partners.\n\n---\n\n## Conducting the Research: Key Principles\n\n**Ask, don't assume.** Never assume what accommodations a participant needs. Ask in advance: \"Are there any accommodations that would help you participate comfortably?\" Then follow through.\n\n**Respect participant expertise.** People with disabilities are experts in their own experience. Approach interviews as a learner, not a problem-solver. Your job is to understand, not to fix.\n\n**Use people-first or identity-first language according to participant preference.** Some communities prefer \"person with a disability\" (people-first); others prefer \"disabled person\" (identity-first, common in Deaf and autistic communities). When in doubt, follow the participant's lead.\n\n**Build in more time.** Accessible research sessions often take longer. Build this into your project timeline and never rush a participant.\n\n**Compensate fairly.** Participation in research takes time and effort that may be disproportionately taxing for participants with certain disabilities. Compensation should reflect this.\n\n**Protect privacy.** Disability status is sensitive personal information. Treat accessibility data with the same care as any other sensitive data. Obtain explicit consent before recording, storing, or sharing information about participants' disabilities.\n\n---\n\n## Analyzing Accessibility Research Data\n\nAccessibility research data is qualitative by nature. Analysis approaches include:\n\n**Task analysis:** For usability-focused accessibility research, record which tasks participants completed, where they struggled, and what workarounds they developed. Map findings to specific UI elements or interaction patterns.\n\n**Severity rating:** Not all accessibility barriers are equal. Rate findings by frequency (how many participants encountered it), impact (how severely it affected task completion), and persistence (whether it blocked participants entirely or just slowed them).\n\n**Thematic analysis:** For attitudinal and experience-focused research, use thematic analysis to identify patterns across participant narratives. Koji automatically surfaces themes from interview data, making cross-participant pattern recognition fast even with smaller sample sizes.\n\n**Comparison with non-disabled participants:** Where possible, compare findings from participants with and without disabilities. Differences reveal accessibility-specific barriers; similarities reveal general usability problems.\n\n---\n\n## How Koji Supports Accessible Research\n\nKoji's platform has structural features that make it well-suited for accessibility research:\n\n- **Async format** removes real-time pressure, letting participants engage when their capacity is highest\n- **Text and voice modes** give participants agency over how they communicate\n- **No moderator presence** removes the social pressure that can be taxing for participants with anxiety or communication-related disabilities\n- **Flexible session duration** means participants can pause and return without losing progress\n- **Automatic transcription** creates a text record of voice interviews for participants who want to review or correct their responses\n- **AI follow-up questions** probe naturally, reducing the cognitive load of figuring out what to say next\n\nFor teams serious about inclusive research, Koji makes it practical to run accessible studies as a matter of course — not just for dedicated accessibility projects.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide)\n- [How to Set Up AI Voice Interviews: A Researcher's Complete Guide](/docs/setting-up-voice-interviews)\n- [Text Interview Experience](/docs/text-interview-experience)\n- [Remote Interview Best Practices for Qualitative Research](/docs/remote-interview-best-practices)\n- [Screening Research Participants Effectively](/docs/screening-participants-effectively)\n- [Research Incentive Strategies: What to Pay and How](/docs/incentive-strategies)\n\n## Further reading on the blog\n\n- [User Research: The Complete Guide to Understanding Your Users (2025)](/blog/user-research-the-complete-guide-to-understanding-your-users-2025) — Learn what user research is, why it matters, and how to conduct it effectively. Discover how AI tools like Koji are transforming the researc\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [B2C User Research: How to Understand Consumer Behavior at Scale (2026)](/blog/b2c-user-research-guide-2026) — B2C user research is systematically underinvested at most consumer companies. While B2B teams run structured customer discovery as a matter \n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Accessibility Research: How to Include Users with Disabilities in Your Studies","metaDescription":"A practical guide to inclusive user research — how to recruit participants with disabilities, adapt your methods, and use async AI interviews to remove participation barriers and build products that work for everyone.","keywords":["accessibility research","inclusive user research","research with disabled users","accessible UX research","disability research","inclusive research design"],"aiSummary":"Accessibility research involves intentionally including users with disabilities in research studies. This guide covers why most research programs exclude users with disabilities (and how to fix it), how to adapt recruitment and methods for different disability types, key principles for inclusive facilitation, and how async AI interview platforms like Koji remove structural participation barriers.","aiPrerequisites":["Basic familiarity with user research methods"],"aiLearningOutcomes":["Understand why standard research programs often exclude users with disabilities","Know which disability types to include and how each affects research participation","Adapt recruitment, methods, and facilitation for accessibility","Use async AI interviews to lower participation barriers","Analyze accessibility research findings effectively"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 minutes"},{"type":"documentation","id":"80a9d77c-44a3-4c1c-b159-aa44f0995055","slug":"formative-vs-summative-research","title":"Formative vs. Summative Research: When to Use Each Method (And Why It Matters)","url":"https://www.koji.so/docs/formative-vs-summative-research","summary":"Formative research shapes a product during design (5-10 participants, qualitative, weekly cadence) and outputs a list of issues. Summative research evaluates a shipped or near-shipped product (20-40+ participants, quantitative, rare cadence) and outputs a benchmark score. Confusing the two is the most common research-budget mistake. The \"5 users\" rule applies only to formative qualitative work — quantitative measurement requires 20-40+. SUS benchmark is mean 68 (SD 12.5) across 500 studies and 5,000 users. Koji runs both modes in the same platform: AI-moderated conversational interviews for formative discovery, structured questions and quality scoring for summative measurement.","content":"The single most expensive mistake in user research is running the wrong type of study at the wrong stage of a project. Teams routinely commission a 40-person summative benchmark to evaluate an unfinished prototype, or a 5-person formative usability test to \"prove\" that a shipped product is performing well. Both fail — not because the methods are bad, but because they were designed for a different question.\n\nFormative research and summative research are the two foundational modes of evaluating any design, product, or service. Understanding the distinction — and knowing which one your current question requires — is the difference between research that drives decisions and research that decorates them.\n\n## The Core Distinction\n\n**Formative research** is research that *forms* the product. It is conducted *during* design and development to identify what is working, what is broken, and what to change next. The output is a list of specific issues and concrete recommendations.\n\n**Summative research** is research that *sums up* the product. It is conducted *after* a design is largely complete to measure overall performance, often against a benchmark or competitor. The output is a number, a comparison, or a verdict.\n\nThe Nielsen Norman Group puts it cleanly: \"Formative evaluations are used in an iterative process to make improvements before production. Summative evaluations are used to evaluate a shipped product in comparison to a benchmark.\" Get this framing right and most other decisions — sample size, methodology, metrics — follow naturally.\n\n## Side-by-Side Comparison\n\n| Dimension | Formative | Summative |\n|-----------|-----------|-----------|\n| **When in lifecycle** | Early to mid (during design and iteration) | Late (post-launch or pre-launch benchmark) |\n| **Primary question** | *What's wrong and what should we fix?* | *How well does it perform?* |\n| **Approach** | Mostly qualitative — observation, think-aloud, interviews | Mostly quantitative — task success rates, SUS, time-on-task |\n| **Typical sample size** | 5–10 participants per iteration | 20–40+ participants for statistical confidence |\n| **Output** | A prioritized list of issues + design recommendations | A score, a comparison, a pass/fail verdict |\n| **Frequency** | Often — every sprint or design iteration | Rarely — pre-launch, post-launch, annual benchmark |\n| **Decision it informs** | What to change in the next iteration | Whether to ship, whether you improved, how you compare |\n| **Failure mode if used wrong** | Mistakes a noisy benchmark for a real signal | Spends 40-participant budget on issues 5 users would surface |\n\n## When to Run Formative Research\n\nFormative research is your default mode during active design and development. Use it when:\n\n- A prototype, mock-up, or early build needs to be vetted before you invest more engineering time\n- You are between design iterations and need to know what to change\n- You're running a discovery study where the goal is to surface unknown problems, not measure known ones\n- A specific feature, flow, or copy block is suspected of underperforming and you need to understand *why*\n- You're in continuous discovery — talking to users weekly to keep design decisions grounded\n\nThe canonical formative method is qualitative usability testing with 5 participants. Jakob Nielsen and Tom Landauer's 1993 mathematical model showed that 5 qualitative participants typically uncover around 85% of usability issues in an interface — assuming you run multiple rounds and fix what you find between them. Critically, the value of testing with 5 users *only* holds for qualitative formative work. Quantitative summative measurements need substantially more.\n\nOther common formative methods:\n- **Think-aloud protocols** — participants narrate their thoughts as they complete tasks\n- **Cognitive walkthroughs** — experts simulate user decision-making at each step\n- **Heuristic evaluation** — design audit against established usability principles\n- **Diary studies** — longitudinal observation during real use\n- **Concept testing** — early feedback on an idea before prototyping\n\nAll share the same underlying goal: figure out what is wrong while changing it is still cheap.\n\n## When to Run Summative Research\n\nSummative research is your evaluation tool. It is expensive, slower, and statistically rigorous. Use it when:\n\n- You need a defensible number to share with stakeholders or executives\n- You're comparing a new version against an old one (\"did the redesign actually help?\")\n- You're benchmarking against a competitor or industry standard\n- You're measuring whether a product meets a usability threshold before launch\n- You're running a regulatory or audit-grade evaluation\n\nThe canonical summative instrument is the **System Usability Scale (SUS)** — a 10-item questionnaire developed by John Brooke in 1986. SUS has been validated across more than 500 published studies and over 5,000 participants. The benchmark from that body of work: an average SUS score of 68 (SD 12.5). Scores above 68 are above average; below 68, below average.\n\nSUS requires at least 20–30 participants to produce a statistically reliable score. The Nielsen Norman Group's recent guidance for quantitative user testing is around 40 users, depending on the effect size you're trying to detect. This is the math that makes summative research expensive — and the math that makes it appropriate only when a precise number actually matters.\n\nOther common summative methods:\n- **Task success rate measurement** at scale\n- **Time-on-task benchmarking** versus prior version or competitor\n- **A/B tests** comparing two designs in production\n- **NPS, CSAT, CES** for overall product perception (with proper sample sizes)\n- **Large-scale surveys** measuring satisfaction or attitude shifts\n\n## The \"Five Users\" Confusion\n\nNo principle in user research is more misquoted than \"you only need 5 users.\" It is true *for qualitative formative research only*. If your goal is to find usability issues to fix, 5 users per iteration is well-supported by the data. If your goal is to *measure* anything quantitatively — task success rate, time, SUS score, NPS — 5 users will produce numbers with such wide confidence intervals that the result is statistically meaningless.\n\nA 100% task success rate from 5 users has a 95% confidence interval that stretches from roughly 48% to 100%. That is not a benchmark. That is a guess with extra steps.\n\nThe right mental model: **formative research finds problems, summative research measures them.** Finding needs few people. Measuring needs many.\n\n## How to Sequence Formative and Summative Together\n\nIn a healthy research program, formative and summative work in cycles:\n\n1. **Discovery (formative).** Qualitative interviews, contextual inquiry, opportunity mapping. 5–15 participants.\n2. **Ideation + prototyping.** Designers and PMs translate insight into options.\n3. **Iterative testing (formative).** Round 1 with 5 users → fix → Round 2 with 5 users → fix → Round 3. Repeat until the prototype stabilizes.\n4. **Pre-launch benchmark (summative).** SUS, task success, time-on-task with 30–40 participants. Establishes a baseline you can compare future versions against.\n5. **Post-launch monitoring (summative).** Periodic re-runs of the same instruments to track drift over time.\n6. **Back to formative.** When the benchmark dips or the team plans a major change, return to discovery.\n\nThe failure mode in immature research orgs is skipping straight to step 4 with no formative work — producing a precise score on a design full of issues that 5 users would have surfaced in week one.\n\n## How Koji Supports Both Modes\n\nMost research tools force a choice: a survey platform optimizes for quantitative summative work; a usability testing platform optimizes for qualitative formative work. Running both modes means stitching together separate tools, recruitment funnels, and analysis workflows.\n\nKoji is designed to run both modes in the same platform. For **formative discovery**, Koji's AI moderator runs adaptive, conversational interviews with 5–15 participants, probing for specific issues, surfacing unexpected friction, and producing a thematic analysis automatically. The structured questions can be edited or removed; the AI follows the brief. Methodology presets — *exploratory*, *mom_test*, *jtbd*, *discovery* — pre-configure the interview around classic formative frameworks.\n\nFor **summative measurement**, the same Koji study can scale to 100+ respondents, with [structured questions](/docs/structured-questions-guide) of six types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) producing the kind of quantitative data SUS-style benchmarks require. Quality scoring (1–5 scale) flags low-effort responses automatically. The same study can produce both a qualitative theme summary *and* a quantitative benchmark, sidestepping the usual tool-fragmentation tax.\n\nThe operational benefit: teams using AI-assisted research report meaningfully faster time-to-insight, because the boundary between \"formative interview round\" and \"summative benchmark wave\" collapses into the same workflow. You stop choosing between depth and scale.\n\n## A Quick Decision Heuristic\n\nBefore commissioning any study, write down the *exact* decision the research will inform:\n\n- *\"Should we ship this redesign or wait?\"* → summative\n- *\"What's making people drop off in onboarding?\"* → formative\n- *\"Did our v2 actually improve over v1?\"* → summative\n- *\"Why does usage spike in week three then crash?\"* → formative\n- *\"How do we compare to our biggest competitor on usability?\"* → summative\n- *\"What jobs are users trying to get done?\"* → formative\n\nIf the decision needs a number, you're looking at summative. If it needs an explanation, you're looking at formative. Match the method to the question and the budget follows.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six structured question types that make summative measurement possible inside an AI-moderated study\n- [Generative vs Evaluative Research](/docs/generative-vs-evaluative-research) — a related but distinct framing for the same lifecycle question\n- [Research Synthesis Guide](/docs/research-synthesis-guide) — how to convert formative findings into shareable insight\n- [How Many User Interviews](/docs/how-many-user-interviews) — practical guidance on sample size at each stage\n- [Continuous Discovery User Research](/docs/continuous-discovery-user-research) — embedding formative research into a weekly cadence\n- [UX Research Process](/docs/ux-research-process) — the end-to-end research lifecycle\n\n## Sources\n\n- Nielsen Norman Group, *Formative vs. Summative Evaluations*\n- Nielsen Norman Group, *Why 5 Participants Are Okay in a Qualitative Study, but Not in a Quantitative One*\n- Brooke, J. (1986). *SUS — A Quick and Dirty Usability Scale*\n- MeasuringU, *Measuring Usability with the System Usability Scale (SUS)* — 500-study benchmark of mean 68 (SD 12.5)\n- Nielsen, J. & Landauer, T. (1993). *A mathematical model of the finding of usability problems*","category":"Research Methods","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Formative vs. Summative Research: When to Use Each (2026 Guide)","metaDescription":"Formative research shapes a product during design with 5-10 participants. Summative research evaluates after launch with 30-40+. Learn when to use each, and why mixing them up wastes research budgets.","keywords":["formative research","summative research","formative vs summative","formative evaluation","summative evaluation","usability testing","SUS score","sample size usability","qualitative research","quantitative research"],"aiSummary":"Formative research shapes a product during design (5-10 participants, qualitative, weekly cadence) and outputs a list of issues. Summative research evaluates a shipped or near-shipped product (20-40+ participants, quantitative, rare cadence) and outputs a benchmark score. Confusing the two is the most common research-budget mistake. The \"5 users\" rule applies only to formative qualitative work — quantitative measurement requires 20-40+. SUS benchmark is mean 68 (SD 12.5) across 500 studies and 5,000 users. Koji runs both modes in the same platform: AI-moderated conversational interviews for formative discovery, structured questions and quality scoring for summative measurement.","aiPrerequisites":["ux-research-process"],"aiLearningOutcomes":["Distinguish formative from summative research by question, sample size, method, and output","Pick the right sample size for the question you're asking (5-10 for formative, 20-40+ for summative)","Sequence formative and summative research into a coherent research program","Recognize when \"5 users\" is the right answer and when it's wildly insufficient","Use SUS and similar instruments correctly for summative benchmarking"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"5bec4544-f763-4956-9113-2038bc1c5f3e","slug":"content-testing-guide","title":"Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)","url":"https://www.koji.so/docs/content-testing-guide","summary":"Content testing measures whether users understand and act correctly on interface copy, and differs from A/B testing by explaining why a variant confused people rather than only which one converted. Six methods cover most cases: cloze tests (delete every nth word, roughly 60% restoration indicates comprehensible prose), highlighter tests (mark confusing and reassuring phrases), comprehension checks (restatement in the users own words, never a yes/no did-you-understand), term-choice or nomenclature tests (single_choice plus ranking with probing on the mental model evoked), expectation tests (what will this button do, with a confidence scale), and label first-click for findability. Readability scores are a drafting aid, not evidence, because they cannot detect unfamiliar vocabulary or correct-sounding copy that describes the wrong behaviour. Never test on colleagues or power users. Sample 20-30 for clearly different term candidates and 50+ for close calls or segment comparisons. AI-moderated conversational studies remove the scheduling constraint: text mode for reading tasks, all six Koji structured question types in one pass, and automatic follow-up probing to capture the reasoning behind each choice.","content":"# Content Testing: How to Test Microcopy, Labels, and UX Writing With Real Users (2026)\n\n**The bottom line:** Content testing measures whether people understand and act correctly on your words — button labels, error messages, form hints, empty states, pricing copy, onboarding instructions. It is not proofreading, and it is not A/B testing: A/B tells you which variant converted, content testing tells you *why* one variant confused people, which is the part you can generalise to the next hundred strings. Six methods cover almost every case, and with an AI-moderated platform like Koji you can run them as a single conversational study across a hundred participants instead of scheduling twelve sessions.\n\n---\n\n## Why words deserve their own test\n\nCopy is the interface. A user does not experience your information architecture; they experience the eleven words on the screen that tell them what will happen if they click. Yet copy is usually the least-tested layer of a product: it changes late, it changes often, and it is written by people who have spent months internalising the exact vocabulary the user has never heard.\n\nThat last point is the core problem. Content testing exists because **the author cannot un-know the meaning.** You cannot self-assess comprehension of your own terminology, and neither can your colleagues — insiders share the company-specific assumptions that make unclear copy feel obvious. Testing with representative users is not a nice-to-have here; it is the only way to obtain the signal at all.\n\nThree symptoms suggest you have a content problem rather than a design problem:\n\n- Support tickets quote your own UI text back at you as the source of confusion.\n- Users complete a task but describe what they did incorrectly afterwards.\n- Two internal teams argue about a label and both cite \"clarity\" as the reason.\n\n---\n\n## The six methods\n\n### 1. Cloze test — does the copy hold together?\n\nThe classic comprehension measure. Take a passage, delete every *n*th word, and ask participants to fill in the blanks from context. For long-form prose, *n* = 6 is conventional; for microcopy you need a much smaller interval, and you want roughly 25 to 50 blanks in total to get a stable score.\n\n**Scoring:** count exact and near-synonym matches. A commonly used rule of thumb is that copy is comprehensible when participants restore around 60% or more of the missing words. Below that, the text is not predictable enough — usually because sentence structure is convoluted or terminology is unfamiliar.\n\n**Use it for:** legal-adjacent explanations, policy text, onboarding paragraphs, security or billing copy where misunderstanding is costly.\n\n**Do not use it for:** three-word button labels. There is nothing to restore.\n\n### 2. Highlighter test — where does confidence break?\n\nShow the copy and ask participants to mark anything confusing, and separately anything reassuring. Digitally this is a two-pass exercise: first the friction, then the trust.\n\nThe output is a heat map of specific strings rather than an overall score, which makes it the most directly actionable method for a writer. In a conversational study you replicate it by asking participants to quote back the exact phrase that felt unclear, then probing why. Koji's follow-up probing does the \"why\" automatically, which is where the rewrite instruction actually comes from.\n\n**Use it for:** pricing pages, plan comparison tables, consent and permission dialogs, anything where hesitation kills conversion.\n\n### 3. Comprehension check — did the meaning survive?\n\nShow the copy, remove it, then ask what it meant and what would happen next. The rule is to ask for a restatement in the participant's own words, never a yes/no \"did you understand?\" — everyone says yes.\n\nStrong variants:\n- \"In your own words, what does this setting do?\"\n- \"If you click this, what happens to your existing data?\"\n- \"Who is this plan for?\"\n\n**Use it for:** destructive actions, permissions, data-sharing explanations, plan selection.\n\n### 4. Term-choice test — which word wins?\n\nPresent two to five candidate terms for the same concept and ask which one participants would expect to lead where. This is nomenclature research and it is the highest-leverage content test in most products, because the term propagates across navigation, docs, support macros, and sales collateral.\n\nRun it as a `single_choice` question for the winner plus a `ranking` question when you need the full preference order — then probe the choice: \"what did you expect *Workspace* to contain that *Project* would not?\" The reasoning matters more than the vote count, because it tells you what mental model the term evoked.\n\n### 5. Expectation test — what do they think the button does?\n\nShow the control in isolation, before the click. Ask what the participant expects to happen, how confident they are, and what would make them hesitate. A `scale` question captures confidence; the probed open-ended answer captures the reason.\n\nThis catches the most expensive class of copy failure: labels that are perfectly clear and describe the wrong thing.\n\n### 6. Label first-click — can they find it at all?\n\nPresent the navigation or menu labels and ask where they would go to accomplish a specific task. Findability failures are frequently vocabulary failures in disguise; if 40% pick the wrong label, the problem is rarely the layout.\n\nPair it with the term-choice test: first learn what people call the thing, then verify that the label you chose is where they look.\n\n---\n\n## Readability scores are not content testing\n\nFlesch-Kincaid, grade-level scores, and their cousins measure sentence and word length. They do not know whether *reconciliation*, *entitlement*, or *seat* means anything to your audience, and they cannot detect a perfectly readable sentence that describes the wrong behaviour. Use readability tools as a drafting aid and a floor, never as evidence. Comprehension is a property of the reader, not of the text — which is why it has to be measured with readers.\n\n---\n\n## Running content tests at conversational scale\n\nThe traditional constraint on content testing is throughput. Each method above is cheap to run once and painful to run twenty times: you schedule sessions, read copy aloud, take notes, and hand-code the answers. In practice teams test the redesign and skip the two hundred strings that ship every quarter.\n\nAn AI-moderated study removes the scheduling layer. A few specifics that matter for content work:\n\n- **Use text mode for reading tasks.** Copy has to be *seen*. Text interviews let participants read the string and respond to it; voice is better suited to the expectation and reasoning parts of the study. Koji supports both from a single link, so a mixed design works.\n- **Mix structured and open in one pass.** Term choice as `single_choice`, preference order as `ranking`, confidence as `scale`, comprehension restatement as `open_ended` with probing, and a quick `yes_no` on whether the participant would take the action. The six [structured question types](/docs/structured-questions-guide) mean one study returns both the vote counts and the reasoning — no second round.\n- **Let the AI ask the \"why\".** The value in content testing sits entirely in the follow-up: not \"which label did you pick\" but \"what did you expect that label to contain?\" Koji probes each answer up to three times automatically, which is exactly the interviewer behaviour a form cannot replicate and a busy team rarely sustains by hand.\n- **Recruit outsiders.** Never test copy on colleagues, and be careful with power users, who have already learned your vocabulary. Screen for the segment whose comprehension you actually care about — often new or prospective users.\n- **Watch for order effects.** Showing variant A before variant B primes the comparison. Randomise where you can and keep the sequence identical across participants where you cannot, so the bias is at least constant. See [question order bias](/docs/question-order-bias-guide).\n\n### How many participants?\n\nMore than qualitative usability work, less than a survey. Comprehension is a proportion, and proportions need a bit of sample: 20 to 30 participants gives a usable read on a term-choice test between two clearly different candidates; 50 or more if the options are close or you need to compare segments. Cloze tests are more forgiving because each participant contributes dozens of data points. For the reasoning layer, the usual qualitative saturation logic applies — see [how many user interviews you need](/docs/how-many-user-interviews).\n\n---\n\n## A ready-made study structure\n\nA single 8-to-12 minute conversation can cover an entire feature's copy:\n\n1. **Warm-up** (`open_ended`) — What do you use a tool like this for today?\n2. **Expectation** (`open_ended` + `scale`) — Here is the button. What happens if you click it? How confident are you, 1 to 7?\n3. **Comprehension** (`open_ended`, probed) — Here is the explanation text. In your own words, what does it mean for your data?\n4. **Highlight** (`open_ended`, probed) — Which exact phrase, if any, made you pause? Why that one?\n5. **Term choice** (`single_choice` + probe) — Which of these would you click to find your saved work?\n6. **Preference order** (`ranking`) — Rank these four headings by how clearly they describe the page.\n7. **Action** (`yes_no`, probed) — Based only on this copy, would you turn the setting on?\n\nKoji aggregates the structured items into distributions automatically and themes the open-ended answers with supporting quotes, so the writer receives \"62% expected *Archive* to delete the file, and here are the eleven quotes explaining why\" rather than a folder of recordings.\n\n---\n\n## What to do with the results\n\n- **Rewrite against the failure, not the score.** A 45% cloze score tells you the passage failed; the highlighter quotes tell you which clause did it.\n- **Fix the term everywhere at once.** A nomenclature change that lands in the UI but not in docs, emails, and support macros creates a new comprehension problem.\n- **Keep a decision log.** Record the tested alternatives and the winning rationale. It ends the recurring internal argument and it is the fastest onboarding artefact a new writer can get.\n- **Re-test after the rewrite.** Content changes are cheap enough that a second wave is realistic, and comprehension gains are the easiest research win to demonstrate to stakeholders.\n\n---\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types used throughout this guide\n- [Cognitive Interviews](/docs/cognitive-interview-guide) — the sibling method for testing whether your *questions* are understood\n- [Open-Ended vs Closed-Ended Questions](/docs/open-ended-vs-closed-ended-questions) — when to count and when to probe\n- [Choice and Ranking Questions in AI Interviews](/docs/choice-ranking-questions-guide) — running term-choice and preference-order tests\n- [Question Order Bias](/docs/question-order-bias-guide) — avoiding priming when comparing copy variants\n- [AI Usability Testing](/docs/ai-usability-testing-guide) — where content testing fits inside a broader usability study\n- [How Many User Interviews Do You Need?](/docs/how-many-user-interviews) — sample size for the qualitative layer\n","category":"Research Methods","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Content Testing: Test Microcopy & UX Writing (2026 Guide)","metaDescription":"Six ways to test whether your UX copy works — cloze tests, highlighter tests, comprehension checks, term-choice tests, expectation tests, and label first-click — plus how to run them conversationally at scale.","keywords":["content testing","ux writing research","microcopy testing","copy testing method","cloze test ux","comprehension testing","terminology testing","label testing","ux content research","test button labels"],"aiSummary":"Content testing measures whether users understand and act correctly on interface copy, and differs from A/B testing by explaining why a variant confused people rather than only which one converted. Six methods cover most cases: cloze tests (delete every nth word, roughly 60% restoration indicates comprehensible prose), highlighter tests (mark confusing and reassuring phrases), comprehension checks (restatement in the users own words, never a yes/no did-you-understand), term-choice or nomenclature tests (single_choice plus ranking with probing on the mental model evoked), expectation tests (what will this button do, with a confidence scale), and label first-click for findability. Readability scores are a drafting aid, not evidence, because they cannot detect unfamiliar vocabulary or correct-sounding copy that describes the wrong behaviour. Never test on colleagues or power users. Sample 20-30 for clearly different term candidates and 50+ for close calls or segment comparisons. AI-moderated conversational studies remove the scheduling constraint: text mode for reading tasks, all six Koji structured question types in one pass, and automatic follow-up probing to capture the reasoning behind each choice.","aiPrerequisites":["Draft copy or candidate labels ready to test","Access to participants outside your company"],"aiLearningOutcomes":["Choose between cloze, highlighter, comprehension, term-choice, expectation, and first-click tests","Score a cloze test and interpret the result","Write comprehension questions that avoid false yes answers","Run a nomenclature test that produces reasoning, not just votes","Size a content test sample","Structure a single conversational study covering a whole features copy"],"aiDifficulty":"beginner","aiEstimatedTime":"10 min read"},{"type":"documentation","id":"3ada4f04-2cec-4395-a059-2dcf6ea1b790","slug":"anonymizing-customer-interview-data","title":"Anonymizing Customer Interview Data: A Practical Guide for Privacy-Safe Research","url":"https://www.koji.so/docs/anonymizing-customer-interview-data","summary":"A 5-technique operational playbook for anonymizing customer interview data: minimize intake collection, use participant codes, strip PII from quotes, control transcript access, and set retention windows. Distinguishes pseudonymization from true anonymization and clarifies what Koji handles vs what stays on the research team.","content":"**TL;DR:** Customer interview data is full of PII — names, emails, employer, role, and verbatim stories that can identify a participant. Anonymizing this data before it leaves the research team is now table stakes for privacy-conscious B2B teams. The 5 practical techniques are (1) collect only what you need at intake, (2) use participant codes instead of real names in synthesis, (3) review and strip PII from quotes before sharing, (4) keep transcripts behind access controls, and (5) set explicit data retention windows. Done right, anonymization improves research quality — participants speak more candidly when they know they''re not being personally identified.\n\n## Why anonymization matters now\n\nCustomer research has always lived in tension with privacy. The most useful interview data is verbatim, specific, and emotional — exactly the data most likely to identify a participant. As privacy regulation (GDPR in the EU, CCPA in California, PIPL in China, and a growing set of state-level US laws) has tightened, the consequences of mishandling interview data have shifted from \"embarrassing\" to \"legally and financially material.\"\n\nThere are three forces pushing this to the top of the research-ops agenda:\n\n1. **Regulation.** Under GDPR Article 4, any data that can identify a \"natural person\" is personal data — and an interview transcript almost always qualifies, even without the participant''s name attached.\n2. **B2B procurement.** Enterprise buyers now ask vendors how they handle research data in security reviews. \"We anonymize before sharing\" is the answer that closes deals.\n3. **Participant trust.** The 2024 Pew Research Center survey on AI and privacy showed 81% of Americans believe companies collect more data than they need. Participants speak more freely when they''re assured of anonymity — which means better research.\n\nThis guide focuses on the *operational* side: what to do at each step of the research process to keep PII contained. For the legal/compliance lens, pair this with the [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research) doc.\n\n## The 5-technique playbook\n\n### Technique 1: Minimize collection at intake\n\nThe cheapest PII to protect is the PII you never collect. Before designing your intake form (the screener questions at the start of an interview), ask: **do I need this field to do my research?**\n\n- **Email** — needed if you''re sending follow-up incentives. Otherwise, skip.\n- **Full name** — almost never needed. A first name or chosen pseudonym is enough for \"Hi {{name}}!\" personalization.\n- **Employer** — only needed if you''re segmenting by company. If segmenting by industry, ask \"What industry are you in?\" instead.\n- **Job title** — needed for B2B segmentation. Ask for it.\n- **Phone number** — almost never needed. Skip.\n\nKoji''s [intake form configuration](/docs/intake-forms-and-consent) lets you toggle every field. Default to **off** and turn on only what serves the research.\n\nA real-world example: a 2025 study on developer tooling collected 200 interviews using only \"first name + chosen pseudonym + role.\" That dataset has effectively zero PII risk while still supporting full segmentation.\n\n### Technique 2: Use participant codes in synthesis\n\nOnce interviews are in, switch from real names to participant codes for all downstream synthesis. A code is a short label like `P01`, `P02`, ... `P25` — or for more memorable codes, `developer_remote_03`, `designer_inhouse_07`.\n\nIn Koji, you can:\n\n- Pull the participant list with their internal IDs\n- Map each to a sequential code in your synthesis doc (Notion, Figma, Miro)\n- From that point forward, refer to interviews only by their code\n\nThis is good hygiene even for non-regulated research. It prevents stakeholders from over-indexing on \"what would Sarah think?\" (the well-known *single anecdote* bias) and forces the conversation onto themes.\n\n### Technique 3: Strip PII from quotes before sharing\n\nThe riskiest moment in research workflows is quote sharing — pasting a verbatim quote into a Slack channel, a PRD, or an investor deck. Three things commonly leak:\n\n- **Names mentioned in the answer** (\"...I asked Maria from sales to help me...\")\n- **Employer mentioned in the answer** (\"...at Acme Corp we have this exact problem...\")\n- **Unique role + geography combinations** (\"...I''m the only DevOps engineer at a 50-person fintech in Munich...\")\n\nBefore sharing any quote externally, do a quick scrub:\n\n| Before | After |\n|---|---|\n| \"...I asked Maria from sales...\" | \"...I asked a colleague from sales...\" |\n| \"...at Acme Corp we have...\" | \"...at our company we have...\" |\n| \"...I''m the only DevOps engineer at a 50-person fintech in Munich...\" | \"...I''m on a small DevOps team at an EU-based fintech...\" |\n\nFor high-volume quote-sharing workflows, use Koji''s AI report features to generate scrubbed quote summaries — the AI can rewrite quotes to preserve the insight while removing identifying details. Always do a human review before publishing.\n\n### Technique 4: Keep transcripts behind access controls\n\nFull transcripts are the highest-risk artifact in your research repository. They contain everything — the participant''s name, voice (in voice interviews), employer, and stories. Treat them like production data.\n\nConcrete practices:\n\n- **Limit transcript access to the research team.** Stakeholders see themes and quotes, not raw transcripts, unless they have a specific need.\n- **Use Koji''s access controls** to scope who in your workspace can read transcripts.\n- **Don''t paste full transcripts into shared channels.** Use the [share link feature](/docs/sharing-your-interview-link) to send view-only access instead.\n- **Avoid downloading transcripts to local drives.** If you must (for backup or offline analysis), encrypt the drive and delete after use.\n- **For voice interviews**, audio is even higher-risk than text. Treat recordings as the most sensitive artifact you have.\n\n### Technique 5: Set explicit data retention windows\n\nPII you''ve already deleted can''t be breached. Set and enforce retention windows for raw interview data:\n\n- **30 days** — for incentive-fulfillment use (after sending the gift card, the email can be deleted)\n- **6 months** — for active research projects (you may want to re-interview)\n- **12 months** — for historical reference (after this, archive themes + anonymized quotes only and delete raw data)\n- **Forever** — never, for raw transcripts. There''s no business value that justifies indefinite retention.\n\nDocument the retention policy in your research operations doc, and either run a quarterly manual cleanup or schedule automated deletion if your platform supports it.\n\nGDPR''s \"right to be forgotten\" (Article 17) means a participant can request deletion at any time. Having a documented retention policy and a clear deletion process is essential.\n\n## What \"anonymized\" actually means (and doesn''t)\n\nIt''s worth being precise. There are two related but distinct standards:\n\n- **Pseudonymized data** — direct identifiers (name, email) are replaced with a code, but the mapping still exists somewhere. Re-identification is possible if the mapping leaks. This is the most common state of \"anonymized\" research data, and it''s still personal data under GDPR.\n- **Truly anonymized data** — no mapping exists, and even combined with other data, the person can''t be re-identified. This is the gold standard but rarely achieved with verbatim qualitative data, because the content itself can identify (e.g., \"I''m the CTO at the only seed-stage AI dental practice in Lisbon\").\n\nBe honest about which you''re achieving. Most research operations land at *pseudonymized* — that''s fine, as long as you treat the code mapping like a secret.\n\n## How Koji helps (and where the responsibility is still yours)\n\nKoji is built to make privacy-safe research realistic:\n\n- **Configurable intake** — collect only what you need\n- **Per-study access controls** — limit who in your workspace can read transcripts\n- **BYOK option** — if you bring your own AI provider key, transcripts are processed via your LLM provider account, keeping your data inside your provider relationship\n- **Webhook control** — you decide which downstream tools receive interview data\n- **Consent collection** at intake — capture explicit research consent before the interview starts\n\nBut anonymization is ultimately a workflow discipline, not a feature toggle. The platform makes it easy; the team has to do it. Build the 5 techniques above into your study setup checklist.\n\n## Anonymization vs. survey-only alternatives\n\nSome teams try to dodge anonymization complexity by sticking with surveys (Typeform, SurveyMonkey, Google Forms) and never doing qualitative interviews. The problem: surveys collect PII too (just less of it), and they give you 10× less insight per response. The tradeoff isn''t \"privacy vs. research\" — it''s \"discipline vs. shortcut.\"\n\nA well-run Koji study with proper anonymization gives you the depth of qualitative interviews *and* a defensible privacy posture. The structured-question portion ([6 question types](/docs/structured-questions-guide)) gives you the quantification you''d get from a survey, in the same conversation, with the same anonymization treatment.\n\n## A starter checklist for your next study\n\n- [ ] Intake collects only essential fields\n- [ ] Consent language at intake is clear and timestamped\n- [ ] Each participant is assigned a code for synthesis\n- [ ] Transcript access is limited to the research team\n- [ ] Shared quotes are scrubbed of names/employers/unique identifiers\n- [ ] A retention window is documented and on the calendar to enforce\n- [ ] For voice interviews, audio handling matches text-transcript controls\n\nRun through this list before publishing each new study. Most violations happen because the team didn''t pause to check.\n\n## Frequently Asked Questions\n\n**Is \"anonymized\" the same as \"GDPR-compliant\"?**\nNo. Anonymization is one *technique* within GDPR compliance. Full GDPR compliance also requires lawful basis, consent records, data subject rights handling, breach notification, and DPA agreements with vendors. See the [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research) doc for the full picture.\n\n**Can AI-moderated interviews be more anonymous than human-moderated?**\nOften, yes. There''s no human moderator who could later recognize a participant. With proper intake-time anonymization, AI-moderated interviews can offer stronger anonymity than a Zoom call with a researcher.\n\n**Should I anonymize before AI synthesis runs, or after?**\nGenerally after. The AI synthesis benefits from full context to spot themes, and runs inside Koji''s controlled environment. The critical step is anonymizing *outputs* (themes, quotes, reports) before they leave that controlled environment.\n\n**What if a participant explicitly wants attribution?**\nSome participants (especially in B2B advocacy contexts) explicitly want their name attached to a quote. Get this consent in writing, scoped narrowly (\"you may use this quote with my name on your website\"), and document it. Default is still anonymity.\n\n**How long can I keep interview audio recordings?**\nFor voice interviews, treat audio as the highest-risk artifact. A 30–90 day window is common for active research; beyond that, transcribe and delete the audio. Document the policy.\n\n**Does anonymization reduce research quality?**\nNo — done well, it *improves* quality. Participants speak more candidly when they trust their anonymity is respected, and team discussions stay theme-focused rather than anecdote-focused.\n\n## Related Resources\n\n- [Structured Questions Guide: 6 Question Types Every Koji Study Needs](/docs/structured-questions-guide)\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research)\n- [Research Consent Form Templates](/docs/research-consent-form-templates)\n- [Research Ethics Guide](/docs/research-ethics-guide)\n- [Intake Forms and Consent](/docs/intake-forms-and-consent)\n- [Research Operations Guide](/docs/research-ops-guide)","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Anonymizing Customer Interview Data: Privacy-Safe Research (2026)","metaDescription":"Five practical techniques for handling PII in AI customer interviews — from intake to stakeholder-safe quotes — without sacrificing research signal.","keywords":["anonymize customer interview data","pii in research transcripts","customer interview privacy","anonymize research participants","research participant anonymity","redact pii interviews","research data privacy","interview data protection","b2b research privacy"],"aiSummary":"A 5-technique operational playbook for anonymizing customer interview data: minimize intake collection, use participant codes, strip PII from quotes, control transcript access, and set retention windows. Distinguishes pseudonymization from true anonymization and clarifies what Koji handles vs what stays on the research team.","aiPrerequisites":["Basic familiarity with running customer interviews","Understanding of what PII means at a high level"],"aiLearningOutcomes":["Configure an intake form that collects only essential PII","Switch from real names to participant codes in synthesis","Scrub quotes of identifying details before sharing","Apply transcript access controls in Koji","Set and enforce a documented data retention window"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"eaee5706-e75d-4e8a-9fcc-9fd073e9fadd","slug":"hipaa-compliant-ai-user-research","title":"HIPAA-Compliant AI User Research: A Practical Playbook for Healthcare and HealthTech","url":"https://www.koji.so/docs/hipaa-compliant-ai-user-research","summary":"HIPAA-compliant AI customer research is possible by scoping studies to avoid the 18 HIPAA identifiers entirely — name, DOB, MRN, email, voice biometrics, etc. The default Koji pattern uses anonymous mode (no email or demographics collected), structured screener questions, and an AI moderator briefed to steer away from PHI. Voice mode is avoided for regulated studies because voiceprints are themselves identifiers. When PHI is genuinely required, the Enterprise tier supports a Business Associate Agreement plus BYOK (Bring Your Own Key) so LLM calls route through the customer's own Anthropic, OpenAI, or Vertex contract. Per-study retention, workspace-scoped access, AES-256 at rest, and TLS 1.2+ in transit are standard. The compliance lever Koji uniquely provides is structured questions: six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that let teams capture sensitive content as categorical buckets instead of free-text PHI. Most product, churn, onboarding, and brand research questions can be answered without any PHI at all — making the de-identified path faster, cheaper, and avoidant of the BAA negotiation for the majority of healthcare research programs.","content":"# HIPAA-Compliant AI User Research: A Practical Playbook for Healthcare and HealthTech\n\n**Answer first:** You can run AI-moderated customer research in HIPAA-regulated contexts without ever touching Protected Health Information (PHI) — and that is almost always the right design. Most patient and provider research questions (about workflow friction, brand perception, app usability, billing experience, even symptom journeys) can be answered through interviews that are scoped to *avoid* the 18 HIPAA identifiers entirely. When PHI is genuinely required, the same study can be re-run on the [Enterprise tier](/enterprise) with a signed Business Associate Agreement, customer-controlled LLM keys (BYOK), and per-study retention controls. The BAA is **available on request and executed during onboarding** — see [/compliance/hipaa](/compliance/hipaa) for the full HIPAA-ready posture, scope, and the controls applied. Koji is designed so that the default research path produces *de-identified* qualitative data: anonymous-mode interviews, no demographic identifiers required, transcript-level redaction in reports, and a per-study retention setting. With tools like Koji, healthcare teams get the speed of AI customer research without inheriting the compliance overhead of a clinical system of record.\n\nIf you're a product, UX, or marketing team at a payer, provider, EHR vendor, digital therapeutics company, pharmacy, or any HealthTech startup, this guide walks through how to design AI customer research that holds up under your privacy office's review.\n\n## What HIPAA actually requires (in research terms)\n\nHIPAA's Privacy Rule applies to **Covered Entities** (health plans, healthcare providers that bill electronically, healthcare clearinghouses) and their **Business Associates** (vendors that handle PHI on a Covered Entity's behalf). It governs *Protected Health Information* — health information tied to one of 18 specific identifiers (name, address, dates more granular than year, phone, email, MRN, etc.).\n\nFor customer research, three rules dominate:\n\n1. **De-identification removes the obligation.** If your interview data does not contain any of the 18 HIPAA identifiers and there is no reasonable basis to re-identify a participant, the data is not PHI and HIPAA does not apply.\n2. **Authorization is needed to use PHI for research** outside treatment, payment, and operations — and most Covered Entities require IRB review for any study that touches PHI.\n3. **Business Associates need a BAA.** Any vendor that *receives* PHI must sign a Business Associate Agreement that flows down HIPAA obligations.\n\nThe practical implication for product teams: **scope your research so you never need to collect PHI in the first place.** Almost every product-discovery, churn, onboarding, pricing, and brand question can be answered without it.\n\n## The default Koji healthcare research pattern\n\nFor 90% of healthcare and HealthTech research, run the study with the following configuration. This produces interview data that is not PHI under HIPAA's de-identification standard and does not require a BAA.\n\n- **Anonymous mode on.** No email, name, phone, or address collected at intake. Participants are assigned a stable but opaque respondent ID. ([Anonymizing customer interview data](/docs/anonymizing-customer-interview-data) covers the controls in detail.)\n- **Screening avoids the 18 identifiers.** Use structured screener questions on role and behavior (\\\"how often do you log medication intake?\\\") rather than identifiable demographics (\\\"what is your date of birth?\\\"). The [research screener questions](/docs/research-screener-questions) doc has a full pattern library.\n- **The AI interviewer is briefed not to elicit PHI.** Add a short company-context instruction (see [company context guide](/docs/company-context-guide)) telling the AI moderator: *\\\"Do not ask for, and do not record, the participant's name, date of birth, address, medical record number, specific diagnoses, or specific treatment dates. If a participant volunteers this information, acknowledge it briefly and steer the conversation back to the research topic.\\\"*\n- **Per-study retention configured.** Set transcripts to auto-delete after 90 days (or whatever your privacy office approves). Reports and themes remain; raw transcripts are purged.\n- **Reports use anonymized quotes.** Koji's AI-generated reports surface themes and pull representative quotes. Configure the export to scrub any inadvertent identifiers before sharing outside your team.\n\nThis pattern lets a digital health PM run 30 patient interviews in a week, surface the friction themes, and ship the fix — without putting their company through a BAA negotiation for every study.\n\n## When you genuinely need PHI: the Enterprise path\n\nSome research questions cannot be answered without PHI. Examples: post-discharge care research that requires linking back to the original encounter; clinical trial recruitment screening; a payer study comparing benefit utilization across named members; provider research that requires NPI-level tracking.\n\nFor those studies:\n\n- **Move to the Enterprise tier** and request a Business Associate Agreement before any PHI flows. Standard self-serve plans (Insights, Interviews) do not include a BAA.\n- **Use BYOK (Bring Your Own Key)** so the LLM calls hit *your* Anthropic, OpenAI, or Google Vertex account on a contract you've already negotiated with HIPAA terms. See the [Bring Your Own Key](/docs/bring-your-own-key) doc for setup. BYOK means the LLM provider is *your* sub-processor, not Koji's, and the conversation content never enters the default shared inference pipeline.\n- **Restrict the participant audience** with [personalized interview links](/docs/personalized-interview-links) tied to an internal participant ID rather than a name or email.\n- **Tighten retention and access.** Restrict the study to specific workspace members, set retention to the shortest period your protocol allows, and disable export to third-party tools like Slack and Notion unless those vendors are also under BAA.\n\nThis is the same pattern enterprise healthcare buyers expect from any AI vendor in 2026 — and it's the reason Koji's architecture separates the application layer (Vercel + Supabase, both SOC 2 Type II) from the inference layer (where BYOK lets you bring your own contracts).\n\n## The 18 HIPAA identifiers your interviews must avoid\n\nIf your study is running under the de-identification path, the AI moderator and your screener should never elicit or store:\n\n1. Names\n2. Geographic subdivisions smaller than a state (street, city, ZIP — except first 3 digits of ZIP in some cases)\n3. Dates more granular than year (DOB, admission date, discharge date)\n4. Phone numbers\n5. Fax numbers\n6. Email addresses\n7. Social Security numbers\n8. Medical record numbers\n9. Health plan beneficiary numbers\n10. Account numbers\n11. Certificate / license numbers\n12. Vehicle identifiers\n13. Device identifiers and serial numbers\n14. URLs that identify the individual\n15. IP addresses\n16. Biometric identifiers (fingerprints, voiceprints)\n17. Full-face photos or comparable images\n18. Any other unique identifying number, characteristic, or code\n\nNumber 16 is worth a pause for voice-mode interviews. A voiceprint is a HIPAA identifier. For studies in regulated contexts, prefer text-mode interviews ([voice vs text interviews](/docs/voice-vs-text-interviews)), or use Enterprise + BAA + BYOK if voice is required.\n\n## Structured questions as a compliance lever\n\nKoji's six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — give you a way to capture sensitive content as *categorical buckets* instead of free-text PHI. Examples:\n\n- Instead of asking \\\"what medications are you on?\\\" (free-text → likely PHI), use a `multiple_choice` question with anonymized therapeutic categories.\n- Instead of \\\"when were you diagnosed?\\\" (date → identifier), use a `single_choice` question with year ranges.\n- Instead of \\\"how often do you experience symptoms?\\\" (open-text invitation to overshare), use a `scale` (1-5) or a `single_choice` with frequency buckets.\n\nThe [structured questions guide](/docs/structured-questions-guide) is the canonical reference; the [scale questions guide](/docs/scale-questions-guide) covers Likert and frequency patterns specifically.\n\nThis is one of the most important Koji differentiators for regulated industries. Traditional survey tools force you into a binary choice: free-text (sensitive overshare risk) or rigid multiple choice (low signal). Koji's AI-moderated structured questions let the AI follow up *within* the bucket boundary — depth without unbounded free text.\n\n## A safer participant intake pattern\n\nFor any healthcare research study, your intake form should:\n\n- **Not require email.** Use anonymous mode and let participants land via a generic study link rather than a personalized one.\n- **Include an explicit research-only consent.** Re-using the [research consent form templates](/docs/research-consent-form-templates) library, add a sentence: *\\\"This research is conducted under our internal research policy, not as a clinical activity. We do not need or want your PHI. Please do not share names, dates of birth, MRNs, or specific clinical details.\\\"*\n- **Set expectations on AI moderation.** Disclose that an AI is moderating, that no human will listen to recordings in real time, and that the data will be used only for product/service improvement.\n- **Provide a contact for withdrawal.** Even with anonymous mode, give participants a way to request deletion of their interview by quoting their respondent ID (visible at the end of every Koji interview).\n\n## What Koji's architecture gives you out of the box\n\n- **TLS 1.2+ in transit, AES-256 at rest** for all participant content.\n- **No model training on customer data** — Koji has contractual no-train clauses with its LLM providers, and on Enterprise BYOK the contract is yours directly.\n- **SOC 2 Type II infrastructure** via Vercel and Supabase as primary sub-processors.\n- **Per-study retention controls** so transcripts can be purged on a schedule.\n- **Workspace-scoped access** so a study is only visible to the people you explicitly add.\n- **Audit logs** of who accessed which study and when (Enterprise tier).\n\nKoji is not, by default, a HIPAA Covered Entity or Business Associate. The Enterprise tier supports a BAA on request for customers who need to send PHI through the platform. For most healthcare research programs, the de-identified pattern above is faster to launch, faster to ship insights from, and avoids the BAA path entirely.\n\n## Comparison: HIPAA on Koji vs. traditional research tools\n\n- **SurveyMonkey and Typeform** require enterprise plans to sign a BAA — and even then, free-text responses are an overshare hazard. Without AI-moderated probing, you can't coach participants away from PHI in the moment.\n- **Qualtrics** has a HIPAA-eligible XM tier, but the platform is built around quantitative survey logic, not conversational depth. The cost and complexity is enterprise-only.\n- **Manual moderator interviews** can be HIPAA-aligned but they're slow and expensive. A single in-depth provider interview is typically $150–$400 in incentive plus 60-90 minutes of moderator time.\n- **Koji** lets you run de-identified AI interviews at scale on self-serve plans, and graduates to Enterprise + BAA + BYOK when a study genuinely needs PHI. The same study design, two compliance modes.\n\n## Common pitfalls to avoid\n\n- **Treating anonymized as anonymous.** A combination of role + employer + ZIP can re-identify a single person in a small market. Be careful with combinations.\n- **Forgetting voice biometrics.** Voice recordings are themselves an identifier. Default to text in HIPAA-regulated studies unless you have a BAA in place.\n- **Exporting raw transcripts to non-BAA tools.** If you push transcripts to a Notion workspace or a Slack channel, those vendors are now in scope. Use anonymized themes and quotes instead, or restrict integration use to BAA-covered destinations.\n- **Letting screener questions identify rare conditions.** A screener that filters for \\\"adults with a rare disease in a small ZIP code\\\" produces an effectively re-identifiable sample even before the interview starts.\n- **Skipping the IRB conversation when PHI is involved.** Most Covered Entities require IRB review for any research that touches PHI, even when an external vendor handles the moderation.\n\n## A 5-step checklist before you launch\n\n1. **Scope the question.** Can you answer it without any of the 18 identifiers? If yes, default path. If no, Enterprise + BAA path.\n2. **Configure anonymous mode and disable demographic intake.** Use behavioral screeners only.\n3. **Brief the AI interviewer** with explicit PHI-avoidance instructions in company context.\n4. **Set retention** to the shortest acceptable window (often 30-90 days for transcripts).\n5. **Review the first 3 transcripts** before scaling. If participants are oversharing, tighten the AI's steering instructions and re-launch.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that let you capture sensitive content as categorical buckets.\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — sister guide for EU privacy law.\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — the field-level controls Koji provides for de-identification.\n- [Bring Your Own Key (BYOK)](/docs/bring-your-own-key) — how to route LLM calls through your own contracted account on the Enterprise tier.\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — starter templates you can adapt for healthcare contexts.\n- [Research Screener Questions](/docs/research-screener-questions) — screening patterns that don't collect identifiers.\n- [AI Research for Healthcare](/docs/ai-research-for-healthcare) — broader playbook for patient and provider research.\n- [Research Ethics Guide](/docs/research-ethics-guide) — informed consent, incentives, and minimizing harm in qualitative research.","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"HIPAA-Compliant AI User Research: Healthcare & HealthTech Playbook","metaDescription":"Run AI-moderated customer research in healthcare contexts without putting PHI at risk. Anonymous-mode patterns, BYOK, BAA path, and the 18 HIPAA identifiers your interviews must avoid.","keywords":["hipaa compliant user research","hipaa research tool","hipaa survey tool","healthtech user research","patient research compliance","ai user research healthcare","hipaa de-identification","healthcare customer research","baa user research"],"aiSummary":"HIPAA-compliant AI customer research is possible by scoping studies to avoid the 18 HIPAA identifiers entirely — name, DOB, MRN, email, voice biometrics, etc. The default Koji pattern uses anonymous mode (no email or demographics collected), structured screener questions, and an AI moderator briefed to steer away from PHI. Voice mode is avoided for regulated studies because voiceprints are themselves identifiers. When PHI is genuinely required, the Enterprise tier supports a Business Associate Agreement plus BYOK (Bring Your Own Key) so LLM calls route through the customer's own Anthropic, OpenAI, or Vertex contract. Per-study retention, workspace-scoped access, AES-256 at rest, and TLS 1.2+ in transit are standard. The compliance lever Koji uniquely provides is structured questions: six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that let teams capture sensitive content as categorical buckets instead of free-text PHI. Most product, churn, onboarding, and brand research questions can be answered without any PHI at all — making the de-identified path faster, cheaper, and avoidant of the BAA negotiation for the majority of healthcare research programs.","aiPrerequisites":["Basic understanding of HIPAA Privacy Rule","A research question that touches healthcare context","A Koji account (self-serve for de-identified studies; Enterprise for PHI)"],"aiLearningOutcomes":["Understand which research questions need PHI and which do not","Design AI interview studies that stay outside HIPAA scope","Configure Koji anonymous mode, screeners, and AI steering for healthcare research","Know when to escalate to Enterprise + BAA + BYOK","Avoid the 5 most common HIPAA pitfalls in customer research"],"aiDifficulty":"intermediate","aiEstimatedTime":"18 min read"},{"type":"documentation","id":"a41f650d-ff03-4fe7-96c2-176b70de4f97","slug":"irb-approval-user-research","title":"Do You Need IRB Approval for User Research? A 2026 Decision Guide","url":"https://www.koji.so/docs/irb-approval-user-research","summary":"Most commercial product and UX research does not require IRB approval because it is not designed to produce generalizable knowledge - the hinge test under the Common Rule (45 CFR 46). Four situations require review: peer-reviewed publication intent, FDA or regulatory submissions, federal funding, and institutional policy including vulnerable populations. Exempt status (45 CFR 46.104, especially Category 2 for surveys and interviews) must be determined by an IRB and cannot be self-declared. Publication intent must be decided before data collection because retroactive approval is generally unavailable.","content":"## The short answer\n\n**Most commercial product and UX research does not need IRB approval. Most academic research does.** The dividing line is not the method you use — interviews, surveys, and usability tests all sit on both sides of it. The line is drawn by federal regulation (the Common Rule, 45 CFR 46) and turns on one hinge question: **is your study designed to produce generalizable knowledge?**\n\nDiscovery interviews, usability tests, NPS programs, and churn research run to improve your own product are almost never \"human subjects research\" in the regulatory sense. They are business operations that happen to involve talking to people.\n\nFour situations flip the answer to yes:\n\n1. You intend to **publish in a peer-reviewed venue**\n2. The work supports an **FDA or other regulatory submission**\n3. The study is **federally funded** under an agency that applies the Common Rule\n4. Your **institution's own policy** requires review — common with universities, health systems, school districts, and studies involving vulnerable populations\n\n> This guide is written for research and product teams, not lawyers, and it is not legal advice. Critically: **only an IRB can determine that a study is exempt.** You cannot self-certify exemption.\n\n## What an IRB actually is\n\nAn Institutional Review Board is a committee, registered and operating under 45 CFR 46, that reviews research involving human subjects before it begins. Its job is participant protection, not research quality policing.\n\nThe approval criteria are spelled out in [45 CFR 46.111](https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-A/part-46/subpart-A/section-46.111). A board must find that:\n\n- Risks to participants are minimized\n- Risks are reasonable relative to anticipated benefits\n- Participant selection is equitable\n- Informed consent is obtained and properly documented\n- There are adequate provisions for monitoring data\n- There are adequate protections for privacy and confidentiality\n\nRead that list again with a research-ops eye: five of the six are documentation problems. That is why teams with clean consent flows and data-management plans clear review quickly, and teams without them stall.\n\n## The two-part regulatory test\n\nUnder the Common Rule, you need review only if the activity is **research** *and* involves **human subjects**. Both must be true.\n\n**Is it \"research\"?** The regulation defines it as \"a systematic investigation, including research development, testing, and evaluation, designed to develop or contribute to generalizable knowledge.\"\n\nThe operative phrase is *generalizable knowledge*. A study designed to tell you whether **your** onboarding flow confuses **your** users is not designed to contribute generalizable knowledge — it is designed to improve a specific product. That is the single most important distinction in this entire guide.\n\n**Does it involve \"human subjects\"?** A living individual about whom an investigator obtains information through intervention or interaction, or obtains identifiable private information.\n\nThere is also a threshold question that trips people up: **is your organization even covered?** A private company with no federal funding and no Federalwide Assurance is generally not subject to the Common Rule at all. That does not mean ethics are optional — it means the *regulatory* mechanism of IRB review does not apply, and your obligations run through privacy law, contracts, and professional ethics instead.\n\n## Decision table\n\n| Your scenario | IRB review typically needed? |\n|---|---|\n| Discovery interviews to shape your roadmap | No |\n| Usability test of your own checkout flow | No |\n| Ongoing NPS, CSAT, or churn interview program | No |\n| Blog post or conference talk about what you learned | Usually no — but confirm the venue's policy |\n| Peer-reviewed journal or academic conference paper | **Yes** |\n| FDA submission (device human factors, labeling comprehension) | **Yes** |\n| NIH, NSF, or other federally funded study | **Yes** |\n| Collaboration with a university co-author | **Yes — through their institution** |\n| Research with children recruited through a school | **Often, plus district and school approval** |\n| Study of your own employees intended for publication | **Yes** |\n\nThe pattern: **audience and purpose decide, not method.** The same 30-minute interview guide can be exempt business research on Monday and reviewable human subjects research on Tuesday, purely because you decided to publish.\n\n## The exempt categories worth knowing\n\nThe 2018 revised Common Rule (45 CFR 46.104) defines eight exempt categories. Two matter most for research teams:\n\n- **Category 1** — research conducted in established educational settings involving normal educational practices\n- **Category 2** — research involving educational tests, **survey procedures, interview procedures**, or observation of public behavior\n\nCategory 2 covers a great deal of ordinary UX work. But two caveats matter enormously:\n\n**Exempt is a determination, not a self-declaration.** You submit; the IRB decides. Skipping this step and later claiming exemption is the most common way teams lose the ability to publish.\n\n**Exempt does not mean paperwork-free.** Many institutions require registration, and some exempt categories trigger a \"limited IRB review\" focused specifically on privacy and confidentiality protections.\n\n## The four triggers, in detail\n\n**1. Publication intent.** Journals and academic conferences routinely require an IRB approval statement. Retroactive approval is generally **not** available — a board cannot approve a study that already happened. Decide whether you might publish *before* you collect a single response. This is the mistake that cannot be fixed later.\n\n**2. Regulated submissions.** FDA-regulated human factors and usability work for medical devices, and labeling comprehension studies for drugs, sit under their own regulatory regime with its own review expectations.\n\n**3. Federal funding.** If an agency applying the Common Rule funds the work, review follows the money.\n\n**4. Vulnerable populations and institutional policy.** Children, prisoners, pregnant people, and individuals with cognitive impairments attract additional scrutiny and, often, additional subparts of the regulation. Schools and health systems layer their own gatekeepers on top — a school district's approval is a separate hurdle from the IRB's, and neither substitutes for the other. If you are researching minors, start with [user research with children and teens](/docs/user-research-with-children-teens).\n\n## If you do need review: how to move faster\n\n**Pick the right track.** Exempt determination, expedited review (minimal-risk studies reviewable by a single member), or full board review. Full board review is tied to a meeting calendar, which is usually what makes timelines feel unpredictable — expect turnaround to vary substantially by institution, and ask for the submission calendar up front rather than guessing.\n\n**Assemble the packet before you start writing prose:**\n\n- Protocol and research questions\n- Recruitment materials and [screener questions](/docs/screener-questions-guide)\n- Informed consent document\n- The full question list or interview guide\n- A data-management plan: where data lives, who can access it, how long you keep it, how you de-identify\n- Incentive structure and amounts\n\nIn practice the biggest source of delay is not the science — it is an incomplete consent document or a missing data-management plan. Those are exactly the artifacts a disciplined research operation already produces.\n\n## Why AI-moderated interviews produce cleaner IRB submissions\n\nAn AI interviewer does not change whether you need review. It does change how good your documentation is — and documentation is what review actually examines.\n\n- **The protocol is delivered identically to every participant.** A board reviewing a human-moderated study has to trust that eight different moderators asked questions consistently. With AI moderation, the protocol *is* the instrument, and every participant receives it the same way. See [AI vs human moderators](/docs/ai-vs-human-moderators) for where each approach wins.\n- **Consent is captured before the conversation starts**, as a structured step rather than a verbal aside — see [intake forms and consent](/docs/intake-forms-and-consent).\n- **Verbatim transcripts** give reviewers and auditors the complete record, rather than a moderator's notes.\n- **Structured questions make your instrument reviewable.** Koji supports six question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — so the exact instrument an IRB wants to see is the same object you configured. The [structured questions guide](/docs/structured-questions-guide) walks through building one.\n- **De-identification is a design choice you make up front.** Scope studies to avoid collecting identifiers at all; see [anonymizing customer interview data](/docs/anonymizing-customer-interview-data).\n\nTwo honest caveats. First, AI moderation does not reduce the regulatory need for review — it improves the quality of what you submit. Second, **disclose in your protocol that the interviewer is an AI system.** Reviewers will ask, participants deserve to know, and in the EU the AI Act imposes its own transparency duty on AI interactions independent of anything an IRB requires.\n\n## A worked example\n\nA healthtech team runs 40 AI-moderated interviews about how patients describe medication side effects. Purpose: improve their symptom-logging feature. No federal funding, no publication plan, no PHI collected. **No IRB review needed** — this is product research.\n\nSix months later, a clinician advisor wants to co-author a paper on the findings with a university affiliation. Now the calculus inverts: the university is engaged, the intent is generalizable knowledge, and the data was collected without approval. The paper is likely unpublishable from that dataset. The fix would have been a 30-minute conversation with the university's IRB office *before* fieldwork — and, if needed, a fresh round of consented interviews under an approved protocol.\n\n## Common mistakes\n\n1. **Assuming \"it's just interviews\" means exempt.** Category 2 often applies, but only the IRB can say so.\n2. **Deciding to publish after collecting data.** The most expensive and least fixable error in this guide.\n3. **Treating \"exempt\" as \"no process.\"** Registration and limited review are common.\n4. **Forgetting the second gatekeeper.** Schools, employers, and health systems approve separately from the IRB.\n5. **Submitting a consent form with no data-management plan.** The fastest route to a revise-and-resubmit.\n6. **Confusing ethics with regulation.** No IRB requirement does not mean no obligations — see the [research ethics guide](/docs/research-ethics-guide).\n\n## Related Resources\n\n- [Research Ethics and Informed Consent](/docs/research-ethics-guide) — the ethical foundation that applies whether or not an IRB does\n- [Structured Questions Guide](/docs/structured-questions-guide) — build the reviewable instrument behind your protocol\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — starting points for your consent documentation\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — de-identification as a design choice\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — the data-management plan reviewers ask for\n- [User Research With Children and Teens](/docs/user-research-with-children-teens) — when vulnerable-population rules apply\n- [Trauma-Informed User Research](/docs/trauma-informed-user-research) — minimizing risk on sensitive topics","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Do You Need IRB Approval for User Research? (2026 Guide)","metaDescription":"Most commercial UX research does not need IRB approval - most academic research does. Learn the Common Rule test, the exempt categories, the four triggers that require review, and how to prepare a fast-clearing submission.","keywords":["irb approval user research","do i need irb approval","irb ux research","common rule 45 cfr 46","human subjects research","irb exempt category 2","irb user experience research","institutional review board product research","irb submission checklist","generalizable knowledge test"],"aiSummary":"Most commercial product and UX research does not require IRB approval because it is not designed to produce generalizable knowledge - the hinge test under the Common Rule (45 CFR 46). Four situations require review: peer-reviewed publication intent, FDA or regulatory submissions, federal funding, and institutional policy including vulnerable populations. Exempt status (45 CFR 46.104, especially Category 2 for surveys and interviews) must be determined by an IRB and cannot be self-declared. Publication intent must be decided before data collection because retroactive approval is generally unavailable.","aiPrerequisites":["Basic familiarity with running user interviews or surveys","An understanding of your organization funding sources and publication plans"],"aiLearningOutcomes":["Apply the two-part Common Rule test to decide whether your study needs IRB review","Identify the four situations that require review regardless of method","Understand why exempt status must be determined by an IRB rather than self-declared","Assemble an IRB submission packet that clears review without revise-and-resubmit","Design AI-moderated studies whose documentation satisfies reviewer expectations"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"939c276e-d0cd-4bc2-b4c0-9a9694103c78","slug":"dpia-user-research","title":"DPIA for User Research: When You Need One and How to Write It (2026)","url":"https://www.koji.so/docs/dpia-user-research","summary":"A DPIA is required under GDPR Article 35 when research is likely to result in high risk to participants. Score your study against the WP29 nine criteria; two or more means do the DPIA. AI-moderated interviews at scale typically hit three or four (large scale, innovative technology, sensitive data, vulnerable subjects). Article 35(7) requires four elements: a systematic description of processing, a necessity and proportionality assessment, a risk assessment from the participant point of view, and mitigations. Follow seven steps: screen, describe, consult, assess necessity, score risks, mitigate, sign off and review. If high residual risk remains, Article 36 requires prior consultation with the supervisory authority. Koji reduces DPIA risk at the design stage: six structured question types minimise collected data, no human moderator removes incidental disclosure, and automatic reporting makes short retention realistic.","content":"## The short answer\n\n**Most one-off user research does not require a DPIA. Research programs that record people, run at scale, use AI to moderate or analyse, or touch sensitive topics usually do.** A Data Protection Impact Assessment (DPIA) is the GDPR's structured way of asking \"could this processing hurt the people in it, and what have we done about that?\" Article 35 makes it mandatory whenever processing is \"likely to result in a high risk to the rights and freedoms of natural persons.\"\n\nThe practical test used across the EU comes from the Article 29 Working Party (now the EDPB): score your study against nine risk criteria. **Hit two or more, and you should assume a DPIA is required.** A recorded, AI-moderated interview study with EU participants routinely hits three or four — which is why research teams that adopted AI interviewing in 2025 and 2026 are being asked for DPIAs they have never written before.\n\nThis guide gives you the triggers, the required contents, a seven-step process, and a worked example you can adapt.\n\n## What a DPIA actually is\n\nA DPIA is a documented risk assessment carried out **before** processing begins. It is not a legal opinion, a security review, or a privacy policy. It is a written record showing that you identified the risks to participants, weighed them against your purpose, and reduced them.\n\nThree things follow from that definition, and teams get all three wrong:\n\n- **It is prospective.** Article 35 and the data-protection-by-design principle both require the assessment in advance. A DPIA written after launch is remediation, not compliance.\n- **It is about the participant, not the company.** The risk you assess is risk to the person being interviewed — re-identification, exposure of a sensitive disclosure, loss of control over their words. Not your brand risk.\n- **It is a living document.** If you change methodology, add voice recording, or start using a new analysis vendor, you revisit it.\n\n## When a DPIA is required\n\nArticle 35(3) names three situations where a DPIA is always required:\n\n1. **Systematic and extensive evaluation** of personal aspects based on automated processing, including profiling, where decisions produce legal or similarly significant effects.\n2. **Large-scale processing of special category data** (Article 9) — health, race, religion, political opinion, sex life, biometrics — or criminal conviction data.\n3. **Systematic monitoring of a publicly accessible area** on a large scale.\n\nBeyond those, the WP29 guidelines give nine criteria that indicate high risk:\n\n| # | Criterion | Common research trigger |\n|---|---|---|\n| 1 | Evaluation or scoring | Scoring or ranking participants, quality-rating responses |\n| 2 | Automated decision-making with legal or significant effect | Auto-screening applicants or customers out |\n| 3 | Systematic monitoring | Always-on feedback capture, session monitoring |\n| 4 | Sensitive data or data of a highly personal nature | Health, finances, workplace grievances, trauma |\n| 5 | Data processed on a large scale | Hundreds or thousands of participants |\n| 6 | Matching or combining datasets | Joining interview data to CRM or product analytics |\n| 7 | Data concerning vulnerable subjects | Children, patients, employees, asylum seekers |\n| 8 | Innovative use or new technological solutions | **AI-moderated interviews, voice analysis, LLM synthesis** |\n| 9 | Processing that prevents subjects exercising a right | Barriers to withdrawing consent or requesting deletion |\n\n**The rule of thumb regulators apply: two or more criteria means do the DPIA.** Where you are genuinely unsure, do it anyway — it is cheaper than defending the decision not to.\n\n### Four research scenarios that almost always trigger one\n\n- **AI-moderated interviews at scale with EU participants.** Criteria 5 and 8 at minimum; add 4 if the topic is personal. This is the single most common trigger in 2026.\n- **Employee research.** Employees are vulnerable subjects (criterion 7) because consent is rarely freely given in an employment relationship. Add sensitive topics and you are at three.\n- **Health, financial, or minor-facing research.** Special category data or children — often mandatory under 35(3) alone.\n- **Research joined to behavioural data.** Matching interview transcripts to product analytics or CRM records is criterion 6.\n\nConversely: a five-person moderated usability test on a checkout flow, unrecorded, notes anonymised immediately, is not high risk. Do not burn a DPIA on it. Document the screening decision instead.\n\n## What goes in a DPIA\n\nArticle 35(7) sets four mandatory elements. Everything else is your own structure.\n\n1. **A systematic description of the processing** — purposes, categories of data, categories of participants, recipients, retention periods, international transfers, and your legal basis.\n2. **An assessment of necessity and proportionality** — why you need this data to answer this question, and why a less intrusive method would not do.\n3. **An assessment of the risks** to participants' rights and freedoms.\n4. **The measures envisaged to address those risks**, including safeguards and security measures.\n\nA workable research DPIA runs 4-8 pages. Longer usually means you have pasted in vendor marketing.\n\n### The necessity-and-proportionality section is where DPIAs fail\n\nThis is the section auditors actually read, and the one teams skip. It has to answer: could you have learned this without collecting this data? Concretely — do you need full-name identifiers, or would a pseudonymous participant ID do? Do you need video, or is audio enough? Do you need audio retained, or just the transcript? Every \"no, we do not need that\" you write here removes risk from every later section.\n\n## The seven-step process\n\n1. **Screen.** Score against the nine criteria. Record the score and the decision even when the answer is \"no DPIA needed\" — the screening record is itself evidence of accountability.\n2. **Describe the processing.** Data flow from recruitment through to deletion. Name every system that touches participant data, including transcription and analysis.\n3. **Consult.** Article 35(9) says seek the views of data subjects \"where appropriate.\" For research, a short question in your pilot round covers this cheaply. Involve your DPO if you have one — that consultation is mandatory under 35(2).\n4. **Assess necessity and proportionality.** Justify each field you collect against the research question.\n5. **Identify and score risks.** Likelihood x severity, from the participant's point of view.\n6. **Identify mitigations** and record the residual risk after each.\n7. **Sign off and review.** Named owner, date, and a review trigger tied to methodology changes.\n\n### Risk and mitigation table you can reuse\n\n| Risk to participant | Likelihood | Severity | Mitigation | Residual |\n|---|---|---|---|---|\n| Re-identification from quoted transcript | Medium | High | Pseudonymous IDs; strip names, employers, rare job titles before sharing; review quotes before they enter reports | Low |\n| Voice recording reused beyond stated purpose | Low | High | Purpose limitation in consent; audio deleted at 30-90 days; transcripts retained instead | Low |\n| Sensitive disclosure captured incidentally | Medium | High | Topic guardrails in the interview script; participants told they may skip any question; redaction pass before analysis | Low |\n| Participant cannot withdraw | Low | Medium | Withdrawal contact in the consent screen; documented deletion path with SLA | Low |\n| Cross-border transfer without safeguards | Medium | Medium | Documented transfer mechanism and processing locations | Low |\n\n## Where AI-moderated research changes the analysis\n\nCriterion 8 — \"innovative use of new technological solutions\" — is not a formality. When an AI moderates the conversation, two things genuinely change:\n\n**The interview is adaptive, not fixed.** A human-written script is auditable in advance; an AI that generates follow-up questions is not. Your DPIA should describe the *guardrails* rather than the exact questions: the topic scope the AI is instructed to stay within, whether it is permitted to probe on sensitive areas, and what it does when a participant volunteers something out of scope.\n\n**Analysis is automated.** If a model summarises, themes, or scores responses, say so, and say whether any of that produces a decision about the individual. In most product research it does not — the output is aggregate insight, not a decision about the participant — and stating that plainly is often what closes the review.\n\nThis is also where the EU AI Act now interacts with your DPIA. Article 26 obligations for deployers of certain AI systems and the transparency duty to tell people they are interacting with an AI sit alongside, not instead of, your GDPR work. The clean pattern is one paragraph in the DPIA cross-referencing your AI Act position — see [the EU AI Act and user research](/docs/eu-ai-act-user-research-compliance) for the classification detail.\n\n## How Koji makes this section easy to write\n\nMost of a research DPIA is a description of *what the tool does with participant data*. That is far easier to write when the platform is built for research rather than adapted from a meeting tool.\n\nWith Koji, the description section writes itself along these lines: participants join a link-based [AI-moderated interview](/docs/how-ai-interviewers-work) by text or voice; the AI asks a defined set of questions and probes within a scoped topic; responses are transcribed and analysed automatically into an aggregate report; no human moderator is present, which removes an entire category of incidental disclosure to staff.\n\nTwo design choices in particular reduce risk before you mitigate anything:\n\n- **Structured questions constrain what gets collected.** Koji supports six question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Every question you convert from open-ended to `scale`, `single_choice`, or `yes_no` collects a bounded value instead of free text that might contain anything. That is data minimisation implemented in the study design, not promised in a policy. See the [structured questions guide](/docs/structured-questions-guide).\n- **No moderator means no incidental audience.** In a traditional study, a moderator, a notetaker, and often observers hear every word. Koji's AI runs the session and produces aggregate analysis, so fewer humans see raw disclosures.\n\nBecause studies run asynchronously and analysis is automatic, teams also stop the worst privacy habit in research: keeping recordings \"just in case\" because nobody has had time to analyse them. When [reports generate automatically](/docs/how-to-automate-user-research), short retention becomes realistic — and short retention is the mitigation that does the most work in any DPIA.\n\n## Prior consultation: the step nobody plans for\n\nIf, after mitigation, residual risk is still high, Article 36 requires you to consult your supervisory authority **before** processing. Expect that to add weeks. In practice, research studies almost never reach this point — but the way to guarantee you avoid it is to mitigate down to acceptable residual risk in step 6, not to quietly downgrade your own risk scores.\n\n## Five common mistakes\n\n1. **Treating the DPIA as a vendor questionnaire.** Your vendor's SOC 2 report is evidence, not an assessment. The DPIA is about *your* study.\n2. **Writing it after fieldwork starts.** It must be prospective.\n3. **Assessing company risk instead of participant risk.** \"Reputational damage to us\" is not an Article 35 risk.\n4. **Skipping the screening record for low-risk studies.** The record of deciding you did not need one is itself accountability evidence.\n5. **Never reviewing it.** Adding voice to a text-only study, or a new analysis vendor, changes the assessment.\n\n## Related Resources\n\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — lawful basis, participant rights, and the wider GDPR picture\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) — how AI Act duties sit alongside your DPIA\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — building the retention schedule your DPIA references\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — the practical mitigation behind most risk rows\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how they minimise collected data\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — consent language that matches what your DPIA promises\n- [Do You Need IRB Approval for User Research?](/docs/irb-approval-user-research) — the parallel question for ethics review\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — the vendor-side evidence you attach","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"DPIA for User Research: When You Need One + Template (2026)","metaDescription":"When does user research need a Data Protection Impact Assessment? The Article 35 triggers, the WP29 nine criteria, required contents, a 7-step process, and a reusable risk table.","keywords":["dpia user research","data protection impact assessment","gdpr article 35","dpia template research","when is a dpia required","dpia ai interviews","research privacy risk assessment","wp29 nine criteria","dpia checklist","user research gdpr compliance"],"aiSummary":"A DPIA is required under GDPR Article 35 when research is likely to result in high risk to participants. Score your study against the WP29 nine criteria; two or more means do the DPIA. AI-moderated interviews at scale typically hit three or four (large scale, innovative technology, sensitive data, vulnerable subjects). Article 35(7) requires four elements: a systematic description of processing, a necessity and proportionality assessment, a risk assessment from the participant point of view, and mitigations. Follow seven steps: screen, describe, consult, assess necessity, score risks, mitigate, sign off and review. If high residual risk remains, Article 36 requires prior consultation with the supervisory authority. Koji reduces DPIA risk at the design stage: six structured question types minimise collected data, no human moderator removes incidental disclosure, and automatic reporting makes short retention realistic.","aiPrerequisites":["Basic familiarity with GDPR terminology","A defined research question and study plan"],"aiLearningOutcomes":["Decide whether your study requires a DPIA using the WP29 nine criteria","Write each of the four sections Article 35(7) requires","Score participant risk and document proportionate mitigations","Recognise when AI-moderated research changes the assessment","Know when Article 36 prior consultation is triggered"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"127ae94e-e987-490d-8047-e82fb1abe0c3","slug":"accessibility-compliance-research","title":"Accessibility Compliance Research: What WCAG, the ADA, and the European Accessibility Act Require You to Test","url":"https://www.koji.so/docs/accessibility-compliance-research","summary":"No accessibility law currently requires user testing with disabled people, yet research shows only about 50% of barriers real users encounter map to any WCAG success criterion. This guide covers the four regimes in force (EAA/EN 301 549, ADA Title II, Section 504, Section 508), the 2026 US deadline extensions that most published guidance has not caught up with, the difference between conformance artefacts (audit, VPAT/ACR) and research evidence, and a six-step method for running accessibility research that finds the barriers audits miss.","content":"## The short answer\n\n**No accessibility law in force today requires you to test with disabled users — and that is exactly why conformance audits keep passing products that disabled people cannot use.**\n\nEvery current regime (the European Accessibility Act, ADA Title II, Section 504, Section 508) points at the same technical target: the Web Content Accessibility Guidelines, at Level AA. WCAG is a checklist of testable success criteria. It is excellent at catching missing alt text, insufficient contrast, and unlabelled form fields. It was never designed to tell you whether a screen reader user can actually complete your checkout.\n\nThe evidence on that gap is unusually clean. In a study of 32 blind users across 16 websites, researchers logged 1,383 instances of real user problems — and **only 50.4% of them mapped to any WCAG 2.0 success criterion**. Roughly half the barriers people actually hit were invisible to the standard. The W3C's own Web Accessibility Initiative says the same thing in plainer language: evaluating with users with disabilities identifies usability issues that conformance evaluation alone does not discover.\n\nSo the compliance answer and the research answer are two different jobs. This guide covers both: what each regime legally obliges you to do, and what user research you should run on top of it.\n\n## The 2026 deadline reset most guidance has not caught up with\n\nTwo US deadlines moved in 2026. A great deal of published advice — and a fair number of internal compliance trackers — still cite the old dates.\n\n- **ADA Title II.** The DOJ's 2024 rule set WCAG 2.1 Level AA as the standard for state and local government web content and mobile apps. In April 2026 the DOJ issued an interim final rule pushing both compliance dates back one year: **26 April 2027** for public entities serving populations of 50,000 or more, and **26 April 2028** for smaller entities and special district governments.\n- **Section 504 (HHS).** The HHS Office for Civil Rights followed on **7 May 2026** — four days before the original deadline — with its own interim final rule. Recipients of HHS federal financial assistance with 15 or more employees now have until **11 May 2027**; those with fewer than 15 employees have until **10 May 2028**. The standard is still WCAG 2.1 Level AA, and the rule reaches hospitals, health plans, digital health companies, and clinical research organisations.\n\nTwo cautions. First, the extensions delay the *technical conformance dates only*. They do not suspend the underlying nondiscrimination and effective-communication duties, which have supported accessibility claims for years without any technical standard attached. Private plaintiffs can and do sue during the extension window. Second, litigation volume is not slowing: UsableNet's tracking put 2025 at roughly 5,000 digital accessibility lawsuits, with more than 1,400 filed against companies that had already been sued once before.\n\n## The four regimes at a glance\n\n| Regime | Who it binds | Technical standard | Key date |\n|---|---|---|---|\n| European Accessibility Act (Directive (EU) 2019/882) | Private businesses selling covered products/services to EU consumers — e-commerce, banking, transport, telecoms, e-books | EN 301 549 (currently v3.2.1, which incorporates WCAG 2.1 AA) | Applicable since **28 June 2025** |\n| ADA Title II (US) | State and local government entities | WCAG 2.1 Level AA | 26 Apr 2027 / 26 Apr 2028 |\n| Section 504 (US, HHS rule) | Recipients of HHS federal financial assistance | WCAG 2.1 Level AA | 11 May 2027 / 10 May 2028 |\n| Section 508 (US) | Federal agencies and their vendors | EN-aligned; WCAG-based | In force |\n\nA note on version drift that trips up EU teams: **the EAA does not currently mandate WCAG 2.2.** The harmonised standard cited in the Official Journal is EN 301 549 v3.2.1, which incorporates WCAG 2.1 Level AA. Version 4.1.1, which brings in WCAG 2.2 Level AA and its six additional success criteria, is expected to be referenced in the Official Journal around late 2026. Building to 2.2 now is sensible future-proofing, not a current legal requirement — and anyone telling you 2.2 is mandatory in the EU today is ahead of the citation.\n\nThe other EAA subtlety: EN 301 549 is broader than WCAG. It covers hardware, mobile apps, documentation, video players, and third-party content embedded in your service. A WCAG-only audit does not fully answer the EAA.\n\n## Conformance evidence vs. research evidence\n\nProcurement and legal will ask you for conformance artefacts. Product will ask you whether the thing works. Keep the two straight:\n\n**Conformance artefacts** — an accessibility audit against WCAG 2.1 AA, an Accessibility Conformance Report (usually authored on the VPAT template), and a documented remediation plan with dates. These are what you hand to a buyer's security and legal review, and what you request from every vendor you buy.\n\n**Research evidence** — task-based sessions with disabled participants using their own assistive technology, on their own devices, with their own settings. This is what tells you the checkout is completable. It is not a substitute for the audit, and the audit is not a substitute for it.\n\nThe failure mode is treating the ACR as the finish line. A product can be fully conformant and still strand a user, because WCAG cannot express \"the error message is technically announced but appears 40 seconds after the field it refers to.\"\n\n## How to run accessibility research that actually finds barriers\n\n**1. Recruit for assistive technology, not for diagnosis.** The useful screening variable is the technology stack and how fluently the person uses it — JAWS, NVDA, VoiceOver, Dragon, switch access, screen magnification, high-contrast modes. \"Has a visual impairment\" is too coarse to plan a session around. Structured screener questions do this cleanly: use `single_choice` for primary assistive technology, `multiple_choice` for the full stack, and `scale` for self-rated proficiency.\n\n**2. Never make the research method itself a barrier.** This is the most common own-goal in accessibility research. Scheduled video calls with screen sharing, unfamiliar moderation software, and \"can you install this plugin\" all add access friction that has nothing to do with your product. If a participant spends the first ten minutes fighting your research tool, you have measured your research tool.\n\n**3. Let participants choose their modality.** Some people are far faster typing with a screen reader than speaking; others are the reverse. Forcing one channel biases who can take part and how much they say.\n\n**4. Test tasks, not pages.** \"Find and change your billing address\" surfaces barriers that a page-by-page audit sweep will not. Sequence matters — most real barriers appear at handoffs between steps.\n\n**5. Budget properly.** Assistive technology users routinely take longer on tasks. Sessions run long, and incentives should reflect the expertise participants bring.\n\n**6. Log every barrier against both frames.** For each problem, record the task, the assistive technology, the WCAG success criterion if one applies — and mark the ones where none does. That \"no matching criterion\" column is the most valuable output of the whole study, because it is the part your audit will never produce.\n\n## Where Koji fits\n\nKoji runs AI-moderated interviews over a link. Participants open it when they want, on their own device, with their own assistive technology already configured — there is nothing to install, no meeting to join, no screen to share, and no moderator waiting while someone gets set up. For accessibility research specifically, that removes the exact friction that suppresses participation.\n\nParticipants can respond by **text or by voice**, and Koji's AI interviewer asks follow-up questions either way — so when someone says a flow was \"confusing\", the AI probes for what specifically broke, in the moment, rather than leaving you a one-line note to chase later. Because sessions are asynchronous, you are not constrained to the handful of slots a moderator can run in a day, which matters when your recruiting pool is narrower than usual.\n\nKoji's six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — let you carry a quantitative spine through the study (task success, severity ratings, assistive technology used) while the AI handles the qualitative probing. Reports aggregate both automatically, so severity distributions and the verbatim explanation of each barrier land in the same place.\n\nOne thing to do regardless of platform, including this one: **ask your research vendor for its own Accessibility Conformance Report.** A research tool that is not itself accessible cannot credibly collect accessibility feedback. Treat that request the same way you treat a DPA — a standard, unremarkable part of vendor review.\n\n## Common mistakes\n\n- **Treating the deadline extensions as a pause.** The nondiscrimination duty never moved, and neither did the litigation trend.\n- **Auditing to WCAG 2.2 and assuming the EAA is satisfied.** EN 301 549 covers more than web content; check the full standard against your product surfaces.\n- **Recruiting one blind participant and calling it accessibility research.** Screen reader users are not a proxy for cognitive, motor, or low-vision users, whose barriers differ completely.\n- **Running the study through inaccessible research software.** You will get a clean-looking dataset from the subset of people who could get through your tooling.\n- **Filing barriers only against WCAG criteria.** The uncategorisable half is where the real product work is.\n\n## Frequently asked questions\n\n**Does WCAG conformance make me legally compliant?**\nIt makes you conformant to the technical standard each of these regimes cites, which is the bulk of what regulators ask for. It does not discharge broader nondiscrimination and effective-communication duties, and it does not guarantee a disabled user can complete your tasks.\n\n**Do the laws require testing with disabled users?**\nNo current regime mandates it as a compliance step. It is nonetheless the only method that surfaces the roughly half of real barriers that WCAG success criteria do not describe, and regulators and courts look favourably on documented user involvement.\n\n**Which WCAG version applies to the European Accessibility Act right now?**\nEN 301 549 v3.2.1, incorporating WCAG 2.1 Level AA, is the harmonised standard currently cited. Version 4.1.1 with WCAG 2.2 Level AA is expected to be cited around late 2026.\n\n**Did the US accessibility deadlines really move?**\nYes. DOJ issued an interim final rule in April 2026 moving ADA Title II to 26 April 2027 and 26 April 2028; HHS issued its own on 7 May 2026 moving Section 504 to 11 May 2027 and 10 May 2028. Both kept WCAG 2.1 Level AA.\n\n**How many accessibility research participants do I need?**\nPlan per assistive technology group rather than in aggregate. Five to eight participants within a single technology profile will surface most severe barriers for that profile; a study spanning screen reader, magnification, voice control, and cognitive accessibility needs that depth in each.\n\n**What is a VPAT, and do I need one?**\nA VPAT is the template; the completed document is an Accessibility Conformance Report. If you sell to government, education, healthcare, or large enterprises, expect to be asked for one — and expect to ask your own vendors for theirs.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types, and how to use them for screeners and severity ratings\n- [Accessibility Research Guide](/docs/accessibility-research-guide) — the methods side: how to include users with disabilities in your studies\n- [Research Ethics and Informed Consent](/docs/research-ethics-guide) — consent practice for studies involving disability data\n- [Trauma-Informed User Research](/docs/trauma-informed-user-research) — adjacent guidance for sensitive-topic and vulnerable-participant studies\n- [Research Screener Questions](/docs/research-screener-questions) — writing screeners that find the right assistive technology profiles\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — the vendor-review companion to requesting an ACR\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — handling the health and disability data these studies touch","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Accessibility Compliance Research: WCAG, the ADA, and the EAA (2026)","metaDescription":"WCAG conformance misses about half of real accessibility barriers. What the EAA, ADA Title II, and Section 504 require in 2026 — and the research that closes the gap.","keywords":["accessibility compliance research","wcag 2.2 user testing","european accessibility act compliance","ada title ii deadline 2027","section 504 web accessibility","en 301 549","vpat accessibility conformance report","testing with disabled users","accessibility user research","wcag 2.1 aa"],"aiSummary":"No accessibility law currently requires user testing with disabled people, yet research shows only about 50% of barriers real users encounter map to any WCAG success criterion. This guide covers the four regimes in force (EAA/EN 301 549, ADA Title II, Section 504, Section 508), the 2026 US deadline extensions that most published guidance has not caught up with, the difference between conformance artefacts (audit, VPAT/ACR) and research evidence, and a six-step method for running accessibility research that finds the barriers audits miss.","aiDifficulty":"intermediate","aiEstimatedTime":"12 min"},{"type":"documentation","id":"1ac10273-7718-4a87-9321-4c04d26d8398","slug":"dsar-research-data","title":"DSARs for Research Data: Handling Access, Deletion, and Portability Requests from Participants","url":"https://www.koji.so/docs/dsar-research-data","summary":"A DSAR gives you one month under GDPR (extendable by two with notice) or 45 days under CCPA/CPRA. For research teams the difficulty is scoping what counts as a participant's data across transcripts, audio, screeners, incentive records, repositories, exports, and verbatims in decks. The Article 89 research derogation rarely helps commercial product research because it depends on Member State implementation, is necessity-bounded, and targets scientific research. The genuinely useful provision is Article 11(2): where you can no longer identify a data subject, most rights do not apply, so anonymisation takes data out of scope. Includes a seven-step runbook and the hard cases: published quotes, multi-voice audio, aggregate analysis, and third-party data in transcripts.","content":"## The short answer\n\n**A data subject access request is a deadline, not a project.** Under GDPR you have **one month** from receipt, extendable by a further two months for complex or numerous requests — but only if you tell the requester about the extension, and why, inside the original month. Under CCPA/CPRA you have **45 calendar days**, extendable by another 45 with notice, and consumers may make the request free twice a year.\n\nFor research teams the hard part is almost never the deadline. It is knowing what \"their data\" means when someone spent forty minutes talking to you: a transcript, an audio file, a screener response, a scheduling record, an incentive payment record, a set of tagged quotes, and possibly a verbatim already sitting in a slide deck that went to the exec team last quarter.\n\nRequest volumes are climbing. Privacy vendors tracking their own platforms reported DSAR volumes up roughly 43% between 2023 and 2024, and CCPA-route requests growing several-fold since 2021. A research programme that has never received one should assume it will.\n\n## The rights that actually land on research data\n\n| Right | GDPR | What it means for an interview study |\n|---|---|---|\n| Access | Art. 15 | A copy of their personal data plus context: purposes, recipients, retention period, source |\n| Erasure | Art. 17 | Delete transcript, audio, and identifiers — and tell sub-processors |\n| Portability | Art. 20 | Structured, machine-readable export of data they provided, where processing rests on consent or contract |\n| Rectification | Art. 16 | Correct inaccurate personal data (rarely a transcript; often a profile field) |\n| Objection | Art. 21 | Stop processing based on legitimate interests |\n| Restriction | Art. 18 | Freeze processing while a dispute is resolved |\n\nCCPA/CPRA maps loosely onto access (right to know), deletion, and correction, with its own 45-day clock and its own verification rules.\n\n## Why the research exemption probably does not save you\n\nArticle 89 permits derogations from several data subject rights where personal data is processed for **scientific or historical research or statistical purposes** — and Article 17(3)(d) contains a matching erasure carve-out. Research teams reach for this constantly. Three reasons to be careful:\n\n1. **The derogations are national law, not self-executing.** Article 89(2) lets Member States legislate derogations. Implementation varies significantly across the EU/EEA, so \"the research exemption\" means different things in Germany and Ireland. You cannot invoke a derogation that your governing Member State has not enacted.\n2. **They are necessity-bounded.** A derogation applies only so far as the right would render the research purpose impossible or seriously impair it. \"It would be inconvenient to re-run the analysis\" does not meet that bar.\n3. **Commercial product research is not the archetype.** The provision was written with scientific and statistical research in mind, and it comes bundled with Article 89(1) safeguards — data minimisation and, critically, using identifiable data only where the purpose cannot be achieved with anonymous or pseudonymous data.\n\nThat third clause is the useful one, and it points at the real answer.\n\n## Anonymisation is the off-switch\n\nArticle 11(2) is the most operationally valuable provision in this whole area: **where you can no longer identify a data subject, the access, rectification, erasure, restriction, and portability rights do not apply** — provided you can demonstrate you cannot identify them, and the requester cannot supply information that lets you.\n\nGenuinely anonymised research data is out of scope. Pseudonymised data — a participant code with a key table sitting in someone's spreadsheet — is still personal data, and every right still applies.\n\nThis is the same structural point that governs your retention schedule: anonymisation stops the clock. Building anonymisation into synthesis rather than treating it as a cleanup task means most of your historical corpus is simply not DSAR-addressable, which is both cheaper and better privacy practice. Our guide to [anonymising customer interview data](/docs/anonymizing-customer-interview-data) covers the techniques.\n\n## The seven-step runbook\n\n**1. Log receipt and start the clock the day it arrives.** A DSAR does not have to say \"DSAR\", does not have to be in writing, does not have to go to a privacy inbox, and does not have to cite a law. A reply to a research invitation saying \"actually, delete whatever you have on me\" is a valid erasure request. Brief whoever monitors your participant inbox.\n\n**2. Verify identity proportionately.** You must be reasonably satisfied the requester is who they claim. You must not use verification as a stalling tactic, and you must not collect more identity data than the request warrants. If your only link to a person is the email address they interviewed from, a reply from that address is usually sufficient.\n\n**3. Scope against a data map you wrote in advance.** Under time pressure is the wrong moment to discover where interview data lives. Your map should cover: the research platform, transcripts and audio, screener and recruitment records, the incentive payment trail, the analysis repository, exports sitting in a warehouse or BI tool, decks and documents containing verbatims, and every sub-processor in the chain.\n\n**4. Decide identifiability per artefact, not per study.** One study routinely contains raw identifiable transcripts, pseudonymised working notes, and genuinely anonymous aggregate findings. The first two are in scope; the third is not.\n\n**5. Handle third-party data in transcripts.** Participants name colleagues, managers, and customers. Those references are other people's personal data. For an access request, redact third-party identifiers unless you have that person's consent or it is reasonable to disclose without it. This is the single most common place research DSARs go wrong.\n\n**6. Propagate to sub-processors.** An erasure request reaches your processors and their sub-processors — transcription providers, model providers, storage, analytics. Your DPA should already oblige them to act on it. If you cannot name the chain, you cannot complete the request; see [reviewing a research vendor's DPA](/docs/research-vendor-dpa-review).\n\n**7. Respond in writing, and record what you did.** Say what you disclosed or deleted, what you withheld and why, and what remains. Keep the record — proving you handled a request correctly is itself an accountability obligation, and that record is one of the few things you keep after deleting everything else.\n\n## The hard cases\n\n**A quote already published in a report.** If the quote is attributed or reasonably identifiable, it is personal data and erasure reaches it. If it is genuinely anonymous in context, it is not. The lesson runs backwards into your process: attribute quotes to participant codes in decks, never to names and job titles at named employers, which are frequently identifying in combination.\n\n**Audio with more than one voice.** A recording containing a participant and a colleague contains two people's personal data. Deleting for one may mean deleting the artefact.\n\n**Aggregate analysis derived from their interview.** Once findings are aggregated to the point that no individual is identifiable, the analysis survives an erasure request. You do not have to re-run your study; you do have to remove the identifiable inputs.\n\n**Incentive payment records.** These usually have an independent legal basis — tax and accounting retention obligations — that survives an erasure request. Say so explicitly in your response rather than silently keeping them. See [research participant incentives and taxes](/docs/research-participant-incentive-taxes).\n\n## Where Koji helps\n\nKoji is built so that the awkward parts of a DSAR are lookups rather than archaeology. Interview data is organised per participant and per study rather than scattered across recordings, notes, and calendar entries, so scoping a request is a query instead of a search party. Data can be exported in structured form, which is what a portability request wants, and deleted when a participant asks.\n\nKoji's six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — help more than they look like they should. Structured responses carry stable question IDs from interview plan through to report, so a participant's quantitative answers are addressable fields rather than prose to be re-read. When you need to produce \"everything we hold about this person\" in a month, typed data is the difference between an afternoon and a fortnight.\n\nBecause Koji's AI interviewer runs asynchronously over a link, there is no third-party meeting recorder holding a second copy of the conversation — one fewer sub-processor to chase on an erasure request, and one fewer place for a copy to survive.\n\nFor enterprise data-handling arrangements, contact the Koji team.\n\n## Common mistakes\n\n- **Waiting for the word \"DSAR\".** Most requests arrive as ordinary sentences in ordinary replies.\n- **Assuming Article 89 covers commercial product research.** It is national-law dependent, necessity-bounded, and written for scientific research.\n- **Treating pseudonymised data as out of scope.** If a key exists anywhere, every right still applies.\n- **Forgetting the exports.** Warehouse tables, BI dashboards, and CSVs in someone's downloads folder are in scope.\n- **Disclosing third-party names in an access response.** Redact colleagues mentioned in transcripts.\n- **Missing the extension notice.** If you need the extra two months under GDPR, you must say so within the first month, with reasons.\n\n## Frequently asked questions\n\n**How long do I have to respond to a DSAR?**\nOne month from receipt under GDPR, extendable by two further months for complex or numerous requests if you notify the requester within the first month. Under CCPA/CPRA it is 45 calendar days, extendable by another 45 with notice.\n\n**Does a participant have to use the word \"DSAR\" or cite a law?**\nNo. A request is valid however it is phrased, whoever receives it, and whether or not it is in writing. Train anyone who monitors participant communications to recognise and escalate one.\n\n**Can I refuse deletion because the data is research data?**\nRarely. Article 89 derogations and the Article 17(3)(d) erasure carve-out depend on Member State implementation, are limited to what is necessary to avoid making the research impossible or seriously impaired, and were written with scientific research in mind. Do not assume they cover commercial product research.\n\n**Does an erasure request mean I have to redo my analysis?**\nNo. Genuinely anonymous aggregate findings are not personal data and survive the request. You must remove the identifiable inputs — transcript, audio, identifiers, and attributed quotes.\n\n**What about anonymised interview data?**\nIf you can demonstrate you can no longer identify the individual, Article 11(2) means the access, rectification, erasure, restriction, and portability rights do not apply. Pseudonymised data with a key does not qualify.\n\n**Do I have to delete incentive payment records?**\nUsually not. Tax and accounting retention obligations provide an independent legal basis that survives an erasure request. State this explicitly in your response instead of quietly retaining them.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and the addressable, typed data they produce\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the legal bases and participant rights underneath every DSAR\n- [CCPA/CPRA Compliance for Customer Research](/docs/ccpa-user-research-compliance) — the US route, its 45-day clock, and verification rules\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — the schedule that determines how much data a DSAR can even reach\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — the techniques that take data out of scope entirely\n- [Reviewing a Research Vendor DPA](/docs/research-vendor-dpa-review) — the contract clauses that make sub-processors act on your requests\n- [Research Participant Incentives and Taxes](/docs/research-participant-incentive-taxes) — why payment records outlive an erasure request","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"DSARs for Research Data: Access, Deletion, and Portability (2026)","metaDescription":"One month under GDPR, 45 days under CCPA. What counts as a participant's data in an interview study, why Article 89 rarely saves you, and a seven-step DSAR runbook.","keywords":["dsar research data","data subject access request research","participant deletion request","right to erasure interview data","gdpr article 89 research exemption","dsar runbook","ccpa 45 day response","data portability research","anonymisation article 11","participant data rights"],"aiSummary":"A DSAR gives you one month under GDPR (extendable by two with notice) or 45 days under CCPA/CPRA. For research teams the difficulty is scoping what counts as a participant's data across transcripts, audio, screeners, incentive records, repositories, exports, and verbatims in decks. The Article 89 research derogation rarely helps commercial product research because it depends on Member State implementation, is necessity-bounded, and targets scientific research. The genuinely useful provision is Article 11(2): where you can no longer identify a data subject, most rights do not apply, so anonymisation takes data out of scope. Includes a seven-step runbook and the hard cases: published quotes, multi-voice audio, aggregate analysis, and third-party data in transcripts.","aiDifficulty":"intermediate","aiEstimatedTime":"12 min"},{"type":"documentation","id":"fa8ad1c4-4c2a-4b77-a148-824c0a9c4029","slug":"consumer-duty-customer-research","title":"FCA Consumer Duty Customer Research: How to Evidence Consumer Understanding and Fair Value","url":"https://www.koji.so/docs/consumer-duty-customer-research","summary":"The FCA Consumer Duty (PRIN 2A, effective 31 July 2023 for open products and 31 July 2024 for closed products) requires firms to evidence that customers can make informed decisions. Its 13 March 2026 consumer understanding review named weak evidence of communication testing, inaccessible or overly complex communications, insufficient consideration of diverse customer needs, weak monitoring and unclear accountability as areas for improvement, and stated that reliance on sales data or the absence of complaints provides no reliable assurance. Good practice included analysing insight from call listening, complaints, chat transcripts, analytics, drop-off data and surveys; testing communications with real customers before and after changes using short surveys, comprehension checks, callbacks and A/B testing; documenting what changed, why and the effect; and setting comprehension targets such as at least 80% correct recall of key points and redrafting until met. The 16 April 2026 board-report observations asked firms to move from MI dashboards to analysis that draws conclusions, document board challenge, monitor distribution-chain outcomes and deepen consumer understanding and support evidence. A comprehension test defines key points, sets pass criteria in advance, presents the real artefact, tests recall with single_choice items, captures self-rated clarity with scale items, probes wrong answers with open_ended items and AI follow-up, then redrafts and retests. Evidence should be split by segment and vulnerability characteristic, offer voice and text participation, and retain transcripts and exports for lineage.","content":"**Answer first:** The FCA's Consumer Duty does not ask whether your communications were sent. It asks whether customers *understood* them, and it expects you to have tested that with real customers and documented the result. In its 13 March 2026 review of the consumer understanding outcome, the FCA identified weak evidence of communication testing as a leading area for improvement and stated plainly that firms relying on sales data or the absence of complaints have no reliable assurance of understanding. One month later, in its 16 April 2026 observations on second-year board reports, it told firms to move beyond management-information dashboards to analysis that draws conclusions — and specifically to deepen the evidence on consumer understanding and support. In practice this makes structured comprehension testing a recurring research obligation for every UK regulated firm, not a project.\n\n## What the Duty demands, evidentially\n\nThe Consumer Duty (PRIN 2A) took effect on 31 July 2023 for new and existing open products and 31 July 2024 for closed products. It sets a consumer principle, three cross-cutting rules — act in good faith, avoid causing foreseeable harm, and enable and support customers to pursue their financial objectives — and four outcomes. Each outcome implies a different research question:\n\n| Outcome | The regulatory question | What research has to produce |\n|---|---|---|\n| **Products and services** | Does the product meet the needs of the identified target market? | Evidence that real customers in that segment have the need, and that distribution is reaching them |\n| **Price and value** | Is the price reasonable relative to the benefits? | Evidence of what customers believe they are buying and which benefits they actually value or never use |\n| **Consumer understanding** | Can customers make informed decisions? | Tested comprehension of the actual communications, with pass criteria and a record of redrafting |\n| **Consumer support** | Can customers realise the benefits and act without unreasonable friction? | Evidence on where support journeys stall, for whom, and what the customer did next |\n\nNotice what all four have in common: the evidence has to come from customers, not from process documentation. That is the shift most firms have still not fully made.\n\n## The 2026 reviews: what the FCA praised and criticised\n\n**13 March 2026 — consumer understanding review.** The FCA published findings on how firms approach the consumer understanding outcome, split into good practice and areas for improvement.\n\n| Good practice observed | Areas for improvement |\n|---|---|\n| Analysing insight from multiple sources — call listening, complaints, chat transcripts, website analytics, drop-off data and surveys | Weak evidence of communication testing |\n| Testing communications with real customers both before and after changes, using proportionate methods including short surveys, comprehension checks, callbacks, A/B testing and feedback during digital trials | Inaccessible or overly complex communications |\n| Documenting what changed, why, and what effect it had | Insufficient consideration of diverse customer needs |\n| Setting explicit comprehension targets — in one case at least 80% correct recall of key points — and redrafting until the target was met | Weak monitoring, and unclear accountability for who decides what and how |\n\n**16 April 2026 — board report observations.** Reviewing second-year reports, the FCA asked firms to move past MI dashboards to analysis that draws conclusions, to document board challenge, to monitor outcomes delivered through distribution chains, and to deepen the evidence on consumer understanding and support.\n\nThe pattern across both is consistent: the regulator is no longer satisfied by activity. It wants a testable claim, the test, the result, and the decision that followed.\n\n## How to actually run a comprehension test\n\nA comprehension test is not a satisfaction survey. You are not asking whether customers liked the letter; you are measuring whether they can correctly answer questions about what it means for them.\n\n**1. Choose the communication and the key points.** Take the real artefact — the annual statement, the renewal notice, the arrears letter, the fee disclosure — and list the three to six key points a customer must take away. If you cannot list them, the communication has no defined purpose and that is the first finding.\n\n**2. Define pass criteria before fielding.** The FCA highlighted a firm using at least 80% correct recall of key points. Pick your threshold, write it down, and treat it as a gate rather than a metric to report.\n\n**3. Show the artefact, then test.** Present the actual text or screen, then ask closed questions with one correct answer per key point, plus open questions on interpretation.\n\n**4. Probe the wrong answers.** The score tells you the communication failed; only the reasoning tells you which clause caused it. This is where a static survey stops and an interview keeps going.\n\n**5. Redraft and retest.** The evidence the FCA wants is the *pair*: the before result, the change, the after result. One-shot testing produces a number, not a demonstration of improvement.\n\n**6. Record the decision trail.** Who approved the redraft, on what evidence, and when. Unclear accountability was named as an area for improvement.\n\n### Mapping the test to structured question types\n\nKoji's six [structured question types](/docs/structured-questions-guide) let one instrument produce the comprehension score and the diagnosis together:\n\n| Element | Question type | Purpose |\n|---|---|---|\n| Key-point recall | `single_choice` | One correct option per key point — this is what generates the pass rate |\n| Action comprehension | `yes_no` | \"Based on this letter, do you need to do anything before 30 September?\" |\n| Confidence | `scale` | Self-rated clarity, which usefully diverges from actual score and exposes false confidence |\n| Perceived relevance of terms | `multiple_choice` | Which listed features the customer believes apply to them |\n| Priority of information | `ranking` | Which points customers think matter most, versus which you intended to foreground |\n| Interpretation | `open_ended` | Their reading in their own words, with AI follow-up probing on any incorrect answer |\n\nThe last row is the differentiator. Platforms like Koji ask the follow-up automatically: when a customer selects the wrong answer about a fee, the AI interviewer asks what they thought the fee covered and where in the document they looked. Ten minutes of conversation produces both the 72%-versus-80% pass rate and the sentence that tells the drafting team which clause to rewrite. A survey tool gives you the first and leaves you guessing at the second; a moderated session gives you both but at twenty times the cost per participant and only during business hours.\n\nThe mechanics of comprehension and wording tests generalise beyond financial services — see [content testing](/docs/content-testing-guide) for the underlying method.\n\n## Diverse customer needs and vulnerability\n\n\"Insufficient consideration of diverse customer needs\" was one of the FCA's named improvement areas, and it is the hardest to evidence with conventional research, because the customers least likely to complete a twelve-minute online survey are precisely the ones the Duty is most concerned with.\n\nPractical steps:\n\n- **Sample deliberately for characteristics of vulnerability** — health conditions, low resilience, low capability, negative life events — rather than hoping a general panel covers them. See [researching hard-to-reach audiences](/docs/hard-to-reach-participants-research).\n- **Offer voice as well as text.** [Voice interviews](/docs/ai-voice-interviews) remove the reading and typing burden, which matters for customers with low literacy, dyslexia, visual impairment, or motor difficulty. Offering both modes on the same study is itself evidence of accommodating diverse needs.\n- **Make participation async.** Fixed appointment slots systematically exclude shift workers, carers and people managing health conditions.\n- **Report outcomes split by vulnerability characteristic.** The Duty asks whether outcomes differ between groups. An aggregate pass rate hides exactly the disparity the regulator is looking for.\n- **Follow accessible-research practice throughout.** See [accessibility research](/docs/accessibility-research-guide) and, for the wider compliance picture, [accessibility compliance research](/docs/accessibility-compliance-research).\n\n## Evidence for price and value, and for consumer support\n\n**Price and value.** Fair value assessments are usually built from cost, margin and benchmark data. What they typically lack is the customer's side: which benefits customers know they have, which they use, and which they would not miss. A short annual study using `multiple_choice` for feature awareness, `ranking` for value ordering, and `open_ended` for the reasoning gives your assessment a customer-evidence limb it probably does not have — and directly addresses whether customers understand what they are paying for.\n\n**Consumer support.** The support outcome fails at friction points, not in aggregate satisfaction. Interview customers immediately after a support interaction, a claim, a cancellation attempt, or a failed self-service journey, and ask what they were trying to achieve, where they stopped, and what they did next. Triggering an interview from the event is what makes this feasible: Koji studies can be launched from [webhooks](/docs/research-automation-webhooks) or a CRM record so the interview arrives while the experience is fresh, rather than in the next quarterly wave.\n\n## The board-report evidence pack\n\nGiven the April 2026 observations, a defensible pack for the consumer understanding and support outcomes contains:\n\n1. **The inventory** — which communications and journeys were tested this period, and which were not, with a reason.\n2. **The pass criteria and results** — before-and-after comprehension rates against a stated threshold, split by segment and by vulnerability characteristic.\n3. **The verbatim evidence** — customer quotes showing *how* a communication was misread. This is what turns a dashboard into analysis.\n4. **The decision trail** — what was redrafted, by whose decision, on what evidence, and what the retest showed.\n5. **The residual risk** — communications that still fail the threshold, with owner and remediation date.\n6. **Distribution-chain evidence** — outcomes for customers acquired through intermediaries, not just direct.\n7. **Method and data lineage** — instrument, sample, dates, and where raw responses are retained. See [research data retention and deletion](/docs/research-data-retention-deletion).\n\nKoji supports the lineage requirement directly: every interview retains its full transcript alongside the structured answers, and [data export](/docs/exporting-research-data) in CSV or JSON gives you an auditable record to attach to the assessment. Composite [quality scores](/docs/understanding-quality-scores) from 1 to 5 let you show that low-effort responses were excluded from the pass-rate calculation rather than quietly diluting it.\n\nOn data handling: interviews are async and link-based, so no third-party meeting recorder sits in the chain, and a data processing agreement is available to business customers. For firm-specific requirements on hosting, retention or sub-processors — which any FCA-regulated firm should put to every vendor in writing — contact the Koji team, and run a [DPIA](/docs/dpia-user-research) covering the processing before you field a study on vulnerable customers.\n\n## Cadence and anti-patterns\n\nTest before a change and after it, then monitor continuously. Annual-only testing produces evidence that is eleven months stale for most of the year, and the FCA specifically criticised weak monitoring. Set a refresh trigger on any communication that is materially redrafted, any product change, and any spike in complaints or drop-off — see [insight decay and when to re-run a study](/docs/research-refresh-cadence).\n\nAvoid these:\n\n- **Treating sales volume or complaint absence as evidence of understanding.** Named explicitly by the FCA as providing no reliable assurance.\n- **Testing satisfaction instead of comprehension.** \"Was this letter clear?\" is not a comprehension test; customers routinely rate unclear documents as clear.\n- **Testing with staff or a general panel.** Comprehension is population-specific. Test with holders of that product.\n- **Reporting an aggregate pass rate only.** Split by segment and vulnerability characteristic or you have hidden the finding.\n- **Stopping at the score.** Without the reasoning behind wrong answers, you cannot fix the wording — and the retest will fail too.\n- **No documented owner.** Unclear accountability was one of the FCA's named improvement areas.\n\n## Frequently asked questions\n\n**Does the Consumer Duty require customer research?**\nIt does not name research as an activity, but it requires firms to evidence that customers can make informed decisions and are achieving good outcomes. In its 13 March 2026 review the FCA identified weak evidence of communication testing as an area for improvement and said reliance on sales data or an absence of complaints gives no reliable assurance — so in practice, testing with real customers is how the consumer understanding outcome gets evidenced.\n\n**What comprehension target should we set?**\nThere is no prescribed figure. The FCA's review highlighted a firm using at least 80% correct recall of key points and redrafting until it was met. What matters is that you set a threshold in advance, apply it as a gate, and record the before-and-after result rather than reporting a single unbenchmarked score.\n\n**How is a comprehension test different from a customer satisfaction survey?**\nA satisfaction survey asks whether customers found a communication clear; a comprehension test measures whether they can correctly answer questions about what it means for them. The two frequently disagree, which is why self-rated clarity should be captured alongside actual recall rather than instead of it.\n\n**How do we evidence consideration of diverse customer needs?**\nSample deliberately for characteristics of vulnerability rather than relying on a general panel, offer both voice and text participation, keep the study async so it does not exclude shift workers and carers, and report results split by vulnerability characteristic. An aggregate pass rate conceals exactly the disparity the Duty asks about.\n\n**How often should communications be tested?**\nBefore and after any material change, plus continuous monitoring on live communications, with a refresh triggered by product changes or spikes in complaints and drop-off. The FCA criticised weak monitoring in 2026, and annual-only testing leaves your evidence stale for most of the year.\n\n**What should go in the board report on consumer understanding?**\nA tested-communications inventory with reasons for any omissions, stated pass criteria and before-and-after results split by segment and vulnerability, verbatim evidence showing how communications were misread, the decision trail for each redraft, residual risks with owners and dates, distribution-chain outcomes, and the method and data lineage. The April 2026 observations asked firms to move from dashboards to analysis that draws conclusions.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how each aggregates into a report\n- [Content Testing](/docs/content-testing-guide) — the underlying method for testing wording and comprehension\n- [AI Customer Research for Banking and Financial Services](/docs/ai-research-for-banking) — sector playbook\n- [AI-Powered Customer Research for Insurance Companies](/docs/ai-research-for-insurance) — renewal, claims and fair value research\n- [Accessibility Research](/docs/accessibility-research-guide) — including customers with disabilities in your testing\n- [DPIA for User Research](/docs/dpia-user-research) — the assessment to complete before researching vulnerable customers\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — retention and lineage for audit evidence\n- [Insight Decay and When to Re-Run a Study](/docs/research-refresh-cadence) — setting a monitoring cadence","category":"Research Operations","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"FCA Consumer Duty Customer Research: Evidencing Consumer Understanding","metaDescription":"The FCA says sales data and absent complaints prove nothing about consumer understanding. How to run comprehension tests, hit an 80% target, and build a board-report evidence pack.","keywords":["consumer duty customer research","FCA consumer understanding","consumer understanding outcome testing","comprehension testing financial services","fair value assessment research","vulnerable customer research","consumer duty board report evidence","communication testing FCA"],"aiSummary":"The FCA Consumer Duty (PRIN 2A, effective 31 July 2023 for open products and 31 July 2024 for closed products) requires firms to evidence that customers can make informed decisions. Its 13 March 2026 consumer understanding review named weak evidence of communication testing, inaccessible or overly complex communications, insufficient consideration of diverse customer needs, weak monitoring and unclear accountability as areas for improvement, and stated that reliance on sales data or the absence of complaints provides no reliable assurance. Good practice included analysing insight from call listening, complaints, chat transcripts, analytics, drop-off data and surveys; testing communications with real customers before and after changes using short surveys, comprehension checks, callbacks and A/B testing; documenting what changed, why and the effect; and setting comprehension targets such as at least 80% correct recall of key points and redrafting until met. The 16 April 2026 board-report observations asked firms to move from MI dashboards to analysis that draws conclusions, document board challenge, monitor distribution-chain outcomes and deepen consumer understanding and support evidence. A comprehension test defines key points, sets pass criteria in advance, presents the real artefact, tests recall with single_choice items, captures self-rated clarity with scale items, probes wrong answers with open_ended items and AI follow-up, then redrafts and retests. Evidence should be split by segment and vulnerability characteristic, offer voice and text participation, and retain transcripts and exports for lineage.","aiPrerequisites":["A UK FCA-regulated firm subject to the Consumer Duty (PRIN 2A)","The actual customer communications or journeys you intend to test"],"aiLearningOutcomes":["Map the four Consumer Duty outcomes to specific research questions","Run a comprehension test with pre-set pass criteria rather than a satisfaction survey","Apply the FCA March 2026 good-practice and improvement findings to your testing programme","Evidence consideration of diverse customer needs and vulnerability characteristics","Assemble a board-report evidence pack that moves from dashboards to analysis","Set a before-and-after testing cadence with monitoring triggers"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"c490c421-ae61-4676-9a0f-2021401b4a9d","slug":"patient-experience-survey-guide","title":"How to Build Patient Experience Surveys That Improve Care Quality","url":"https://www.koji.so/docs/patient-experience-survey-guide","summary":"Comprehensive guide to patient experience surveys for healthcare. Covers HCAHPS alignment, study design, sensitivity considerations, accessibility, and how Koji produces higher response rates and deeper feedback than traditional healthcare survey instruments.","content":"# How to Build Patient Experience Surveys That Improve Care Quality\n\nPatient experience is now a core quality measure, financial driver, and competitive differentiator in healthcare. In the US, HCAHPS scores directly impact Medicare reimbursement. Globally, patient experience scores correlate with clinical outcomes, staff retention, and institutional reputation.\n\nYet healthcare faces a unique challenge: patients are vulnerable. They may be anxious, in pain, or navigating complex medical situations. Standard survey instruments feel impersonal and clinical. Response rates for mail-based HCAHPS surveys average just 25-30%. The patients who do respond skew toward the extremes, creating data that's noisy and non-representative.\n\nKoji addresses this by offering a conversational approach that feels more like speaking with a compassionate listener than filling out a government-mandated form. The result: higher response rates, more nuanced feedback, and insights that go beyond standardized metrics.\n\n## Patient Experience Measurement Frameworks\n\n### HCAHPS (Hospital Consumer Assessment of Healthcare Providers and Systems)\nThe US national standard. 29 questions covering:\n- Communication with nurses and doctors\n- Responsiveness of hospital staff\n- Cleanliness and noise levels\n- Pain management\n- Communication about medications\n- Discharge information\n- Overall rating and recommendation\n\n### NHS Friends and Family Test\nUK standard. One question: \"How likely are you to recommend this service?\" with free-text follow-up.\n\n### CAHPS (Consumer Assessment of Healthcare Providers and Systems)\nFamily of surveys for different settings: outpatient, home health, hospice, health plans.\n\n### Press Ganey / NRC Health\nProprietary instruments used by many healthcare systems for benchmarking.\n\n## Building Patient Experience Studies with Koji\n\n### Core Patient Experience Study\n\n**Q1: Overall Experience (Scale, 0-10)**\n\"Overall, how would you rate your experience with us?\"\n- Labels: 0 = \"Worst possible\", 10 = \"Best possible\"\n- Anchor probing: \"What most influenced your rating?\"\n\n**Q2: Communication Quality (Scale, 1-5)**\n\"How well did your care team communicate with you about your condition and treatment?\"\n- Labels: 1 = \"Very poorly\", 5 = \"Very well\"\n- Probing depth: 2\n- AI instruction: \"Explore specific communication moments. Were explanations clear? Did they feel heard? Were questions welcomed?\"\n\n**Q3: Respect and Dignity (Scale, 1-5)**\n\"How well were you treated with courtesy and respect?\"\n- Probing on low scores: AI sensitively explores specific incidents\n\n**Q4: Care Coordination (Open-ended)**\n\"How smoothly did the different parts of your care work together?\"\n- Probing depth: 2\n- AI explores handoffs between departments, staff, and settings\n\n**Q5: Information and Education (Open-ended)**\n\"Did you receive clear information about your condition, treatment options, and what to expect?\"\n- Probing depth: 2\n- Captures health literacy gaps and education opportunities\n\n**Q6: Emotional Support (Scale, 1-5)**\n\"How supported did you feel emotionally during your care?\"\n- Probing: \"Was there a moment where you felt particularly supported or unsupported?\"\n\n**Q7: Environment (Single Choice)**\n\"How would you describe the cleanliness and comfort of the facility?\"\n- Options: Excellent / Good / Acceptable / Poor\n- Probing on negative responses\n\n**Q8: Wait Time (Scale, 1-5)**\n\"How reasonable were the wait times you experienced?\"\n- Probing: \"Where did you experience the longest wait? How was that communicated?\"\n\n**Q9: Discharge/Follow-up (Open-ended)**\n\"How clear were the instructions you received about what to do after your visit/discharge?\"\n- Probing depth: 2\n- Critical for readmission prevention\n\n**Q10: Recommendation (Yes/No)**\n\"Would you recommend this facility to a friend or family member?\"\n- Probing: \"Why or why not?\"\n\n### Specialized Studies\n\n**Pre-Procedure Anxiety Study:**\nInterview patients before procedures to understand anxiety levels, information needs, and preparation quality.\n\n**Post-Discharge Follow-Up:**\nInterview 48-72 hours after discharge to assess transition quality, medication understanding, and early recovery experience.\n\n**Chronic Care Experience:**\nOngoing conversational check-ins with chronic disease patients about their care journey, medication management, and quality of life.\n\n## Why Conversational AI Works for Patient Feedback\n\n### Sensitivity\nHealthcare feedback is inherently emotional. Patients may feel vulnerable discussing their care. Koji's AI is trained to be empathetic, patient, and non-judgmental, creating a safe space for honest feedback.\n\n### Accessibility\nPatients with limited literacy, visual impairments, or language barriers struggle with paper/online forms. Koji's voice interview option and multi-language support (30+ languages) make feedback accessible to all patients.\n\n### Depth\n\"Rate your communication with nurses: 1-5\" tells you nothing actionable. \"Tell me about your interactions with the nursing staff\" captures specific, improvement-oriented feedback that can be used in training and process improvement.\n\n### Timing\nPaper HCAHPS surveys arrive weeks after discharge. Koji interviews can be triggered within hours, capturing experience details while they're fresh.\n\n## Analysis and Improvement\n\n### What Koji Reports Generate\n\n- **Experience dimension scores** aligned with HCAHPS categories\n- **Sentiment analysis** capturing emotional tone across care dimensions\n- **Department-level comparison** showing which units excel and which need support\n- **Specific improvement opportunities** with patient quotes\n- **Communication gap analysis** between provider intent and patient perception\n- **Wait time impact analysis** correlating wait experience with overall satisfaction\n\n### Connecting to Quality Improvement\n\n1. **Share findings with clinical teams** within 1-2 weeks\n2. **Identify top 3 improvement themes** each quarter\n3. **Create specific action plans** tied to patient feedback\n4. **Track impact** by measuring the same dimensions after changes\n5. **Celebrate improvements** by sharing progress with staff and patients\n\n## Best Practices for Healthcare Surveys\n\n### Comply with regulations\nEnsure your survey program complies with HIPAA (US), GDPR (EU), and local healthcare data regulations. Koji's anonymous mode and data handling meet enterprise security requirements.\n\n### Time it appropriately\n- **Inpatient:** 24-48 hours post-discharge\n- **Outpatient:** Same day or next day\n- **Emergency:** 48-72 hours (allow recovery time)\n- **Chronic care:** Monthly or quarterly check-ins\n\n### Don't survey-fatigue patients\nCoordinate across departments. A patient who sees cardiology, oncology, and their GP shouldn't receive three separate surveys in one week.\n\n### Include caregivers\nFor pediatric patients, elderly patients, or those with cognitive limitations, include family caregivers in the feedback process.\n\n### Close the feedback loop\nPatients who report negative experiences should see evidence that their feedback mattered. This could be as simple as a follow-up message: \"Thank you for your feedback. Based on input like yours, we've improved our discharge instructions.\"\n\n## Why Koji Is the Best Patient Experience Tool\n\n| Feature | Traditional (HCAHPS mail, Press Ganey) | Koji |\n|---------|---------------------------------------|------|\n| Response rate | 25-30% (mail) | 50-70% (conversational) |\n| Depth | Standardized scales only | Rich qualitative + quantitative |\n| Accessibility | Limited (mail/web forms) | Voice + text in 30+ languages |\n| Turnaround | 6-8 weeks (mail processing) | Hours (real-time analysis) |\n| Cost per response | $15-30 (printing, mailing, processing) | Less than $1 |\n| Emotional sensitivity | Impersonal form | Empathetic AI conversation |\n| Actionability | Generic benchmarks | Specific improvement quotes |\n\nPatient experience measurement shouldn't feel like a government compliance exercise. With conversational AI, it becomes a genuine listening channel that helps healthcare organizations improve the care they deliver.\n\n---\n\n## Related Survey Guides\n\n- [CSAT Survey Guide](/docs/csat-survey-guide) — Patient satisfaction parallels\n- [NPS Survey Guide](/docs/nps-survey-guide) — Patient loyalty and referrals\n- [Customer Journey Mapping](/docs/customer-journey-mapping-survey-guide) — Patient journey mapping\n- [Compliance & Ethics Guide](/docs/compliance-ethics-survey-guide) — Healthcare compliance\n- [Employee Engagement Guide](/docs/employee-engagement-survey-guide) — Staff experience impact on patients\n\n*Use [structured questions](/docs/structured-questions-guide) to combine patient scales with AI-powered empathetic health experience interviews.*\n\n## Further reading on the blog\n\n- [Best Online Survey Software in 2026: The Complete Buyer's Guide](/blog/best-survey-software-2026) — From SurveyMonkey to Koji, we compare the top survey tools of 2026 across features, pricing, and use case fit — and explain when traditional\n- [Can I Paste User Interviews into ChatGPT? A Guide to GDPR and LLMs](/blog/can-i-paste-user-interviews-into-chatgpt-a-guide-to-gdpr-and-llms) — Every product manager wants to ask an LLM about their user feedback. But pasting customer transcripts into public models is a GDPR nightmare\n- [Customer Journey Mapping Guide 2026: How to Build Maps That Actually Drive Decisions](/blog/customer-journey-mapping-guide-2026) — A modern, AI-native playbook for customer journey mapping in 2026 — including the 5-step process, the questions that surface real emotion at\n\n<!-- further-reading:blog -->\n","category":"Survey & Study Templates","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"Patient Experience Survey Guide: Build Healthcare Surveys That Improve Care | Koji","metaDescription":"Complete guide to patient experience surveys. Learn how to design HCAHPS-aligned surveys with conversational AI that produces higher response rates and more actionable feedback than traditional healthcare survey instruments.","keywords":["patient experience survey","patient satisfaction survey","HCAHPS survey","healthcare survey","patient feedback","hospital survey","patient experience template","healthcare quality survey"],"aiSummary":"Comprehensive guide to patient experience surveys for healthcare. Covers HCAHPS alignment, study design, sensitivity considerations, accessibility, and how Koji produces higher response rates and deeper feedback than traditional healthcare survey instruments."},{"type":"documentation","id":"cc3b67dd-d9a2-4be4-b852-458df941eb83","slug":"ai-research-for-banking","title":"AI Customer Research for Banking & Financial Services","url":"https://www.koji.so/docs/ai-research-for-banking","summary":"Retail banks, credit unions, and wealth firms use Koji's AI-moderated voice and text interviews to understand the why behind transactional data — onboarding drop-off, primary-bank attrition, trust, channel preferences, and product fit. Koji fields studies in days, probes every answer automatically, supports six structured question types for tracked metrics, and produces real-time reports. Compliance practices include data minimization, neutral questioning, complaint escalation paths, AES-256/TLS 1.2+ encryption, a DPA, and in-flow consent. Distinct from the fintech and insurance guides.","content":"## The Bottom Line\n\nBanks, credit unions, and wealth management firms sit on mountains of transactional data but rarely understand the *why* behind it — why a customer abandoned an account application, why they keep most of their money elsewhere, what would earn their trust with a new product. Koji lets financial services teams run AI-moderated voice and text interviews with customers and prospects, fielded in days, with automatic analysis that turns hundreds of conversations into clear themes. The result is the qualitative depth of customer interviews at the scale and speed a regulated, multi-segment institution needs — roughly 10x faster than booking moderated sessions, without the cost of a research agency for every study.\n\nThis guide covers where AI interviews fit in banking, how to design studies for a regulated environment, and what makes Koji a fit. (For fintech startups and embedded-finance products, see the fintech guide; for carriers, see the insurance guide — this article focuses on banks, credit unions, and wealth/investment firms.)\n\n## Where AI interviews fit in financial services\n\nFinancial institutions face research questions across the entire customer relationship:\n\n- **Account opening and onboarding** — Where do applicants drop off, and why? What friction or trust concern stops them?\n- **Primary-bank status** — Why do customers keep their direct deposit and balances with you, or with a competitor? What would make you their primary institution?\n- **Digital experience** — How do customers actually use the mobile app and online banking, and where does it frustrate them?\n- **Trust and security perception** — How do customers feel about fraud protection, data use, and AI-driven features? Trust is the core currency of banking.\n- **Product fit** — Will customers adopt a new card, savings product, lending option, or advisory service, and what would justify switching?\n- **Branch and channel preferences** — How do segments differ in their preference for branch, phone, chat, and self-service?\n- **Wealth and advisory** — What do clients value in an advisor relationship, and where are robo and human models winning or losing?\n\nAcross all of these, the institutions that win are the ones that understand customer motivation — not just behavior.\n\n## Why surveys and panels fall short\n\nMost banks default to NPS surveys and occasional panels. Both have limits:\n\n- **Surveys** capture a score but not the reason. An NPS detractor box tells you a customer is unhappy; it never asks the follow-up that explains why or what would fix it.\n- **Moderated research** gives depth but is slow and expensive, so it gets reserved for the biggest initiatives and skips the everyday decisions.\n- **Online panels** can be unrepresentative and prone to low-quality responses.\n\nThe gap is a method that asks the follow-up question — at scale, fast, and across every customer segment. That is what AI-moderated interviews provide.\n\n## How Koji works for banking teams\n\n**AI-moderated voice and text interviews.** Write a brief and Koji conducts the interview, asking your questions and generating intelligent follow-ups. A customer can answer by voice from their phone or by text on their own schedule — no moderator, no scheduling, no agency fee per study.\n\n**Automatic follow-up probing.** When a customer says \"I just trust my old bank more,\" Koji probes what trust means to them, what would change it, and what specifically gives the competitor an edge. This is the laddering that turns a vague sentiment into an actionable insight.\n\n**Structured questions for tracking metrics.** Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Track a satisfaction or trust scale, a single_choice on primary-bank status, and a ranking of decision factors (rates, fees, app quality, branch access, trust) alongside open-ended narrative — all in one interview. Reports visualize scales as distributions and choices as frequency charts, so you get a quant dashboard and qual story together. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Real-time, analysis-ready reports.** Koji synthesizes themes, surfaces representative quotes, and quantifies structured answers as interviews arrive — so a product or CX team reads insights the same week.\n\n**Quality gate.** Only conversations scoring 3 or higher count toward your plan, protecting your dataset from low-effort responses — useful when running incentivized studies at scale.\n\n## Designing studies for a regulated environment\n\nFinancial services research carries compliance weight. A few practices:\n\n1. **Protect personal and financial data.** Collect only what the study needs — use structured questions and screeners to avoid gathering account numbers or sensitive identifiers. Koji encrypts data in transit (TLS 1.2+) and at rest (AES-256), offers a DPA, and supports anonymization and retention controls.\n2. **Mind fair-treatment and complaint-handling rules.** If a research conversation surfaces a complaint or potential harm, have an escalation path — coordinate study design with your compliance team so issues route correctly.\n3. **Keep questions neutral.** Avoid anything that resembles a sales pitch or could imply a guarantee. Koji's probing follows the respondent rather than steering them, which supports clean, unbiased data.\n4. **Capture consent in the flow.** Because interviews are link-based and asynchronous, consent and disclosures appear at the start, creating an auditable trail.\n5. **Segment deliberately.** Banking customers differ sharply by life stage, balance tier, and channel preference. Use screeners to interview each segment and compare with structured questions.\n\n## A practical example\n\nA regional bank wants to fix a high drop-off in its online account-opening flow. With Koji they would:\n\n- Recruit recent applicants — both completers and abandoners — and screen by segment.\n- Build a brief with open-ended questions on the application experience, a scale on how easy it felt, a single_choice on where they stopped, and a ranking of what would have helped most.\n- Let Koji run voice and text interviews and probe every point of friction automatically.\n- Read a real-time report that quantifies drop-off points and clusters the open-ended frustrations into prioritized themes with quotes.\n\nA study that an agency would scope over six weeks becomes a few days of fielding and same-week synthesis the product team can act on.\n\n## Getting started\n\nStart with one costly, well-defined problem — onboarding drop-off, primary-bank attrition, or trust in a new digital feature — and design a short, segmented study around it. Lean on structured questions for the metrics you will track over time and open-ended questions plus AI probing for the why. Bring compliance in early, and you will have a fast, repeatable research engine that keeps pace with how quickly banking customer expectations now change.\n\n## From one study to an always-on program\n\nThe highest-performing financial institutions do not treat research as a once-a-year project. They build a continuous signal:\n\n- **Trigger interviews from key moments.** Invite customers to a short interview right after onboarding, a failed application, a fraud flag, or a product upgrade — when the experience is fresh and the feedback is most actionable.\n- **Standardize tracked metrics.** Reuse the same scale and single_choice questions (trust, satisfaction, primary-bank status) across studies so you can trend them over quarters and across segments.\n- **Close the loop with frontline teams.** Route themes to the product, CX, and branch teams that own them, and confirm changes back to customers — the feedback loop that turns research into loyalty.\n- **Segment continuously.** Keep separate study streams for retail, wealth, small business, and digital-only customers so each segment's distinct needs stay visible rather than averaging out.\n\nBecause Koji interviews are AI-moderated and fast to field, an always-on program is realistic for a lean team — you get a steady stream of qualitative signal without standing up a full research department.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types for mixed-methods banking studies\n- [AI Research for Fintech](/docs/ai-research-for-fintech) — guidance for fintech and embedded-finance products\n- [AI Research for Insurance](/docs/ai-research-for-insurance) — adjacent guidance for carriers\n- [Voice of Customer Research Program](/docs/voice-of-customer-research-program) — always-on listening across segments\n- [Customer Journey Mapping Guide](/docs/customer-journey-mapping) — mapping onboarding and channel journeys\n- [Generating Research Reports](/docs/generating-research-reports) — turning interviews into analysis-ready reports","category":"Use Cases","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"AI Customer Research for Banking & Financial Services | Koji","metaDescription":"How retail banks, credit unions, and wealth firms use AI voice and text interviews to understand onboarding friction, trust, channel preferences, and product fit — at scale and in days.","keywords":["banking customer research","financial services research","ai interviews banking","credit union research","wealth management research","bank onboarding research","customer experience banking","retail banking insights","financial services voice of customer","bank customer interviews"],"aiSummary":"Retail banks, credit unions, and wealth firms use Koji's AI-moderated voice and text interviews to understand the why behind transactional data — onboarding drop-off, primary-bank attrition, trust, channel preferences, and product fit. Koji fields studies in days, probes every answer automatically, supports six structured question types for tracked metrics, and produces real-time reports. Compliance practices include data minimization, neutral questioning, complaint escalation paths, AES-256/TLS 1.2+ encryption, a DPA, and in-flow consent. Distinct from the fintech and insurance guides.","aiPrerequisites":["Familiarity with financial services customer experience","Awareness of your institution's compliance and complaint-handling requirements"],"aiLearningOutcomes":["Identify where AI interviews fit across the banking customer relationship","Design segmented, mixed-methods banking studies with structured questions","Apply compliance-aware practices for a regulated environment","Cut research time-to-insight from weeks to days"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 minutes"},{"type":"documentation","id":"e4642de1-8d5d-47bd-b185-89fa94a1971d","slug":"ai-research-for-insurance","title":"AI-Powered Customer Research for Insurance Companies (2026)","url":"https://www.koji.so/docs/ai-research-for-insurance","summary":"Insurers can run policyholder satisfaction, claims-experience, retention, product concept, and pricing research at scale with AI interviews. Koji conducts voice or text conversations with policyholders, probes claims stories with adaptive follow-ups, captures NPS and CSAT with scale questions and coverage preferences with choice and ranking questions, and aggregates everything into a compliance-aware research report - no moderator or panel agency required.","content":"Insurance is one of the hardest industries to research well - and one with the most to gain from doing it. Policyholders only think about their insurer at a few emotionally charged moments (buying, renewing, and claiming), engagement with traditional surveys is low, and the journeys are long and regulated. AI interviews change the economics: with a platform like Koji, an insurer can conduct hundreds of voice or text conversations with policyholders, probe the *story* behind every claim or cancellation, and get an automatically analyzed report - without a panel agency, a moderator, or a six-week timeline.\n\nThis guide covers the highest-value research use cases across the insurance lifecycle and how to run each one.\n\n## Why Insurance Research Is Different\n\nFour things make insurance research uniquely challenging:\n\n- **Emotionally charged moments.** A claim is often filed during a stressful life event. A static survey cannot read tone or follow up on a painful detail; a conversation can.\n- **Low baseline engagement.** Most policyholders ignore email surveys. Response rates are poor, and the people who do respond skew to the extremes.\n- **Long, multi-touch journeys.** Quote, bind, service, claim, renewal - friction at any stage drives churn months later.\n- **Regulation and sensitivity.** Insurance data is sensitive, and health insurance touches PHI, so compliance has to be designed in, not bolted on.\n\nThe result: insurers often fly blind on the *why* behind their NPS and retention numbers. AI interviews close that gap.\n\n## High-Value Research Use Cases for Insurers\n\n### 1. Claims-experience research\nThe claim is the moment of truth. Run interviews with recent claimants - approved and denied - to learn where the process felt slow, opaque, or unfair. Voice mode captures the emotion; the AI probes \"What would have made that easier?\" so you get actionable friction, not just a score.\n\n### 2. Policyholder satisfaction and NPS drivers\nStop guessing why your NPS moved. Pair a **scale** question (0-10 likelihood to recommend) with adaptive probing so every detractor and promoter explains their rating. See the [NPS Survey Guide](/docs/nps-survey-guide).\n\n### 3. Retention and churn research\nInterview policyholders who cancelled or switched. Was it price, a bad claim, a competitor offer, or a life change? The AI digs into the real trigger - the kind of insight that a checkbox exit survey never surfaces.\n\n### 4. Product and coverage concept testing\nBefore launching a new rider, add-on, or coverage tier, test it. Use **single_choice** and **multiple_choice** questions for coverage preferences and a **ranking** question to prioritize features, all alongside open-ended reactions.\n\n### 5. Pricing and willingness-to-pay\nUnderstand price sensitivity for new products with conversational pricing research - the AI explores not just *what* a policyholder would pay but *why* a price feels fair or excessive.\n\n### 6. Digital onboarding and quote-flow usability\nWhere do prospects abandon the online quote? Run task-based interviews on your quote and bind flow to find the drop-off points.\n\n### 7. Agent and broker channel feedback\nFor intermediated lines, interview agents and brokers about tooling, commissions, and the support they need to sell more effectively.\n\n## How Koji Powers Insurance Research\n\n- **Voice or text, on the policyholder's schedule.** Participants join via a link 24/7 - no call center, no calendar coordination. Voice mode is ideal for emotional claims stories; text suits quick coverage or pricing checks.\n- **Adaptive AI follow-ups.** Koji probes hesitation and vague answers in real time, the way a skilled researcher would, surfacing the root cause behind a complaint.\n- **Structured questions for hard numbers.** Combine qualitative depth with chartable data: **scale** for NPS/CSAT, **single_choice** and **multiple_choice** for coverage and channel preferences, **ranking** for feature priorities, and **yes_no** for claim resolution. See the [Structured Questions Guide](/docs/structured-questions-guide).\n- **Methodology built in.** Studies run on real frameworks like Customer Discovery and Jobs to be Done, so the AI asks disciplined, non-leading questions.\n- **Automatic, segmentable reporting.** Koji aggregates every interview into a report with themes, verbatim quotes, and distribution charts, and you can segment by policy type, tenure, or claim status.\n\n## Compliance and Data Handling\n\nInsurance research must respect strict data rules. Koji supports GDPR-aligned consent and **transcript anonymization**, and PrimeClub members can run on their own model keys (BYOK) so conversations are never used to train third-party models. For health-insurance research where protected health information may surface, follow the HIPAA-focused guidance and practice data minimization - collect only the personal detail your analysis truly needs. See [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) and [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research).\n\n## A Simple Way to Start\n\n1. **Pick one moment of truth** - usually the claims experience or a recent-churn cohort.\n2. **Create a Koji study** and choose Customer Discovery as the methodology.\n3. **Add 4-6 questions**, mixing one open-ended claims story, an NPS scale question, and a coverage-preference choice question.\n4. **Choose voice** for emotional depth, **text** for speed.\n5. **Import or share the link** with the cohort and let interviews run in parallel.\n6. **Read the report** - themes and NPS drivers are ready as soon as interviews complete.\n\nWith new accounts getting 10 free credits, you can field a first claims-experience study before lunch. That is the 10x advantage over commissioning a panel study: insight in days, at a fraction of the cost.\n\n## Worked Example: A Claims-Experience Study\n\nHere is what a full study looks like end to end, so you can copy the shape.\n\n**Goal:** Understand why claims satisfaction dipped last quarter among auto policyholders.\n\n**Audience:** 30 recent claimants - a mix of approved and denied claims, filed in the last 60 days.\n\n**Interview plan (voice mode, ~7 minutes):**\n\n1. *Open-ended:* \"Walk me through what happened from the moment you needed to file a claim.\" (The AI probes for the emotional and practical friction points.)\n2. *Scale (0-10):* \"How likely are you to recommend us to a friend?\" with anchor probing on the rating.\n3. *Single choice:* \"At which stage was the experience most frustrating?\" (Filing / Waiting for a decision / Communication / Payout / None)\n4. *Yes/No:* \"Did you always know the status of your claim?\"\n5. *Open-ended:* \"If you could change one thing about the process, what would it be?\"\n\n**What you get back:** Within a day or two, Koji aggregates all 30 interviews into a report. You see the NPS distribution split by approved vs denied, a bar chart of the most frustrating stage, and the recurring themes - say, \"unclear status updates\" and \"slow adjuster callbacks\" - each backed by verbatim policyholder quotes. That is a board-ready insight in days, not a six-week panel engagement.\n\n## Which Use Case Fits Each Line of Business\n\n| Line of business | Highest-value first study |\n|---|---|\n| Auto | Claims-experience and FNOL friction |\n| Home/property | Claims-experience and renewal-price reaction |\n| Health | Onboarding/navigation and care-access experience (HIPAA-aware) |\n| Life | Application and underwriting-friction research |\n| Commercial/SME | Broker channel feedback and coverage concept testing |\n| Insurtech/D2C | Quote-flow usability and trial-to-policy conversion |\n\nStart with the moment that most directly drives churn or complaints for your line, prove the value with one study, then expand into an always-on voice-of-customer program. Because Koji runs interviews in parallel and analyzes them automatically, scaling from one study to a continuous program does not require hiring a research team - the platform absorbs the operational load that traditionally capped how much research an insurer could do.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types for insurance research\n- [Build a Voice-of-Customer Program](/docs/voice-of-customer-research-program) - make policyholder listening continuous\n- [NPS Survey Guide](/docs/nps-survey-guide) - measure and explain loyalty\n- [Customer Discovery Interviews at Scale](/docs/customer-discovery-interviews-at-scale) - talk to 100 policyholders in a week\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) - run compliant studies\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) - for health-insurance research","category":"Use Cases","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"AI Customer Research for Insurance Companies (2026)","metaDescription":"A practical guide to AI-powered customer research for insurers: claims experience, policyholder retention, product concept testing, and pricing research - run with voice or text AI interviews and analyzed automatically.","keywords":["insurance customer research","policyholder research","claims experience research","insurance market research","AI interviews insurance","insurance churn research","insurance product research"],"aiSummary":"Insurers can run policyholder satisfaction, claims-experience, retention, product concept, and pricing research at scale with AI interviews. Koji conducts voice or text conversations with policyholders, probes claims stories with adaptive follow-ups, captures NPS and CSAT with scale questions and coverage preferences with choice and ranking questions, and aggregates everything into a compliance-aware research report - no moderator or panel agency required.","aiPrerequisites":["Familiarity with your policyholder journey (quote, bind, service, claim, renewal)","A Koji account (free tier includes 10 credits)","Awareness of your data-handling and compliance requirements"],"aiLearningOutcomes":["Identify the highest-value research use cases across the insurance lifecycle","Design claims-experience and retention interviews that capture emotion and detail","Choose the right structured question types for NPS, coverage, and pricing research","Run policyholder research at scale without a panel agency","Handle insurance research data in a compliance-aware way"],"aiDifficulty":"intermediate","aiEstimatedTime":"9 minutes"},{"type":"documentation","id":"c653c7a7-8d48-4a19-b964-e6d8ae676563","slug":"ai-research-for-pharma-life-sciences","title":"AI Customer Research for Pharma & Life Sciences","url":"https://www.koji.so/docs/ai-research-for-pharma-life-sciences","summary":"Pharma, biotech, and medical device teams use Koji's AI-moderated voice and text interviews to gather HCP and patient insights about 10x faster than manual moderated research. Koji probes every answer automatically, supports six structured question types for mixed-methods rigor, and produces real-time reports. Compliance practices include non-promotional framing, adverse event handling, data minimization, AES-256/TLS 1.2+ encryption, a DPA, and in-flow consent.","content":"## The Bottom Line\n\nPharma and life sciences teams need qualitative depth — why a physician prescribes one therapy over another, how a patient experiences a treatment journey, what a payer values in a formulary decision — but traditional market research in this industry is slow and expensive, often taking weeks to recruit and field, with moderator costs that limit sample size. Koji lets pharma, biotech, and medical device teams run AI-moderated voice and text interviews that field in days instead of weeks, probe every answer automatically, and produce analysis-ready reports — while you keep tight control over what data is collected. The result is roughly 10x faster time-to-insight than manual moderated research, at a fraction of the cost per conversation.\n\nThis guide covers where AI interviews fit in pharma research, how to design compliant studies, and the specific Koji capabilities that make it work.\n\n## Where AI interviews fit in life sciences research\n\nLife sciences research spans a wide set of audiences and decisions. AI-moderated interviews are a strong fit wherever you need structured qualitative depth at scale:\n\n- **HCP insights**: understand prescribing rationale, unmet clinical needs, treatment switching triggers, and reactions to new data or indications.\n- **Patient journey research**: map diagnosis-to-treatment journeys, adherence barriers, and quality-of-life impacts in the patient's own words.\n- **Message and concept testing**: test non-promotional educational concepts, positioning, and value propositions before committing budget.\n- **Market access and payer research**: explore how decision-makers weigh evidence, cost, and outcomes.\n- **Medical device usability and adoption**: capture how clinicians and patients actually use a device and where friction lives.\n- **Patient advisory and voice-of-patient programs**: run always-on listening that scales beyond a handful of advisory board members.\n\nIn each case, the qualitative *why* is what drives the decision — and that is exactly what surveys miss and what AI interviews capture.\n\n## Why traditional pharma research falls short\n\nThe classic approaches each carry a tax:\n\n- **Moderated interviews and focus groups** deliver depth but cost thousands per session and take weeks to schedule across busy clinicians.\n- **Surveys** scale and field fast, but force respondents into pre-written boxes and never ask the obvious follow-up. You learn *what* but rarely *why*.\n- **Advisory boards** give you deep relationships with a few experts, but the sample is tiny and the cadence is occasional.\n\nThe gap is a method that combines the depth of an interview with the scale and speed of a survey. That is the category Koji is built for.\n\n## How Koji works for life sciences teams\n\n**AI-moderated voice and text interviews.** You write a research brief, and Koji's AI conducts the interview — asking your questions and generating intelligent follow-ups in real time. An HCP can complete a voice interview from their phone between patients; a patient can respond by text on their own schedule. No moderator to book, no calendar Tetris.\n\n**Automatic follow-up probing.** When a respondent gives a thin answer — \"I switched because of side effects\" — Koji probes: which side effects, how severe, what would have changed the decision. This is the laddering a skilled human moderator does, applied consistently across every interview.\n\n**Structured questions for mixed-methods rigor.** Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — so you can capture a clean satisfaction or likelihood-to-prescribe scale alongside rich open-ended narrative in the same interview. Reports visualize scales as distributions and choices as frequency charts, giving you quant and qual in one study. See the [structured questions guide](/docs/structured-questions-guide) for details.\n\n**Real-time, analysis-ready reports.** As interviews come in, Koji synthesizes themes, surfaces representative verbatim quotes, and quantifies structured responses — so you are reading insights, not transcribing recordings.\n\n**Quality gate.** Only conversations that score 3 or higher on Koji's quality scale count toward your plan, which protects against low-effort responses — a real concern when incentivizing busy professionals.\n\n## Designing compliant pharma studies\n\nPharma research carries obligations that consumer research does not. A few principles for using AI interviews responsibly:\n\n1. **Keep it non-promotional.** Use research to learn, not to detail. Frame concepts and educational materials neutrally and avoid leading questions — Koji's probing is designed to follow the respondent, not to steer.\n2. **Plan for adverse event (AE) handling.** Any research touching marketed products needs an AE reporting process. Build a screening note and a clear escalation path into your study design, and brief your pharmacovigilance team before fielding.\n3. **Minimize and protect personal data.** Use structured questions and screeners to collect only what the study needs. Koji encrypts data in transit (TLS 1.2+) and at rest (AES-256), offers a DPA, and supports anonymization workflows — important when handling HCP or patient information. For regulated patient data, review the HIPAA and GDPR guidance below.\n4. **Capture consent in the flow.** Because Koji interviews are link-based and asynchronous, consent and study information can be presented at the start of the interview, creating a clean, auditable trail.\n5. **Document your methodology.** Koji's research brief becomes your study documentation — the questions, the audience, and the analysis approach in one place.\n\n## A practical example\n\nSuppose a biotech team wants to understand why oncologists hesitate to adopt a newly approved therapy. With Koji they would:\n\n- Write a brief targeting practicing oncologists, with open-ended questions on adoption barriers plus a scale question on likelihood to prescribe and a ranking question on decision factors (efficacy, safety, cost, guidelines).\n- Share a personalized interview link with a recruited HCP panel.\n- Let Koji conduct voice interviews and probe every hesitation automatically.\n- Read a real-time report that quantifies the ranking, charts the likelihood scale, and clusters the open-ended barriers into themes with supporting quotes.\n\nWhat used to be a six-week moderated study becomes a few days of fielding and same-day synthesis.\n\n## Getting started\n\nStart with one high-stakes question — an adoption barrier, a patient adherence gap, or a positioning test — and design a short study around it. Use a tight screener to reach the right audience, lean on structured questions for the metrics you need to track over time, and let open-ended questions plus AI probing surface the why. Bring your compliance and pharmacovigilance teams in early, and you will have a repeatable, defensible, and dramatically faster research engine.\n\n## Common pitfalls to avoid\n\nLife sciences teams new to AI interviews tend to stumble on the same few issues — all avoidable with planning:\n\n- **Treating it like a survey.** The value of an interview is the follow-up. Write fewer, deeper open-ended questions and let Koji probe, rather than porting a 30-item survey into an interview format.\n- **Skipping the screener.** Reaching the wrong specialty or patient cohort wastes the study. Use a tight screener so every completed interview is in-segment.\n- **Forgetting longitudinal value.** Reuse the same structured scale questions across studies so likelihood-to-prescribe, satisfaction, or adherence become trackable over time, not one-off readings.\n- **Underestimating compliance lead time.** Loop in compliance and pharmacovigilance before fielding, not after, so adverse event handling and consent language are settled up front.\n- **Over-collecting personal data.** Gather only what the research question requires; minimization is both a compliance win and a trust signal to respondents.\n\nAvoiding these keeps studies fast, defensible, and genuinely insightful — the combination that makes AI interviews a durable part of the life sciences research toolkit rather than a one-time experiment.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types for mixed-methods pharma studies\n- [AI Research for Healthcare](/docs/ai-research-for-healthcare) — adjacent guidance for care delivery and health systems\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) — handling regulated patient data\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — research under GDPR\n- [Concept Testing Guide](/docs/ai-concept-testing-guide) — testing educational concepts and positioning\n- [Generating Research Reports](/docs/generating-research-reports) — turning interviews into analysis-ready reports","category":"Use Cases","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"AI Customer Research for Pharma & Life Sciences | Koji","metaDescription":"How pharma, biotech, and medical device teams use AI voice and text interviews for HCP and patient insights — faster fielding, deeper qualitative depth, and compliance-aware study design.","keywords":["pharma market research","life sciences research","hcp insights","patient journey research","ai interviews pharma","biotech customer research","medical device research","payer research","voice of patient","pharma qualitative research"],"aiSummary":"Pharma, biotech, and medical device teams use Koji's AI-moderated voice and text interviews to gather HCP and patient insights about 10x faster than manual moderated research. Koji probes every answer automatically, supports six structured question types for mixed-methods rigor, and produces real-time reports. Compliance practices include non-promotional framing, adverse event handling, data minimization, AES-256/TLS 1.2+ encryption, a DPA, and in-flow consent.","aiPrerequisites":["Familiarity with pharma or life sciences market research","Awareness of your organization's compliance and pharmacovigilance requirements"],"aiLearningOutcomes":["Identify where AI interviews fit across HCP, patient, and payer research","Design mixed-methods pharma studies with structured questions","Apply compliance-aware practices including AE handling and data minimization","Cut time-to-insight from weeks to days"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 minutes"},{"type":"documentation","id":"bd390597-36c2-4894-8543-7ddd0d514f98","slug":"ai-research-for-healthcare","title":"AI-Powered Patient and Provider Research for Healthcare","url":"https://www.koji.so/docs/ai-research-for-healthcare","summary":"Koji enables healthcare organizations to conduct patient experience and provider satisfaction research at scale through async AI voice interviews. The format overcomes healthcare scheduling challenges, achieves 3-4x higher participation than traditional methods, and delivers actionable depth that HCAHPS surveys cannot provide.","content":"## The Bottom Line\n\nHealthcare research faces unique constraints: strict privacy requirements, hard-to-reach provider populations, and the need to capture nuanced clinical experiences that surveys cannot convey. Koji's AI voice interviews let healthcare organizations conduct patient experience research, provider satisfaction studies, and clinical workflow analysis at the scale needed for meaningful improvement — without the scheduling nightmares and costs of traditional healthcare research.\n\n## Why Healthcare Needs Better Research Methods\n\n### The Patient Experience Gap\nHCAHPS surveys measure satisfaction on standardized scales, but they cannot explain why patients feel the way they do. When a hospital's \"communication with nurses\" score drops from the 75th to the 60th percentile, HCAHPS tells you there is a problem. It does not tell you whether the issue is response time, bedside manner, discharge instructions, or something else entirely.\n\n### The Provider Burnout Blind Spot\nHealthcare organizations know burnout is a crisis — but annual engagement surveys produce the same generic findings year after year. \"Providers are stressed about administrative burden\" is not actionable. Understanding the specific administrative tasks, the moments of peak frustration, and the workarounds providers have invented is actionable.\n\n### The Research Logistics Problem\nHealthcare professionals are among the hardest populations to schedule for research:\n- Physicians have 15-minute schedule blocks with no flex time\n- Nurses work rotating shifts across day, evening, and night\n- Administrators are in back-to-back meetings\n- Patients have varying energy levels and availability\n\nTraditional interview-based research in healthcare takes 8-12 weeks just for the scheduling and moderation — before analysis even begins.\n\n## How Koji Solves Healthcare Research Challenges\n\n### Async Format for Busy Clinicians\nKoji's voice interviews can be completed at any time — during a lunch break, after a shift, during a commute. A 15-minute voice interview fits into healthcare schedules in ways that a 60-minute scheduled Zoom call never will. Participation rates among clinicians are 3-4x higher with async voice interviews than with scheduled sessions.\n\n### Depth Without Observation Bias\nWhen a researcher observes clinical workflows, the Hawthorne effect kicks in — providers modify their behavior because they are being watched. Voice interviews capture self-reported workflow experiences without observational interference, often revealing workarounds and frustrations that providers would not demonstrate in front of an observer.\n\n### Scale for Statistical Credibility\nHealthcare improvement initiatives require evidence that is credible to clinician audiences. A study based on 12 interviews can be dismissed as anecdotal. A study based on 100+ provider interviews across departments and specialties carries the weight needed to drive change in evidence-driven organizations.\n\n## Healthcare Research Use Cases\n\n### Patient Experience Deep-Dives\n\n**Beyond HCAHPS**: Interview patients about their complete care journey — from first symptom through treatment and follow-up. Capture the emotional experience, communication quality, care coordination gaps, and moments that mattered most.\n\n**Post-Discharge Research**: Interview patients 48-72 hours after discharge about their understanding of care instructions, medication changes, and follow-up plans. Identify where communication failed before readmissions occur.\n\n**Chronic Condition Management**: Interview patients managing chronic conditions about their daily experience with treatment protocols, medication adherence, and quality of life. Understand the patient perspective that clinical metrics miss.\n\n### Provider Experience Research\n\n**Burnout Driver Identification**: Move beyond generic burnout surveys to understand specific triggers. Which EHR tasks consume the most time? Where do documentation requirements interfere with patient care? What administrative processes create the most frustration?\n\n**Technology Adoption**: When implementing new clinical systems, interview providers about their experience with training, workflow integration, and workaround development. Identify adoption barriers before they become entrenched resistance.\n\n**Interdepartmental Collaboration**: Interview providers across departments about handoff quality, communication breakdowns, and coordination challenges. Surface the workflow friction that patient safety incidents often trace back to.\n\n### Operational Research\n\n**Wait Time Experience**: Surveys tell you patients are unhappy about wait times. Voice interviews reveal what specifically about the wait is frustrating — is it the length, the lack of communication, the uncomfortable environment, or the uncertainty about when they will be seen?\n\n**Telehealth Optimization**: Interview patients and providers about their telehealth experiences. Understand where virtual care excels and where it falls short from both perspectives.\n\n**Staff Retention**: Interview nursing staff about career satisfaction, manager effectiveness, scheduling flexibility, and what would make them stay. The insights from 75 nurse interviews are more actionable than any engagement survey.\n\n## Healthcare-Specific Discussion Guide Templates\n\n### Patient Experience Interview (12 minutes)\n1. Tell me about your most recent visit. What brought you in?\n2. Walk me through your experience from arrival to departure\n3. How would you describe the communication you received from your care team?\n4. Was there a moment during your visit that stood out — positively or negatively?\n5. How well did you understand your care plan when you left?\n6. If you could change one thing about your experience, what would it be?\n\n### Provider Workflow Interview (15 minutes)\n1. Describe a typical shift or clinic day for you\n2. Where do you spend time on tasks that feel like they should be faster or easier?\n3. Tell me about the documentation requirements in your workflow\n4. How well do the tools and technology support your clinical work?\n5. When was the last time you felt frustrated by a system or process?\n6. What would free up more of your time for patient care?\n\n### Post-Implementation Interview (12 minutes)\n1. How has the new [system/process] affected your daily work?\n2. What aspects have been easier than expected? Harder than expected?\n3. Have you developed any workarounds for limitations?\n4. How does this compare to what you were using before?\n5. What training or support would help you use it more effectively?\n\n## Privacy and Compliance Considerations\n\n### HIPAA-Aware Research Design\n- Design interview questions that capture experiences without requiring protected health information (PHI)\n- Focus on process quality, communication, and workflow — not clinical details\n- Configure studies to avoid collecting identifying health information\n- Ensure data handling meets organizational IRB and compliance requirements\n\n### Data Security\n- Koji implements encryption for data in transit and at rest\n- Access controls limit who can view study data\n- Data retention policies can be configured per organizational requirements\n- No voice biometric identification is performed\n\n### Informed Consent\n- Digital consent capture before interview begins\n- Clear explanation of how responses will be used\n- Participant right to stop the interview at any time\n- Anonymity protections for all respondents\n\n## Implementation Roadmap for Healthcare Organizations\n\n### Phase 1: Pilot (Month 1-2)\n- Select one department or service line for initial study\n- Run 40-50 patient or provider interviews\n- Compare insights to existing survey data\n- Validate approach with quality and compliance stakeholders\n\n### Phase 2: Expand (Month 3-4)\n- Roll out to 2-3 additional departments\n- Create healthcare-specific discussion guide templates\n- Train quality improvement teams on Koji\n- Integrate findings into existing improvement workflows\n\n### Phase 3: Scale (Month 5+)\n- Organization-wide availability\n- Continuous pulse programs for patient experience and provider satisfaction\n- Integration with quality metrics dashboards\n- Longitudinal tracking of improvement initiative impact\n\n## Frequently Asked Questions\n\n### Is Koji HIPAA compliant?\nKoji implements enterprise-grade security measures including encryption, access controls, and configurable data retention. Healthcare organizations should design studies to minimize PHI collection and work with their compliance teams to ensure research protocols meet organizational HIPAA requirements.\n\n### Can Koji interview patients with varying health literacy levels?\nYes. The AI interviewer adapts its language complexity to match the participant. Discussion guides can be designed with plain language principles, and the AI uses conversational prompting rather than clinical terminology.\n\n### How do you get physicians to participate in a 15-minute interview?\nThe async format is key — physicians can participate between patients, during lunch, or after hours. Incentives appropriate to the physician population ($100-200) and framing the research as contributing to workflow improvement (not just another survey) drive participation rates of 25-35%.\n\n### Can voice interviews replace HCAHPS?\nNot as a regulatory compliance mechanism — HCAHPS is required for CMS reporting. But Koji provides the depth layer that HCAHPS lacks, explaining the drivers behind scores and identifying specific improvement opportunities.\n\n### How do healthcare organizations handle negative findings about specific providers?\nReport findings at the department or service line level, never at the individual provider level. If patterns suggest serious quality or safety concerns, established organizational protocols for peer review and quality assurance should be followed.\n\n---\n\n## Related Resources\n\n- [Patient Experience Guide](/docs/patient-experience-survey-guide) — Patient feedback research\n- [Compliance & Ethics Guide](/docs/compliance-ethics-survey-guide) — Healthcare compliance\n- [NPS Survey Guide](/docs/nps-survey-guide) — Patient loyalty tracking\n- [Customer Journey Mapping](/docs/customer-journey-mapping-survey-guide) — Patient journey mapping\n- [CSAT Survey Guide](/docs/csat-survey-guide) — Patient satisfaction\n\n*Explore [structured questions](/docs/structured-questions-guide) for healthcare-specific patient and provider research.*\n\n## Further reading on the blog\n\n- [Koji vs Microsoft Forms: AI-Powered Research vs Enterprise Form Builder (2026)](/blog/koji-vs-microsoft-forms-2026) — Microsoft Forms is free with every M365 subscription, which makes it the default for a lot of teams. But it's a form builder — not a researc\n- [Koji vs UserTesting: AI-Powered vs Panel-Based Research (2026)](/blog/koji-vs-usertesting-2026) — UserTesting averages $36,265/year for SMB teams. Koji starts at €99/month. Here's an honest comparison of AI-native interviews vs panel-base\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n\n<!-- further-reading:blog -->\n","category":"Use Cases","lastModified":"2026-07-31T03:23:50.07297+00:00","metaTitle":"AI-Powered Patient and Provider Research for Healthcare | Koji","metaDescription":"How healthcare organizations use Koji for patient experience research, provider satisfaction studies, and clinical workflow analysis at scale with AI voice interviews.","keywords":["healthcare research","patient experience","provider research","healthcare AI","HCAHPS alternative","patient interviews","clinical workflow","provider burnout","healthcare quality","patient satisfaction","telehealth research","healthcare voice interviews"],"aiSummary":"Koji enables healthcare organizations to conduct patient experience and provider satisfaction research at scale through async AI voice interviews. The format overcomes healthcare scheduling challenges, achieves 3-4x higher participation than traditional methods, and delivers actionable depth that HCAHPS surveys cannot provide.","aiPrerequisites":["Healthcare operations or quality improvement experience","Understanding of patient experience measurement"],"aiLearningOutcomes":["Design healthcare research programs with AI voice interviews","Navigate privacy and compliance considerations","Conduct provider workflow and satisfaction research","Move beyond HCAHPS scores to actionable patient insights"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 minutes"},{"type":"documentation","id":"5468597f-7c3a-457b-8bc0-b35a30b13980","slug":"vulnerable-customer-research","title":"Vulnerable Customer Research: How to Evidence Good Outcomes Under FCA Rules","url":"https://www.koji.so/docs/vulnerable-customer-research","summary":"FG21/1, the FCA guidance on the fair treatment of vulnerable customers, is issued under the Principles and sets expectations in four areas: understanding the needs of vulnerable customers, staff skills and capability, responding through product design, customer service and communications, and monitoring and assessing outcomes. It frames vulnerability through four drivers: health, life events, resilience and capability. These are situational states rather than permanent traits, so annual point-in-time research cannot monitor them. The FCA published its review of firms treatment of customers in vulnerable circumstances on 7 March 2025, updated December 2025, based on interviews with 29 firms across 12 markets, a survey of 725 firms, commissioned consumer research with around 1,500 participants, and analysis of the 2020, 2022 and 2024 Financial Lives waves. It found that consumers in vulnerable circumstances continue to report poorer outcomes than other consumers, with the widest gap for those with multiple characteristics, and that firm-level progress has not yet appeared in UK-wide data. Firms asked for sector-specific case studies, better guidance on outcome monitoring methods, approaches for customers who do not disclose, and recognition of the intersection with Equality Act protected characteristics; the FCA published good practice case studies rather than revising the guidance. The central methodological problem is non-disclosure: disclosed vulnerability flags are a biased sample, so research should measure the drivers rather than the label using multiple_choice circumstance questions, behavioural scale questions on confidence and comprehension, and open_ended AI follow-ups on low scores. Outcomes should be monitored per Consumer Duty outcome, triggered by life events rather than by calendar, and always reported split by driver and by number of characteristics rather than in aggregate. Research methods themselves must not exclude low-capability customers: asynchronous participation, voice as a first-class option, short resumable sessions, plain language, no app install and deliberate recruiting for characteristics.","content":"If you rely on your vulnerability flags to evidence outcomes for customers in vulnerable circumstances, **your evidence is drawn from a biased sample and your monitoring is measuring the wrong population.** Most customers in vulnerable circumstances never disclose. The FCA has explicitly named both problems — how to monitor outcomes, and how to reach customers who do not tell you — and neither is solved by better reporting on the flags you already hold. They are solved by primary research designed to find vulnerability rather than waiting for it to be declared.\n\nThis guide is about that research programme. It sits alongside the [Consumer Duty guide](/docs/consumer-duty-customer-research), which covers comprehension testing for the consumer understanding outcome; here the subject is vulnerability as a **sampling and monitoring** problem across all four Duty outcomes.\n\n## What the FCA actually expects\n\nFG21/1, *Guidance for firms on the fair treatment of vulnerable customers*, is issued under the Principles rather than as Handbook rules — which is why firms so often under-invest in it and then find it applied firmly at supervision. It sets expectations in four areas:\n\n1. Understanding the needs of vulnerable customers in your target market and customer base\n2. Making sure staff have the skills and capability to recognise and respond to those needs\n3. Responding to those needs through product and service design, flexible customer service and communications\n4. Monitoring and assessing whether you are meeting and responding to those needs\n\nAreas 1 and 4 are research obligations in everything but name. Area 3 cannot be evidenced without them.\n\nThe guidance frames vulnerability through **four drivers**, and the practical value of the model is that it is situational — most of these are states people move through, not permanent traits.\n\n| Driver | What it covers | Typical research implication |\n|---|---|---|\n| Health | Physical or mental health conditions, cognitive impairment | Journey length, memory load, ability to complete in one sitting |\n| Life events | Bereavement, job loss, relationship breakdown, caring responsibilities | Timing-sensitive research; consent and duty of care |\n| Resilience | Irregular income, over-indebtedness, low savings | Financial stress changes decision-making under any communication |\n| Capability | Low literacy or numeracy, limited digital skills, low financial knowledge | Method itself can exclude the population you need |\n\nBecause vulnerability is situational, **a single annual study is structurally incapable of monitoring it.** Someone bereaved in March is a different research subject in September.\n\n## The 2025 review verdict, and what it means for your research\n\nThe FCA published its review of firms' treatment of customers in vulnerable circumstances on **7 March 2025**, updated in December 2025. It was substantial work: interviews with 29 firms across 12 markets, a survey of 725 firms, commissioned quantitative and qualitative consumer research with around 1,500 participants, and analysis of the 2020, 2022 and 2024 Financial Lives waves.\n\nThe verdict was uncomfortable. Consumers in vulnerable circumstances **continue to report poorer outcomes than other consumers**, and the gap is widest for people with multiple vulnerability characteristics. The Consumer Duty has visibly renewed firms' focus, but the progress seen in firm-level assessments has not yet shown up in the UK-wide data.\n\nFirms themselves asked the FCA for four things: more sector-specific case studies, better guidance on **outcome monitoring methods**, approaches for **customers who do not disclose vulnerability**, and recognition of how vulnerability intersects with Equality Act protected characteristics. The FCA chose not to rewrite FG21/1, publishing good-practice case studies instead — which means the expectation is unchanged and the burden of designing the monitoring sits with you.\n\nThree things follow directly for a research programme:\n\n- **Aggregate outcome metrics are not enough.** A firm-level satisfaction or complaints number cannot show a gap it is not split by.\n- **Disclosed flags are the wrong denominator.** They measure who told you, which is a different question from who is affected.\n- **Multiple characteristics need to be visible.** If your analysis treats vulnerability as one binary, you cannot see the group with the worst outcomes.\n\n## Problem one: the customers who never tell you\n\nNon-disclosure is the central methodological problem. People do not disclose because they do not recognise the label, because they fear consequences for credit or access, because the channel gave them no natural opening, or because they were in a hurry and the disclosure question sat behind a menu.\n\nThe fix is not to ask harder. It is to **measure the drivers rather than the label**:\n\n- Ask about circumstances, not categories. \"In the last twelve months, have you experienced any of the following?\" with a `multiple_choice` list of concrete life events outperforms any question containing the word vulnerable.\n- Capture capability behaviourally. A `scale` question on how confident someone felt completing the last step of a journey tells you more than a self-declared digital skills rating.\n- Probe the ones that matter. In Koji, an `open_ended` follow-up fires automatically when a confidence score is low — \"you said you were not sure what would happen next; what did you do at that point?\" — which is where the actual failure story lives.\n- Let people answer by voice. Typing is itself a capability filter, and a spoken answer from someone with low literacy carries detail no form field would have captured.\n\nRun this in the research instrument, and you get a vulnerability signal for **every participant**, not only the ones already flagged. That is the denominator the FCA is asking about, and it lets you do the analysis that matters: comparing outcomes between customers you had flagged, customers you had not flagged but who show characteristics, and customers who show none.\n\nAlmost every firm that runs this comparison for the first time finds the same thing — the undisclosed group looks materially worse than the flagged group, because the flagged group has been receiving support.\n\n## Problem two: monitoring outcomes rather than reporting activity\n\nOutcomes monitoring fails in a predictable way: firms report what they did (calls handled, flags recorded, training completed) instead of what happened to customers. Activity is easy to count and proves nothing.\n\nA workable monitoring design measures each Duty outcome, split by vulnerability characteristic:\n\n| Outcome | What to measure with customers | Method |\n|---|---|---|\n| Products and services | Did the product still fit after circumstances changed? | Triggered study after a life-event signal |\n| Price and value | Did the customer understand what they were paying and why? | Comprehension test plus `scale` value perception |\n| Consumer understanding | Can the customer correctly state what happens next? | Recall test with pre-set pass criteria |\n| Consumer support | Could the customer get help, first time, through their channel of choice? | Post-interaction study across channels |\n\nTwo design rules carry most of the weight. First, **trigger studies on events rather than a calendar** — a bereavement notification, a missed payment, a power of attorney registration, a complaint about a communication. Vulnerability is situational and your evidence should be too. Second, **always report split, never only aggregate.** An 84% satisfaction figure that hides 61% among customers with three or more characteristics is worse than no figure, because it creates false assurance — exactly the pattern the FCA criticised when it said reliance on sales data or an absence of complaints provides no reliable assurance.\n\n## Problem three: your research method is probably excluding them\n\nThis is the failure that quietly invalidates everything above. Standard research methods select against precisely the customers you need:\n\n- Scheduled video interviews exclude shift workers, carers and anyone without reliable connectivity or a quiet room.\n- Panel recruitment over-represents confident, digitally fluent, repeat participants.\n- Long written surveys select for literacy and stamina.\n- App downloads and account creation exclude low-capability users at the first step.\n- Incentives paid only by digital transfer exclude the unbanked and underbanked.\n\nAn inclusive design does the opposite. Make participation **asynchronous** so it fits around caring responsibilities and irregular shifts. Offer **voice as a first-class option**, not a fallback — for someone with low literacy or a visual impairment, speaking is not an accommodation, it is the usable path. Keep sessions short and allow them to be resumed. Write at a genuinely plain reading level. Offer the study in the languages your customer base actually speaks. And recruit deliberately for the characteristics rather than hoping a general panel contains them.\n\nThis is the strongest structural argument for AI-moderated research in this domain: **a study that runs by voice or text, asynchronously, in any language, with no moderator to schedule and no software to install, removes most of the access barriers that make vulnerable customers hard to research at all.** Koji runs exactly that shape of study, and because the AI probes automatically, a short session still reaches the depth a moderator would have needed a booked hour for.\n\n## Building the evidence pack\n\nWhat a board or supervisor needs to see, in order:\n\n1. **Who you researched, and how you found them** — including how you identified characteristics beyond disclosed flags.\n2. **Outcome results split by driver and by number of characteristics**, with the comparison against the non-vulnerable group stated explicitly.\n3. **Verbatim evidence** of where the journey failed, quoted, with the transcript retained.\n4. **What changed as a result**, with dates and owners — the decision trail, not just the finding.\n5. **Re-test results after the change**, which is what converts a finding into evidence of improvement.\n6. **Residual gaps**, named, with dates. Boards get more credit for a known gap with an owner than for an unblemished dashboard.\n7. **Method and data lineage** — how participants were sampled, what was asked, where transcripts and exports live.\n\nAdd the intersection the FCA called out: report where vulnerability characteristics overlap with Equality Act protected characteristics, because that is where both regulatory and reputational risk concentrate.\n\n## Common mistakes\n\n- Treating vulnerability as a permanent customer attribute rather than a situational state.\n- Monitoring only customers who disclosed, then reporting the result as coverage of vulnerable customers.\n- Collapsing all vulnerability into one binary flag, which hides the multiple-characteristic group with the worst outcomes.\n- Reporting activity metrics as outcome evidence.\n- Running the research with methods that structurally exclude low-capability customers, and never noticing.\n- Testing once a year, when the drivers are events that occur continuously.\n- Collecting health and financial hardship data without a lawful basis, a retention limit and a deletion route — vulnerability data is often special category data and deserves a data protection impact assessment.\n\n## Frequently asked questions\n\n**Does the FCA require firms to do research with vulnerable customers?**\nFG21/1 does not name research as an activity, but it requires firms to understand the needs of vulnerable customers in their customer base and to monitor whether those needs are being met. Neither can be evidenced from internal activity metrics alone, and the FCA's March 2025 review found that firms themselves asked for better guidance on outcome monitoring methods. Primary research is how those two expectations get satisfied in practice.\n\n**How do we research customers who never disclose their vulnerability?**\nStop asking for the label and measure the drivers instead. Ask about concrete circumstances in the last twelve months with a multiple_choice list, capture confidence and comprehension behaviourally with scale questions, and probe low scores with open_ended follow-ups. That produces a vulnerability signal for every participant, letting you compare outcomes across flagged customers, unflagged customers with characteristics, and everyone else.\n\n**What are the four drivers of vulnerability in FG21/1?**\nHealth, life events, resilience and capability. Health covers physical and mental health conditions; life events covers bereavement, job loss, relationship breakdown and caring responsibilities; resilience covers irregular income, over-indebtedness and low savings; capability covers low literacy, numeracy, digital skills or financial knowledge. Most are situational states rather than permanent traits, which is why point-in-time annual research cannot monitor them.\n\n**How often should we research vulnerable customer outcomes?**\nContinuously, triggered by events rather than by the calendar — bereavement notifications, missed payments, power of attorney registrations, complaints about communications — plus before and after any material change to a product, journey or communication. An annual study leaves the evidence stale for most of the year and cannot capture a situational driver.\n\n**How do we stop our research method from excluding the customers we need?**\nMake it asynchronous so it fits around shifts and caring responsibilities, offer voice as a first-class option rather than a fallback, keep sessions short and resumable, write at a genuinely plain reading level, offer the languages your customer base speaks, avoid app installs and account creation, and recruit deliberately for the characteristics instead of relying on a general panel. Then check the achieved sample against your customer base and report the gap honestly.\n\n**What should we report to the board on vulnerable customer outcomes?**\nOutcome results split by driver and by number of characteristics with an explicit comparison to non-vulnerable customers, how the sample was identified beyond disclosed flags, verbatim evidence of journey failures, the decision trail for what changed, re-test results after those changes, named residual gaps with owners and dates, and the intersection with Equality Act protected characteristics. Aggregate satisfaction figures without splits create false assurance rather than evidence.\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types behind driver measurement and comprehension tests\n- [FCA Consumer Duty Customer Research](/docs/consumer-duty-customer-research) — comprehension testing for the consumer understanding outcome\n- [Accessibility Research Guide](/docs/accessibility-research-guide) — including users with disabilities in your studies\n- [AI Customer Research for Banking & Financial Services](/docs/ai-research-for-banking) — the sector view\n- [AI-Powered Customer Research for Insurance Companies](/docs/ai-research-for-insurance) — insurance-specific outcome monitoring\n- [DPIA for User Research](/docs/dpia-user-research) — assessing risk before you collect health and hardship data\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) — retention limits for special category data","category":"Research Operations","lastModified":"2026-07-31T03:21:09.13565+00:00","metaTitle":"Vulnerable Customer Research: Evidencing Outcomes Under FCA Rules","metaDescription":"How to design research that evidences good outcomes for customers in vulnerable circumstances: measuring the four FG21/1 drivers, reaching customers who never disclose, and monitoring outcomes by characteristic.","keywords":["vulnerable customer research","FCA vulnerable customers","FG21/1 guidance","characteristics of vulnerability","vulnerable customer outcomes monitoring","financial vulnerability research","inclusive customer research","Consumer Duty vulnerability"],"aiSummary":"FG21/1, the FCA guidance on the fair treatment of vulnerable customers, is issued under the Principles and sets expectations in four areas: understanding the needs of vulnerable customers, staff skills and capability, responding through product design, customer service and communications, and monitoring and assessing outcomes. It frames vulnerability through four drivers: health, life events, resilience and capability. These are situational states rather than permanent traits, so annual point-in-time research cannot monitor them. The FCA published its review of firms treatment of customers in vulnerable circumstances on 7 March 2025, updated December 2025, based on interviews with 29 firms across 12 markets, a survey of 725 firms, commissioned consumer research with around 1,500 participants, and analysis of the 2020, 2022 and 2024 Financial Lives waves. It found that consumers in vulnerable circumstances continue to report poorer outcomes than other consumers, with the widest gap for those with multiple characteristics, and that firm-level progress has not yet appeared in UK-wide data. Firms asked for sector-specific case studies, better guidance on outcome monitoring methods, approaches for customers who do not disclose, and recognition of the intersection with Equality Act protected characteristics; the FCA published good practice case studies rather than revising the guidance. The central methodological problem is non-disclosure: disclosed vulnerability flags are a biased sample, so research should measure the drivers rather than the label using multiple_choice circumstance questions, behavioural scale questions on confidence and comprehension, and open_ended AI follow-ups on low scores. Outcomes should be monitored per Consumer Duty outcome, triggered by life events rather than by calendar, and always reported split by driver and by number of characteristics rather than in aggregate. Research methods themselves must not exclude low-capability customers: asynchronous participation, voice as a first-class option, short resumable sessions, plain language, no app install and deliberate recruiting for characteristics.","aiPrerequisites":["A UK FCA-regulated firm subject to FG21/1 and the Consumer Duty","Access to customer records or journeys where vulnerability characteristics may be present","A lawful basis and DPIA for collecting health and financial hardship data"],"aiLearningOutcomes":["Map the four FG21/1 drivers of vulnerability to concrete research questions","Measure vulnerability characteristics without asking customers to self-label","Compare outcomes across flagged, unflagged-with-characteristics and non-vulnerable customers","Design event-triggered outcome monitoring instead of annual studies","Remove the method-level barriers that exclude low-capability customers from research","Assemble a board evidence pack that reports outcomes split by characteristic"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"b31ac21a-572e-49b8-8fa1-1452076905dc","slug":"ai-incident-postmortem-user-research","title":"AI Incident Postmortems: How to Investigate Model Failures with User Evidence (2026)","url":"https://www.koji.so/docs/ai-incident-postmortem-user-research","summary":"An AI incident postmortem is a written, blameless investigation of a model or agent failure that records impact, timeline, contributing causes, and owned corrective actions. AI incidents differ from ordinary software incidents because they usually produce no error code - the system returns a confident, well-formed, wrong answer - so telemetry cannot scope the harm. This guide covers what counts as an incident, how to classify severity, the EU AI Act Article 73 reporting clocks, the postmortem template adapted for probabilistic systems, and how to gather user evidence fast enough to meet those clocks using Koji.","content":"# AI Incident Postmortems: How to Investigate Model Failures with User Evidence (2026)\n\n**Short answer:** An AI incident postmortem is a blameless written investigation of a model or agent failure that records impact, timeline, contributing causes, and owned corrective actions. It differs from an ordinary software postmortem in one decisive way: AI systems usually fail without an error. There is no exception, no 500, no alert - just a confident, fluent, wrong answer returned as a success. That means your telemetry can tell you what the model said, but only your users can tell you what it cost.\n\nMost teams discover this the hard way, mid-incident, when someone asks \"how many customers were affected?\" and the honest answer is that nobody knows.\n\n## What counts as an AI incident\n\nDraw the line explicitly, before you need it. An AI incident is any deployed model behaviour that:\n\n- causes or nearly causes harm to a person, financially, physically, reputationally, or psychologically;\n- produces a materially wrong result that a user **acted on**;\n- violates a stated policy, a permission boundary, or a legal obligation;\n- degrades an agreed quality metric beyond a defined threshold for a defined period.\n\nNote what is absent from that list: downtime, exceptions, and error rates. The Responsible AI Collaborative's AI Incident Database - which indexes real-world AI harms and near-harms and has now issued incident IDs past 1,600 - is instructive here. Read through it and the pattern is unmistakable: very few entries are outages. They are systems working exactly as built, on inputs nobody anticipated, for users nobody profiled.\n\n**The near-harm is worth logging too.** Aviation safety culture, from which this practice descends, treats the near miss as the cheapest possible lesson. An AI system that produced a dangerous recommendation which a user happened to catch is a free incident - all of the signal, none of the damage.\n\n## Why AI incidents need user evidence\n\nIn a conventional incident, the blast radius is a query: which requests failed, between which timestamps, for which accounts. In an AI incident, the same query returns everything and tells you nothing, because every one of those requests succeeded.\n\nWhether a wrong output became a harm depends entirely on the human step that followed it:\n\n| Model behaviour | Log signature | Actual outcome | Determined by |\n|---|---|---|---|\n| Wrong dosage guidance | 200 OK | User noticed, ignored it | The user |\n| Wrong dosage guidance | 200 OK | User followed it | The user |\n| Fabricated citation | 200 OK | Caught in review | The reviewer |\n| Fabricated citation | 200 OK | Published | The reviewer |\n| Over-refusal | 200 OK | User rephrased, succeeded | The user |\n| Over-refusal | 200 OK | User abandoned the task | The user |\n\nIdentical telemetry, opposite severity. This is the structural reason an AI postmortem that consults only logs will systematically under-report harm - and why the phrase \"we saw no elevated error rate\" is not evidence of anything.\n\nThe corroborating industry signal: the 2025 Stack Overflow Developer Survey found **66% of developers name AI output that is \"almost right, but not quite\" as their single largest frustration, and 45% report that debugging AI-generated code takes longer** than writing it themselves. Near-misses are the dominant failure mode, and near-misses are precisely what monitoring cannot see.\n\n## Blameless, properly understood\n\nGoogle's SRE practice popularised the blameless postmortem, and it is probably the most misread term in the discipline. Blameless does **not** mean no consequences, no accountability, or no follow-up. It means assuming everyone involved had good intentions and acted on the best information available to them at the time, and pointing the investigation at the system rather than the person.\n\nThe reason this is operational rather than sentimental: removing blame gives people the confidence to escalate early. In AI incidents, where detection depends on a human noticing something subtly wrong, the cost of a culture where people hesitate to raise a concern is measured in weeks of undetected harm.\n\nTwo adaptations for AI systems:\n\n1. **The cause is rarely a commit.** Expect a combination - a prompt revision, a retrieval index rebuild, an upstream model version change, and a shift in the input distribution, none individually sufficient. Resist the pressure to name one.\n2. **\"The model did it\" is not a root cause.** It is a restatement of the incident. Push through to the design decision that let a probabilistic component take an irreversible action without a check.\n\n## The reporting clock\n\nThis is no longer purely voluntary practice. Under **Article 73 of the EU AI Act**, providers of high-risk AI systems must report serious incidents to the market surveillance authority of the member state where the incident occurred, as soon as a causal link to their system is established or reasonably likely. The deadlines are tiered by severity:\n\n| Situation | Deadline from awareness |\n|---|---|\n| Serious incident (general) | 15 days |\n| Death may have been caused | 10 days |\n| Widespread infringement, or serious disruption to critical infrastructure | 2 days |\n\nArticle 73(5) explicitly permits an **initial incomplete report** with fuller information to follow - which is the mechanism that makes these deadlines survivable. Authorities may order market surveillance measures within days of receiving a report, so the quality of your initial evidence directly shapes what happens next.\n\nA two-day clock is the part worth internalising. Two days is not enough time to recruit, schedule, moderate, transcribe, and synthesise a conventional research study. It is enough time to run an AI-moderated one.\n\n## The AI postmortem template\n\n| Section | What goes in it | AI-specific note |\n|---|---|---|\n| Summary | Two sentences: what happened, who was affected | Quantify affected users, not affected requests |\n| Impact | Harm distribution by severity | Requires user evidence, not logs |\n| Detection | How and when you found out | Record whether a user or a monitor found it - if it was always users, that is itself a finding |\n| Timeline | First bad output to full mitigation | Include the silent period before detection |\n| Contributing causes | Plural, always | Prompt, retrieval, model version, distribution shift, missing guardrail |\n| What went well | Genuinely - what limited the damage | Usually a human check somewhere |\n| Action items | Owned, dated, split mitigative vs preventative | Add an evaluation-set item to every AI postmortem |\n| Evidence appendix | Sample outputs, participant quotes, severity counts | This is your regulatory filing material |\n\nSplit action items into **mitigative** (closes this specific gap) and **preventative** (addresses the whole class of failure). AI postmortems should almost always generate a third kind: an **evaluation** item - the failure, converted into permanent test cases and added to your golden set so the regression can never silently return.\n\n## Severity classification\n\nScore every confirmed harm on one scale, defined in advance:\n\n| Level | Definition | Reporting implication |\n|---|---|---|\n| S1 - Harmful | User suffered material harm; irreversible or costly | Regulatory reporting likely; executive notification |\n| S2 - Blocking | User could not complete a critical task; no workaround | Full postmortem required |\n| S3 - Degraded | Task completed but with wasted effort or lost trust | Postmortem if systemic |\n| S4 - Cosmetic | Noticed, no consequence | Log and aggregate |\n\nThe count that matters in the summary is **S1 and S2 users**, not total requests. Executives and regulators both ask the same first question, and \"0.3% of requests\" is not an answer to it.\n\n## Running the user-evidence study with Koji\n\nThe traditional path - recruit affected users, schedule interviews, moderate 12 of them, transcribe, tag, synthesise - takes two to three weeks. That is longer than every deadline in the table above and longer than the patience of any executive during an active incident. It is also why most postmortems quietly substitute a support-ticket sample for real evidence, which biases everything toward the users who complained loudly.\n\nThe AI-native approach:\n\n**1. Launch within hours, not weeks.** Define the exposed population from your incident window, invite a representative sample, and run an always-on AI-moderated study. Participants respond on their own schedule, which removes the scheduling bottleneck entirely - the single largest source of delay in conventional research.\n\n**2. Probe what the user actually did.** This is the question a survey cannot ask well, because the useful follow-up depends on the answer. Koji's AI moderator asks it: *\"You said the recommendation looked off - what did you do next?\"* and then follows the answer wherever it goes. That branch is the difference between knowing an output was wrong and knowing whether it caused harm.\n\n**3. Produce a comparable severity distribution.** Koji's six structured question types turn testimony into countable evidence:\n\n| Question type | Incident use |\n|---|---|\n| `single_choice` | Which of these did you experience? |\n| `yes_no` | Did you act on the output before realising it was wrong? |\n| `scale` | How much impact did this have on you? (1-5) |\n| `ranking` | Order these consequences by how much they mattered |\n| `multiple_choice` | What did you do after you noticed? |\n| `open_ended` | Describe what happened in your own words |\n\nThe `yes_no` acted-on-it question is the one that converts a quality problem into a harm count - and it is the number your postmortem summary and any regulatory filing both need. See the [structured questions guide](/docs/structured-questions-guide) for how to sequence these without leading participants.\n\n**4. Cluster harms automatically.** Thematic analysis groups hundreds of open-ended accounts into named harm categories, each linked back to verbatim quotes and specific participants. Manual tagging of 60 transcripts is roughly a week of researcher time; this is the step that makes a real evidence base compatible with a two-day clock.\n\n**5. Keep it as a regression study.** The same study, re-run after the fix, is your verification that the mitigation worked from the user's side rather than the dashboard's.\n\nBecause Koji is AI-native rather than a scheduling layer over human moderators, a 50-participant incident study runs in the window where it can still change the outcome. That is the whole argument: research that arrives after the postmortem is filed is not evidence, it is history.\n\n## Sample sizes during an incident\n\n| Goal | Participants | Turnaround |\n|---|---|---|\n| Confirm the failure is real and characterise it | 10-15 | Same day |\n| Scope the harm distribution defensibly | 40-60 | 2-3 days |\n| Verify the mitigation from the user side | 30-40 | Post-fix |\n| Track trust recovery | 40+ | Monthly for a quarter |\n\n## Five mistakes to avoid\n\n1. **Reporting requests instead of people.** Percentages of traffic conceal concentrated harm on a small, specific group.\n2. **Sampling only complainants.** Support tickets over-represent articulate, high-engagement users and under-represent the ones who silently left.\n3. **Naming a single root cause.** AI incidents are almost always a conjunction. A single named cause usually means the investigation stopped early.\n4. **Closing without an evaluation item.** If the failure did not become a permanent test case, you have licensed its return.\n5. **Waiting for a complete picture before filing.** Article 73(5) exists precisely so you do not have to. File the initial report, then complete it.\n\n## Frequently asked questions\n\n**What counts as an AI incident?**\nAny deployed model behaviour that causes or nearly causes harm, produces a materially wrong result a user acted on, violates a policy or legal obligation, or degrades a quality metric past an agreed threshold. It does not require an error, an exception, or downtime - most AI incidents return HTTP 200.\n\n**How is an AI incident postmortem different from a normal one?**\nThree ways: there is usually no error signal, so detection depends on users rather than alerts; the cause is rarely a single commit but a conjunction of prompt, retrieval, model version, and distribution shift; and the blast radius cannot be read from logs, because whether a wrong output caused harm depends on what the user did with it.\n\n**What does blameless actually mean?**\nAssuming everyone involved acted with good intent on the best information available at the time, and directing the investigation at systems rather than individuals. It does not mean no consequences. Its practical purpose is that people escalate early instead of hiding problems.\n\n**What are the EU AI Act reporting deadlines for serious incidents?**\nUnder Article 73, providers of high-risk systems report to the market surveillance authority of the member state where the incident occurred: 15 days generally, 10 days where a death may have been caused, and 2 days for widespread infringement or serious disruption to critical infrastructure. Article 73(5) permits an initial incomplete report.\n\n**How do you scope who was affected when telemetry cannot tell you?**\nYou sample. Identify the population plausibly exposed during the incident window, draw a representative sample, and ask them what they saw and what they did next. A structured study of 40-60 affected users characterises the harm distribution far better than any log query.\n\n**How does Koji help during an incident?**\nSpeed is the binding constraint when a 15-day or 2-day clock is running. Koji launches an AI-moderated study to affected users within hours, probes what each person actually did with the wrong output, produces a comparable severity distribution through structured questions, and clusters responses into named harm categories - evidence in days rather than weeks.\n\n## Related Resources\n\n- [Root Cause Analysis for Customer Research](/docs/root-cause-analysis-guide) - the general RCA method this adapts\n- [The Five Whys Technique](/docs/five-whys-technique-user-research) - drilling past the first plausible cause\n- [Human Evaluation of AI Outputs](/docs/human-evaluation-ai-outputs) - the ongoing quality practice that catches incidents earlier\n- [AI Red Teaming with Real Users](/docs/ai-red-teaming-with-users) - finding these failures before they become incidents\n- [The EU AI Act and User Research](/docs/eu-ai-act-user-research-compliance) - the wider compliance picture\n- [AI Governance for Customer Research](/docs/ai-governance-frameworks-research) - ISO 42001 and the NIST AI RMF\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - turning testimony into countable evidence\n\n---\n\n**Sources:** EU AI Act Article 73 (Reporting of Serious Incidents) and the European Commission draft guidance on serious incident reporting; Google SRE, *Postmortem Culture: Learning from Failure*; AI Incident Database, Responsible AI Collaborative; 2025 Stack Overflow Developer Survey.","category":"Research Operations","lastModified":"2026-07-31T03:19:37.807161+00:00","metaTitle":"AI Incident Postmortems: Investigating Failures with User Evidence","metaDescription":"How to run a blameless postmortem for an AI product failure: what counts as an incident, why telemetry alone cannot scope harm, the EU AI Act reporting clocks, and how to gather user evidence in days with Koji.","keywords":["ai incident postmortem","blameless postmortem","ai incident response","model failure investigation","ai incident reporting","eu ai act article 73","ai regression","root cause analysis ai"],"aiSummary":"An AI incident postmortem is a written, blameless investigation of a model or agent failure that records impact, timeline, contributing causes, and owned corrective actions. AI incidents differ from ordinary software incidents because they usually produce no error code - the system returns a confident, well-formed, wrong answer - so telemetry cannot scope the harm. This guide covers what counts as an incident, how to classify severity, the EU AI Act Article 73 reporting clocks, the postmortem template adapted for probabilistic systems, and how to gather user evidence fast enough to meet those clocks using Koji.","aiPrerequisites":["Familiarity with incident response or root cause analysis","Basic understanding of how AI models fail in production"],"aiLearningOutcomes":["Distinguish an AI incident from an ordinary software bug","Classify AI incident severity and scope affected users","Apply blameless postmortem principles to probabilistic systems","Meet EU AI Act Article 73 reporting deadlines with defensible evidence","Run a rapid user-evidence study during an active incident with Koji"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"},{"type":"documentation","id":"4094868a-9721-4be6-b49e-96e8a725b20b","slug":"hcp-research-guide","title":"HCP Research: How to Recruit and Interview Physicians, Nurses and Other Healthcare Professionals","url":"https://www.koji.so/docs/hcp-research-guide","summary":"Research with healthcare professionals carries four constraints beyond ordinary qualitative research: low access, regulated payment, transparency reporting and adverse event obligations. Median physician response rates in research are around 18 percent with a range of roughly 10 to 60 percent, which is why asynchronous AI-moderated interviewing outperforms scheduled calls for this audience. Recruiting channels ranked by yield are own customer lists, professional associations, specialist HCP panels, referral and snowball recruiting, and congress intercepts, and every channel requires verification through credential checks plus clinical screener questions a non-clinician cannot answer. Honoraria must be set on a documented fair market value basis using an hourly rate by specialty and seniority and a record-to-recruit ratio against the qualified universe; published ranges put five-minute polls at 5 to 15 dollars and 60-minute in-depth interviews at 200 to 500 dollars and above, with oncology, immunology and rare-disease specialists at 100 to 300 dollars per hour and higher. Under the US Physician Payments Sunshine Act, payments in genuinely double-blinded market research are not reportable to Open Payments, and the exclusion also applies to single-blind designs where only the manufacturer is blinded; in a partially blinded study only the blinded portion is excluded, and blinding cannot be retro-fitted after identities are shared. Research touching a marketed medicinal product requires an adverse event process: trained reviewers, a defined reporting route and timeline set by the client SOP, capture of the four case elements of identifiable reporter, identifiable patient, suspect product and event, re-contact consent, and a reminder script. EphMRA maintains the Code of Conduct and Adverse Event Reporting Guidelines revised in September 2025. Interviews should cap participant time at around 20 minutes, lead with clinical scenarios, offer voice, stay non-promotional and use structured question types to produce countable data.","content":"Research with healthcare professionals is ordinary qualitative research wrapped in four constraints that do not apply anywhere else: **access is brutal, payment is regulated, transparency reporting may be triggered, and anything a participant mentions about a product may create a pharmacovigilance obligation.** Get those four right and the research itself is familiar. Get them wrong and you have a compliance incident rather than a study.\n\nThe headline access problem is real. **Median physician response rates in research sit around 18%**, with reported ranges of roughly 10–60% depending on specialty, incentive and follow-up. A specialist oncologist has perhaps twenty minutes of uncommitted time in a working day, and it is not at 2pm on a Tuesday when your moderator is free. This is the single strongest argument for asynchronous, AI-moderated interviewing in healthcare: **a study a clinician can complete by voice at 22:40 after clinic, in six minutes of speaking, converts at rates a calendar invite never will.**\n\n## Who counts as an HCP, and why the label matters\n\n\"HCP\" covers physicians, nurses and nurse practitioners, pharmacists, physician assistants, dentists, allied health professionals and, in most compliance frameworks, anyone in a position to prescribe, dispense, purchase or recommend a medical product. Two adjacent groups are often confused with them and carry different rules:\n\n- **Payers and formulary decision-makers** — commercially sensitive, rarely subject to transparency reporting, usually recruited through specialist networks.\n- **Practice and hospital administrators** — a purchasing audience, not a clinical one, and frequently the actual buyer for medtech and health IT.\n\nScope the audience before you scope the method. A study that mixes prescribers and administrators without separating the samples produces findings that neither group would recognise.\n\n## The four constraints, in order of how often they bite\n\n| Constraint | What it means in practice | Who it applies to |\n|---|---|---|\n| Access and verification | Low response rates; must prove the participant is genuinely the specialist claimed | Everyone |\n| Fair market value honoraria | Payment must be defensible against specialty, time and expertise — not \"whatever it takes\" | Anyone paid by or on behalf of a manufacturer |\n| Transparency reporting | US Open Payments reporting of transfers of value to covered recipients | Applicable manufacturers of covered products |\n| Adverse event reporting | Any mention of a suspected adverse reaction may trigger a reporting obligation | Any research touching a marketed medicinal product |\n\n## Recruiting: five channels, ranked by what they actually deliver\n\n1. **Your own customer or user list.** Highest response, lowest cost, most bias. Perfect for product research; useless for market landscape work. Import the list and send personalised interview links so each participant arrives pre-identified and you never ask a question you already know the answer to.\n2. **Professional associations and society lists.** Credible and specialty-accurate; slow, and often requires sponsorship framing that colours responses.\n3. **Specialist HCP panels.** Fast and verified, and the default for pharma and medtech market research. Expensive, and heavy panel users are not representative of the specialty.\n4. **Referral and snowball recruiting.** The only reliable route to genuinely rare specialties and to key opinion leaders. Slow to start, then compounds.\n5. **Congress and conference intercepts.** Excellent for breadth in a short window, terrible for depth — nobody gives you thirty focused minutes in an exhibition hall. Capture short async studies there instead and let people complete them on the train home.\n\nWhichever channel you use, **verify**. Ask for a licence or registration number where lawful, cross-check specialty against claimed practice setting, and build screener questions that a non-clinician cannot pass. A well-designed clinical screener asks about workflow specifics rather than credentials: which formulation you reach for first in a specific presenting scenario, what your unit's protocol says about a particular threshold, who signs off on a given order. Professional survey-takers fail those; real clinicians answer them without thinking.\n\n## Honoraria and fair market value\n\nHCP payments are not incentives in the consumer sense. They are compensation for professional time, and they must be defensible.\n\nThe mechanics most agencies use:\n\n- **Hourly-rate model.** A per-hour figure set by specialty, seniority and geography, then multiplied by the true participation time — including any pre-work. Life sciences clients typically maintain internal FMV bands and will not approve anything above them.\n- **Record-to-recruit ratio.** Divide the size of the qualified universe by the required sample size. A large universe means a lower honorarium clears; a rare specialty with a universe of a few hundred needs a much higher one.\n- **Published market ranges.** Short polls of around five minutes commonly sit in the $5–$15 range, while 60-minute in-depth interviews commonly run $200–$500 and above. Oncology, immunology and rare disease specialists command the highest effective rates, often $100–$300+ per hour.\n\nThree rules keep this clean. Pay for time, not for outcome or opinion. Document the FMV basis before fielding, not after. And never let honoraria vary by what a participant says.\n\n**Shorter interviews save real money here.** If an AI interviewer can extract the same depth in eighteen minutes that a human moderator needs forty-five minutes to reach — because it never spends time on rapport-building preamble, never re-asks something already answered, and probes only where the answer was thin — the honorarium falls proportionally and the response rate rises. That is the strongest commercial argument for AI moderation in HCP work.\n\n## Transparency reporting and the blinding exclusion\n\nIn the US, the Physician Payments Sunshine Act requires applicable manufacturers to report transfers of value to covered recipients through Open Payments. Market research is one of the few areas with a practical carve-out.\n\n**Payments made in genuinely double-blinded market research are not reportable**, because the manufacturer does not know the identity of the participant and the participant does not know the sponsor. The exclusion also applies where only the manufacturer is blinded — single-blind research where the sponsor never learns who participated. If a study is only partially blinded, only the blinded portion falls outside reporting.\n\nThe operational consequence: **decide the blinding design before recruitment, not after.** If your sponsor needs to know exactly which KOLs participated — which is often legitimate for advisory work — accept that it is reportable and budget the compliance process. If the research genuinely does not need identities, keep the fieldwork with an agency or platform that holds them and deliver de-identified results. Retro-fitting blinding to a study that already exchanged names does not work.\n\nThis is one reason to keep participant identity separated from analysis output. Koji lets you run studies where the sponsor sees aggregated findings and anonymised quotes while identity stays with whoever recruited — and de-identification technique is covered in depth in the anonymisation guide linked below.\n\n## Adverse event reporting: the obligation nobody plans for\n\nIf your research touches a marketed medicinal product, a participant may mention a suspected adverse reaction. That mention can create a pharmacovigilance obligation for the marketing authorisation holder, and the obligation exists whether or not you were looking for it.\n\nEphMRA maintains the standard framework here — a Code of Conduct alongside dedicated **Adverse Event Reporting Guidelines, most recently revised in September 2025** — and it complements the ICC/ESOMAR code. Practically, a compliant HCP study needs five things in place before it fields:\n\n1. **Trained interviewers or a trained review process.** Everyone touching the data must be able to recognise an adverse event mention.\n2. **A defined reporting route and timeline** to the client's pharmacovigilance function. Reporting windows are tight and are set by the client's SOP — confirm the exact window contractually before fielding.\n3. **The four case elements.** A reportable case classically requires an identifiable reporter, an identifiable patient, a suspect product and an event. Your process must be able to capture all four, or explain why it cannot.\n4. **Re-contact consent.** Pharmacovigilance follow-up may need to go back to the reporter. If your consent form forbids re-contact, you have built a dead end.\n5. **A reminder script.** The interviewer must remind a respondent who mentions an adverse event to report it through normal channels.\n\nAsynchronous AI interviews do not remove this obligation — they change where it sits. Because every session is transcribed in full, adverse event review moves from \"did the moderator notice?\" to a systematic pass over complete transcripts, which is a genuinely stronger control than relying on live recall. Build the review step into your process and give it a named owner.\n\n## Designing an interview a clinician will actually finish\n\n- **Cap it at 20 minutes of participant time.** Publish the cap and honour it. Clinicians abandon studies that overrun more readily than any other audience.\n- **Lead with the clinical scenario, not the product.** \"Walk me through the last patient where you considered switching therapy\" outperforms any attitudinal opener.\n- **Use structured questions where you need to count.** Koji supports six types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no`. Rank treatment attributes with `ranking`, capture confidence on a `scale`, take the free-text reasoning with `open_ended`, and let the AI probe every low score automatically.\n- **Offer voice.** Speaking is faster than typing for a clinician between patients, and voice answers run considerably longer and richer than typed ones.\n- **Stay non-promotional.** Research must not drift into product promotion, and it must never discuss unapproved uses. Have medical or legal review the discussion guide when the sponsor is a manufacturer.\n- **Localise properly.** Multi-market HCP studies need genuine language coverage, not a translated survey — clinical terminology carries local convention.\n\n## Mistakes that cost studies\n\n- Recruiting \"physicians\" without verifying specialty, then discovering half the sample is out of scope at analysis.\n- Setting honoraria by negotiation rather than by documented FMV basis.\n- Deciding blinding after names have already reached the sponsor.\n- Fielding a study touching a marketed product without an adverse event process or trained reviewers.\n- Booking 60-minute moderated calls when a 20-minute async study answers the same questions with three times the sample.\n- Treating patients, caregivers and clinicians as one study. They need separate designs, separate consent and — where patients are involved — a much closer look at whether ethical review applies.\n\n## How Koji fits HCP research\n\nKoji runs AI-moderated interviews by voice or text that clinicians complete on their own schedule, in their own language, with automatic follow-up probing that reaches the depth a static survey never gets to. Personalised interview links let you field to a verified HCP list without re-screening people you already know, CSV import handles the list itself, and every session produces a full transcript plus structured answers that aggregate into a report immediately — so a study that would have taken six weeks of scheduling closes in days.\n\nUsed properly, that is not just faster. Shorter sessions lower honoraria, higher completion improves sample quality, and full transcription makes adverse event review systematic rather than dependent on a moderator's memory.\n\n## Frequently asked questions\n\n**Do market research payments to physicians have to be reported under the Sunshine Act?**\nNot when the research is genuinely double-blinded — the manufacturer does not know who participated and the participant does not know the sponsor. The exclusion also covers single-blind designs where only the manufacturer is blinded. If a study is partially blinded, only the blinded portion falls outside reporting. Decide the blinding design before recruiting, because it cannot be retro-fitted once identities have been shared.\n\n**What honorarium should we offer a physician for a 60-minute interview?**\nPublished market ranges commonly put 60-minute in-depth interviews at $200–$500 and above, with oncology, immunology and rare-disease specialists at the top. The right figure comes from your fair market value basis: specialty, seniority, geography and true participation time, sized against how rare the qualified universe is. Document that basis before fielding, and pay for time rather than for opinions.\n\n**Do we need ethics or IRB review for HCP market research?**\nMarket research that produces business decisions rather than generalisable knowledge usually falls outside the human-subjects research definition, but the answer depends on your jurisdiction, on whether patients or patient data are involved, and on whether you intend to publish. Studies touching patients, clinical outcomes or identifiable health information should be assessed properly rather than assumed exempt.\n\n**How do we handle an adverse event mentioned during an interview?**\nCapture it, remind the respondent to report it through normal channels, and pass it to the client's pharmacovigilance function within the window their SOP specifies. A reportable case classically needs an identifiable reporter, an identifiable patient, a suspect product and an event. Make sure your consent permits the re-contact that follow-up may require, and give the review step a named owner rather than leaving it to whoever reads the transcripts.\n\n**How do you verify that a research participant is really the specialist they claim to be?**\nCombine credential checks where lawful with clinical screener questions a non-clinician cannot answer — the protocol threshold their unit uses, who signs off a particular order, which formulation they reach for first in a specific presenting scenario. Credential checks catch impostors; workflow questions catch people who hold the credential but do not do the work in question.\n\n**Is asynchronous AI interviewing acceptable for HCP research?**\nYes, and it usually outperforms scheduled calls for this audience. Median physician response rates around 18% reflect calendar friction more than unwillingness. An async voice study removes the scheduling problem, shortens participation time, lowers the honorarium accordingly, and produces complete transcripts that make compliance review systematic.\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and when to use each with time-poor experts\n- [Expert Interviews Guide](/docs/expert-interviews-guide) — the general playbook for interviewing specialists\n- [AI Customer Research for Pharma & Life Sciences](/docs/ai-research-for-pharma-life-sciences) — the wider industry view\n- [Research Participant Incentives](/docs/research-participant-incentives) — incentive design outside regulated audiences\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — how to deliver findings without identities\n- [IRB Approval for User Research](/docs/irb-approval-user-research) — when ethical review actually applies\n- [Multi-Language User Research](/docs/multilingual-research-guide) — running multi-market clinical studies properly","category":"Participant Recruitment","lastModified":"2026-07-31T03:19:12.482477+00:00","metaTitle":"HCP Research: Recruiting and Interviewing Physicians and Nurses","metaDescription":"How to recruit healthcare professionals for research: verification screeners, fair market value honoraria, Sunshine Act blinding rules, adverse event reporting, and interview design for time-poor clinicians.","keywords":["HCP research","physician interviews","recruiting healthcare professionals for research","KOL interviews","nurse research participants","fair market value honoraria","Sunshine Act market research","medical affairs research","pharma market research recruiting"],"aiSummary":"Research with healthcare professionals carries four constraints beyond ordinary qualitative research: low access, regulated payment, transparency reporting and adverse event obligations. Median physician response rates in research are around 18 percent with a range of roughly 10 to 60 percent, which is why asynchronous AI-moderated interviewing outperforms scheduled calls for this audience. Recruiting channels ranked by yield are own customer lists, professional associations, specialist HCP panels, referral and snowball recruiting, and congress intercepts, and every channel requires verification through credential checks plus clinical screener questions a non-clinician cannot answer. Honoraria must be set on a documented fair market value basis using an hourly rate by specialty and seniority and a record-to-recruit ratio against the qualified universe; published ranges put five-minute polls at 5 to 15 dollars and 60-minute in-depth interviews at 200 to 500 dollars and above, with oncology, immunology and rare-disease specialists at 100 to 300 dollars per hour and higher. Under the US Physician Payments Sunshine Act, payments in genuinely double-blinded market research are not reportable to Open Payments, and the exclusion also applies to single-blind designs where only the manufacturer is blinded; in a partially blinded study only the blinded portion is excluded, and blinding cannot be retro-fitted after identities are shared. Research touching a marketed medicinal product requires an adverse event process: trained reviewers, a defined reporting route and timeline set by the client SOP, capture of the four case elements of identifiable reporter, identifiable patient, suspect product and event, re-contact consent, and a reminder script. EphMRA maintains the Code of Conduct and Adverse Event Reporting Guidelines revised in September 2025. Interviews should cap participant time at around 20 minutes, lead with clinical scenarios, offer voice, stay non-promotional and use structured question types to produce countable data.","aiPrerequisites":["A defined clinical audience and specialty scope","A documented fair market value basis for honoraria if a manufacturer is funding the research","An adverse event reporting route agreed with the client if a marketed product is in scope"],"aiLearningOutcomes":["Rank HCP recruiting channels by yield, cost and bias","Write screener questions that a non-clinician cannot pass","Set and document defensible fair market value honoraria","Apply the Sunshine Act blinding exclusion correctly before recruitment starts","Build an adverse event capture and escalation process into a research study","Design a 20-minute asynchronous interview that time-poor clinicians complete"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"},{"type":"blog","id":"22a10776-8c88-4c60-b4cf-c7bc0af2063e","slug":"maze-pricing-2026","title":"Maze Pricing in 2026: What It Actually Costs (Seats, Credits & the Enterprise Gate)","url":"https://www.koji.so/blog/maze-pricing-2026","summary":"Maze pricing in 2026: maze.co/pricing publishes no dollar figures, only a feature comparison. Vendr benchmark data across 33 purchases shows an average annual contract of $29,000, range $12,000-$62,682. Maze public feature table marks AI Study Builder, Maze AI, Interview Studies, mobile testing and SSO as Enterprise-only. Panel recruitment is billed separately per response with credits that expire after 12 months and do not roll over. Koji publishes EUR 29-79/month pricing with AI-moderated interviews on every plan and no seat minimum.","content":"## The Short Answer\n\n**Maze does not publish pricing.** As of 31 July 2026, maze.co/pricing is a feature-comparison table with no dollar figures on it — every tier resolves to a form or a sales conversation.\n\nWhat buyers actually pay is documented. [Vendr data across 33 Maze purchases](https://www.vendr.com/marketplace/maze) shows an **average and median annual contract of $29,000**, with a range from **$12,000 to $62,682**.\n\nThird-party trackers report list prices, but they **disagree with each other** — a predictable consequence of unpublished pricing. The most commonly cited figures are a **Starter tier somewhere between $49 and $99 per month** (lower when billed annually), a **Team tier around $249/month for three seats**, and an **Organization tier frequently cited near $15,000/year**, with per-seat estimates of **$200+** at that level. Treat all of these as directional, not quotable.\n\nThe number you can verify yourself is the feature table — and it tells you something more useful than a price: **Maze's AI capabilities are largely Enterprise-gated**. AI Study Builder, Maze AI, Interview Studies, mobile testing and SSO are all marked Enterprise on the public comparison.\n\nKoji does the opposite on both counts. Pricing is on the website — **€29/month (Insights)** or **€79/month (Interviews)** — and **AI-moderated interviews are the product, not the top tier**.\n\n---\n\n## What Maze's Pricing Page Actually Tells You\n\nSince there are no prices, the comparison table is the real document. Reading it top to bottom:\n\n| Capability | Availability per Maze's public table |\n|---|---|\n| Prototype testing | All plans (limits apply) |\n| Panel recruitment | All plans (limits apply) |\n| Automated reporting | All plans |\n| **AI Study Builder** | **Enterprise** |\n| **Maze AI** | **Enterprise** |\n| **Interview studies** | **Enterprise** |\n| **Mobile testing** | **Enterprise** |\n| **Security / SSO** | **Enterprise** |\n\nThe pattern is clear: **the unmoderated testing core is available broadly; the AI layer and the interview layer are not.** If you are evaluating Maze because you want AI-assisted research or moderated-style interviews, you are evaluating the Enterprise tier — and Enterprise means a quote, a procurement cycle, and, based on the benchmark data, a five-figure annual commitment.\n\n\"Limits apply\" on prototype testing and panel recruitment is doing quiet work too. Free and entry tiers cap studies and responses; the free plan is generally reported as roughly **one study per month with Maze branding on it**.\n\n---\n\n## What Buyers Actually Pay\n\nThe [Vendr benchmark across 33 Maze contracts](https://www.vendr.com/marketplace/maze):\n\n- **Average / median annual contract: $29,000**\n- **Range: $12,000 to $62,682**\n- **Volume discounts (5+ seats):** 10–25% off list\n- **Multi-year commitments (2–3 years):** 15–30% lower per-seat pricing\n- **Annual prepayment:** 5–15% versus quarterly billing\n\nFor context, that median is **higher than the equivalent [Dovetail benchmark](/blog/dovetail-pricing-2026) ($21,600)** and higher than the [SurveyMonkey median](/blog/surveymonkey-pricing-2026) ($16,482) from the same dataset family — see the full [user research tool pricing comparison](/blog/user-research-tool-pricing-2026). Maze is not a cheap tool, and the contracts cluster well above what the reported Starter list price would suggest — because the tiers people actually buy are the ones with the AI and interview features in them.\n\n## The Credits Problem\n\nMaze recruitment is billed **separately from the subscription**. Panel responses are charged per response collected, and reporting on the credit system consistently notes two things worth budgeting for:\n\n- **Credits expire 12 months from purchase**\n- **Unused credits do not roll over**\n\nThat is a use-it-or-lose-it clause on a line item that is already unpredictable. Research volume is lumpy — a quarter with a big launch burns credits fast, a quiet quarter burns none — and the expiry rule converts that natural variance into waste.\n\nAdvanced screening is also gated: entry tiers get basic recruitment at roughly one credit per participant but cannot use the more precise screening criteria, which is exactly the feature that determines whether the participants you paid for are the ones you needed.\n\n---\n\n## The Seat Math\n\nPer-seat pricing has a specific failure mode in research tools, and Maze sits right in it.\n\nA research tool has one or two people who **build studies** and a much larger group who **need to see results**. Charge per seat, and the second group becomes a budget line — so teams buy fewer seats, results stay locked inside the tool, and the research does not reach the people making decisions.\n\nIndustry benchmarks put [SaaS licence utilisation at roughly 49–54%](https://zylo.com/blog/how-much-wasted-on-saas-spend) — about half of purchased seats go unopened in a given month. At an estimated $200+ per seat on Organization tier, each dormant seat is **$2,400 a year** for a login nobody uses.\n\nIf you are trying to justify or defend that spend, [proving research ROI](/docs/research-roi-guide) and our [ResearchOps guide](/docs/research-ops-guide) cover how to audit seats and consolidate before a renewal lands.\n\n---\n\n## Where Maze Is Genuinely Good\n\nMaze earns its place for a specific job:\n\n- **Rapid unmoderated usability testing** on prototypes, especially Figma flows\n- **Quantitative usability metrics** at speed — task success, time on task, misclick rates\n- **Design-team velocity**: ship a prototype, get numbers back the same day\n- **Automated reporting** that non-researchers can read\n\nIf your question is *\"can users complete this flow, and how long does it take\"*, Maze answers it well. Our [unmoderated usability testing guide](/docs/unmoderated-usability-testing-guide) and [usability metrics guide](/docs/usability-metrics-guide) cover what that method does and does not tell you.\n\n## Where It Stops\n\nUnmoderated testing measures **behaviour on a task**. It cannot ask why someone hesitated, what they expected instead, or what they were actually trying to accomplish. When the click-path data shows a 40% drop-off, Maze tells you *that* it happened. It does not tell you *why* — and the why is what changes the design.\n\nMaze's answer to that is Interview Studies. Which is Enterprise-gated.\n\n---\n\n## Maze vs Koji: The Pricing Model Compared\n\n| | Maze | Koji |\n|---|---|---|\n| **Published price** | None — quote only | €29/mo (Insights), €79/mo (Interviews) |\n| **Benchmark contract** | $29,000/year average | Self-serve, monthly or annual |\n| **Seats** | Per seat | No seat minimum, no per-user fee |\n| **AI features** | Enterprise tier | Every plan |\n| **Interview studies** | Enterprise tier | Core product |\n| **Participant billing** | Credits, expire in 12 months, no rollover | Credits included in plan; prepaid packs for overage |\n| **Unused-credit policy** | Expire after 12 months | Quality gate — only conversations scoring 3+ consume credits |\n| **Follow-up questions** | Not in unmoderated tests | AI probes every answer in real time |\n| **Analysis** | Automated usability reporting | Automatic thematic analysis across transcripts |\n\nTwo differences do most of the work here.\n\n**First: what the AI costs.** At Maze, the AI layer is the reason to move to Enterprise. At Koji, [AI-moderated interviews](/docs/ai-moderated-interviews) — [voice or text](/docs/ai-voice-interviews) — are what you get on a €79/month plan. There is no tier where the AI turns on.\n\n**Second: what you pay for.** Maze charges for seats and for panel credits that expire. Koji charges for research that actually happened: **text conversations cost 1 credit, voice interviews 3, a report refresh 5**, and a quality gate means **only conversations scoring 3 or above consume credits at all**. A junk response does not bill you.\n\nKoji also keeps quantitative rigour inside the conversation rather than making you choose between numbers and reasons. **Six structured question types** — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — produce chartable data in the same session where the AI is probing for the why behind each answer ([structured questions guide](/docs/structured-questions-guide)). For usability work specifically, [AI usability testing](/docs/ai-usability-testing-guide) covers how an AI moderator handles a task-based study.\n\nThe full side-by-side lives at [Koji vs. Maze](/docs/koji-vs-maze).\n\n---\n\n## How to Negotiate a Maze Contract\n\n1. **Anchor on $29,000.** It is the published average across 33 real contracts. Sales will open higher.\n2. **Ask which tier unlocks AI in writing.** The public table says Enterprise. Confirm before you scope a pilot around features you cannot access.\n3. **Negotiate credit rollover.** The 12-month expiry with no rollover is a policy, and policies are negotiable at contract time. Ask for rollover or a longer window.\n4. **Separate builders from viewers.** Push for low-cost or free read-only access so results actually circulate.\n5. **Prepay annually, but check the escalator.** The 5–15% prepayment discount is real; multi-year uplift clauses often claw it back.\n\n---\n\n## Frequently Asked Questions\n\n### How much does Maze cost in 2026?\n\nMaze does not publish prices. Benchmark data across 33 purchases shows an average annual contract of $29,000, ranging from $12,000 to $62,682. Third-party sources report a Starter tier between roughly $49 and $99/month and an Organization tier often cited near $15,000/year, but these figures conflict and should be verified directly with Maze.\n\n### Does Maze have a free plan?\n\nYes, with tight limits — generally reported as about one study per month, restricted blocks and responses, and Maze branding on your study. It is suitable for trying the product, not for running ongoing research.\n\n### Are Maze's AI features available on all plans?\n\nNo. Maze's public feature comparison marks AI Study Builder, Maze AI, Interview Studies, mobile testing and SSO as Enterprise. The unmoderated prototype-testing core and automated reporting are available more broadly, with limits.\n\n### How do Maze credits work?\n\nPanel recruitment is billed separately from the subscription, per response collected. Credits expire 12 months after purchase and unused credits do not roll over, so unpredictable research volume can turn into wasted spend.\n\n### Why does Maze not publish pricing?\n\nUnpublished pricing lets a vendor quote by segment and keeps buyers from benchmarking against a list price. The practical effect is that every real purchase becomes a sales cycle, and third-party estimates diverge — which is why negotiated-contract data is more reliable than any published estimate.\n\n### What is a cheaper alternative to Maze?\n\nFor unmoderated prototype testing, several [Maze alternatives](/blog/maze-alternatives-2026) compete below its benchmark. If you need the why behind the behaviour, Koji runs AI-moderated interviews from €29/month with no seat minimum and no Enterprise gate on the AI features.\n\n---\n\n## The AI Should Not Be the Upsell\n\nPaying five figures a year to reach the tier where the AI turns on is a 2023 pricing model applied to 2026 expectations.\n\nKoji ships AI moderation, automatic thematic analysis, and one-click reports on a plan you can buy with a card in about a minute. No seat minimums, no expiring credits, no sales call to learn the price.\n\n**[Start free with Koji](https://www.koji.so)** — from question to insight in hours, not weeks, with no research expertise required.","category":"Comparisons","lastModified":"2026-07-31T03:18:58.417069+00:00","metaTitle":"Maze Pricing 2026: Real Costs, Credits & the Enterprise AI Gate","metaDescription":"Maze publishes no prices. See what buyers actually pay (average $29,000/year), how expiring panel credits work, and which AI features sit behind the Enterprise tier.","keywords":["maze pricing","maze cost","maze pricing 2026","how much does maze cost","maze.co pricing","maze credits","maze enterprise plan","maze starter plan price"],"aiSummary":"Maze pricing in 2026: maze.co/pricing publishes no dollar figures, only a feature comparison. Vendr benchmark data across 33 purchases shows an average annual contract of $29,000, range $12,000-$62,682. Maze public feature table marks AI Study Builder, Maze AI, Interview Studies, mobile testing and SSO as Enterprise-only. Panel recruitment is billed separately per response with credits that expire after 12 months and do not roll over. Koji publishes EUR 29-79/month pricing with AI-moderated interviews on every plan and no seat minimum.","aiKeywords":["maze pricing","usability testing cost","per-seat pricing","research budget","unmoderated testing","ai moderated interviews"],"aiContentType":"comparison","faqItems":[{"answer":"Maze does not publish prices. Benchmark data across 33 purchases shows an average annual contract of $29,000, ranging from $12,000 to $62,682. Third-party sources report a Starter tier between roughly $49 and $99/month and an Organization tier often cited near $15,000/year, but these figures conflict and should be verified directly with Maze.","question":"How much does Maze cost in 2026?"},{"answer":"Yes, with tight limits — generally reported as about one study per month, restricted blocks and responses, and Maze branding on your study. It is suitable for trying the product, not for running ongoing research.","question":"Does Maze have a free plan?"},{"answer":"No. Maze's public feature comparison marks AI Study Builder, Maze AI, Interview Studies, mobile testing and SSO as Enterprise. The unmoderated prototype-testing core and automated reporting are available more broadly, with limits.","question":"Are Maze's AI features available on all plans?"},{"answer":"Panel recruitment is billed separately from the subscription, per response collected. Credits expire 12 months after purchase and unused credits do not roll over, so unpredictable research volume can turn into wasted spend.","question":"How do Maze credits work?"},{"answer":"Unpublished pricing lets a vendor quote by segment and keeps buyers from benchmarking against a list price. The practical effect is that every purchase becomes a sales cycle and third-party estimates diverge, which is why negotiated-contract data is more reliable than published estimates.","question":"Why does Maze not publish pricing?"},{"answer":"For unmoderated prototype testing, several tools compete below Maze's benchmark. If you need the why behind the behaviour, Koji runs AI-moderated interviews from EUR 29/month with no seat minimum and no Enterprise gate on the AI features.","question":"What is a cheaper alternative to Maze?"}],"relatedTopics":["maze","pricing","usability testing","research budget","per-seat pricing"]},{"type":"blog","id":"bff97a22-b1f1-41bc-a6e2-5d0b386fe7ac","slug":"dovetail-pricing-2026","title":"Dovetail Pricing in 2026: What a Research Repository Actually Costs","url":"https://www.koji.so/blog/dovetail-pricing-2026","summary":"Dovetail pricing in 2026: the public pricing page lists only a Free plan ($0, one channel, one project) and a custom-quoted Enterprise plan — the previously reported Professional tier at roughly $39/editor/month is no longer self-serve. Vendr benchmark data across 65 contracts shows a median of $21,600/year, range $10,800-$61,200, with multi-year discounts of 15-25% often carrying 3-7% annual escalators. A repository stores research but does not produce it, so recruiting, moderating and transcription costs sit outside the subscription. Koji publishes EUR 29-79/month pricing with no seat minimum and runs the interviews itself.","content":"## The Short Answer\n\nAs of 31 July 2026, **Dovetail no longer publishes per-seat pricing**. Its pricing page lists exactly two options: a **Free plan at $0 with no card required**, and an **Enterprise plan at \"custom pricing available\"**. The mid-market self-serve tier that buyers used to budget against is not on the page.\n\nThis matters, because most third-party guides still quote a **Professional plan at roughly $39 per editor per month** (sources range from $29 to $49), plus a **Channels add-on from around $50/month**. Those figures reflect an earlier packaging. If you are building a 2026 budget from a comparison article, **you may be budgeting against a plan you can no longer self-serve**.\n\nThe reliable number is what buyers actually sign — and it sits alongside [UserTesting](/blog/usertesting-pricing-2026), [Qualtrics](/blog/qualtrics-pricing-2026) and [Maze](/blog/maze-pricing-2026) in our [user research tool pricing comparison](/blog/user-research-tool-pricing-2026). [Vendr data across 65 Dovetail deals](https://www.vendr.com/marketplace/dovetail) puts the **median contract at $21,600 per year**, with a range of **$10,800 to $61,200**.\n\nKoji publishes its pricing and always has: **€29/month (Insights)** or **€79/month (Interviews)**, no seat minimum, no sales call required to see a number.\n\n---\n\n## What Dovetail's Pricing Page Shows Today\n\n| Plan | Price | What you get |\n|---|---|---|\n| **Free** | $0, no card required | One channel, one project, basic AI Chat, basic AI summaries, integrations, calendar sync, search, contacts |\n| **Enterprise** | Custom quote | Unlimited Agents, unlimited channels, projects and Docs, advanced AI summaries and semantic search, unlimited chat queries in Slack and Teams, folders, global tags, templates, SSO, access control, data redaction, HIPAA add-on, dedicated CSM, priority support |\n\nThe Free plan is genuinely usable for evaluation — **one channel and one project** is enough to see how the repository feels. It is not enough to run a research practice: a single project means one study at a time, and the AI features are the basic tier.\n\nEverything else — every feature a team of more than one person needs — sits behind a quote.\n\n## What This Packaging Change Means for Buyers\n\nMoving from \"published tiers plus Enterprise\" to \"Free plus Enterprise\" is a familiar move in B2B SaaS, and it has three practical consequences:\n\n**1. You cannot benchmark yourself.** With no list price, you have no anchor. The only defence is third-party spend data — which is why the $21,600 median matters more than any blog's list price.\n\n**2. Every real purchase becomes a sales cycle.** A team that wants two seats and three projects now enters the same motion as a 200-seat enterprise. Budget approval takes longer, and procurement gets involved earlier.\n\n**3. Price discovery moves to renewal.** You learn what expansion costs when you need it, not when you are choosing.\n\nIf you are heading into that conversation, our [ResearchOps guide](/docs/research-ops-guide) covers consolidating tooling before a renewal, and [proving research ROI](/docs/research-roi-guide) helps build the internal case for whatever number comes back.\n\n---\n\n## What Buyers Actually Pay\n\nThe [Vendr benchmark of 65 Dovetail contracts](https://www.vendr.com/marketplace/dovetail) is the most useful public dataset:\n\n- **Median annual contract: $21,600**\n- **Range: $10,800 to $61,200**\n- **Small teams (5–10 seats):** reported 5–15% below list\n- **Mid-sized teams (10–25 seats):** 10–20% below list\n- **Larger teams (25+ seats):** 20–30% off initial Enterprise quotes\n- **Multi-year (2–3 years):** 15–25% discount versus annual — but frequently with **3–7% annual price escalation clauses** that erode the saving\n\nThat last point is the one buyers miss. A 20% multi-year discount with a 7% annual escalator is not a 20% saving by year three. Read the uplift clause before you sign the discount.\n\n---\n\n## The Cost a Repository Does Not Cover\n\nHere is the structural thing about repository pricing that no plan comparison shows: **a repository does not produce research. It stores it.**\n\nDovetail is an analysis and storage layer. To fill it, you still need:\n\n- A way to **recruit participants** — a panel, a recruiting tool, or your own list\n- A way to **run the sessions** — Zoom, a moderated testing platform, or a researcher's calendar\n- A **person to moderate** — the single largest cost in most research programmes\n- **Transcription**, if it is not bundled\n- Then, finally, **the repository seat** to tag and organise what came out\n\nWhen teams compare \"Dovetail vs alternatives\" on price alone, they are comparing one line of a five-line budget. The repository seat is rarely the expensive part. The **moderator's time is**, and every hour of it is spent before a single tag gets applied.\n\nThis is the point our [research repository guide](/docs/research-repository-guide) makes at length, and why [insight repository methodology](/docs/insight-repository-methodology) argues that storage without activation is a filing cabinet with a subscription fee.\n\n### The seat-utilisation problem\n\nRepositories are also where per-seat pricing goes wrong fastest. A repository has two user types: **the few people who tag and synthesise**, and **the many who occasionally read**. Per-editor pricing charges for the first group correctly and then makes the second group expensive to include — which is precisely backwards, because a repository nobody reads has failed at its only job.\n\nIndustry benchmarks put [SaaS licence utilisation around 49–54%](https://zylo.com/blog/how-much-wasted-on-saas-spend) — roughly half of purchased seats go unopened in a given month. In a repository, that ratio tends to be worse, not better.\n\n---\n\n## Dovetail vs Koji: What You Are Buying\n\n| | Dovetail | Koji |\n|---|---|---|\n| **Published price** | Free tier only; everything else custom | €29/mo (Insights), €79/mo (Interviews) |\n| **Seats** | Per editor, quote-based | No seat minimum, no per-user fee |\n| **Collects data?** | No — you bring transcripts | Yes — runs the interviews itself |\n| **Moderator required** | Yes, for interviews | No — AI moderates every session |\n| **Analysis** | AI summaries and tagging on stored data | Automatic thematic analysis on every study |\n| **Structured quant** | Not a survey tool | Six question types built in |\n| **Time to first insight** | However long recruiting and moderating takes | Hours |\n| **Contract** | Annual, quote-based | Monthly or annual (2 months free) |\n\nDovetail is a good repository. The question is whether a repository is what you are short of.\n\n**Koji covers the whole loop.** It recruits into a study, runs the interview with an AI moderator — [voice or text](/docs/ai-voice-interviews) — probes every answer with follow-ups a form cannot ask, then produces the [thematic analysis](/docs/thematic-analysis-guide) and a one-click report at the end. There is no separate moderating step, so there is no moderator cost, and no gap between the session and the synthesis.\n\nIt also does something no repository can: **structured questions inside a conversation**. Six types — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — so you get clean, chartable quantitative data *and* the reasoning behind each answer in the same session ([structured questions guide](/docs/structured-questions-guide)). A repository can only analyse what you managed to collect elsewhere.\n\nAnd the pricing model matches the work: **text conversations cost 1 credit, voice interviews 3, a report refresh 5**, with a quality gate that means **only conversations scoring 3 or above consume credits at all**. You pay for research that happened, not for people who have logins.\n\nIf you want the direct feature-level comparison, we maintain [Koji vs. Dovetail](/docs/koji-vs-dovetail).\n\n---\n\n## How to Negotiate a Dovetail Contract\n\n1. **Anchor on the median, not the ask.** $21,600/year across 65 deals is a real number. Bring it.\n2. **Separate editors from viewers.** Push for read-only access at low or no cost. If everyone who should read insights needs a paid seat, the repository will not get read.\n3. **Get the escalator in writing.** A multi-year discount with a 7% annual uplift is a different deal than it looks.\n4. **Price the whole stack.** Ask what the total is once recruiting, moderating, and transcription are included. Compare *that* to alternatives, not the seat price.\n5. **Run the free tier first.** One channel and one project is enough to test whether the team will actually adopt it before you commit to a quote.\n\n---\n\n## Frequently Asked Questions\n\n### How much does Dovetail cost in 2026?\n\nDovetail publishes only a Free plan ($0) and a custom-quoted Enterprise plan on its pricing page as of July 2026. Third-party spend data across 65 contracts puts the median at $21,600 per year, ranging from $10,800 to $61,200.\n\n### Does Dovetail still have a Professional plan?\n\nDovetail's public pricing page currently shows only Free and Enterprise. Many comparison articles still quote a Professional tier at roughly $39 per editor per month, which reflects earlier packaging — verify directly with Dovetail rather than budgeting from those figures.\n\n### Is Dovetail's free plan enough for a small team?\n\nIt is enough to evaluate the product, not to run a practice. The free tier is limited to one channel and one project with basic AI features, so you cannot run parallel studies or use the advanced semantic search and summarisation.\n\n### What is the Channels add-on and what does it cost?\n\nChannels is Dovetail's automated feedback-ingestion feature, pulling in sources like support tickets and reviews. Third-party sources have reported it starting around $50/month, though it now appears bundled into Enterprise packaging rather than sold as a published add-on.\n\n### Does Dovetail run interviews or just analyse them?\n\nDovetail is an analysis and storage layer. You recruit participants, moderate the sessions, and bring transcripts in. The moderating cost — usually the largest line in a research budget — sits outside the subscription.\n\n### What is a cheaper alternative to Dovetail?\n\nIf you need storage only, several [Dovetail alternatives](/blog/dovetail-alternatives-2026) compete on price. If the actual bottleneck is running the research, Koji starts at €29/month with no seat minimum and handles recruiting through analysis in one tool, removing the separate moderator cost entirely.\n\n---\n\n## The Repository Is Not the Bottleneck\n\nMost teams do not have an insight-storage problem. They have an insight-*production* problem: research takes too long, needs a trained moderator, and by the time it is tagged the decision has been made.\n\nKoji closes that loop. AI-moderated interviews that run while you sleep, thematic analysis the moment the last participant finishes, and a report your stakeholders can read without translation. No seat minimums. No custom quote to find out the price.\n\n**[Start free with Koji](https://www.koji.so)** — from question to insight in hours, not weeks.","category":"Comparisons","lastModified":"2026-07-31T03:18:58.417069+00:00","metaTitle":"Dovetail Pricing 2026: Real Costs, Seat Fees & What Buyers Pay","metaDescription":"Dovetail now publishes only a Free and custom Enterprise plan. See what buyers actually pay (median $21,600/year), how discounts work, and the research costs a repository does not cover.","keywords":["dovetail pricing","dovetail cost","dovetail pricing 2026","how much does dovetail cost","dovetail per user pricing","dovetail enterprise cost","research repository pricing"],"aiSummary":"Dovetail pricing in 2026: the public pricing page lists only a Free plan ($0, one channel, one project) and a custom-quoted Enterprise plan — the previously reported Professional tier at roughly $39/editor/month is no longer self-serve. Vendr benchmark data across 65 contracts shows a median of $21,600/year, range $10,800-$61,200, with multi-year discounts of 15-25% often carrying 3-7% annual escalators. A repository stores research but does not produce it, so recruiting, moderating and transcription costs sit outside the subscription. Koji publishes EUR 29-79/month pricing with no seat minimum and runs the interviews itself.","aiKeywords":["dovetail pricing","research repository cost","per-seat pricing","research budget","ux research tools","ai moderated interviews"],"aiContentType":"comparison","faqItems":[{"answer":"Dovetail publishes only a Free plan ($0) and a custom-quoted Enterprise plan on its pricing page as of July 2026. Third-party spend data across 65 contracts puts the median at $21,600 per year, ranging from $10,800 to $61,200.","question":"How much does Dovetail cost in 2026?"},{"answer":"Dovetail's public pricing page currently shows only Free and Enterprise. Many comparison articles still quote a Professional tier at roughly $39 per editor per month, which reflects earlier packaging — verify directly with Dovetail rather than budgeting from those figures.","question":"Does Dovetail still have a Professional plan?"},{"answer":"It is enough to evaluate the product, not to run a practice. The free tier is limited to one channel and one project with basic AI features, so you cannot run parallel studies or use advanced semantic search and summarisation.","question":"Is Dovetail's free plan enough for a small team?"},{"answer":"Channels is Dovetail's automated feedback-ingestion feature, pulling in sources like support tickets and reviews. Third-party sources have reported it starting around $50/month, though it now appears bundled into Enterprise packaging rather than sold as a published add-on.","question":"What is the Channels add-on and what does it cost?"},{"answer":"Dovetail is an analysis and storage layer. You recruit participants, moderate the sessions, and bring transcripts in. The moderating cost — usually the largest line in a research budget — sits outside the subscription.","question":"Does Dovetail run interviews or just analyse them?"},{"answer":"If you need storage only, several repositories compete on price. If the actual bottleneck is running the research, Koji starts at EUR 29/month with no seat minimum and handles recruiting through analysis in one tool, removing the separate moderator cost.","question":"What is a cheaper alternative to Dovetail?"}],"relatedTopics":["dovetail","pricing","research repository","research budget","per-seat pricing"]},{"type":"blog","id":"ce743c54-c666-4ad7-ba53-09417faa37fa","slug":"surveymonkey-pricing-2026","title":"SurveyMonkey Pricing in 2026: Plans, Per-User Costs & What You Actually Pay","url":"https://www.koji.so/blog/surveymonkey-pricing-2026","summary":"SurveyMonkey 2026 pricing: Team Advantage EUR 30/user/month and Team Premier EUR 75/user/month, both with a three-seat minimum, plus custom Enterprise on five seats. Individual plans run $46-$139/month. Response overage is EUR 0.10 each. Median negotiated contract is $16,482/year across 193 purchases with 18.31% average savings. Koji charges EUR 29-79/month with no seat minimum, pricing research volume rather than headcount.","content":"## The Short Answer\n\nSurveyMonkey sells on a **per-user, per-month basis with a three-seat minimum** on every team plan. As of 31 July 2026, the published team pricing is **€30/user/month for Team Advantage** (50,000 responses per year) and **€75/user/month for Team Premier** (100,000 responses per year), both billed annually, with **Enterprise quoted custom on a five-seat minimum**. Additional responses beyond your plan allowance are billed at **€0.10 each**.\n\nThat means the real floor for a team is not €30 — it is **€90/month (€1,080/year)**, because you cannot buy fewer than three seats.\n\nIndividual plans still exist and are usually the cheaper route for one person: **Standard Monthly at $99/month** is the only no-commitment option, while **Advantage Annual runs roughly $46/month** and **Premier Annual roughly $139/month**, both billed as a full year up front.\n\nWhat buyers actually pay is different again. Spend-benchmark data from [Vendr covering 193 SurveyMonkey purchases](https://www.vendr.com/marketplace/momentive) puts the **median contract at $16,482 per year**, with a range of **$6,491 to $41,275** and **average savings of 18.31%** off list for buyers who negotiate.\n\nKoji takes the opposite approach: **€29/month (Insights)** or **€79/month (Interviews)**, credits included, **no seat minimum, no per-user fee, and no annual commitment**.\n\n> **A note on currency:** SurveyMonkey geo-prices its site. The figures above are what surveymonkey.com/pricing serves in euros. US list pricing is reported at **$30/user/month for Team Advantage** and higher for Team Premier, with third-party trackers disagreeing on the exact Premier figure ($75 to $92). Always check the price your own region is served before budgeting.\n\n---\n\n## SurveyMonkey Team Plans, 2026\n\n| Plan | Price (annual billing) | Seat minimum | Responses included | Notable gates |\n|---|---|---|---|---|\n| **Team Advantage** | €30/user/month | 3 | 50,000/year | No white-label branding, no advanced logic, no phone support |\n| **Team Premier** | €75/user/month | 3 | 100,000/year | Adds white-label, advanced logic, block randomisation, multilingual surveys, crosstabs, phone support |\n| **Enterprise** | Custom quote | 5 | Custom | Adds SSO, governance controls, centralised billing, data-centre choice (US/CA/EU), custom subdomain, HIPAA |\n\nThe jump from Advantage to Premier is **2.5x the price per seat**, and it is where several things most research teams consider basic live: **advanced survey logic, randomisation, multilingual surveys, and crosstab analysis**. If you need cross-tabulated results — the single most common analysis request from a stakeholder — you are on Premier, not Advantage.\n\n## SurveyMonkey Individual Plans, 2026\n\n| Plan | Reported US price | Billing |\n|---|---|---|\n| **Standard Monthly** | $99/month | Monthly, cancel anytime |\n| **Advantage Annual** | ~$46/month | Billed annually (~$552/year) |\n| **Premier Annual** | ~$139/month | Billed annually |\n\nThe pattern here is deliberate and worth naming: **the only plan you can cancel any month is also the most expensive per month**. Monthly flexibility carries roughly a 2x premium over the annual equivalent.\n\n---\n\n## The Three Costs the Pricing Page Does Not Show\n\n### 1. The seat minimum is a floor, not a suggestion\n\nTeam plans start at three seats whether or not three people will use the tool. In practice, teams buy three, one person builds every survey, and two seats sit idle. That is not a hypothetical: industry SaaS benchmarks consistently find **roughly half of purchased licences go unused** — [Zylo and comparable trackers report 49–54% utilisation](https://zylo.com/blog/how-much-wasted-on-saas-spend), meaning about half of provisioned seats are never opened in a given month.\n\nOn Team Premier at €75/seat, two dormant seats cost **€1,800 a year** for nothing.\n\n### 2. Response overage is metered\n\n50,000 responses per year sounds generous until you run a few large-sample studies. Past the allowance, every response is **€0.10**. A single 20,000-response study that tips you over the limit adds €2,000 to the bill at renewal-time reconciliation.\n\n### 3. Feature gates push you up a tier, not out a door\n\nMultilingual surveys, crosstabs, and advanced logic sit on Premier. SSO, HIPAA, and data-residency controls sit on Enterprise. The common upgrade trigger is not survey volume — it is a **compliance or analysis requirement arriving mid-contract**, at which point you are negotiating from a weak position because you already depend on the tool.\n\n---\n\n## What Real Buyers Actually Pay\n\nList price and contract price are different numbers. The [Vendr dataset of 193 SurveyMonkey purchases](https://www.vendr.com/marketplace/momentive) shows:\n\n- **Median annual contract: $16,482**\n- **Range: $6,491 (low) to $41,275 (high)**\n- **Average savings achieved: 18.31%**\n\nBuyers who apply competitive pressure and prepare a benchmark reportedly land **15–30% off list on team plans**. The leverage points are seat count, multi-year commitment, and timing your renewal against the vendor's quarter-end.\n\nIf you are running that negotiation, our guide to [proving research ROI](/docs/research-roi-guide) is useful for building the internal case, and [ResearchOps](/docs/research-ops-guide) covers how to consolidate tooling before renewal rather than after.\n\n---\n\n## Where SurveyMonkey Is Genuinely the Right Tool\n\nBe fair to it. SurveyMonkey is excellent when:\n\n- You need **large-sample quantitative measurement** — NPS tracking, CSAT, market sizing\n- You already know **exactly what to ask** and only need to count answers\n- You need a survey live in **fifteen minutes** with a familiar interface\n- Non-researchers across the business need to self-serve simple forms\n\nIf your question is *\"what percentage of customers prefer A over B\"*, a survey is the correct instrument and SurveyMonkey does it well.\n\n## Where the Per-Response Model Breaks Down\n\nThe limitation is structural, not a bug. A survey can only collect the answers to questions you already thought to ask. When a respondent selects \"Other\" or writes two words in a free-text box, the interesting part — *why* — is gone and unrecoverable.\n\nThat is the gap that pushes teams to look past survey pricing entirely. We cover the trade-off in depth in [AI interviews vs. surveys](/docs/ai-interviews-vs-surveys) and [conversational surveys](/docs/conversational-survey-guide), and if you want to keep your existing questionnaire, [converting a SurveyMonkey survey into an AI interview](/docs/convert-surveymonkey-to-ai-interview) takes about as long as duplicating it.\n\n---\n\n## SurveyMonkey vs Koji: The Pricing Model Compared\n\n| | SurveyMonkey | Koji |\n|---|---|---|\n| **Entry price** | €30/user/month, 3-seat minimum (€90/mo floor) | €29/month, no seat minimum |\n| **Full research tier** | €75/user/month (Team Premier) | €79/month (Interviews) |\n| **Seats** | Charged per user | Unlimited — you pay for research volume, not headcount |\n| **Contract** | Annual for best rate | Monthly or annual (2 months free) |\n| **Overage** | €0.10 per response | €1 per credit, prepaid packs |\n| **Follow-up questions** | None — static form | AI probes every answer automatically |\n| **Analysis** | Crosstabs on Premier only | Automatic thematic analysis on every plan |\n| **Published enterprise price** | No | Custom, but self-serve tiers are public |\n\nThe important difference is not the headline number — it is **what the number is attached to**. SurveyMonkey charges for *people who log in*. Koji charges for *research you actually run*: text conversations cost 1 credit, voice interviews cost 3, and a report refresh costs 5. A quality gate means **only conversations scoring 3 or above consume credits at all**, so junk responses do not bill.\n\nFor a team of five where one person builds studies, that is the difference between paying for five seats and paying for none.\n\n### What Koji does that a survey tool structurally cannot\n\n- **AI-moderated voice interviews** that adapt in real time and ask the follow-up a form cannot\n- **Six structured question types** — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — so you keep clean quantitative data *and* get the reasoning behind it in the same session ([structured questions guide](/docs/structured-questions-guide))\n- **Automatic thematic analysis** across every transcript, not manual coding of an export ([thematic analysis](/docs/thematic-analysis-guide))\n- **One-click reports** stakeholders can read without a researcher translating them\n- **No moderator bias** — every participant gets the same neutral interviewer\n\nYou do not need research training to run it, and you go from question to insight in hours rather than weeks.\n\n---\n\n## How to Cut Your SurveyMonkey Bill\n\n1. **Audit seats before renewal.** If half your licences are dormant — the industry norm — drop to the three-seat floor.\n2. **Check whether you actually need Premier.** If crosstabs are the only reason, ask whether the analysis belongs in your BI tool instead.\n3. **Time the renewal.** Quarter-end and fiscal year-end are where the reported 15–30% discounts come from.\n4. **Benchmark with real data.** Walk in with the $16,482 median rather than the list price — our [user research tool pricing comparison](/blog/user-research-tool-pricing-2026) sets it against UserTesting, Qualtrics, Dovetail, dscout and Maze.\n5. **Split the workload.** Keep surveys for measurement, move the *why* to interviews. Many teams find a smaller survey seat count plus a research tool costs less than upgrading everyone to Premier.\n\n---\n\n## Frequently Asked Questions\n\n### How much does SurveyMonkey cost per month in 2026?\n\nTeam plans are €30/user/month (Team Advantage) and €75/user/month (Team Premier) billed annually, both with a three-seat minimum — so the practical floor is €90/month. Individual plans run from about $46/month billed annually up to $99/month for month-to-month Standard.\n\n### Does SurveyMonkey have a free plan?\n\nYes. The free tier allows unlimited surveys but caps responses you can view and blocks most analysis, export, and logic features. It is suitable for testing the interface, not for running research you need to report on.\n\n### Why is there a three-seat minimum on SurveyMonkey team plans?\n\nIt is a packaging decision, not a technical one. Every team plan bills a minimum of three users (five on Enterprise), so a solo researcher on a team plan pays for two seats nobody uses. If only one person builds surveys, an individual plan is usually cheaper.\n\n### What does SurveyMonkey charge for extra responses?\n\n€0.10 per response beyond your plan allowance — 50,000/year on Team Advantage, 100,000/year on Team Premier. Enterprise allowances are negotiated.\n\n### What do companies actually pay for SurveyMonkey?\n\nBenchmark data across 193 purchases puts the median annual contract at $16,482, ranging from $6,491 to $41,275, with buyers averaging 18.31% off list price when they negotiate.\n\n### Is there a cheaper alternative to SurveyMonkey for customer research?\n\nFor measurement-only work, plenty of [SurveyMonkey alternatives](/blog/surveymonkey-alternatives-2026) undercut it, and [Typeform](/blog/typeform-pricing-2026) and [Qualtrics](/blog/qualtrics-pricing-2026) sit either side of it on price. For understanding *why* customers answer the way they do, Koji starts at €29/month with no seat minimum and runs AI-moderated interviews rather than static forms — a different instrument at a lower entry price.\n\n---\n\n## Stop Paying for Seats. Start Paying for Insight.\n\nSurveyMonkey prices your headcount. Koji prices your research.\n\nIf you are staring at a renewal quote and wondering why three seats cost more than the answers are worth, run one study the other way: same questions, but with an AI interviewer that asks the follow-up your form never could, and a thematic analysis waiting when the last participant finishes.\n\n**[Start free with Koji](https://www.koji.so)** — no seat minimums, no annual contract, no sales call. From question to insight in hours, not weeks.","category":"Comparisons","lastModified":"2026-07-31T03:18:58.417069+00:00","metaTitle":"SurveyMonkey Pricing 2026: Plans, Costs & What You Actually Pay","metaDescription":"SurveyMonkey costs EUR 30-75/user/month on team plans with a 3-seat minimum. Full 2026 pricing breakdown, individual plan costs, response overage fees, and what real buyers negotiate.","keywords":["surveymonkey pricing","surveymonkey cost","surveymonkey plans","surveymonkey pricing 2026","how much does surveymonkey cost","surveymonkey team advantage price","surveymonkey enterprise pricing"],"aiSummary":"SurveyMonkey 2026 pricing: Team Advantage EUR 30/user/month and Team Premier EUR 75/user/month, both with a three-seat minimum, plus custom Enterprise on five seats. Individual plans run $46-$139/month. Response overage is EUR 0.10 each. Median negotiated contract is $16,482/year across 193 purchases with 18.31% average savings. Koji charges EUR 29-79/month with no seat minimum, pricing research volume rather than headcount.","aiKeywords":["surveymonkey pricing","survey tool cost","per-seat pricing","research budget","survey software comparison","ai interviews"],"aiContentType":"comparison","faqItems":[{"answer":"Team plans are EUR 30/user/month (Team Advantage) and EUR 75/user/month (Team Premier) billed annually, both with a three-seat minimum — so the practical floor is EUR 90/month. Individual plans run from about $46/month billed annually up to $99/month for month-to-month Standard.","question":"How much does SurveyMonkey cost per month in 2026?"},{"answer":"Yes. The free tier allows unlimited surveys but caps the responses you can view and blocks most analysis, export, and logic features. It is suitable for testing the interface, not for running research you need to report on.","question":"Does SurveyMonkey have a free plan?"},{"answer":"It is a packaging decision. Every team plan bills a minimum of three users (five on Enterprise), so a solo researcher on a team plan pays for two seats nobody uses. If only one person builds surveys, an individual plan is usually cheaper.","question":"Why is there a three-seat minimum on SurveyMonkey team plans?"},{"answer":"EUR 0.10 per response beyond your plan allowance — 50,000/year on Team Advantage and 100,000/year on Team Premier. Enterprise allowances are negotiated.","question":"What does SurveyMonkey charge for extra responses?"},{"answer":"Benchmark data across 193 purchases puts the median annual contract at $16,482, ranging from $6,491 to $41,275, with buyers averaging 18.31% off list price when they negotiate.","question":"What do companies actually pay for SurveyMonkey?"},{"answer":"For measurement-only work, several survey tools undercut it. For understanding why customers answer the way they do, Koji starts at EUR 29/month with no seat minimum and runs AI-moderated interviews rather than static forms.","question":"Is there a cheaper alternative to SurveyMonkey for customer research?"}],"relatedTopics":["surveymonkey","pricing","survey tools","research budget","per-seat pricing"]},{"type":"documentation","id":"8731c80e-c766-49bd-a302-d24543f0d008","slug":"root-cause-analysis-guide","title":"Root Cause Analysis for Customer Research: The Complete Guide","url":"https://www.koji.so/docs/root-cause-analysis-guide","summary":"Root cause analysis (RCA) is a structured method for finding the underlying cause of a customer problem — churn, complaints, abandoned signups — so teams fix the source instead of patching symptoms. This guide explains the iceberg of symptoms vs. causes, the five-step RCA process, the 5 Whys, fishbone (Ishikawa) diagrams and their six categories, Pareto analysis, how to apply RCA to qualitative interview and survey data, common pitfalls like blame culture and single-cause bias, and how AI-moderated research with Koji surfaces root causes at scale.","content":"Root cause analysis (RCA) is a structured method for finding the underlying cause of a customer problem — the churn, the complaint, the abandoned signup — so you fix the source instead of repeatedly patching the symptom. Done well, it is the difference between a team that ships the same bug fix three quarters in a row and one that eliminates the problem for good.\n\nThis guide covers the core RCA techniques — the 5 Whys, fishbone diagrams, and Pareto analysis — and shows how to apply them to qualitative customer research, where the real cause is almost never stated outright.\n\n## What Is Root Cause Analysis?\n\nRoot cause analysis is a problem-solving discipline for identifying the deepest cause of a fault or problem — the factor that, if removed, prevents the problem from recurring. It originated in manufacturing and quality management (Toyota, Six Sigma, total quality management) and has since become standard practice in product, customer success, and research teams.\n\nThe central idea is the **iceberg**: visible symptoms — a spike in cancellations, a flood of support tickets, a low activation rate — sit above the waterline. The systemic causes that generate them sit below it. Treating the symptom (a faster support queue, a discount to a churning customer) feels like progress, but the problem regenerates because the cause is untouched.\n\nThe economic case is blunt. The IBM System Science Institute's widely cited research found that fixing a defect in the maintenance phase costs roughly **100 times** more than fixing it during design. The Consortium for Information and Software Quality (CISQ) estimated that poor software quality cost U.S. businesses **$607 billion in 2022**. And because acquiring a new customer is **four to five times** more expensive than retaining one (Forbes), every churn you trace to its root and prevent compounds in value.\n\n## Symptoms vs. Root Causes: The Iceberg Problem\n\nMost dissatisfaction is invisible. Zendesk research found that **56% of consumers rarely complain about a negative experience — they quietly switch instead**. Qualtrics' 2026 Consumer Experience Trends Report found that out of every ten poor experiences, five result in reduced or eliminated spending. The implication for researchers is uncomfortable: the complaints you can see are a tiny, unrepresentative sample of the problems that are actually costing you revenue.\n\nThis is exactly why RCA belongs in the research toolkit, not just the engineering one. You cannot fix what customers will not tell you directly — you have to infer the root cause from patterns across many conversations.\n\n## The Core RCA Process\n\nWhatever technique you use, the underlying process is the same five steps:\n\n1. **Define the problem precisely.** \"Activation is down\" is a symptom, not a problem statement. \"Users who sign up on mobile reach the first key action 40% less often than desktop users\" is something you can investigate.\n2. **Collect evidence.** Pull the relevant interviews, support tickets, session data, and open-ended survey responses. RCA is only as good as the data behind it.\n3. **Identify possible causal factors.** Brainstorm broadly before narrowing — this is where the fishbone diagram earns its place.\n4. **Isolate the root cause(s).** Trace each candidate cause until you reach one that is both actionable and explanatory.\n5. **Implement and verify.** Apply a corrective action and confirm the symptom actually disappears. An unverified \"root cause\" is a hypothesis, not a finding.\n\n## Technique 1: The 5 Whys\n\nThe 5 Whys, developed at Toyota by Sakichi Toyoda and embedded in the Toyota Production System by Taiichi Ohno, is the simplest RCA tool. You state the problem and ask \"Why?\" — then ask why of each answer, chaining them until you reach a systemic cause.\n\nAs Taiichi Ohno put it: *\"By repeating why five times, the nature of the problem as well as its solution becomes clear.\"*\n\nA research example:\n\n- **Problem:** Trial users on the team plan cancel within two weeks.\n- **Why?** They never invited teammates.\n- **Why?** They could not find the invite flow.\n- **Why?** It is buried in account settings.\n- **Why?** Onboarding never prompts collaboration.\n- **Why?** We assumed the first user would explore on their own.\n\nThe root cause is not \"the invite button is hard to find\" — it is an onboarding model that assumes self-guided exploration. The fix is structural, not cosmetic.\n\nThe 5 Whys has one well-known weakness: **single-path bias**. Because you follow one chain, you can miss parallel causes. Use it to go deep on a hypothesis, not to map the whole problem.\n\n## Technique 2: Fishbone (Ishikawa) Diagrams\n\nWhen a problem likely has multiple contributing causes, the fishbone diagram (also called the Ishikawa or cause-and-effect diagram) forces breadth. You draw the symptom as the \"head\" of the fish and branch possible causes off the spine, grouped into categories. In manufacturing these are the classic **six Ms**:\n\n- **Machine** — tools and technology (your app, infrastructure, integrations)\n- **Method** — processes and workflows (onboarding, support handoffs)\n- **Material** — inputs and content (documentation, data quality)\n- **Manpower / Mindpower** — people and skills\n- **Measurement** — how success is tracked and whether the metric itself misleads\n- **Milieu / Environment** — external context (pricing, competitors, seasonality)\n\nFor software and service teams, the categories are often adapted to People, Process, Policy, and Technology. The value is the same: by spreading causes across categories, you surface candidates the 5 Whys alone would never reach.\n\n## Technique 3: Pareto Analysis\n\nNot all causes are worth fixing. Pareto analysis applies the 80/20 principle: a small number of causes (the \"vital few\") typically drive the majority of the problem. After you have identified candidate causes, count how often each one appears across your evidence — interviews, tickets, churn reasons — and rank them. This turns RCA from an opinion contest into a prioritization based on frequency and impact, so engineering effort goes to the causes that actually move the number.\n\n## RCA in Qualitative Customer Research\n\nApplying RCA to research data means systematically coding qualitative sources — interview transcripts, open-ended survey answers, support conversations — to trace a recurring theme back to its driver. The workflow:\n\n1. **Tag** every mention of a problem with a consistent code (see our [thematic analysis guide](/docs/thematic-analysis-guide)).\n2. **Cluster** the codes to find recurring symptoms.\n3. **Drill** into each cluster with the 5 Whys, using direct quotes as evidence for each link in the chain.\n4. **Triangulate** across sources so a single vocal customer does not masquerade as a root cause.\n\nThe hard part is volume. Treating symptoms is fast; finding root causes requires reading widely enough to see the pattern, and most teams never have the bandwidth to code hundreds of conversations by hand.\n\n## Common Pitfalls\n\n- **Stopping at the symptom.** The most frequent failure. A faster fix is not the same as a permanent one.\n- **Blame culture.** When RCA turns into \"who messed up,\" people hide information and the analysis dies. W. Edwards Deming's estimate is the antidote: *\"94% belongs to the system (responsibility of management), 6% special.\"* Most problems are systemic, not individual.\n- **Single-cause bias.** Real failures are usually multi-causal. Resist the urge to declare victory at the first plausible cause.\n- **Confirmation bias.** Choosing the \"why\" that confirms what you already believed. Let the evidence, not the hypothesis, pick the branch.\n\n## The Modern Approach: Root Cause Analysis with AI\n\nTraditional RCA on qualitative data is slow because a human has to conduct each interview, probe each \"why\" in the moment, and then code every transcript. While legacy survey tools like SurveyMonkey can tell you *that* customers are unhappy, they cannot ask the follow-up question that reveals *why*. AI-native platforms like Koji close that gap.\n\n### How Koji Helps\n\n- **AI-moderated interviews that ask \"Why?\" automatically.** Koji's adaptive interviewer detects a meaningful answer and probes deeper in real time — running the 5 Whys at scale across every conversation, not just the handful a researcher can personally moderate.\n- **Automatic thematic analysis.** Instead of hand-coding transcripts, Koji clusters recurring causes across hundreds of interviews and ranks them — effectively a Pareto analysis of your customers' own words.\n- **Structured questions for clean quantification.** Koji supports six [structured question types](/docs/structured-questions-guide) — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — so you can pair open-ended root-cause probing with rankable, countable data in the same study.\n- **Voice and text interviews** that capture the nuance behind a complaint, plus real-time reporting that turns raw conversations into a root-cause map in minutes.\n- **Customizable AI consultants** tuned to your domain, so the interviewer knows which threads to pull.\n\nTeams using AI-assisted research tools report dramatically faster time-to-insight — minutes instead of the days it takes to manually transcribe, code, and synthesize. That speed is what makes continuous root cause analysis, rather than the occasional fire drill, actually possible.\n\n## Related Resources\n\n- [The 5 Whys Technique for User Research](/docs/five-whys-technique-user-research) — go deeper on the foundational RCA technique\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) — how to code qualitative data before tracing causes\n- [How to Analyze Qualitative Data](/docs/how-to-analyze-qualitative-data) — turn raw transcripts into findings\n- [Research Synthesis Guide](/docs/research-synthesis-guide) — move from individual causes to a coherent story\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that quantify your research\n- [Jobs to Be Done Framework](/docs/jobs-to-be-done-framework) — understand the underlying need behind the behavior","category":"Analysis & Synthesis","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"Root Cause Analysis for Customer Research (5 Whys, Fishbone) | Koji","metaDescription":"Learn root cause analysis for customer and product research. Master the 5 Whys, fishbone (Ishikawa) diagrams, and Pareto analysis to fix the cause of churn and complaints — not just the symptom.","keywords":["root cause analysis","root cause analysis customer research","5 whys","fishbone diagram","ishikawa diagram","rca method","root cause analysis churn","pareto analysis"],"aiSummary":"Root cause analysis (RCA) is a structured method for finding the underlying cause of a customer problem — churn, complaints, abandoned signups — so teams fix the source instead of patching symptoms. This guide explains the iceberg of symptoms vs. causes, the five-step RCA process, the 5 Whys, fishbone (Ishikawa) diagrams and their six categories, Pareto analysis, how to apply RCA to qualitative interview and survey data, common pitfalls like blame culture and single-cause bias, and how AI-moderated research with Koji surfaces root causes at scale.","aiPrerequisites":["how-to-analyze-qualitative-data","thematic-analysis-guide"],"aiLearningOutcomes":["Distinguish surface symptoms from systemic root causes","Run a 5 Whys analysis and avoid single-path bias","Build a fishbone (Ishikawa) diagram using the six cause categories","Apply Pareto analysis to prioritize the vital few causes","Use AI-moderated interviews to find root causes behind churn and complaints"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"7e6acbba-43be-4db7-8a25-d3f58bca53ba","slug":"research-repository-guide","title":"How to Build a UX Research Repository: The Complete Guide","url":"https://www.koji.so/docs/research-repository-guide","summary":"A research repository is a centralized, searchable system for storing and retrieving qualitative insights across studies. This guide covers taxonomy design, tool selection, intake rituals, and how AI-native platforms like Koji automate the repository-building process. Organizations with mature research repositories report 2.7x better business outcomes.","content":"A research repository is a centralized system for storing, organizing, and retrieving qualitative insights from past studies — so institutional knowledge doesn't live in scattered Notion pages, personal drives, or researchers' heads. When built well, a research repository transforms individual study findings into a compounding organizational asset.\n\nThe business case is clear: according to the User Interviews State of Research Operations 2025 report, organizations that embed research into their strategy report 2.7x better business outcomes, including 3.6x more active users and 2.8x increased revenue. But research that isn't findable might as well not exist.\n\n## What Is a Research Repository?\n\nA research repository — sometimes called an insights repository or research library — is a structured database of past research: studies, transcripts, themes, quotes, participant data, and synthesized insights, organized so anyone on the team can find relevant findings quickly.\n\nIt's different from a shared drive or a folder full of reports. A well-designed repository is:\n- **Searchable by topic, theme, or user segment** — not just by study name or date\n- **Cross-linked** — so a finding from last year's usability study surfaces when someone searches for a relevant topic today\n- **Living** — updated after every study, not just when someone has bandwidth\n- **Accessible** — open to PMs, designers, engineers, and leadership, not gated behind researcher access\n\nWithout a repository, teams repeatedly research questions that have already been answered, or make decisions that contradict findings they don't know exist.\n\n## Why Research Repositories Matter\n\n**The insight reuse problem**: Research findings have a long shelf life. A study on onboarding friction from 18 months ago may be directly relevant to a decision being made today — but only if someone can find it. Without a repository, insights decay in email threads and presentation decks that nobody revisits.\n\n**The scale problem**: As research volume grows, synthesis becomes impossible without infrastructure. A single researcher can keep track of 10 studies. At 50, you need a system. At 200, you need search and AI-assisted retrieval.\n\n**The democratization problem**: According to the State of Research Operations 2025, 35% of organizations have at least one dedicated Research Operations professional — but in most companies, insights are still locked in researcher-controlled systems. A good repository lets product managers and designers find relevant research without filing a research request.\n\n**The AI opportunity**: The same report found that 80% of research professionals now use AI in their research workflow — a 24-point increase from the prior year. AI-native repositories can automatically tag, cross-reference, and synthesize insights in ways that manual systems cannot.\n\n## What to Store in a Research Repository\n\n| Content Type | What to Include |\n|-------------|----------------|\n| Studies | Research plan, methodology, participant details, date |\n| Transcripts | Full session transcripts, timestamped |\n| Insights | Synthesized findings with supporting evidence |\n| Themes | Cross-study patterns with evidence from multiple sources |\n| Quotes | Tagged, searchable participant quotes |\n| Participant profiles | Anonymized participant data for cross-study analysis |\n| Reports | Final deliverables distributed to stakeholders |\n\nThe most valuable layer is **insights** — not raw transcripts. Transcripts are evidence; insights are the conclusions drawn from that evidence. Build your taxonomy and search around insights, not raw data.\n\n## How to Build a Research Repository: Step by Step\n\n### Step 1: Agree on a Taxonomy\n\nBefore touching any tooling, define how you'll categorize insights. Common taxonomic dimensions:\n- **Product area** (onboarding, checkout, notifications, settings)\n- **User segment** (enterprise, SMB, consumer; new vs. experienced users)\n- **Research type** (discovery, evaluative, generative)\n- **Theme** (mental models, friction points, motivations, workarounds)\n- **Sentiment** (positive, neutral, negative)\n\nResist the urge to build a perfect taxonomy upfront. Start with 4–6 dimensions and refine as content accumulates. Over-engineered taxonomies don't get maintained.\n\n### Step 2: Choose the Right Tool\n\nYou don't need dedicated software to start. Many teams begin with Notion or Airtable before migrating to purpose-built tools. The right choice depends on:\n- How many researchers are contributing\n- Whether stakeholders need direct self-serve access\n- Whether you need semantic search vs. tag-based search\n- Your budget\n\n**For small teams (fewer than 2 studies per month)**: Notion or Airtable with a consistent tagging convention.\n\n**For mid-size teams (2–8 studies per month)**: Purpose-built tools like Dovetail, Condens, Looppanel, or EnjoyHQ offer automatic tagging, transcript analysis, and insight synthesis.\n\n**For AI-native teams**: Platforms like Koji automatically generate themes and insights from every interview session — building the repository as research happens, rather than requiring manual post-study intake.\n\n### Step 3: Establish an Intake Ritual\n\nThe most common reason repositories fail is that they never get populated. Build intake into your research process, not as an afterthought. After every study, add:\n- The research brief and methodology\n- Deidentified transcripts or session notes\n- 3–5 top-level insights with supporting evidence (quotes, timestamps)\n- Tags across your taxonomy dimensions\n\nThis should take 30–60 minutes per study. If it takes longer, your intake process is too complex — simplify the taxonomy or the template.\n\n### Step 4: Make It Accessible to Non-Researchers\n\nThe repository delivers value only if people outside the research team actually use it. This requires:\n- **Simple, powerful search** — full-text search is the minimum; semantic search (finding conceptually related results even with different terminology) is much more powerful\n- **Insight summaries** — non-researchers don't have time to read full reports; 2–3 sentence summaries with link-to-detail are essential\n- **Proactive sharing** — send relevant insights to stakeholders when a known product decision is in progress\n- **Stakeholder onboarding** — a 15-minute tour of the repository pays dividends in adoption\n\n### Step 5: Audit and Maintain Quarterly\n\nRepositories decay without maintenance. Every quarter:\n- Archive studies older than 3 years (keep the insights, deprecate the raw data)\n- Review and merge duplicate themes\n- Identify insights invalidated by subsequent research and flag them\n- Survey stakeholders: \"Did you find what you needed in the last month?\"\n\n## Common Mistakes to Avoid\n\n1. **Building the taxonomy before you have data**: Start with a few studies, see what patterns emerge, then build your taxonomy around real content. Theoretical taxonomies rarely survive contact with actual research.\n\n2. **Storing raw data instead of insights**: A repository full of unanalyzed transcripts isn't useful — it just creates a larger pile to dig through. The synthesis work is what makes a repository valuable.\n\n3. **Making it researcher-only**: If only researchers can access and update the repository, it becomes a bottleneck rather than a resource. Give PMs and designers read access and contribution rights for their own synthesis work.\n\n4. **Optimizing for completeness over findability**: You don't need every study perfectly tagged — you need the most recent and most relevant studies to be instantly findable. Prioritize accordingly.\n\n5. **Neglecting cross-study synthesis**: Individual study insights are useful; cross-study themes are where the real leverage is. Schedule quarterly synthesis sessions to draw connections across multiple studies.\n\n## The Modern Research Repository: AI-Augmented Insights\n\nLegacy research repositories are passive storage systems — you put things in, and only get them out if you know what to search for. The next generation of research infrastructure is AI-native.\n\nModern AI-powered research platforms can:\n- Automatically tag transcripts with themes and sentiment\n- Surface relevant past insights when you start a new study\n- Generate cross-study synthesis reports on demand\n- Alert stakeholders when new findings touch topics they care about\n\nWhile traditional tools like Dovetail and Condens require manual tagging and structured intake, AI-native platforms like Koji take a different approach: every interview automatically generates themes, sentiment signals, and synthesized insights — creating a continuously updated knowledge base without manual overhead. As research volume scales, the repository grows more intelligent, not just larger.\n\n> \"The goal of an effective research operations program is to magnify the impact and value of UX research in an organization, giving researchers a seat at the table to ensure the voice of the user is at the center of every product release.\" — ResearchOps Community\n\nFor teams running 10+ interviews per month, the AI-native approach isn't just convenient — it's the only way to keep synthesis from becoming a bottleneck.\n\n## Real-World Example\n\nA mid-size SaaS company has been running research for two years. Their researchers have conducted 40+ studies, but findings live in Google Drive folders, Confluence pages, and individual Notion workspaces. When a PM asks \"what do we know about enterprise onboarding friction?\", the answer is \"give me a few days to dig through everything.\"\n\nThey build a research repository in Notion with a simple taxonomy: product area, user segment, and theme. They spend two weeks doing a retroactive intake of the 10 most important past studies. They then commit to 30-minute intake after every new study.\n\nWithin three months, PMs are self-serving research before coming to researchers. The research team spends more time on new studies and less time answering questions that have already been answered.\n\n## Key Takeaways\n\n- A research repository transforms individual findings into a compounding organizational asset that gets more valuable over time\n- The most valuable layer is synthesized insights, not raw transcripts — invest in synthesis before intake\n- Consistent 30–60 minute intake after every study is the key habit; skip it and the repository decays\n- Non-researcher accessibility is what creates ROI — a researcher-only repository is an underused repository\n- AI-native research platforms can automatically build the repository as research happens, eliminating manual intake entirely\n\n## Frequently Asked Questions\n\n**What is the difference between a research repository and a research report?**\nA research report is a deliverable from a single study — a document summarizing what was found. A research repository is infrastructure connecting findings across many studies over time. Think of reports as inputs to the repository.\n\n**What tool should I use for a research repository?**\nIt depends on your stage. Start with Notion or Airtable if running fewer than 2 studies per month. Graduate to purpose-built tools (Dovetail, Condens, Looppanel) when automatic tagging and search become necessary. For teams using AI-moderated interviews, platforms like Koji create the repository automatically as each study runs.\n\n**How do I get stakeholders to actually use the repository?**\nThree tactics work consistently: (1) share proactive insight alerts when relevant decisions are in progress, (2) give PMs and designers self-serve access so they can answer questions without filing a research request, and (3) create a monthly insights digest surfacing the most relevant recent findings.\n\n**How long does it take to build a research repository?**\nYou can have a working repository in a week using Notion. A robust, searchable repository with 50+ cross-linked studies takes 3–6 months to mature. The key discipline is consistent intake after every study — not a single large migration project.\n\n**Should I include external research — published reports, competitor analysis — in the repository?**\nYes. Many teams create a secondary research section. Store summaries and citations rather than full documents to avoid copyright issues. Flag all external research with its source and date so users understand the provenance.\n\n---\n\n## Related Resources\n\n- [How to Analyze Qualitative Data](/docs/how-to-analyze-qualitative-data) — Analysis methodology\n- [Presenting Research Findings](/docs/presenting-research-findings) — Share repository insights\n- [Scaling User Research](/docs/scaling-user-research) — Scale your research practice\n- [Continuous Discovery Guide](/docs/continuous-discovery-user-research) — Ongoing research habits\n- [AI-Generated Insights](/docs/ai-generated-insights) — Automated insight generation\n\n*Explore [structured questions](/docs/structured-questions-guide) for building searchable, structured research repositories.*\n\n## Further reading on the blog\n\n- [Koji vs Dovetail: Which Research Tool Is Right for You?](/blog/koji-vs-dovetail) — Dovetail organizes research data. Koji conducts the research for you. An honest breakdown of both tools to help you decide which one your te\n- [Koji vs Marvin: Full-Stack AI Research vs Analysis Repository (2026)](/blog/koji-vs-marvin-2026) — Marvin organizes research you've already collected. Koji runs new AI-moderated interviews from scratch. Here's how to choose — and when you \n- [AI-Moderated vs Human-Moderated Interviews: Which Should You Choose?](/blog/ai-moderated-vs-human-moderated-interviews) — AI-moderated and human-moderated interviews each have a time and a place. Here is the honest comparison to help you choose the right approac\n\n<!-- further-reading:blog -->\n","category":"Analysis & Synthesis","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"How to Build a UX Research Repository — Koji Research KB","metaDescription":"Learn how to build a research repository that teams actually use. Covers taxonomy, tool selection, intake rituals, and AI-native approaches.","keywords":["research repository","insights repository","ResearchOps","UX research repository","research knowledge base","how to build research repository","research operations"],"aiSummary":"A research repository is a centralized, searchable system for storing and retrieving qualitative insights across studies. This guide covers taxonomy design, tool selection, intake rituals, and how AI-native platforms like Koji automate the repository-building process. Organizations with mature research repositories report 2.7x better business outcomes.","aiPrerequisites":["thematic-analysis-guide","coding-qualitative-data"],"aiLearningOutcomes":["Design a taxonomy for organizing research insights","Choose the right repository tool for your team size and maturity","Establish a consistent intake ritual that keeps the repository current","Enable non-researcher stakeholders to self-serve past findings"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min read"},{"type":"documentation","id":"dd8fea56-0960-441f-93f7-d5f50dfe296e","slug":"structured-questions-guide","title":"Structured Questions in AI Interviews","url":"https://www.koji.so/docs/structured-questions-guide","summary":"Koji supports six question types — open ended, scale, single choice, multiple choice, ranking, and yes/no — embedded in natural AI conversations. Unlike traditional surveys, every structured response is automatically probed by the AI to capture both quantitative data and qualitative reasoning. Features include anchor probing for scale questions, always-available \"prefer not to answer,\" and automatic chart generation in reports.","content":"Structured questions are what make Koji fundamentally different from both traditional surveys and standard interview tools. They let you embed quantitative data collection — scales, ratings, multiple choice, ranking — directly inside a natural AI-powered conversation. Every structured response is automatically followed up by the AI, so you capture both the number and the reasoning behind it in a single interaction.\n\n## The Problem with Traditional Approaches\n\nResearch has traditionally forced a painful choice: run a survey to collect structured, chartable data (but miss the \"why\"), or run an interview to understand motivations and context (but lose the ability to aggregate and compare across participants). Most teams end up running both — doubling the cost, timeline, and participant burden.\n\nKoji's structured questions eliminate that tradeoff. A single AI interview captures quantitative metrics that aggregate into charts and distributions, alongside rich qualitative context from conversational follow-up. You get the benchmarkable numbers your stakeholders want and the deep human insight your research needs — from the same conversation.\n\n## How It Works\n\nWhen you add a structured question to your study, two things happen during every interview:\n\n1. **The participant provides a structured response** — a number on a scale, a selected option, a ranked list, or a yes/no answer. In [text mode](/docs/text-interview-experience), this appears as an interactive widget (slider, radio buttons, checkboxes, drag-to-rank) embedded naturally in the chat. In [voice mode](/docs/voice-interview-experience), the AI handles everything conversationally — it asks the question, hears the answer, and extracts the structured value.\n\n2. **The AI immediately probes for the \"why\"** — based on the response, the AI follows up with contextual probing. A participant who gives a low satisfaction score gets different follow-up questions than one who gives a high score. This probing is what transforms a data point into an insight.\n\nEvery response is captured twice: as a structured value (the number, choice, or ranking) for aggregate reporting, and as qualitative context from the follow-up conversation. When you generate a report, structured questions produce charts — distribution charts for scales, bar charts for choices, ranked lists for rankings — alongside AI-synthesized qualitative insights explaining the patterns in the data.\n\n## The Six Question Types\n\nKoji supports six structured question types, each optimized for a different kind of data:\n\n**Open Ended** — The default question type. Pure qualitative, no structured value. The AI asks the question and probes for depth. Best for exploratory discovery where you don't yet know the right categories or scales.\n\n**Scale** — A numeric rating (e.g., 1–5, 1–10, NPS 0–10). You define the range and can add labels for the endpoints (e.g., \"Very Unlikely\" to \"Very Likely\"). Reports show a distribution chart with average and median. Best for NPS, CSAT, satisfaction scores, and any metric you want to track over time.\n\n**Single Choice** — Pick one option from a list. Reports show a frequency bar chart. Best for segmentation questions like \"which of these describes your role?\" or \"which feature do you use most?\"\n\n**Multiple Choice** — Pick any number of options from a list. You can also enable an \"Other\" free-text option for responses you didn't anticipate. Reports show a stacked frequency chart. Best for \"what challenges do you face?\" (select all that apply).\n\n**Ranking** — Order a list of items by preference or priority. In text mode, participants drag items into their preferred order. Reports show average position for each item. Best for feature prioritization and preference ordering.\n\n**Yes/No** — A binary question. Reports show a pie or donut chart. Best for confirmation questions (\"Have you tried X?\") that open up more nuanced follow-up conversation.\n\n## Anchor Probing: The Power of Scale Questions\n\nScale questions have a unique probing feature called **anchor probing**. When enabled, after a participant gives their rating, the AI asks a question like:\n\n> \"You said 7 — what would need to change to make it a 9 or 10?\"\n\nThis technique is powerful because it:\n- Surfaces specific, actionable improvements rather than vague dissatisfaction\n- Anchors the follow-up in the participant's own evaluation framework\n- Reveals the gap between current experience and ideal experience\n- Produces insights that map directly to product or service improvements\n\nAnchor probing works on any scale question and is especially valuable for NPS, satisfaction, and likelihood-to-recommend metrics. You can enable it in the probing configuration for any scale question in the [brief editor](/docs/editing-the-brief-manually).\n\n## Step-by-Step Guide\n\n1. **Open your study's research brief**\n   From your study dashboard, open the research brief editor and navigate to the Questions tab. For details on the editor, see [Editing the Brief Manually](/docs/editing-the-brief-manually).\n\n2. **Add a new question**\n   Click the Add Question button. Enter your question text and select the question type from the dropdown.\n\n3. **Configure type-specific options**\n   For Scale questions, set the min/max values and optional endpoint labels (e.g., \"Very Unlikely\" to \"Very Likely\"). For Choice or Ranking questions, add your option list. You can enable \"Allow Other\" on choice questions to capture write-in responses.\n\n4. **Set AI probing depth**\n   Each question has a probing configuration. Set the maximum number of follow-up questions the AI will ask (0 = no probing, 1–3 = increasing depth). You can add specific probing instructions, like \"If the participant gives a low score, ask what would need to change.\" For scale questions, consider enabling anchor probing.\n\n5. **Order your questions**\n   Drag questions into the sequence that makes sense for the conversation flow. Open-ended discovery questions work well at the start; structured ratings fit naturally toward the end once rapport is established.\n\n6. **Publish your study**\n   Once you're satisfied with your question mix, [publish the study](/docs/publishing-your-study) as normal. Question IDs are assigned at publish time and remain stable across the life of the study, ensuring accurate data aggregation.\n\n## \"Prefer Not to Answer\"\n\nEvery structured question includes a \"Prefer not to answer\" option. Participants can always skip any question they're uncomfortable answering. Skipped responses are tracked separately from answered responses in your report data, so your aggregate statistics (averages, distributions, frequencies) remain accurate and are not skewed by missing data.\n\nThis is deliberate — research ethics require that participants never feel forced to answer any question. The \"prefer not to answer\" option is always available and cannot be disabled.\n\n## Key Things to Know\n\n- **Text vs. voice behavior**: In [text mode](/docs/text-interview-experience), quantitative questions render as interactive widgets — sliders, radio buttons, checkboxes, or drag-to-rank lists. In [voice mode](/docs/voice-interview-experience), all questions are handled conversationally and the AI extracts structured values from spoken answers. Both modes capture the same data.\n\n- **It's a conversation, not a survey**: Unlike traditional survey tools, structured questions in Koji are embedded within a natural, flowing AI conversation. The AI builds rapport, asks open-ended questions for context, and then introduces structured questions at natural conversation points. Participants experience a conversation with data capture moments, not a form with chat bolted on.\n\n- **Question IDs are stable**: Each question has a stable identifier assigned at [publish time](/docs/publishing-your-study) that flows through from the brief to the interview to the report. This means your reports accurately aggregate data across all interviews, even if question text is lightly edited.\n\n- **AI probing follows structured answers**: After a participant selects their answer via widget (or speaks it in voice mode), the AI adapts its follow-up based on what they selected. A participant who gives a low satisfaction score gets different follow-up questions than one who gives a high score.\n\n- **Reports automatically visualize structured data**: When you generate a research report, each structured question gets the right visualization: distribution charts for scales, bar charts for choices, ranked lists for rankings — no manual chart-building required.\n\n- **Sections help with longer studies**: If your study has many questions, use section grouping to organize them (e.g., \"Background,\" \"Product Experience,\" \"NPS\"). This helps the AI present questions in a coherent, conversational arc.\n\n## Tips and Best Practices\n\n- **Put qualitative questions first**: Build rapport and let participants share naturally before asking them to rate anything. Open-ended questions early in the interview produce richer context for the structured answers that follow.\n- **Use scale questions for benchmarks**: If you're tracking a metric over time (NPS, CSAT, feature satisfaction), use a Scale question with consistent parameters so you can compare across studies.\n- **Don't over-structure**: Resist the urge to turn every question into multiple choice. Structured questions are most powerful when used sparingly — they anchor key metrics while open-ended questions carry the qualitative depth. Most studies benefit from 1–3 structured questions mixed with several open-ended ones.\n- **Always enable probing on quantitative questions**: A scale or choice question without AI follow-up is just a survey. The real value is in the \"why\" — configure at least one follow-up for every quantitative question.\n- **Enable anchor probing on key scales**: For your most important scale questions (NPS, overall satisfaction), turn on anchor probing. The \"what would change your score?\" follow-up produces some of the most actionable insights in any study.\n- **Mix and match for richer reports**: A study that ends with an NPS score and immediately probes the reason behind it gives you both the benchmarkable number and the insight — something no traditional survey tool can deliver.\n- **Keep choice lists focused**: For single and multiple choice questions, aim for 3–7 clear, mutually exclusive options. More than that and participants spend too much time reading instead of thinking. Use \"Allow Other\" as a catch-all.\n\n## Related Articles\n\n- [Understanding the Research Brief](/docs/understanding-the-research-brief) — how structured questions fit into the overall brief\n- [Editing the Brief Manually](/docs/editing-the-brief-manually) — using the question editor to configure structured questions\n- [Text Interview Experience](/docs/text-interview-experience) — how structured questions appear as interactive widgets\n- [Voice Interview Experience](/docs/voice-interview-experience) — how structured questions work in voice mode\n- [Choosing a Methodology](/docs/choosing-a-methodology) — which methodologies pair best with structured data capture\n- [Publishing Your Study](/docs/publishing-your-study) — how question IDs are assigned at publish time\n- [Generating Research Reports](/docs/generating-research-reports) — how structured data becomes charts and visualizations\n- [AI-Generated Insights](/docs/ai-generated-insights) — how qualitative follow-ups are synthesized\n\n## Frequently Asked Questions\n\n**Q: Do structured questions work in both voice and text mode?**\nA: Yes, but they work differently. In text mode, quantitative questions appear as interactive widgets — sliders, radio buttons, checkboxes, or drag-to-rank lists. In voice mode, the AI handles all questions conversationally and extracts structured values from spoken responses. Both modes capture identical data for reporting.\n\n**Q: Can I mix structured and open-ended questions in the same study?**\nA: Absolutely — this is the recommended approach. Open-ended questions explore freely and build rapport; structured questions anchor key metrics. A typical study might start with 2–3 open-ended discovery questions, then include 1–2 scale or choice questions toward the end.\n\n**Q: How does the AI use structured question data in reports?**\nA: When you generate a research report, each structured question produces a chart (distribution for scales, bar chart for choices) alongside AI-synthesized qualitative context from follow-up conversations. You get both the numbers and the reasoning in a single report.\n\n**Q: Can participants skip a structured question?**\nA: Yes. A \"Prefer not to answer\" option is always available on every structured question and cannot be disabled. Skipped questions are tracked separately from answered questions in the report, so your aggregate data stays accurate.\n\n**Q: How many structured questions should I include in a study?**\nA: Most studies benefit from 1–3 structured questions mixed with several open-ended ones. More than 5 structured questions can make the conversation feel survey-like, which reduces the qualitative depth participants share.\n\n**Q: What is the difference between structured questions in Koji and a regular survey?**\nA: In a regular survey, structured questions stand alone — participants select an answer and move on. In Koji, every structured question is embedded in a natural AI conversation and followed by intelligent probing. The AI adapts its follow-up based on the specific response given, capturing both the quantitative data point and the qualitative reasoning behind it. It's the difference between knowing your NPS is 7 and knowing your NPS is 7 because the onboarding was confusing but the support team was excellent.\n\n**Q: What is anchor probing?**\nA: Anchor probing is a feature specific to scale questions. After a participant gives a rating, the AI asks what would need to change to improve that rating (e.g., \"You said 7 — what would make it a 9?\"). This surfaces specific, actionable insights about the gap between current and ideal experience. You can enable it in the probing configuration for any scale question.\n\n## Further reading on the blog\n\n- [The 11 Best AI Survey Tools in 2026 (Ranked & Reviewed)](/blog/best-ai-survey-tools-2026) — We tested 11 AI-powered survey platforms head-to-head — from AI question generation to automatic open-ended analysis to fully AI-moderated c\n- [Koji vs Insight7: AI Customer Research Platforms Compared (2026)](/blog/koji-vs-insight7-2026) — Insight7 analyzes recorded calls. Koji runs the interviews itself — AI-moderated voice conversations from €29/month. Compare features, prici\n- [Koji vs Listen Labs: AI Interview Platforms Compared (2026)](/blog/koji-vs-listen-labs-2026) — Listen Labs raised $69M and powers enterprise research at brands across the Fortune 500. Koji is the accessible AI-native alternative starti\n\n<!-- further-reading:blog -->\n","category":"Study Design","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"Structured Questions — Koji Docs","metaDescription":"Mix NPS, scale ratings, multiple choice, and ranking questions with AI conversational follow-up. Get quantitative data and qualitative insights in one interview.","keywords":["structured questions","NPS survey AI","mixed methods research","scale questions interview","AI survey tool","quantitative qualitative research","interview question types"],"aiSummary":"Koji supports six question types — open ended, scale, single choice, multiple choice, ranking, and yes/no — embedded in natural AI conversations. Unlike traditional surveys, every structured response is automatically probed by the AI to capture both quantitative data and qualitative reasoning. Features include anchor probing for scale questions, always-available \"prefer not to answer,\" and automatic chart generation in reports.","aiPrerequisites":["creating-your-first-study","understanding-the-research-brief"],"aiLearningOutcomes":["Add scale, choice, and ranking questions to an AI interview study","Understand how each question type affects text vs voice mode behavior","Configure AI probing depth for quantitative questions","Generate reports that combine charts and qualitative context"],"aiDifficulty":"intermediate","aiEstimatedTime":"8 min read"},{"type":"documentation","id":"0476aca3-1e55-4c9f-96f9-2298db794258","slug":"five-whys-technique-user-research","title":"The Five Whys Technique: How to Find Root Causes in User Research (with AI)","url":"https://www.koji.so/docs/five-whys-technique-user-research","summary":"The Five Whys is a root-cause analysis technique developed by Sakichi Toyoda at Toyota in the 1930s. By asking \"why\" five times in succession on a single causal thread, researchers move past surface symptoms to system-level causes. Koji's AI interviewer automates Five Whys probing across hundreds of respondents at once.","content":"\n## What Is the Five Whys Technique?\n\nThe Five Whys is a root-cause analysis technique that uncovers the underlying reason behind a problem by asking \"Why?\" five times in succession. Each answer becomes the foundation for the next \"Why?\" — peeling back layers of symptom until you reach the actual cause. In user research, it converts vague complaints (\"the dashboard is confusing\") into specific, fixable problems (\"users can't tell which metric drives their bonus, so they assume the entire dashboard is wrong\").\n\nThe technique was developed by Sakichi Toyoda in the 1930s and codified by Taiichi Ohno as a foundational element of the Toyota Production System. Originally used to debug manufacturing defects, it has since spread to incident postmortems, customer support investigations, and qualitative user research — anywhere a team needs to move past surface symptoms.\n\nThe number five is not magic. The discipline is to keep asking until you reach a cause that, if changed, would actually prevent the problem. Sometimes that takes three whys. Sometimes seven. Five is just a useful target that prevents teams from stopping at the first plausible explanation.\n\n## Why Surface Feedback Fails Product Teams\n\nMost user feedback dies on the surface. A user says \"the onboarding is too long,\" and the team adds a progress bar. A user says \"I don't trust the AI summaries,\" and the team adds a confidence score. Both responses treat the symptom and miss the actual cause — which might be that the user lost a previous account when they trusted an automated decision they didn't understand.\n\nSurface feedback fails for three reasons:\n\n1. **Users describe symptoms, not causes.** They feel the friction, not the mechanism that produces it.\n2. **First answers are socially shaped.** People give the explanation they think you want to hear.\n3. **Product teams default to action.** A reported problem becomes a Jira ticket before anyone validates the diagnosis.\n\nThe Five Whys disrupts all three failure modes. By insisting on five layers of causation, it forces both interviewer and respondent past the first plausible-sounding answer.\n\n## The Classic Five Whys Example\n\nToyota's canonical example involves a stopped welding robot:\n\n1. **Why did the robot stop?** Its circuit overloaded, blowing a fuse.\n2. **Why did the circuit overload?** There was insufficient lubrication on the bearings, so they locked up.\n3. **Why was there insufficient lubrication?** The lubrication pump was not circulating enough oil.\n4. **Why was the pump not circulating enough oil?** The pump intake was clogged with metal shavings.\n5. **Why was the intake clogged?** There is no filter on the pump.\n\nNote where the chain ends: at a missing filter. That is something you can fix. \"Replace the fuse\" — the symptom — would have left the actual problem in place.\n\nThe same shape works for product research. A user says onboarding is confusing. Five whys later, you discover their company uses a single-sign-on provider that doesn't pass through the user's role, so every new account starts in a permission-stripped state, so every welcome screen looks broken. That is something engineering can fix.\n\n## How to Apply Five Whys in Customer Interviews\n\nRunning Five Whys in a live interview takes practice. The technique looks simple on paper but requires three habits most interviewers don't have:\n\n**1. Stay on a single thread.** Each \"why\" must build on the previous answer, not branch into a new topic. If the respondent jumps, gently bring them back: \"Hold that thought — going back to what you said about the export failing, why was that important?\"\n\n**2. Avoid blame-loaded \"whys.\"** \"Why did you do that?\" sounds accusatory. Reframe as: \"What was happening when you decided to do that?\" or \"Walk me through what was on your mind.\" The \"why\" lives inside the question's purpose, not its literal phrasing.\n\n**3. Stop when you hit a system, not a person.** A bad chain ends at \"because the user wasn't paying attention.\" A good chain ends at \"because the email subject line was indistinguishable from spam.\" The first blames a person; the second points at a system you can change.\n\n## Five Whys with Koji's AI Interviewer\n\nTraditional Five Whys requires a skilled human moderator who can listen, hold context across multiple turns, and gently keep the respondent on a single causal thread. That skill is rare and expensive — which is why most teams collect surface-level feedback and never reach the root cause.\n\nKoji's AI interviewer runs Five Whys by default. When you set a question's `maxFollowUps` to 2 or 3 in Koji's structured question framework, the AI:\n\n- Detects vague answers (\"it was confusing\", \"it didn't work well\") and probes for specificity\n- Holds the original question as the anchor across multiple follow-up turns\n- Asks neutral, blame-free \"what was happening\" style probes instead of accusatory \"why did you\" phrasing\n- Stops probing when the respondent reaches a concrete, system-level cause — preventing the over-probing that exhausts participants\n\nBecause the AI runs in parallel across many interviews, you get root-cause depth at survey scale — something a human-moderated study could never reach without weeks of work. Compared to traditional tools like SurveyMonkey or Typeform that capture a single answer and stop, Koji's AI keeps probing until the cause is actionable.\n\n## Five Whys vs. Other Root-Cause Techniques\n\nThe Five Whys is one of several root-cause methods. Each fits a different situation:\n\n| Technique | Best For | Output | Effort |\n|-----------|----------|--------|--------|\n| Five Whys | Linear, single-cause problems | Causal chain | Low |\n| Fishbone (Ishikawa) | Multi-cause problems with several contributing factors | Categorized cause map | Medium |\n| Fault Tree Analysis | Safety-critical, low-frequency failures | Probabilistic failure tree | High |\n| Critical Incident Technique | Specific past events with rich context | Categorized incident bank | Medium |\n| Pareto Analysis | Many problems, need to prioritize the vital few | Frequency ranking | Low |\n\nFor most user research, Five Whys is the right starting point. Move to Fishbone only when you discover that a problem has multiple parallel causes that don't reduce to a single root.\n\n## Common Five Whys Mistakes\n\nTeams that try Five Whys without practice tend to make the same mistakes:\n\n**Stopping too early.** \"It crashed because the API timed out\" sounds like an answer but it is still a symptom. Why did the API time out? Why was the request that slow? Most chains require six or seven whys to reach an actual root.\n\n**Branching the chain.** Each \"why\" must follow from the previous answer, not from your own theory. If you skip ahead, you confirm your bias instead of finding the real cause.\n\n**Interviewing without context.** Five Whys works best when the respondent recently experienced the problem. Memory fades fast — try to interview within 48 hours of the incident.\n\n**Treating opinion as cause.** \"Why is engagement low?\" \"Because users don't see the value.\" That is an opinion, not a root cause. Reframe: \"Walk me through the last time you opened the product and didn't take any action — what was on your mind?\"\n\n## A Five Whys Interview Template for Koji\n\nHere is a battle-tested structure you can drop into a Koji study:\n\n**Anchor question (open_ended):** \"Tell me about a specific recent moment when you felt frustrated using [product].\"\n*AI probing:* maxFollowUps = 3. Probe for specifics. After each answer, ask why that mattered or why it happened. Stop when the respondent describes a concrete system or process cause.\n\n**Severity check (scale, 1–10):** \"How much did that moment affect your decision to keep using [product]?\"\n*AI probing:* maxFollowUps = 1, anchor = true.\n\n**Counterfactual (open_ended):** \"What would have prevented that moment from happening?\"\n*AI probing:* maxFollowUps = 1.\n\n**Pattern check (yes_no):** \"Has something similar happened more than once?\"\n*If yes, AI probes:* \"Tell me about the most recent time.\"\n\nThis four-question structure runs in 8–12 minutes per respondent. Across 30 respondents, Koji's analysis automatically clusters the root causes by frequency, giving you a Pareto-ranked list of system-level fixes — something traditional research tools simply can't produce without weeks of manual coding.\n\n## When Not to Use Five Whys\n\nFive Whys is powerful but it has limits:\n\n- **Statistical questions** (\"how many users churn?\") need quantitative methods, not causal probing.\n- **Aspirational research** (\"what features do users wish existed?\") is better served by generative discovery interviews.\n- **Multi-causal failures** with several parallel contributing factors usually need a Fishbone diagram.\n- **Sensitive incidents** where the respondent may be embarrassed or defensive require trust-building first; aggressive probing will shut them down.\n\nFor everything else — onboarding friction, churn drivers, feature-adoption stalls, support escalations — Five Whys is the highest-leverage interview technique you can learn.\n\n## Getting Started in Koji\n\n1. Create a new study and add your anchor question as `open_ended` with `maxFollowUps: 3`\n2. Set probing instructions: \"Keep asking why or what was happening until you reach a concrete cause\"\n3. Add the severity, counterfactual, and pattern questions from the template above\n4. Send to 30+ respondents who recently experienced the problem\n5. Open the auto-generated report to see causes clustered by frequency\n\nThe whole study runs end-to-end in a single afternoon — including analysis. No moderator hours, no transcript coding, no spreadsheet wrangling.\n\n## Related Resources\n\n- [How to Conduct User Interviews](/docs/how-to-conduct-user-interviews) — foundational interview skills the Five Whys builds on\n- [Critical Incident Technique](/docs/critical-incident-technique) — pair with Five Whys to investigate specific past events\n- [Avoiding Bias in Interviews](/docs/avoiding-bias-in-interviews) — keep the Five Whys chain neutral and non-leading\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) — cluster Five Whys root causes across many respondents\n- [Structured Questions Guide](/docs/structured-questions-guide) — combine open-ended Five Whys probes with quantitative scales\n- [Mom Test Customer Interviews](/docs/mom-test-user-interviews) — bias-free interview discipline that pairs with Five Whys probing\n\n\n## Further reading on the blog\n\n- [How to Conduct Remote User Interviews: The Complete Guide (2026)](/blog/how-to-conduct-remote-user-interviews-2026) — Remote user interviews are now the default for most research teams. This guide covers everything — from recruiting and scheduling to running\n- [AI-Moderated vs Human-Moderated Interviews: Which Should You Choose?](/blog/ai-moderated-vs-human-moderated-interviews) — AI-moderated and human-moderated interviews each have a time and a place. Here is the honest comparison to help you choose the right approac\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n\n<!-- further-reading:blog -->\n","category":"Interview Techniques","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"Five Whys Technique for User Research: Find Root Causes with AI","metaDescription":"Use the Five Whys technique to uncover root causes behind user complaints. Step-by-step playbook plus a Koji AI template for running Five Whys interviews at scale.","keywords":["five whys","five whys technique","root cause analysis","user research methods","interview probing","toyota five whys","5 whys","sakichi toyoda","qualitative root cause","user interview techniques"],"aiSummary":"The Five Whys is a root-cause analysis technique developed by Sakichi Toyoda at Toyota in the 1930s. By asking \"why\" five times in succession on a single causal thread, researchers move past surface symptoms to system-level causes. Koji's AI interviewer automates Five Whys probing across hundreds of respondents at once.","aiPrerequisites":["Basic user interview skills","Familiarity with qualitative research"],"aiLearningOutcomes":["Apply the Five Whys technique to user research interviews","Stay on a single causal thread without branching","Use blame-free probing language","Run Five Whys at scale with Koji's AI interviewer","Combine Five Whys with structured questions for quantitative anchors"],"aiDifficulty":"beginner","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"63e90016-2a92-4edb-9e92-295bda4039b1","slug":"stakeholder-interview-guide","title":"Stakeholder Interviews: How to Align Your Team Before Research Begins","url":"https://www.koji.so/docs/stakeholder-interview-guide","summary":"Stakeholder interviews are structured conversations with internal decision-makers, domain experts, and implementers conducted before user research begins. This guide covers who to interview (decision-makers, domain experts, skeptics, implementers), how to prepare flexible question guides, what to listen for (assumptions, competing hypotheses, constraints), how to synthesize input into a research alignment document, and how AI-powered tools like Koji can run asynchronous stakeholder interviews to save scheduling overhead.","content":"\n# Stakeholder Interviews: How to Align Your Team Before Research Begins\n\n**The bottom line:** Stakeholder interviews — brief conversations with internal decision-makers, subject matter experts, and key team members before user research begins — are the single most effective way to ensure your research asks the right questions, earns organizational support, and produces findings that actually get acted on.\n\nMost research failures aren't methodological. They're political. Research teams design studies without fully understanding what decisions need to be made, what leadership already believes to be true, or what constraints will shape what's actually implementable. The result: findings that surprise no one, or worse, findings that challenge assumptions no one was prepared to question.\n\nStakeholder interviews fix this before it happens. Done well, they transform research from a deliverable into a decision-making tool that the whole team has a stake in.\n\n---\n\n## What Stakeholder Interviews Are (And Aren't)\n\nStakeholder interviews are structured conversations with internal team members — product leaders, designers, engineers, sales, customer success, executives — conducted *before* user research begins.\n\nThey are not:\n- User research (you're talking to colleagues, not customers)\n- A committee approval process (you're gathering input, not seeking permission)\n- A political exercise (you're building alignment, not managing egos)\n\nThey are:\n- A diagnostic tool to understand what questions matter most\n- An opportunity to surface existing assumptions that need testing\n- A way to discover what organizational constraints will shape how findings are used\n- A relationship investment that makes findings easier to act on\n\nThe 30-60 minutes you invest in stakeholder interviews typically saves 10x that time in research re-scoping, stakeholder pushback, and unused deliverables.\n\n---\n\n## Who to Interview\n\nThe right stakeholders depend on your research context, but most projects benefit from conversations across four categories:\n\n### Decision-Makers\nThe people who will act on research findings. If you're researching onboarding friction, this might be the VP of Product or Head of Growth. Understanding what decisions they're facing — and what would make them feel confident taking action — shapes what your research needs to prove.\n\n**Key question:** \"What would you need to see in the research to feel confident making a change?\"\n\n### Domain Experts\nPeople with deep knowledge about the problem area. Customer success managers who talk to customers daily. Sales reps who hear objections in every call. Support teams who see where users get stuck. These people often hold rich informal knowledge that no one has ever systematically documented.\n\n**Key question:** \"What do you think is driving this problem that we might not be seeing in the data?\"\n\n### Skeptics\nPeople who question the research agenda or have competing hypotheses about what users want. Rather than avoiding them, interview them first. Their objections will sharpen your research design, and including them early converts potential blockers into invested participants.\n\n**Key question:** \"What would this research need to show to change your current thinking?\"\n\n### Implementers\nEngineers, designers, or operations teams who will execute any changes the research informs. Understanding their constraints prevents recommendations that sound great but are impossible to build.\n\n**Key question:** \"What constraints should we keep in mind as we design this research?\"\n\n### How Many?\nFor most projects, 4-8 stakeholder interviews is the sweet spot. Fewer risks missing critical perspectives; more produces diminishing returns and delays research launch. If you're working on a large cross-functional initiative, you might interview up to 12-15, but that's an exception.\n\n---\n\n## Preparing for Stakeholder Interviews\n\n### Set the Right Framing\nStakeholders need to understand why you're talking to them. The framing matters:\n\n**Weak framing:** \"I need to check in before we start the research.\"\n**Strong framing:** \"Before we launch this research, I want to make sure we're asking questions that will be genuinely useful to your team. I'd love 30 minutes to understand what decisions you're facing and what you already know.\"\n\nThe strong framing positions the interview as a service to them, not a box-checking exercise for you.\n\n### Prepare a Flexible Guide\nUnlike user research interviews, stakeholder interviews don't benefit from a rigid script. Prepare 5-8 open questions and let the conversation flow. Good universal starter questions:\n\n- \"What are the most important decisions your team is making in the next 90 days?\"\n- \"What do you currently believe to be true about [the topic] that you wish you had better evidence for?\"\n- \"If this research went perfectly, what would it tell you that you don't already know?\"\n- \"What would make these findings easy for your team to act on?\"\n- \"What's the most important thing I should know before we design this research?\"\n\n### Schedule Efficiently\nStakeholder interviews tend to compress around research launch dates. Plan them early and schedule them in the same week if possible — this gives you a coherent picture of organizational context rather than fragments scattered over weeks.\n\nFor distributed or remote teams, 30-minute video calls work well. With tools like Koji, you can also run asynchronous stakeholder interviews — stakeholders complete the interview on their own schedule with an AI moderator, and you receive synthesized insights without scheduling overhead.\n\n---\n\n## Conducting the Interview\n\n### Listen for Assumptions, Not Just Facts\nThe most valuable thing a stakeholder interview surfaces is often what interviewees believe to be true without evidence. Listen for phrases like:\n- \"We know that users...\"\n- \"It's obvious that...\"\n- \"Everyone agrees...\"\n- \"The data shows...\" (then ask to see the data)\n\nThese are hypotheses dressed as facts. They're exactly what user research should test.\n\n### Probe for Specificity\nStakeholders often speak in generalities. Push for concrete examples:\n\n- \"Can you tell me about a specific customer conversation where this came up?\"\n- \"What's the most recent example you can remember?\"\n- \"What did they actually say?\"\n\nSpecific examples are far more useful for research design than broad assertions.\n\n### Surface Competing Hypotheses\nWhen two stakeholders hold different beliefs about why users behave a certain way, you've found gold. These competing hypotheses are exactly what research should resolve. Document them explicitly and design your study to test them.\n\n### Close With Priority and Constraints\nEnd every stakeholder interview with:\n- \"If we could only answer one question with this research, what should it be?\"\n- \"What would make it hard to act on findings, even if they were compelling?\"\n\nThe first reveals priorities. The second reveals implementation constraints that should shape your recommendations from the start.\n\n---\n\n## Synthesizing Stakeholder Input\n\nAfter completing stakeholder interviews, synthesize findings before designing your research:\n\n### Map Assumptions\nCreate a list of all the assumptions stakeholders hold about users and their behavior. These become your research hypotheses — explicitly stated beliefs your study will either confirm or challenge.\n\n### Identify Overlaps and Tensions\nWhere do stakeholders agree? Where do they contradict each other? Overlaps suggest things the research doesn't need to establish (everyone already believes them). Tensions are the most valuable research targets — answering them builds alignment across the organization.\n\n### Extract Decision Criteria\nWhat would stakeholders need to see to change their behavior or make a specific investment? Document these explicitly. When you present findings, you can directly address each decision criterion — dramatically increasing the chance your research gets acted on.\n\n### Identify Constraints\nWhat can't change? What constraints must the research stay within? Documenting these early prevents wasted effort on recommendations that will never be implemented.\n\n---\n\n## Using AI to Scale Stakeholder Interviews\n\nTraditional stakeholder interviews require scheduling, note-taking, and synthesis — all time-consuming. For large organizations or distributed teams, this overhead can delay research launch by weeks.\n\nAI-powered research platforms like Koji can run stakeholder interviews asynchronously:\n\n1. Configure a structured interview with your key questions\n2. Send the interview link to stakeholders — they complete it on their own schedule\n3. The AI probes for specificity, pursues interesting threads, and handles natural conversation\n4. You receive synthesized themes, key quotes, and a structured summary — without sitting in 8 back-to-back 30-minute meetings\n\nThis approach is particularly powerful for gathering input from senior executives (who have limited calendar availability) and for distributed teams across time zones. It also creates a clean record of stakeholder input that can be referenced throughout the project.\n\n---\n\n## Stakeholder Interview Question Bank\n\nUse these questions across your stakeholder interviews, adapting to each person's role:\n\n**On the problem space:**\n- \"What's the biggest unanswered question about how our users experience this part of the product?\"\n- \"If you could observe any 5 customer conversations from the past month, which scenarios would you most want to see?\"\n\n**On existing beliefs:**\n- \"What's your current hypothesis about what's driving this problem?\"\n- \"What would prove that hypothesis wrong?\"\n\n**On organizational context:**\n- \"Who else should I talk to before designing this research?\"\n- \"Is anyone actively building something in this area that research findings could affect?\"\n\n**On success criteria:**\n- \"What would make this research genuinely useful to you?\"\n- \"What format would make findings easiest to act on — a presentation, a one-pager, a Slack post?\"\n\n**On history:**\n- \"Has research been done on this before? What did it show?\"\n- \"Was there a time when research findings weren't acted on? What happened?\"\n\n---\n\n## Documenting and Sharing Stakeholder Input\n\nAfter synthesizing your stakeholder interviews, create a brief stakeholder alignment document:\n\n1. **Research question:** The primary question your study will answer\n2. **Key assumptions:** Beliefs stakeholders hold that your research will test\n3. **Competing hypotheses:** Where stakeholders disagree — what the research should resolve\n4. **Success criteria:** What decision-makers need to see to act on findings\n5. **Constraints:** What's out of scope or non-negotiable\n6. **Recommended methodology:** How you'll conduct the research and why\n\nShare this document with all stakeholders before launching research. This creates accountability, surfaces any remaining misalignments early, and gives everyone a stake in what the research finds.\n\n---\n\n## Common Stakeholder Interview Mistakes\n\n**Skipping the skeptics.** The stakeholders most likely to resist your findings are the most important to interview early. Include them, listen to their objections, and let their challenges sharpen your research design.\n\n**Letting stakeholders design the research.** There's a difference between incorporating stakeholder input and letting stakeholders dictate methodology. You're the researcher. Gather input on what questions matter; own the decisions about how to answer them.\n\n**Treating stakeholder interviews as user research.** Stakeholders are not representative users. Their beliefs about users may be wrong — that's exactly why you're doing user research. Don't let stakeholder interviews substitute for customer conversations.\n\n**Skipping documentation.** Memory is unreliable. Document everything in writing, share it back to stakeholders for confirmation, and reference it throughout the project.\n\n---\n\n## Key Takeaways\n\nStakeholder interviews are the most underinvested part of the research process — and the most leveraged. An hour of stakeholder conversations before research begins saves weeks of work after findings are delivered.\n\nThe goal is simple: understand what decisions need to be made, what assumptions need to be tested, and what constraints will shape implementation. Do that well, and your research will ask the right questions, earn organizational trust, and drive the kind of action that justifies your team's investment in customer understanding.\n\nFor teams with distributed stakeholders or limited scheduling bandwidth, AI-powered platforms like Koji can run asynchronous stakeholder interviews — gathering rich input without the calendar overhead, and synthesizing findings automatically so you can launch research faster.\n\n---\n\n## Related Resources\n\n- [User Interview Guide Template](/docs/user-interview-guide-template) — How to plan, run, and analyze interviews\n- [Research Brief Template](/docs/research-brief-template) — Define your research before you start\n- [UX Research Plan Template](/docs/ux-research-plan-template) — Structure any research project\n- [Avoiding Bias in Research Interviews](/docs/avoiding-bias-in-interviews) — Keep your questions neutral\n- [Presenting Research Findings to Stakeholders](/docs/presenting-research-findings) — Turn insights into action\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — How Koji structures research conversations\n\n\n## Further reading on the blog\n\n- [Getting Started with Customer Research: A Beginner's Guide](/blog/getting-started-with-customer-research-a-beginner-s-guide) — A practical, step-by-step guide for Product Managers, UX Researchers, and Founders who want to start doing customer research today and build\n- [User Research Budget Template: How to Plan and Justify Research Spending in 2026](/blog/user-research-budget-template-2026) — Build a research budget that actually gets approved. Real benchmarks, line-item templates, ROI arguments, and stage-appropriate guidance — f\n\n<!-- further-reading:blog -->\n","category":"Interview Techniques","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"Stakeholder Interview Guide: Align Your Team Before Research Begins","metaDescription":"Learn how to conduct stakeholder interviews before user research: who to interview, how to ask the right questions, surface assumptions, build alignment, and use AI tools to scale the process.","keywords":["stakeholder interview","stakeholder interview guide","research alignment","stakeholder questions","research planning","internal interviews","research stakeholders"],"aiSummary":"Stakeholder interviews are structured conversations with internal decision-makers, domain experts, and implementers conducted before user research begins. This guide covers who to interview (decision-makers, domain experts, skeptics, implementers), how to prepare flexible question guides, what to listen for (assumptions, competing hypotheses, constraints), how to synthesize input into a research alignment document, and how AI-powered tools like Koji can run asynchronous stakeholder interviews to save scheduling overhead.","aiPrerequisites":["Active research project or study in planning","Access to key internal stakeholders"],"aiLearningOutcomes":["Identify the right stakeholders for any research project","Conduct stakeholder interviews that surface assumptions and competing hypotheses","Synthesize stakeholder input into a research alignment document","Build organizational buy-in for research findings before research begins","Use AI tools to scale stakeholder interviews asynchronously"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 minutes"},{"type":"documentation","id":"c94d9695-7919-4551-9244-8ba57b21875d","slug":"thematic-analysis-guide","title":"The Complete Guide to Thematic Analysis","url":"https://www.koji.so/docs/thematic-analysis-guide","summary":"This guide walks through Braun and Clarke's six-phase thematic analysis framework in detail, covering familiarization, coding, theme generation, reviewing, defining, and reporting. It includes practical examples, comparison tables for coding approaches, and guidance on AI-assisted analysis to accelerate the process.","content":"# The Complete Guide to Thematic Analysis\n\nYou've conducted your interviews. You have hours of recordings and hundreds of pages of transcripts. Now what?\n\nThis is the moment where many research projects stall. The gap between raw data and actionable insights is where thematic analysis comes in — it's the most widely used method for making sense of qualitative data, and this guide will walk you through it step by step.\n\n## What Is Thematic Analysis?\n\nThematic analysis is a method for identifying, analyzing, and reporting patterns (themes) within qualitative data. It was formalized by Virginia Braun and Victoria Clarke in their landmark 2006 paper, which has been cited over 200,000 times according to Google Scholar — making it one of the most influential methodology papers in social science.\n\nUnlike content analysis (which counts occurrences) or grounded theory (which builds theory from data), thematic analysis offers a flexible approach that works across different epistemological positions. You don't need to subscribe to a specific theoretical framework to use it effectively.\n\nAt its core, thematic analysis answers the question: **What are the recurring patterns across my data, and what do they mean?**\n\n## When to Use Thematic Analysis\n\nThematic analysis is appropriate when you:\n\n- Have qualitative data from [user interviews](/docs/user-interview-guide), focus groups, open-ended surveys, or similar sources\n- Want to identify patterns across multiple participants or data sources\n- Need to produce findings that are accessible to non-researcher stakeholders\n- Want a flexible method that doesn't require deep expertise in a specific theoretical tradition\n\nIt's less suitable when you need to preserve individual narratives (use narrative analysis), when you're building theory from scratch (use grounded theory), or when you need to analyze language use itself (use discourse analysis).\n\n## Braun and Clarke's Six-Phase Framework\n\nBraun and Clarke's framework provides a clear, systematic process. These phases are sequential but not strictly linear — you'll often move back and forth between them as your understanding deepens.\n\n### Phase 1: Familiarization with the Data\n\n**Goal:** Immerse yourself in the data so deeply that you begin noticing patterns intuitively.\n\nThis phase is about reading and re-reading your data. If you conducted the interviews yourself, you already have a head start, but don't skip the careful re-reading — you'll notice things you missed in the moment.\n\n**What to do:**\n\n- Transcribe all audio/video recordings if you haven't already\n- Read through every transcript at least twice\n- Take initial notes — jot down things that strike you as interesting, surprising, or recurring\n- Don't start coding yet; stay open and curious\n\nAccording to a 2020 study published in the International Journal of Qualitative Methods, researchers who spend more time on familiarization produce significantly more nuanced and well-supported themes. Rushing this phase is the most common cause of shallow analysis.\n\n**Time investment:** Plan for 1–2 hours per transcript during familiarization. For a study of 10 interviews, that's 10–20 hours before you even begin formal coding.\n\n### Phase 2: Generating Initial Codes\n\n**Goal:** Systematically label meaningful segments of your data.\n\nCoding is the process of attaching short labels to segments of text that capture something relevant to your research questions. A single passage might receive multiple codes.\n\n**What to do:**\n\n- Work through each transcript systematically\n- Highlight data segments that relate to your research questions\n- Assign descriptive codes to each segment\n- Code for as many potential themes as possible — you can always merge or discard later\n- Keep a codebook that defines each code\n\n**Example of coding in practice:**\n\n| Data Excerpt | Codes |\n|---|---|\n| \"I spend half my Monday just trying to figure out what everyone else did last week.\" | Time waste, Status visibility gap, Monday friction |\n| \"The dashboard shows numbers but I don't know what to actually do with them.\" | Data without context, Actionability gap, Dashboard frustration |\n| \"I just ask Sarah because she always knows what's going on.\" | Informal information networks, Key-person dependency, Workaround behavior |\n\n**Coding approaches:**\n\n| Approach | Description | Best For |\n|---|---|---|\n| **Inductive** | Codes emerge from the data itself | Exploratory research, new domains |\n| **Deductive** | Codes come from existing theory or a predefined framework | Testing hypotheses, building on prior research |\n| **Hybrid** | Start with some deductive codes, allow inductive codes to emerge | Most product research scenarios |\n\n### Phase 3: Searching for Themes\n\n**Goal:** Group your codes into potential themes — broader patterns that tell a story about your data.\n\nA theme captures something important about the data in relation to your research question. It's not just a topic (like \"onboarding\") — it's a pattern with a point (like \"users feel abandoned after initial setup because documentation assumes prior expertise\").\n\n**What to do:**\n\n- Gather all codes and their associated data into one view\n- Look for codes that cluster together or relate to a similar concept\n- Create candidate themes by grouping related codes\n- Some codes won't fit anywhere — that's fine; set them aside\n- Consider both semantic themes (surface-level) and latent themes (underlying assumptions or ideologies)\n\nThis is where [affinity mapping](/docs/affinity-mapping) can be incredibly valuable — physically or digitally grouping coded data to see patterns emerge.\n\n### Phase 4: Reviewing Themes\n\n**Goal:** Refine your themes so they are coherent, distinct, and well-supported by data.\n\nThis is a quality-control phase. You're checking that your themes actually work — that they hold together internally and are clearly distinguishable from each other.\n\n**Two levels of review:**\n\n1. **Level 1 — Review coded extracts.** Read all the data segments assigned to each theme. Do they form a coherent pattern? If not, consider splitting the theme, moving some codes to another theme, or discarding codes that don't fit.\n\n2. **Level 2 — Review against the full dataset.** Re-read the entire dataset with your theme map in mind. Do the themes accurately represent the data as a whole? Are there patterns you missed?\n\n**Signs a theme needs work:**\n\n- It's too broad (captures everything but says nothing)\n- It overlaps significantly with another theme\n- It's supported by only one or two data points\n- You can't explain it in one or two sentences\n\n### Phase 5: Defining and Naming Themes\n\n**Goal:** Write clear, concise definitions and give each theme a compelling name.\n\nEach theme should have:\n\n- **A name** that captures the essence (avoid generic labels like \"Communication Issues\" — try \"The Information Black Hole Between Teams\" instead)\n- **A definition** of what the theme captures and what it doesn't\n- **A narrative** that explains the story this theme tells, with supporting data\n\nGood theme names are specific, evocative, and instantly understandable. They help stakeholders grasp the insight without reading the full analysis.\n\n### Phase 6: Producing the Report\n\n**Goal:** Tell a compelling, evidence-based story that answers your research questions.\n\nYour report isn't just a list of themes — it's a narrative that weaves together your findings into a coherent answer to your research questions. Each theme should be illustrated with vivid data extracts that bring the pattern to life.\n\n**A strong thematic analysis report includes:**\n\n- An overview of the research questions and methodology\n- A presentation of each theme with supporting evidence\n- An analysis of how themes relate to each other\n- Implications for design, product, or business decisions\n- Limitations and areas for further research\n\n## The Time Problem — And How to Solve It\n\nLet's be honest: traditional thematic analysis is slow. Research published in Qualitative Research journal estimated that a full thematic analysis of 10 interviews takes between **60 and 120 hours** when done manually — from transcription through reporting.\n\nFor product teams operating on sprint timelines, that's often not feasible. This is where the analysis process has evolved significantly:\n\n**AI-assisted analysis** can reduce the mechanical parts of the process — transcription, initial coding, and code clustering — from days to hours. A 2023 study in the Journal of Medical Internet Research found that AI-assisted qualitative coding achieved 78% agreement with expert human coders, and that human-AI collaborative coding was faster than either alone.\n\nKoji applies this principle to research analysis: after your interviews are complete, AI generates initial codes and theme suggestions that you can review, refine, and build upon. This keeps the researcher's judgment at the center while eliminating the repetitive mechanical work. The result is thematic analysis in hours rather than weeks — without sacrificing rigor.\n\n## Inductive vs. Deductive Thematic Analysis\n\n| Dimension | Inductive | Deductive |\n|---|---|---|\n| **Starting point** | The data itself | A pre-existing framework or theory |\n| **Code generation** | Codes emerge from reading data | Codes derived from theory before analysis |\n| **Flexibility** | High — follows wherever data leads | Lower — constrained by the framework |\n| **Best for** | Exploratory research, new domains | Testing specific hypotheses, building on prior studies |\n| **Risk** | May miss connections to existing knowledge | May force data into ill-fitting categories |\n\nMost product research benefits from a **hybrid approach**: start with a few deductive codes based on your research questions, but remain open to inductive codes that emerge from the data.\n\n## Quality Criteria for Thematic Analysis\n\nHow do you know if your analysis is good? Braun and Clarke outlined a 15-point checklist, but here are the essentials:\n\n1. **Transcriptions are accurate** and retain enough detail for analysis\n2. **Each data item has been given equal attention** during coding\n3. **Themes are not just paraphrases of questions** — they capture patterns\n4. **Data has been analyzed, not just described** — you've interpreted what patterns mean\n5. **There is a good balance between analytic narrative and data extracts**\n6. **Enough time has been allocated** to complete all phases adequately\n\n## Common Pitfalls\n\n- **Using data collection questions as themes.** \"What participants said about onboarding\" is a topic, not a theme. A theme captures a pattern with a point.\n- **Weak or unconvincing themes.** Every theme needs substantial supporting evidence from multiple participants.\n- **Mismatch between data and claims.** Don't claim a theme is prevalent if only two people mentioned it.\n- **Not going beyond description.** Description says what the data contains. Analysis says what the data means.\n\n## Getting Started with Thematic Analysis\n\nIf you're new to thematic analysis, here's a practical starting point:\n\n1. Start with a small dataset — 5 to 6 interviews from a focused study\n2. Use a hybrid coding approach with 3–5 deductive codes and room for inductive codes\n3. Keep a reflexive journal noting your interpretive decisions\n4. Discuss emerging themes with a colleague — external perspective improves quality\n5. Practice naming themes with specificity\n\nFor related techniques, explore our guide to [affinity mapping](/docs/affinity-mapping), which provides a complementary approach to organizing qualitative data into themes.\n\n## Next Steps\n\n- [User Interview Guide](/docs/user-interview-guide) — ensure your data collection supports strong analysis\n- [Affinity Mapping](/docs/affinity-mapping) — a visual approach to theme identification\n- [Writing Interview Questions](/docs/writing-interview-questions) — better questions lead to richer data for analysis\n\n## Further reading on the blog\n\n- [How to Analyze Customer Interview Data: A Complete Guide](/blog/how-to-analyze-customer-interview-data) — You ran the interviews. Now what? Here is a step-by-step process for turning raw transcripts into clear, actionable insights your team will \n- [How to Analyze User Interview Data: A Complete Guide (2026)](/blog/how-to-analyze-user-interview-data) — You ran the interviews. Now what? This step-by-step guide covers how to turn raw interview data into clear, actionable insights — with and w\n- [Best AI Thematic Analysis Tools in 2026: The Complete Buyer's Guide](/blog/best-ai-thematic-analysis-tools-2026) — A side-by-side review of the top AI thematic analysis platforms in 2026 — what each does well, where they fall short, and why AI-native rese\n\n<!-- further-reading:blog -->\n","category":"Research Methods","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"Thematic Analysis Guide — Koji Docs","metaDescription":"Master Braun and Clarke's six-phase thematic analysis framework. Learn to code qualitative data, identify themes, and produce compelling research findings.","keywords":["thematic analysis","qualitative coding","Braun and Clarke","research analysis","qualitative data analysis","coding framework","theme identification"],"aiSummary":"This guide walks through Braun and Clarke's six-phase thematic analysis framework in detail, covering familiarization, coding, theme generation, reviewing, defining, and reporting. It includes practical examples, comparison tables for coding approaches, and guidance on AI-assisted analysis to accelerate the process.","aiPrerequisites":["user-interview-guide"],"aiLearningOutcomes":["Apply Braun and Clarke's six-phase framework to qualitative data","Generate and organize codes systematically","Develop well-defined and well-supported themes","Distinguish between inductive and deductive approaches","Evaluate the quality of thematic analysis using established criteria"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"3130ed02-69e5-4480-8fc5-24f3ea1e5cd7","slug":"eu-ai-act-user-research-compliance","title":"The EU AI Act and User Research: What AI-Moderated Interviews Actually Require (2026)","url":"https://www.koji.so/docs/eu-ai-act-user-research-compliance","summary":"AI-moderated customer research sits in the EU AI Act limited-risk transparency tier under Article 50, applicable from 2 August 2026, requiring only that participants be told they are interacting with an AI at the start. Two escalations matter: inferring emotions from biometric data such as vocal tone (prohibited in workplace and education settings under Article 5, penalties up to EUR 35M or 7% of global turnover), and using AI interviews to screen or evaluate employees or candidates (Annex III high-risk, deferred to 2 December 2027 by the July 2026 AI Omnibus Regulation). Analysing what participants say is not emotion recognition; inferring affect from how they sound is. Koji discloses AI moderation by default, models no vocal affect, and reports in aggregate.","content":"If you run AI-moderated interviews with participants in the EU, here is the short answer: **customer research is almost certainly in the AI Act's limited-risk \"transparency\" tier, not the high-risk tier.** Your core duty is one sentence long — tell participants they are interacting with an AI, before the conversation starts. That obligation, in Article 50, became applicable on **2 August 2026**.\n\nTwo things escalate a study out of that comfortable tier:\n\n1. **Inferring emotions from voice or face.** Emotion recognition is prohibited outright in workplace and education settings under Article 5, and carries disclosure duties everywhere else.\n2. **Using AI interviews to recruit, screen, or evaluate employees.** That is Annex III territory, where the full high-risk regime applies.\n\nMost product and UX teams never touch either. But the teams that do — HR tech, candidate experience, employee listening — are frequently the ones who assume \"it's just a survey\" and get it wrong. This guide draws the line precisely.\n\n> **A note on scope.** The AI Act governs the *AI system*. GDPR governs the *personal data* that flows through it. They are separate regimes with separate penalties, and satisfying one does not satisfy the other. If you have not already worked through lawful basis, retention, and sub-processors, start with our [GDPR-compliant AI user research guide](/docs/gdpr-compliant-ai-user-research) and treat this page as the second layer.\n\n## The four risk tiers, mapped to research\n\nThe AI Act sorts systems by what they do, not by how they are built. Here is where research activities land.\n\n| Tier | What it covers | Typical research example | Your obligation |\n| --- | --- | --- | --- |\n| **Prohibited** (Art. 5) | Unacceptable-risk practices | Inferring employee emotions from voice tone in a workplace study | Do not do it. In force since 2 February 2025 |\n| **High-risk** (Annex III) | Employment, education, essential services, and six other domains | AI interviews used to screen or evaluate job candidates | Full Chapter III regime: risk management, data governance, human oversight, logging, conformity assessment |\n| **Limited-risk** (Art. 50) | Systems that interact directly with people, or generate synthetic content | **AI-moderated customer discovery, VoC, concept testing, usability research** | Disclose that the participant is talking to an AI. Mark synthetic content |\n| **Minimal risk** | Everything else | Thematic analysis run over already-collected, de-identified transcripts | No specific AI Act duty |\n\nThe vast majority of commercial customer research — the discovery calls, the churn interviews, the pricing studies — lives in that third row. The obligation is real but light.\n\n## What Article 50 actually requires of you\n\nArticle 50(1) puts the design duty on the **provider**: a system that interacts directly with natural persons must be built so those people are informed they are dealing with an AI, unless it is already obvious from the context.\n\nArticle 50(3) puts a separate duty on the **deployer**: if you operate an emotion recognition or biometric categorisation system, you must inform the people exposed to it.\n\nThat provider/deployer split matters commercially, because it determines who owes what:\n\n- **If you use a platform like Koji**, the platform is the provider of the AI system. The disclosure has to be engineered into the interview experience, and that is the vendor's job.\n- **You are the deployer.** You choose the purpose, the audience, and the questions. You own the decision about whether your study strays into emotion inference or employment evaluation — and no vendor can make that call for you.\n\nThe practical bar for disclosure is low but specific. It must be **clear, at the first interaction, and not buried**. A line in a privacy policy does not satisfy it. \"You're chatting with an AI interviewer\" on the opening screen does.\n\nThere is also a \"unless it is obvious\" carve-out. Do not lean on it. A participant who clicked an email link labelled \"share your feedback\" has no reason to assume the interviewer is software, and regulators read obviousness narrowly.\n\n## The emotion recognition trap — and the nuance most guides get wrong\n\nThis is the part worth reading twice, because the distinction is genuinely subtle and it decides whether you are doing something regulated, something prohibited, or something entirely unremarkable.\n\nThe Act defines an emotion recognition system as one that identifies or infers the emotions or intentions of natural persons **on the basis of their biometric data**. That last clause does the work:\n\n- **Coding what someone said is not emotion recognition.** If your analysis reads the words \"I was really frustrated when the export failed\" and tags that response as negative sentiment, you are processing language, not biometric data. This is ordinary qualitative analysis and falls outside the definition.\n- **Inferring emotion from how someone sounded is a different matter.** Deriving affect from vocal tone, pitch, or facial expression uses biometric data, and that lands inside the definition.\n\nSo a voice interview is not a compliance problem in itself. A voice interview with a tone-based \"sentiment from audio\" feature is. And in a **workplace or educational** setting, that second thing is not merely regulated — Article 5(1)(f) prohibits it, at the top penalty band.\n\nThe safe posture, and the one Koji is built around: analyse **what participants say**, never how their voice sounds. Thematic analysis, quality scoring, and sentiment in Koji all operate on the transcript. There is no vocal affect model anywhere in the pipeline, which keeps voice studies in the limited-risk tier by design rather than by configuration.\n\n## When research becomes high-risk: the employment line\n\nAnnex III designates AI systems used to recruit, select, and evaluate people as high-risk. The trigger is the **decision the output feeds**, not the interview format.\n\nDraw the line like this:\n\n- **Not high-risk:** anonymous employee engagement research, exit interviews analysed in aggregate to find retention themes, candidate experience studies measuring how your hiring process felt.\n- **High-risk:** an AI interview that scores, ranks, or filters candidates. Any output that influences who gets hired, promoted, or terminated at the individual level.\n\nIf you are in the second bucket, the full Chapter III regime applies — risk management system, data governance, technical documentation, logging, human oversight, accuracy and robustness testing, and a conformity assessment before you go to market. That is a compliance programme, not a checklist, and it needs counsel.\n\nIf you are in the first bucket, keep it there deliberately: report at the cohort level, do not generate per-person scores that feed personnel decisions, and document that design choice. Our guides on [anonymous employee research](/docs/anonymous-employee-research-ai-interviews) and [exit interviews](/docs/exit-interview-survey-guide) both assume aggregate-only reporting for exactly this reason.\n\n## The timeline, after the Omnibus\n\nThe compliance calendar shifted materially in 2026, and a lot of published advice is now stale. The AI Omnibus Regulation entered into force in **July 2026**, deferring the high-risk deadlines that had been set for 2 August 2026.\n\n| Date | What applies |\n| --- | --- |\n| 1 August 2024 | AI Act entered into force |\n| 2 February 2025 | Prohibited practices apply — including workplace and education emotion recognition |\n| **2 August 2026** | **Article 50 transparency obligations apply. This is the date that binds customer research** |\n| 2 December 2026 | Synthetic content marking and watermarking obligations |\n| 2 December 2027 | Annex III high-risk obligations — deferred 16 months by the Omnibus |\n| 2 August 2028 | High-risk AI embedded in products already covered by EU product-safety rules — deferred 12 months |\n\nThe headline for researchers: **the deferral does not help you.** What moved was the high-risk regime. Article 50 — the tier customer research actually sits in — was not postponed. If you are running AI interviews in the EU, your deadline is now, not December 2027.\n\n## What non-compliance costs\n\nPenalties are tiered to match the risk tiers, and they are calculated on **global** turnover, which is why they get executive attention:\n\n- **Prohibited practices** (Art. 5): up to **€35 million or 7%** of worldwide annual turnover, whichever is higher.\n- **High-risk and other obligations**: up to **€15 million or 3%** of worldwide annual turnover for deployers who fail their duties.\n\nPut plainly: switching on a vocal-emotion feature in an employee study is in the same penalty band as the Act's most serious violations. That single configuration choice is worth more scrutiny than most research programmes give it.\n\n## Your compliance checklist\n\nRun this before your next EU study ships:\n\n1. **Classify the study.** Customer research, or employment decision? Write the answer down. This one line determines everything else.\n2. **Disclose the AI up front.** First screen, plain language, before the first question.\n3. **Confirm no biometric emotion inference.** Ask your vendor directly whether any model infers affect from audio or video. Get it in writing.\n4. **Name the AI in your participant notice**, alongside your GDPR disclosures — see [research consent form templates](/docs/research-consent-form-templates).\n5. **Keep employee research aggregate-only.** No per-person scores feeding personnel decisions.\n6. **Log your studies.** Purpose, dates, audience, model used. If a regulator asks, your defence is documentation.\n7. **Check the provider's posture.** Providers carry the design duty — verify yours has actually met it rather than assuming.\n8. **Re-check before 2 December 2026** if you publish AI-generated content externally, when synthetic content marking kicks in.\n\n## How Koji handles this\n\nThe AI Act rewards platforms that made the right architectural decisions early, because the obligations that matter here are engineered in, not toggled on:\n\n- **Disclosure is built into the interview experience.** Every Koji interview identifies itself as AI-moderated at the outset. There is no configuration required and no way to accidentally ship a study that hides it.\n- **No vocal emotion inference anywhere.** Koji's analysis operates on transcripts. [Voice interviews](/docs/ai-voice-interviews) transcribe speech and analyse language — tone is never modelled — which keeps voice studies in the limited-risk tier structurally.\n- **Aggregate-first reporting.** Koji's [report aggregation](/docs/research-repository-guide) rolls findings up to themes and cohorts by default, which is exactly the posture that keeps employee research out of Annex III.\n- **Transcript-level traceability.** Every theme in a Koji report links back to the quotes that produced it, so when someone asks how a conclusion was reached, the answer is a citation rather than a shrug.\n- **Structured questions instead of inference.** This is the underrated compliance advantage. When you need to know how someone feels, the robust move is to *ask them* with a scale question rather than infer it from their voice. Koji's six [structured question types](/docs/structured-questions-guide) — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — give you quantified affect as self-reported data. A 1–7 satisfaction scale is better evidence than a tone model, and it is not regulated as biometric processing. Better methodology and lighter compliance load, from the same design decision.\n\nThe wider point: legacy survey tools were built before any of this existed, and bolt compliance on through settings you have to find and configure correctly. A platform designed in the AI Act era can make the compliant path the default one — which is the difference between a control you have to remember and a property of the system.\n\n## Related Resources\n\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — the data-protection layer that sits underneath AI Act compliance\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types, and why self-reported scales beat inferred affect\n- [AI Interview Data Privacy & Security](/docs/ai-interview-data-privacy-security) — how interview data is stored, encrypted, and retained\n- [Research Ethics Guide](/docs/research-ethics-guide) — the ethical duties that outlast any single regulation\n- [Research Consent Form Templates](/docs/research-consent-form-templates) — copy-ready notices covering AI disclosure\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) — the parallel regime for health-related studies\n- [Anonymous Employee Research with AI Interviews](/docs/anonymous-employee-research-ai-interviews) — aggregate-only design that stays out of Annex III\n\n*Regulatory information current as of July 2026 and reflects the AI Omnibus Regulation. This guide is practitioner orientation, not legal advice — confirm your specific obligations with qualified counsel.*\n","category":"Research Operations","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"EU AI Act & User Research: What AI Interviews Require (2026)","metaDescription":"AI-moderated customer interviews fall under the EU AI Act's Article 50 transparency tier, applicable 2 August 2026. Learn the disclosure rule, the emotion-recognition trap, when research becomes high-risk, and a practical compliance checklist.","keywords":["eu ai act user research","eu ai act ai interviews","ai act article 50","ai act transparency obligations","ai act high risk research","emotion recognition ai act","ai act compliance checklist","ai moderated interviews eu","ai act annex iii employment","ai act penalties"],"aiSummary":"AI-moderated customer research sits in the EU AI Act limited-risk transparency tier under Article 50, applicable from 2 August 2026, requiring only that participants be told they are interacting with an AI at the start. Two escalations matter: inferring emotions from biometric data such as vocal tone (prohibited in workplace and education settings under Article 5, penalties up to EUR 35M or 7% of global turnover), and using AI interviews to screen or evaluate employees or candidates (Annex III high-risk, deferred to 2 December 2027 by the July 2026 AI Omnibus Regulation). Analysing what participants say is not emotion recognition; inferring affect from how they sound is. Koji discloses AI moderation by default, models no vocal affect, and reports in aggregate.","aiPrerequisites":["Basic understanding of user research methods","Familiarity with GDPR fundamentals"],"aiLearningOutcomes":["Classify a research study into the correct EU AI Act risk tier","Meet Article 50 transparency obligations in an AI-moderated interview","Distinguish transcript sentiment analysis from regulated biometric emotion recognition","Recognise when employee or candidate research crosses into Annex III high-risk territory","Apply the post-Omnibus 2026-2028 compliance timeline","Run a pre-launch AI Act compliance checklist"],"aiDifficulty":"intermediate","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"9d1c6645-a534-4134-919d-49a1210da152","slug":"research-ops-guide","title":"ResearchOps: The Complete Guide to Scaling Research Operations","url":"https://www.koji.so/docs/research-ops-guide","summary":"ResearchOps (Research Operations) is the infrastructure layer that makes user research sustainable, consistent, and scalable. This guide covers the eight pillars of research operations: participant recruitment and panel management, consent and compliance, tools and technology stack, research repository and knowledge management, team enablement and democratization, stakeholder engagement, metrics and impact tracking, and governance and prioritization. Includes maturity model, modern AI-powered stack recommendations, and scaling guidance for teams from solo researchers to enterprise programs. Koji is positioned as the AI infrastructure layer that automates moderation, transcription, analysis, and synthesis — enabling continuous research without full-time operational overhead.","content":"\n# ResearchOps: The Complete Guide to Scaling Research Operations\n\n**The bottom line:** ResearchOps (Research Operations) is the infrastructure layer that makes user research sustainable, consistent, and scalable — covering participant recruitment, consent and compliance, tooling, knowledge management, and team enablement. Done well, it transforms research from a series of one-off projects into an always-on organizational capability. AI-powered platforms like Koji have fundamentally changed what's possible, making continuous research accessible without a full-time operations team.\n\nWhen research works well, it looks effortless: researchers spend their time thinking deeply about users, not scrambling for participants or re-entering data into multiple systems. That effortlessness is the product of deliberate operational investment — and ResearchOps is the discipline that creates it.\n\n---\n\n## What Is ResearchOps?\n\nResearchOps is the set of systems, processes, tools, and people that support and enable research practice at scale. It emerged as a formal discipline around 2018 as UX research teams at tech companies grew large enough that coordination overhead was visibly limiting research output.\n\nThe ReOps Community — a global network of research operations professionals — defines ResearchOps as \"the people, mechanisms, and strategies that set user research in motion\" — scaling research reach, impact, and quality.\n\nResearchOps is not:\n- A gatekeeping function that slows research down\n- A purely administrative role\n- Only relevant to large enterprise teams\n\nIt is:\n- The operational infrastructure that makes research faster, more consistent, and more impactful\n- A force multiplier for researchers\n- Increasingly achievable by small teams through AI-powered automation\n\n---\n\n## The Eight Pillars of Research Operations\n\nThe ResearchOps community has identified eight core practice areas. Here's how each works and where AI tools create leverage:\n\n### 1. Participant Recruitment and Panel Management\n\nFinding the right research participants is consistently cited as the top operational bottleneck. Recruitment delays are the most common reason research launches late and findings arrive too late to influence decisions.\n\nEffective ResearchOps builds:\n- **Participant panels:** A pre-screened database of people who have consented to be contacted for research. Panels dramatically reduce time-to-recruit from weeks to days.\n- **Recruitment workflows:** Standardized processes for screening, scheduling, and reminding participants that run with minimal manual effort.\n- **Incentive management:** Scalable systems for compensating participants fairly and efficiently.\n\n**Where AI changes this:** Platforms like Koji eliminate traditional scheduling entirely. Participants receive a link, complete the interview on their own time (voice or text), and you receive a synthesized report — no calendars, no scheduling back-and-forth, no no-shows. One researcher can run 100 interviews in a week with the same effort that previously supported 10.\n\n### 2. Consent, Ethics, and Compliance\n\nResearch data comes with legal and ethical obligations. ResearchOps builds the systems to handle them consistently:\n- Consent form templates that meet legal requirements (GDPR, CCPA, institutional review)\n- Data retention and deletion policies and the systems to execute them\n- Processes for handling sensitive data from vulnerable populations\n- Documentation that supports compliance audits\n\nConsistent consent handling also builds participant trust, improving response rates and data quality. Participants who feel respected and protected are more willing to share honestly.\n\n### 3. Tools and Technology Stack\n\nResearchOps selects, integrates, and maintains the research tooling ecosystem. For most teams, this includes:\n- **Scheduling tools:** Calendly, Doodle, or similar for live interview scheduling\n- **Interview platforms:** Video conferencing for moderated sessions; AI platforms like Koji for unmoderated conversational interviews\n- **Transcription and analysis:** Tools that convert audio to searchable text and identify themes\n- **Repository:** A searchable system for storing and retrieving research artifacts\n- **Synthesis:** Tools that aggregate findings across studies\n\n**The modern stack:** AI-native platforms like Koji collapse several of these tools into one. Interviews are conducted, transcribed, analyzed, themed, and synthesized automatically. This reduces tool sprawl, training overhead, and the manual work of moving data between systems.\n\n### 4. Research Repository and Knowledge Management\n\nA research repository is a searchable, organized system for storing research artifacts — transcripts, recordings, reports, insights, personas, and raw data — so that knowledge accumulates rather than disappearing into email inboxes and individual hard drives.\n\nWithout a repository:\n- The same research questions get asked repeatedly because no one knows prior research exists\n- New team members spend months rebuilding context that previous researchers developed\n- Research impact fades immediately after the findings presentation\n\nAn effective repository includes:\n- Consistent tagging and metadata (research type, date, product area, methodology, participant profile)\n- A clear retention policy (what gets stored, for how long, and in what format)\n- Search that surfaces relevant prior research quickly\n- Connections between insights and the product decisions they informed\n\n**Koji's role:** Koji stores all interview transcripts, AI-generated themes, and reports in a searchable format by study. Studies build on each other — findings from one study inform the brief for the next.\n\n### 5. Team Enablement and Research Democratization\n\nResearchOps builds the capability of everyone who conducts research — from dedicated researchers to product managers running their own customer calls.\n\nThis includes:\n- **Templates and frameworks:** Standardized research plan templates, interview guides, survey designs, and report formats that non-specialists can use without starting from scratch\n- **Training:** Onboarding new researchers, teaching interview technique, explaining methodology choices\n- **Research democratization:** Enabling product managers, designers, and engineers to conduct lightweight research within guardrails established by the research team\n- **Quality standards:** Defining what good research looks like and reviewing work that will inform major decisions\n\n**AI's democratizing effect:** When AI handles moderation (asking questions, probing follow-ups), transcription, and analysis, non-researchers can run high-quality conversational interviews by simply setting up a Koji study. The AI consultant builds the interview guide; the AI moderator conducts the interview; AI synthesis generates the themes and report. Research expertise is increasingly embedded in the tool, not only in the researcher.\n\n### 6. Stakeholder Engagement and Research Socialization\n\nResearch only creates value if findings reach the people who can act on them. ResearchOps builds:\n- Communication channels for sharing research (Slack channels, newsletters, research Slack bots)\n- Presentation templates that make findings accessible to non-research audiences\n- Relationships with product, design, and business stakeholders that ensure research is consulted early and findings are trusted\n- Metrics to demonstrate research impact on decisions and outcomes\n\n### 7. Metrics and Impact Tracking\n\nA research operations function without metrics can't demonstrate its own value or improve over time. Core ResearchOps metrics include:\n\n**Operational efficiency:**\n- Average time from research request to insights delivered\n- Participant recruitment time\n- Interview completion rate\n- Cost per interview\n\n**Research coverage:**\n- Number of unique product areas studied per quarter\n- Percentage of major product decisions informed by research\n- Number of participant touchpoints per month\n\n**Organizational reach:**\n- Number of teams consuming research\n- Proportion of product decisions informed by research\n- Stakeholder satisfaction with research quality and timeliness\n\n### 8. Research Governance and Prioritization\n\nWith multiple teams requesting research and limited researcher bandwidth, ResearchOps creates:\n- A research intake process for capturing and prioritizing requests\n- Criteria for evaluating which research is worth doing (impact, urgency, feasibility)\n- A visible research roadmap that aligns research timing with product planning cycles\n- Standards for when research must be done before a major decision can be made\n\n---\n\n## Building ResearchOps for Different Team Sizes\n\n### Solo Researcher (1 person)\nFocus on the high-leverage operational investments:\n1. Build a simple participant panel (even a spreadsheet with past participants who consented to future contact)\n2. Create 2-3 reusable interview guide templates for your most common research types\n3. Establish a lightweight repository (a shared drive with consistent folder structure and naming)\n4. Set up one consistent report format so findings are easy to consume\n\nUse AI tools like Koji to automate moderation, transcription, and synthesis — freeing your time for the thinking work that benefits from human judgment.\n\n### Small Team (2-5 researchers)\nAdd process and governance:\n1. Formalize a research intake process\n2. Build a proper research repository with consistent tagging\n3. Create a participant panel management system\n4. Establish consent and compliance standards\n5. Define quality review processes for high-stakes research\n6. Build relationships with key stakeholders and establish regular research readouts\n\n### Mid-Size Team (5-15 researchers)\nAdd specialization and infrastructure:\n1. Hire or designate a dedicated ResearchOps specialist\n2. Build a self-service research panel that product managers can recruit from for lightweight studies\n3. Implement a purpose-built research repository tool\n4. Create a research democratization program with training and guardrails\n5. Track metrics and report research impact quarterly\n\n### Enterprise Team (15+ researchers)\nBuild a research operations program:\n1. Multiple ResearchOps specialists with distinct focus areas (recruitment, tools, enablement)\n2. Enterprise-grade compliance and data governance\n3. Vendor management for the full research tooling ecosystem\n4. Formal research democratization program with certification\n5. Research program-level metrics tied to business outcomes\n\n---\n\n## The Modern ResearchOps Stack\n\nThe research operations tooling landscape has shifted dramatically with the rise of AI. Here's what an efficient modern stack looks like:\n\n| Function | Traditional Approach | AI-Powered Approach |\n|----------|---------------------|---------------------|\n| Interview moderation | Researcher moderates live | AI moderates asynchronously (Koji) |\n| Scheduling | Calendly + email back-and-forth | No scheduling (async link) |\n| Transcription | Otter.ai, Rev | Automatic in Koji |\n| Analysis | Manual coding | AI theme synthesis |\n| Reports | Manual writing | Auto-generated, shareable |\n| Repository | Dovetail, Notion | Koji study archive + tags |\n| Recruitment | Respondent, User Interviews | Koji Recruit tab + direct links |\n\nTeams using AI-native platforms consolidate 5-7 tools into 1-2, dramatically reducing tool management overhead, training time, and data transfer errors between systems.\n\n---\n\n## ResearchOps Anti-Patterns to Avoid\n\n**Making ResearchOps a gatekeeper.** Operations should remove friction, not add it. If researchers need to submit multi-week requests to run a 30-minute interview, ResearchOps has become a bottleneck.\n\n**Optimizing for consistency over speed.** Rigid processes designed for large-scale research become obstacles for quick exploratory work. Build tiered processes: lightweight for exploratory, rigorous for high-stakes decisions.\n\n**Neglecting knowledge management.** The most common ResearchOps failure mode. Teams invest in recruitment and tools but don't build the systems to capture and share what's learned. Research knowledge evaporates after every researcher offboarding.\n\n**Under-investing in stakeholder relationships.** Tools and processes mean nothing if stakeholders don't trust or consume research. ResearchOps must invest in research socialization, not just operational efficiency.\n\n**Starting too big.** Don't build an enterprise ResearchOps program before you have the team to use it. Start with the 2-3 highest-leverage operational investments and expand from there.\n\n---\n\n## Measuring ResearchOps Maturity\n\nUse this framework to assess where your research operations currently stands:\n\n**Level 1 — Ad Hoc:** Research happens project by project. No consistent processes. Recruitment is improvised. Findings live in individual researchers' heads.\n\n**Level 2 — Repeatable:** Basic processes exist for common research types. A participant panel or database. Consistent consent forms. Report templates. A basic repository.\n\n**Level 3 — Defined:** Formal research operations function. Standardized intake and prioritization. Research democratization program. Metrics tracked. Knowledge management active.\n\n**Level 4 — Managed:** Research operations measured and optimized. Impact tracked to business outcomes. AI tools automate high-volume operational work. Non-researchers run lightweight research within quality guardrails.\n\n**Level 5 — Optimizing:** Research operations continuously improving. Predictive capacity planning. Research influence measurable across product organization. AI handles all operational work; researchers focus entirely on strategy and synthesis.\n\nMost teams sit at Levels 1-2. The biggest leverage comes from moving to Level 3 — defined processes and a working knowledge management system.\n\n---\n\n## Key Takeaways\n\nResearchOps is the infrastructure that makes research sustainable, scalable, and impactful. It's not a luxury for large teams — even solo researchers benefit enormously from investing in the 3-4 highest-leverage operational foundations.\n\nThe emergence of AI-powered research platforms like Koji has made it possible to achieve Level 3-4 maturity with far less operational investment than previously required. Automated moderation, transcription, and synthesis eliminate the most time-consuming operational bottlenecks. Structured questions provide consistent, comparable data across studies. The AI consultant builds interview guides in minutes. The result is a research operation that can run continuously — not just when there's budget for a dedicated study — keeping your team permanently connected to the people who use your product.\n\n---\n\n## Related Resources\n\n- [How to Automate User Research: Build a Pipeline That Runs 24/7](/docs/how-to-automate-user-research) — Practical automation guide\n- [Continuous Discovery: How to Run Weekly Customer Interviews Without Burning Out](/docs/continuous-discovery-user-research) — The continuous research model\n- [How to Scale Your User Research Practice](/docs/scaling-user-research) — Scaling research team and impact\n- [How to Build a UX Research Repository](/docs/research-repository-guide) — Knowledge management in depth\n- [Proving Research ROI: How to Justify Your Customer Interview Program](/docs/research-roi-guide) — Measuring and communicating research impact\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — How Koji structures research for consistency at scale\n\n\n## Further reading on the blog\n\n- [Research Democratization: How to Scale Insights Beyond the Research Team (2026)](/blog/research-democratization-scaling-insights-2026) — Research demand is growing faster than teams can scale. Learn how to enable non-researchers to run high-quality studies — without sacrificin\n- [The UX Researcher's Guide to Scaling Research with AI (2026)](/blog/ux-researcher-guide-scaling-with-ai-2026) — Demand for user research has never been higher — but researcher headcount hasn't kept pace. Here's how UX researchers are using AI to scale \n- [Why AI Interviewers Are the Future of Customer Research](/blog/why-ai-interviewers-are-the-future-of-customer-research) — AI interviewers are transforming how product teams conduct customer research, enabling conversations at scale without sacrificing depth or q\n\n<!-- further-reading:blog -->\n","category":"Research Operations","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"ResearchOps: The Complete Guide to Scaling Research Operations (2026)","metaDescription":"The complete guide to research operations: how to build participant panels, manage research tooling, create knowledge repositories, democratize research, and use AI platforms to scale continuous customer insight.","keywords":["research ops","researchops","research operations","scaling user research","research infrastructure","ux research operations","research ops guide"],"aiSummary":"ResearchOps (Research Operations) is the infrastructure layer that makes user research sustainable, consistent, and scalable. This guide covers the eight pillars of research operations: participant recruitment and panel management, consent and compliance, tools and technology stack, research repository and knowledge management, team enablement and democratization, stakeholder engagement, metrics and impact tracking, and governance and prioritization. Includes maturity model, modern AI-powered stack recommendations, and scaling guidance for teams from solo researchers to enterprise programs. Koji is positioned as the AI infrastructure layer that automates moderation, transcription, analysis, and synthesis — enabling continuous research without full-time operational overhead.","aiPrerequisites":["Active research practice with recurring research needs","At least one research study completed"],"aiLearningOutcomes":["Understand the eight pillars of research operations","Build a research operations foundation appropriate for your team size","Choose and configure a modern AI-powered research stack","Create participant panels, repository systems, and governance processes","Measure research operations maturity and plan improvements","Use Koji to automate the highest-overhead operational bottlenecks"],"aiDifficulty":"intermediate","aiEstimatedTime":"16 minutes"},{"type":"documentation","id":"822663ee-d4fb-42f5-983f-f9024a994c69","slug":"ux-research-team-structure","title":"UX Research Team Structure: Centralized, Embedded, and Hub-and-Spoke Models Compared (2026)","url":"https://www.koji.so/docs/ux-research-team-structure","summary":"Four research operating models exist: centralized, embedded/decentralized, hub-and-spoke, and ops-enabled democratization. Adoption is nearly even — NN/g found 29% centralized, 32% decentralized, 31% hybrid across 557 professionals in 356 companies — while only 6% of researchers sit on a dedicated research team and nearly a third report into Product (User Interviews). Hybrid structures correlate with larger organisations. NN/g ratio benchmark is roughly 1 researcher : 5 designers : 50 developers, with the caveat that ratios are not maturity scores. AI-moderated research changes the design by lifting the moderation and synthesis constraint, letting a one-to-three-person core support a dozen product teams and making the centre owner of templates, standards, and quality gates rather than individual studies.","content":"## The short answer\n\n**There are four viable research operating models, and the right one is decided by two variables: how fast product teams need answers, and how much standardisation your decisions require.** Centralized wins on rigor and career development; embedded wins on speed and influence; hub-and-spoke is what most organisations converge on as they grow; ops-enabled democratization is what AI-native tooling has made viable below the headcount that used to be required.\n\nThe benchmark data is unusually clear on two points. First, adoption is almost evenly split: Nielsen Norman Group's survey of 557 UX and design professionals across 356 companies in 59 countries found **29% centralized, 32% decentralized, and 31% hybrid/matrix** ([NN/g](https://www.nngroup.com/articles/design-team-statistics/)) — nobody has won this argument. Second, researchers specifically are rarely on a research team at all: **only 6% of researchers sit on a dedicated User Research team, and nearly a third report into Product** ([User Interviews, State of Research Strategy](https://www.userinterviews.com/state-of-research-strategy)).\n\nThat second number is the one that should drive your design. If research is already distributed, the question is not \"centralize or embed\" — it is *what the centre owns* when almost everyone doing research reports somewhere else.\n\n---\n\n## The four models\n\n### 1. Centralized\n\nAll researchers sit on one team under one manager and take intake from product groups.\n\n**Strong when:** methodological rigor matters (regulated products, safety, high-stakes decisions); the team is small enough that context-switching is cheaper than duplication; you need to grow junior researchers.\n**Fails when:** intake becomes a queue. The most common failure is the research team becoming a service desk with a six-week backlog, and product teams routing around it with a survey tool and a weekend.\n\n### 2. Embedded / decentralized\n\nEach researcher sits inside a product team, reporting either into that team or dotted-line to a research lead.\n\n**Strong when:** speed and influence matter more than consistency; researchers need deep domain context; the org runs continuous discovery.\n**Fails when:** it fragments. Among decentralized teams, **43% align to products and 27% to internal departments or business lines** (NN/g) — and without a shared standard, each alignment invents its own screener, its own consent language, and its own repository. Career growth also stalls: an embedded researcher with no research manager has no one who can evaluate their craft.\n\n### 3. Hub-and-spoke (hybrid/matrix)\n\nResearchers are embedded in product teams but report — solid or dotted line — to a central research lead who owns standards, tooling, and development. NN/g's data shows **hybrid structures correlate with larger organisations** (statistically significant), which matches the lived pattern: teams start centralized, embed under pressure, and rebuild a centre once quality problems appear.\n\n**Strong when:** you have more than roughly four researchers and more than one product line.\n**Fails when:** the dotted line has no teeth. If the hub owns standards but not performance reviews, tooling budget, or hiring, it owns nothing.\n\n### 4. Ops-enabled democratization\n\nA small central core (often one to three people) owns standards, tooling, participant operations, and quality review, while PMs, designers, and support staff run most studies themselves on a platform that enforces the method.\n\nThis is the newest model and the fastest-growing. **71% of organisations now have people conducting research who are not researchers**, and the share of organisations where research is essential to strategy at all levels nearly tripled in a year, from 8% to 22% ([Maze, 2026](https://maze.co/blog/future-user-research-2026/)). See the [research democratization playbook](/docs/research-democratization-playbook) for the enablement side.\n\n**Strong when:** demand massively exceeds researcher supply — which is the normal condition.\n**Fails when:** the platform does not enforce quality. Democratization without a method-enforcing tool is just more bad research, faster.\n\n---\n\n## Model comparison\n\n| | Centralized | Embedded | Hub-and-spoke | Ops-enabled |\n|---|---|---|---|---|\n| Speed to answer | Slow (queue) | Fast | Fast | Fastest |\n| Method consistency | High | Low | High | Enforced by tooling |\n| Depth of domain context | Low | High | High | Medium |\n| Career development | Strong | Weak | Strong | Weak for ICs, strong for ops |\n| Repository health | Good | Fragmented | Good | Good |\n| Cost per study | High | High | Medium | Lowest |\n| Typical org size | Under 4 researchers | 2–8 in one product line | 5+ across product lines | Any, especially research-poor orgs |\n| Main failure mode | Backlog and routing-around | Fragmentation and stalled careers | Toothless dotted line | Quality drift |\n\n---\n\n## Where research should report\n\nReporting line determines what research is allowed to influence. NN/g's data shows **27% of design organisations report into product management and 23% directly into the C-suite**; research specifically most often sits under Product.\n\n- **Under Product** — fastest path to roadmap influence; the risk is that research only gets asked questions Product already thought of, and evaluative work crowds out generative work.\n- **Under Design** — protects craft quality and gives researchers a manager who can assess their work; the risk is being scoped to interface questions rather than business ones.\n- **Under a C-level Insights or Strategy function** — the only structure where research routinely shapes decisions above the roadmap; rare below a few hundred employees.\n- **Under Marketing or Growth** — fast funding, but the questions skew to messaging and acquisition.\n\nA practical test: if research reports into the function whose plans it most often contradicts, expect findings to get softened. That is an argument for a dotted line to someone independent, not for a reorg.\n\n---\n\n## Staffing math: how many researchers do you actually need?\n\nNN/g's ratio benchmark is roughly **1 researcher : 5 designers : 50 developers** — improved from 1:5:100 — with the important caveat that ratios describe a state, not a maturity score. NN/g is explicit that \"ratios are not maturity scores.\" The complementary NN/g finding is that headcount tracks organisation size mechanically: on average one additional designer per extra 200 employees.\n\nUse the ratio as a sanity check, not a target. The better question is throughput: how many decisions per quarter require evidence, and what is the largest number of studies a researcher can run well? Human researchers manage roughly four to six in-depth interviews a day before fatigue degrades quality, and a traditional study runs about six weeks end to end.\n\nFor the cost side of this calculation — salary bands by level and region, fully loaded cost, and research budget as a share of ARR — see [UX researcher salary and team cost benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks). For the build-the-team decision, see [hiring a UX researcher](/docs/hiring-ux-researcher-guide) and [agency vs. in-house](/docs/ux-research-agency-vs-in-house).\n\n---\n\n## How AI changes the org design\n\nEvery model above was designed around a constraint that no longer holds: a researcher can only moderate one session at a time, and analysis takes longer than fieldwork.\n\nWhen moderation and synthesis are automated, three things change structurally:\n\n1. **The hub can serve more spokes.** The central function's scarce resource stops being interview hours and becomes study design and quality review — which scales far better. A one-to-three-person core can credibly support a dozen product teams.\n2. **Democratization stops degrading quality.** The historical objection to non-researchers running studies is that they ask leading questions and over-generalise from four conversations. A platform that enforces the method — structured question types, consistent probing, automatic thematic analysis, an interview quality score — removes most of that risk. Koji scores each conversation 1–5 and only counts sessions at 3 or above.\n3. **The centre's real deliverable becomes templates and standards, not studies.** The highest-leverage artefact a central team ships is a validated study template that any PM can launch on Monday.\n\nConcretely with Koji: the central team owns the research briefs and the [six structured question types](/docs/structured-questions-guide) — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — so every team's data is comparable and aggregatable rather than trapped in one team's notes. Spokes launch AI-moderated text or voice interviews without booking a moderator, and a customisable AI consultant carries the organisation's context and tone into every study. Reports are generated in real time, so the hub reviews finished analysis instead of transcribing. Legacy stacks make the opposite trade: SurveyMonkey or Qualtrics lets anyone launch a survey but nobody enforce a method, while a traditional agency enforces method but only for the studies you can afford.\n\nPricing supports the model too: at €29/month (Insights) and €79/month (Interviews), a spoke team can be given a seat for less than the cost of a single recruited participant, and credits are consumed only by conversations that clear the quality gate.\n\n---\n\n## A worked example: a 40-person product org\n\nFour product teams, two researchers, demand for roughly 20 studies a quarter and capacity for six.\n\n**Diagnosis.** Centralized by default. Six-week backlog. Two of the four teams have quietly started running their own unmoderated surveys, and nobody can find last quarter's findings.\n\n**Design chosen: hub-and-spoke with ops-enabled spokes.**\n\n- **Hub (2 researchers, one designated lead):** owns the study template library, screener and consent standards, the repository, quality review of every published finding, and all high-stakes generative work.\n- **Spokes (1 trained PM or designer per product team):** run evaluative and continuous-discovery studies from hub-approved templates on Koji, with the hub reviewing before findings are published to the repository.\n- **Reporting:** researchers report solid-line to the research lead, dotted-line to their product group. The lead owns hiring, craft review, and the tooling budget — the dotted line has teeth.\n- **Quality gate:** no finding enters the repository without hub sign-off; every study uses structured question types so results aggregate across teams.\n\n**Result to expect:** study throughput rises well past the six-per-quarter ceiling because the two researchers stop moderating routine sessions, while method consistency improves rather than degrades — the spokes are running the hub's templates, not inventing their own.\n\n---\n\n## Changing models: a 90-day plan\n\n**Days 1–30 — Diagnose.** Inventory every study run in the last two quarters, including the ones run without the research team. Count decisions that needed evidence and did not get it. Map current reporting lines and where findings live. Interview four to six stakeholders about what they do when they need an answer this week — see [stakeholder buy-in](/docs/stakeholder-buy-in-user-research).\n\n**Days 31–60 — Design and pilot.** Pick the model, write down what the centre owns (standards, tooling, participants, quality review, hiring) and what the spokes own. Publish three study templates. Pilot with one willing product team rather than announcing a reorg.\n\n**Days 61–90 — Institutionalise.** Fix the reporting lines, set the quality gate, agree how research capacity is requested and prioritised, and instrument the outcome: studies per quarter, time from question to answer, and share of decisions with evidence attached. Review the model every two quarters — structures that fit at 40 people break at 120.\n\n---\n\n## Common mistakes\n\n1. **Reorganising instead of fixing intake.** Most \"we need to embed\" conversations are really a prioritisation problem wearing a costume.\n2. **A hub with no authority.** If the centre does not own standards, tooling budget, and craft review, it is a mailing list.\n3. **Embedding a single researcher with no research manager.** Their craft stops developing and they leave within a year.\n4. **Democratizing without a method-enforcing platform.** Volume goes up, evidence quality goes down, trust in research goes with it.\n5. **Using ratios as targets.** They describe the current state of a sample, not what your decisions require.\n6. **Never revisiting.** The model that fits your org today is a snapshot; the NN/g size correlation says you will need a different one after the next two hiring waves.\n\n---\n\n## Frequently asked questions\n\n**What is the most common UX research team structure?**\nThere is no majority model. NN/g's survey of 557 professionals across 356 companies found 29% centralized, 32% decentralized, and 31% hybrid/matrix. For researchers specifically, only 6% sit on a dedicated research team — most are embedded elsewhere, with nearly a third reporting into Product.\n\n**What is a hub-and-spoke research model?**\nResearchers are embedded in product teams for day-to-day work but report to a central research lead who owns standards, tooling, participant operations, hiring, and craft review. It is the model organisations typically converge on past four or five researchers and more than one product line.\n\n**How many researchers should we have?**\nNN/g's benchmark ratio is roughly 1 researcher to 5 designers to 50 developers, improved from 1:5:100 — but NN/g cautions that ratios are not maturity scores. Size the team from decision throughput instead: how many decisions per quarter need evidence, and how many studies a researcher can run well.\n\n**Where should UX research report?**\nMost often Product, which buys roadmap influence at the cost of being scoped to questions Product already asked. Design protects craft quality; a C-level insights function gives the most strategic reach but is rare below a few hundred employees. Whatever the line, keep an independent path for findings that contradict the plan.\n\n**Does democratization mean we do not need researchers?**\nNo. 71% of organisations already have non-researchers running research; the effect is to change what researchers do — from moderating every session to owning standards, templates, quality review, and the hardest generative work.\n\n**How does AI change research team structure?**\nIt lifts the constraint the models were built around. With AI-moderated interviews and automatic thematic analysis, a one-to-three-person core can support a dozen product teams, and the centre's main deliverable becomes validated study templates and a quality gate rather than individually run studies.\n\n---\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types that make findings comparable across teams\n- [Research Democratization Playbook](/docs/research-democratization-playbook) — enabling non-researchers without losing rigor\n- [Hiring a UX Researcher](/docs/hiring-ux-researcher-guide) — when to hire, levels, and the interview loop\n- [UX Researcher Salary and Team Cost Benchmarks](/docs/ux-researcher-salary-team-cost-benchmarks) — the cost side of the staffing model\n- [UX Research Agency vs. In-House](/docs/ux-research-agency-vs-in-house) — the four sourcing models and their TCO\n- [ResearchOps Guide](/docs/research-ops-guide) — the operational layer every model depends on\n- [Stakeholder Buy-In for User Research](/docs/stakeholder-buy-in-user-research) — securing the mandate before you restructure","category":"Research Operations","lastModified":"2026-07-31T03:18:40.779282+00:00","metaTitle":"UX Research Team Structure: Centralized vs Embedded vs Hub-and-Spoke","metaDescription":"Benchmark data on four research operating models — 29% centralized, 32% embedded, 31% hybrid — plus reporting lines, ratios, and a 90-day migration plan.","keywords":["ux research team structure","centralized vs embedded research","hub and spoke research model","research operating model","ux research reporting line","research team size","researcher to designer ratio","research org design","researchops team model","democratized research team"],"aiSummary":"Four research operating models exist: centralized, embedded/decentralized, hub-and-spoke, and ops-enabled democratization. Adoption is nearly even — NN/g found 29% centralized, 32% decentralized, 31% hybrid across 557 professionals in 356 companies — while only 6% of researchers sit on a dedicated research team and nearly a third report into Product (User Interviews). Hybrid structures correlate with larger organisations. NN/g ratio benchmark is roughly 1 researcher : 5 designers : 50 developers, with the caveat that ratios are not maturity scores. AI-moderated research changes the design by lifting the moderation and synthesis constraint, letting a one-to-three-person core support a dozen product teams and making the centre owner of templates, standards, and quality gates rather than individual studies.","aiDifficulty":"intermediate","aiEstimatedTime":"13 min"},{"type":"documentation","id":"a98b59e4-39bd-4578-b57d-0ece22aebaf0","slug":"medical-device-usability-testing","title":"Medical Device Usability Testing: IEC 62366-1 and FDA Human Factors Requirements (2026)","url":"https://www.koji.so/docs/medical-device-usability-testing","summary":"Medical device usability testing is governed by IEC 62366-1 and FDA human factors guidance. The FDA issued its final guidance, Content of Human Factors Information in Medical Device Marketing Submissions, on 29 May 2026, replacing the December 2022 draft. It defines three human factors submission categories: Category 1 for certain modified devices requiring a conclusion and high-level summary, Category 2 requiring documentation of why there are no critical tasks or no new or impacted critical tasks, and Category 3 requiring a full HFE/UE report with validation data. A new Decision Point D weighs user interface history of use, complexity and existing risk controls to determine whether validation data must be submitted, and the guidance allows justification and leveraged prior data in place of new testing. The engineering chain runs use specification, user profiles, known use problems, task analysis, use-related risk analysis, critical tasks, formative evaluation, summative validation and the HFE/UE report. Formative studies have no pass criteria and roughly five users surface 85 percent of usability problems while ten reach 95 percent. Summative validation requires observed simulated use with representative users performing critical tasks on the final design, with an FDA expectation of at least 15 participants per distinct user group and root cause analysis of every use error. AI-moderated research such as Koji can run use specification interviews, known use problem discovery, IFU and labeling comprehension tests with pre-set pass criteria, formative comprehension studies and post-market use surveillance, but cannot replace observed summative validation or hands-on formative interaction studies. There were 111 Class I recall events and early alerts in 2025 and 1,059 US device recall events in 2024, a four-year high.","content":"Medical device usability testing is a regulated engineering process, not a design nicety. Under **IEC 62366-1** and the FDA's human factors guidance, you must identify the tasks where a use error could hurt someone, design the interface to prevent those errors, and then prove with representative users that you succeeded. The FDA issued its final guidance, *Content of Human Factors Information in Medical Device Marketing Submissions*, on **29 May 2026**, replacing the December 2022 draft — and it changed what many manufacturers have to submit.\n\nThe short version for teams under time pressure: **summative validation still requires observed, simulated use with representative users performing critical tasks on the final design.** No AI platform replaces that. But summative validation is the last 10% of the usability engineering file. The other 90% — understanding your user groups, discovering known use problems, testing whether people can understand your instructions for use, gathering formative feedback, and running post-market use surveillance — is interview and comprehension work, and that is exactly where a platform like Koji collapses weeks into days.\n\n## What the regulations actually require\n\nTwo documents drive the work, and they overlap heavily without being identical.\n\n| | IEC 62366-1 | FDA human factors guidance |\n|---|---|---|\n| Status | Harmonised standard; FDA-recognised consensus standard | Agency guidance (recommendations, but reviewers apply them) |\n| Core artifact | Usability engineering file | HFE/UE report in the marketing submission |\n| Sample size for validation | \"Representative users\" — no number given | Minimum **15 participants per distinct user group** |\n| Risk linkage | Via ISO 14971:2019 risk management | Via use-related risk analysis (URRA) |\n| Emphasis | Process conformity and traceability | Evidence that critical-task risk is acceptable |\n\nThe unifying concept is **use error**: not \"user error.\" The standard deliberately avoids blaming the person. A use error is an act or omission that produces a different result than the manufacturer intended — and the assumption is that the interface invited it. Your job is to find the acts and omissions that could cause harm, then engineer them out.\n\nThat chain runs: *use specification → user profiles → known use problems → task analysis → use-related risk analysis → critical tasks → formative evaluation → summative validation → HFE/UE report.*\n\n## What changed in the May 2026 FDA guidance\n\nThe final guidance keeps the risk-based framework but makes it meaningfully less mechanical. Five changes matter for research planning:\n\n1. **Three human factors submission categories.** Category 1 applies to certain modified devices and asks for a conclusion plus a high-level summary of the HF evaluation. Category 2 applies when you can document why there are no critical tasks — or, for a modification, no new or impacted critical tasks. Category 3 requires a full HFE/UE report including validation testing data.\n2. **A new Decision Point D.** The flowchart now asks whether validation data actually needs to be submitted, weighing the user interface's history of use, its complexity, and the adequacy of existing risk controls. Having a critical task no longer automatically forces a new validation study.\n3. **Justification in place of testing.** For modifications and well-understood interfaces with a safe use history, a rigorous justification built on a comprehensive URRA can substitute for new validation data.\n4. **Leveraging existing data.** The FDA explicitly encourages referencing prior submissions and data from similar devices rather than duplicating studies.\n5. **Far more worked examples.** The example section roughly tripled versus the draft, now covering pediatric users, augmented-reality interfaces, and devices with known use-related problems.\n\nThe practical consequence is that **the quality of your use-related risk analysis and your evidence about real-world use now carries more weight than the number of studies you ran.** Justification-based pathways only survive review when the underlying understanding of users, environments and known problems is deep and documented. That raises the value of upstream research considerably.\n\nContext for why reviewers are strict: there were **111 Class I recall events and early alerts in 2025**, and 2024 saw **1,059 US device recall events** — a four-year high, with Class I recalls at their highest level in fifteen years. Design-related causes lead the list. Every use error found in a formative study is one that does not become a field action.\n\n## Step 1 — Write the use specification with users, not about them\n\nThe use specification defines intended medical indication, patient population, intended user profiles, use environment and operating principle. Most teams write it from internal assumptions, then discover in summative testing that a whole user group was missing.\n\nInterview each candidate user group before you write it. The questions that matter are unglamorous: Who actually performs this step in your clinic at 3am? What else is happening in the room? What do you do when the alarm sounds and you are two rooms away? Which steps do experienced staff skip? What did the previous device get wrong?\n\nThis is where AI-moderated interviews earn their keep. Clinicians are hard to schedule and geographically scattered; **median physician response rates in research sit around 18%**, with a range roughly 10–60%. An asynchronous study that a nurse can complete by voice at the end of a shift, in her own language, without a moderator's calendar, converts far better than a booked video call. Koji's AI interviewer probes automatically when an answer is thin — \"you said you usually skip the priming step; walk me through the last time you did that\" — so you get the incident detail that a static survey never surfaces.\n\n## Step 2 — Hunt for known use problems\n\nIEC 62366-1 expects you to consider known use problems with your device and with similar devices. Teams typically satisfy this with a MAUDE search and a literature scan, then stop. That is a thin file.\n\nA stronger approach adds primary evidence:\n\n- Interview users of the predicate or competitor device about workarounds and near misses.\n- Interview your own service and complaints staff — they hold the richest failure narratives in the company.\n- Interview trainers about which steps consistently need re-teaching.\n\nStructure these so they aggregate. In Koji, a `multiple_choice` question that lists candidate problem steps gives you frequency, a `scale` question captures perceived severity, and an `open_ended` follow-up captures the story behind each selection. You end up with a ranked, quotable, traceable input to the URRA rather than a folder of notes.\n\n## Step 3 — Task analysis and the use-related risk analysis\n\nDecompose use into tasks and subtasks, then for each ask what could go wrong perceptually, cognitively, and in action. Which failures could lead to harm? Those are your **critical tasks**, and they define the scope of validation and the depth of the submission.\n\nTwo disciplines separate a strong URRA from a weak one. First, derive tasks from observed practice, not from the draft IFU — the IFU describes intended use, and use errors live in the gap between intended and actual. Second, keep the traceability explicit: every critical task should trace to a hazardous situation in the ISO 14971 file and to a risk control, and every risk control should trace to the evaluation that tested it.\n\n## Step 4 — Formative evaluation: cheap, early, repeated\n\nFormative studies are exploratory. They have no pass/fail criteria, they can use prototypes, and their entire purpose is to find and fix problems before validation. The classic finding is that **about five representative users surface roughly 85% of usability problems and ten reach about 95%** — which is why running three small formative rounds beats running one big one.\n\nFormative work splits cleanly into two kinds:\n\n- **Interaction studies** — someone must handle the device or prototype while an observer watches. Do these in person or over supervised video. This is not something to automate.\n- **Comprehension and expectation studies** — does the label make sense, does the alarm mean what people think it means, does the IFU step read as one action or two, would a user expect the device to do X after Y? These are pure comprehension work, and running them asynchronously with dozens of clinicians instead of six is a strict upgrade.\n\nKoji handles the second category natively. Show the artefact, ask a `single_choice` recall question with one correct answer and plausible distractors, capture self-rated clarity on a `scale`, then let the AI probe every wrong answer with an `open_ended` follow-up to learn *why* it was misread. Set the pass criterion before you field it — for example, 90% correct identification of the dose-confirmation step — and record the before-and-after result across redrafts. That produces exactly the kind of documented, criterion-based evidence a reviewer can follow.\n\n## Step 5 — Summative validation: what cannot be automated\n\nBe clear with yourself and with your quality team about this boundary.\n\n| Activity | Can AI-moderated research do it? |\n|---|---|\n| Use specification and user-profile interviews | Yes — voice or text, async, any language |\n| Known use problems and near-miss discovery | Yes |\n| IFU, labeling and training comprehension testing | Yes, with pre-set pass criteria |\n| Formative feedback on concepts, alarms, wording | Yes |\n| Formative hands-on interaction studies | No — requires observation of use |\n| **Summative human factors validation** | **No — requires observed simulated use with the final design** |\n| Post-market use surveillance and complaint follow-up | Yes |\n\nSummative validation means representative users from each distinct user group performing critical tasks in a realistic simulated-use environment, with observation of what they do, followed by a knowledge-task debrief and root-cause analysis of every use error and difficulty. The FDA expects a minimum of **15 participants per user group**, more for higher-risk products, and a root cause for each observed problem — not a satisfaction score.\n\nWhere AI-moderated research does contribute to summative work is the debrief. Structured post-task interviewing at scale, with consistent probing and automatic transcription, removes a real source of variability between moderators. But the observation itself stays human.\n\n## Step 6 — Post-market use surveillance\n\nBoth frameworks expect production and post-production information to feed back into the usability engineering file. Complaint data tells you that something went wrong; it rarely tells you why. A short quarterly AI interview study with active users — \"walk me through the last time the device did something you did not expect\" — turns a complaint trend into an identified use error with a candidate root cause. That evidence is what supports your next submission's justification-based pathway under the 2026 framework.\n\n## Common mistakes that cost submissions\n\n- **Treating validation as a usability test.** Reviewers want risk evidence, not SUS scores. A benchmark study is not a summative validation.\n- **Missing a user group.** Home caregivers, cleaning and reprocessing staff, and remote monitoring staff are routinely forgotten, and each needs its own 15.\n- **Testing the IFU you wish you had.** Validate the final labeling, not a cleaned-up draft.\n- **No pass criteria in advance.** A comprehension result without a pre-stated threshold reads as post-hoc.\n- **Untraceable findings.** If a use error in the report cannot be traced to a task, a hazard and a risk control, the file does not hold together.\n- **Confusing formative and summative.** Formative studies have no acceptance criteria; using formative data as validation evidence is a predictable rejection.\n\n## How Koji fits a regulated programme\n\nKoji is an AI-native research platform, not a regulatory submission tool. It runs AI-moderated interviews by voice or text, probes follow-ups automatically, supports six structured question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — and produces analysis with quote-level traceability back to transcripts. In a human factors programme that means: faster access to scattered clinicians, comprehension evidence with pre-set criteria, multilingual studies for multi-market submissions, and exports and transcripts you can attach to the usability engineering file as method and data lineage.\n\nWhat it does not do is watch someone use your device. Plan the observed studies properly, and use AI research to make everything around them faster and better evidenced.\n\n## Frequently asked questions\n\n**Can AI-moderated interviews replace summative human factors validation?**\nNo. Summative validation requires representative users performing critical tasks with the final design in a realistic simulated-use environment, observed so that use errors and difficulties can be recorded and root-caused. AI-moderated research covers the work around it: use specification interviews, known use problem discovery, IFU and labeling comprehension testing, formative comprehension studies and post-market use surveillance.\n\n**How many participants does FDA expect in a summative usability study?**\nA minimum of 15 participants per distinct user group, with more expected for higher-risk devices. Distinct user groups are counted separately, so a device used by nurses, home caregivers and reprocessing staff needs 15 of each. IEC 62366-1 itself specifies representative users without naming a number.\n\n**What changed in the FDA human factors guidance issued in May 2026?**\nThe final guidance, published 29 May 2026, replaced the December 2022 draft. It confirms three human factors submission categories, adds Decision Point D — which weighs user interface history of use, complexity and existing risk controls before requiring validation data to be submitted — permits robust justification in place of new testing for well-understood interfaces, encourages leveraging data from prior submissions and similar devices, reorders the HFE/UE report sections, and roughly triples the worked examples.\n\n**What is the difference between formative and summative usability evaluation?**\nFormative evaluation happens during development, uses prototypes, has no pass or fail criteria, and exists to find and fix problems. Summative evaluation is the final validation of the finished design against pre-defined acceptance criteria with representative users performing critical tasks. Presenting formative data as validation evidence is a predictable reason for a deficiency letter.\n\n**How do we test whether users understand our instructions for use?**\nRun a comprehension test rather than a satisfaction survey. Define the key points the IFU must convey, set a pass threshold in advance, present the real artefact, test recall with `single_choice` questions that have plausible distractors, capture self-rated clarity on a `scale` question, and probe every wrong answer with an `open_ended` follow-up. Redraft and retest until the threshold is met, and record the before-and-after result.\n\n**Does a device modification always require a new validation study?**\nNot since the 2026 final guidance. If a modification introduces no new or impacted critical tasks it can fall into Category 2 with documented justification, and Decision Point D allows well-understood interfaces with a safe use history and adequate risk controls to rely on justification and leveraged data. That places more weight on the quality of the use-related risk analysis and on real-world evidence about how the device is actually used.\n\n## Related resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and when each belongs in a comprehension test\n- [Formative vs. Summative Research](/docs/formative-vs-summative-research) — the general distinction behind the regulatory one\n- [Usability Testing Guide](/docs/usability-testing-guide) — the unregulated foundation this builds on\n- [HIPAA-Compliant AI User Research](/docs/hipaa-compliant-ai-user-research) — handling health data when your participants are patients\n- [AI-Powered Patient and Provider Research for Healthcare](/docs/ai-research-for-healthcare) — the wider healthcare research picture\n- [Content Testing Guide](/docs/content-testing-guide) — comprehension testing technique applied to labeling and IFUs\n- [Accessibility Research Guide](/docs/accessibility-research-guide) — including users with disabilities in your user profiles","category":"Research Methods","lastModified":"2026-07-31T03:17:13.730967+00:00","metaTitle":"Medical Device Usability Testing: IEC 62366-1 & FDA Human Factors","metaDescription":"How medical device usability testing works under IEC 62366-1 and the FDA's May 2026 human factors guidance: critical tasks, formative studies, summative validation, and where AI research fits.","keywords":["medical device usability testing","IEC 62366-1","FDA human factors engineering","summative validation testing","formative usability study","use-related risk analysis","critical task analysis","HFE UE report","human factors medical device"],"aiSummary":"Medical device usability testing is governed by IEC 62366-1 and FDA human factors guidance. The FDA issued its final guidance, Content of Human Factors Information in Medical Device Marketing Submissions, on 29 May 2026, replacing the December 2022 draft. It defines three human factors submission categories: Category 1 for certain modified devices requiring a conclusion and high-level summary, Category 2 requiring documentation of why there are no critical tasks or no new or impacted critical tasks, and Category 3 requiring a full HFE/UE report with validation data. A new Decision Point D weighs user interface history of use, complexity and existing risk controls to determine whether validation data must be submitted, and the guidance allows justification and leveraged prior data in place of new testing. The engineering chain runs use specification, user profiles, known use problems, task analysis, use-related risk analysis, critical tasks, formative evaluation, summative validation and the HFE/UE report. Formative studies have no pass criteria and roughly five users surface 85 percent of usability problems while ten reach 95 percent. Summative validation requires observed simulated use with representative users performing critical tasks on the final design, with an FDA expectation of at least 15 participants per distinct user group and root cause analysis of every use error. AI-moderated research such as Koji can run use specification interviews, known use problem discovery, IFU and labeling comprehension tests with pre-set pass criteria, formative comprehension studies and post-market use surveillance, but cannot replace observed summative validation or hands-on formative interaction studies. There were 111 Class I recall events and early alerts in 2025 and 1,059 US device recall events in 2024, a four-year high.","aiPrerequisites":["A device concept or design with an identified intended medical indication","Access to representative users from each intended user group","An ISO 14971 risk management file or the intent to create one"],"aiLearningOutcomes":["Distinguish use error, critical task and hazard-related use scenario as the regulations define them","Apply the three FDA human factors submission categories and the 2026 Decision Point D","Build a use-related risk analysis that traces to risk controls and evaluations","Run formative comprehension studies with pre-set pass criteria","Recognise exactly which activities require observed simulated use and cannot be automated","Feed post-market use surveillance back into the usability engineering file"],"aiDifficulty":"advanced","aiEstimatedTime":"14 min read"},{"type":"documentation","id":"1fb5bb2d-6945-4b01-af26-c2a2a01de79d","slug":"customer-discovery-vs-customer-validation","title":"Customer Discovery vs. Customer Validation: Key Differences & When to Do Each","url":"https://www.koji.so/docs/customer-discovery-vs-customer-validation","summary":"Customer discovery and customer validation are the first two stages of Steve Blank customer development model. Discovery is the learning phase that confirms a real problem and who has it; validation is the confirming phase that tests whether your solution earns commitment and scales. This guide covers the model, what each phase is, a side-by-side comparison, the mistake of skipping discovery, how to know you have moved between phases, mapping each to Koji six structured question types, and how Koji accelerates both with asynchronous AI interviews.","content":"# Customer Discovery vs. Customer Validation: Key Differences and When to Do Each\n\n**Customer discovery and customer validation are the first two stages of Steve Blank's customer development model, and they answer fundamentally different questions. Customer discovery asks \"do we understand a real problem that real people have?\" — it is about learning, listening, and finding a problem worth solving. Customer validation asks \"will people actually buy and use our solution to that problem?\" — it is about testing whether you have a repeatable, scalable way to win customers. Discovery comes first and is exploratory; validation comes second and is confirmatory. Confuse the two, or skip discovery to rush into selling, and you risk building something nobody needs.**\n\nRoughly 42% of failed startups die because there was no market need for what they built, according to CB Insights' analysis of startup post-mortems — making \"we never truly validated the problem\" the single most common cause of failure. Discovery and validation exist precisely to prevent that outcome. This guide explains what each phase is, how they differ, how to know when you have passed from one to the next, and how AI-native research lets you run both far faster than the traditional interview-by-interview slog.\n\n<figure class=\"koji-figure\"><img src=\"https://sybpuenocntpoywqhgkf.supabase.co/storage/v1/object/public/blog-images/customer-discovery-vs-customer-validation/inline-stat-card.webp\" alt=\"Stat: 42% of failed startups die from no market need (CB Insights), the risk customer discovery exists to prevent.\" title=\"42% of startups fail from no market need\" width=\"800\" height=\"522\" loading=\"lazy\" /><figcaption><span class=\"caption-text\">42% of failed startups die from no market need.</span><span class=\"caption-source\">CB Insights, analysis of startup post-mortems</span></figcaption></figure>\n\n## The customer development model in brief\n\nSteve Blank — the entrepreneur and academic whose work became the foundation of the Lean Startup movement — frames the early life of a company as a *search* for a business model, not the execution of a known one. His famous rule captures the whole philosophy: *\"There are no facts inside your building, so get outside.\"* Inside the building you have opinions and assumptions; outside, with customers, you find facts.\n\nCustomer development breaks that search into four stages — discovery, validation, creation, and company-building — but the first two are where research lives and where most of the risk is removed. Discovery and validation form a loop: you discover a problem and a hypothesized solution, you try to validate it, and if validation fails you return to discovery with what you learned. (For the broader frame, see [customer development methodology](/docs/customer-development-methodology) and [lean startup methodology](/docs/lean-startup-methodology).)\n\n## What is customer discovery?\n\nCustomer discovery is the *learning* phase. Its goal is to deeply understand a customer's world — their problems, current workarounds, and the jobs they are trying to get done — and to confirm that a problem you care about is real, painful, and widespread enough to be worth solving.\n\nIn discovery you are testing **problem hypotheses**, not pitching a product. The work is mostly open-ended interviews where you talk far more about the customer's life than about your idea. The Mom Test principle applies: ask about specific past behavior, not hypothetical future enthusiasm. Good discovery outputs include:\n\n- A clear, validated statement of the problem and who has it.\n- An understanding of how customers solve it today and what that costs them.\n- Evidence that the pain is strong enough that people are actively looking for relief.\n- A refined view of which customer segment feels the problem most acutely.\n\nIf discovery reveals the problem is mild, rare, or already well-served, that is a *success* — you just saved months of building the wrong thing. (See [customer discovery call guide](/docs/customer-discovery-call-guide) and [customer pain points research](/docs/customer-pain-points-research).)\n\n## What is customer validation?\n\nCustomer validation is the *confirming* phase. Now that you believe you understand a real problem, you test whether your specific solution earns real commitment — and whether you have a repeatable, scalable model for finding and converting customers.\n\nIn validation you are testing **solution and business-model hypotheses**. The signal you are hunting for is not a polite \"that sounds nice\" but costly action: a pre-order, a signed pilot, a deposit, a sustained usage pattern, a willingness to pay. Good validation outputs include:\n\n- Evidence that customers will adopt — and ideally pay for — your solution.\n- A repeatable sales or acquisition motion that works more than once.\n- Confirmation that the economics (price, cost to acquire) can work.\n- Strong enough signal to justify scaling, or clear signal to pivot.\n\nValidation is where you confirm [product-market fit](/docs/product-market-fit-interviews) is within reach. Crucially, validation can *fail* even after great discovery — you understood the problem but your solution or pricing missed — and that failure sends you back to discovery, not out of business.\n\n## Side-by-side comparison\n\n| Dimension | Customer discovery | Customer validation |\n| --- | --- | --- |\n| Core question | Is there a real problem worth solving? | Will people buy and use our solution? |\n| Stage | First | Second |\n| Mindset | Exploratory, learning | Confirmatory, testing |\n| What you test | Problem hypotheses | Solution + business-model hypotheses |\n| Primary method | Open-ended problem interviews | Pilots, pre-sales, usage, willingness-to-pay |\n| Signal of success | Strong, common, urgent pain | Costly commitment (money, time, adoption) |\n| A \"no\" means | Find a different problem or segment | Revisit the solution or return to discovery |\n\n## The biggest mistake: skipping discovery to start validating\n\nThe most expensive error founders make is jumping straight to validation — building a product and trying to sell it — without ever doing discovery. It feels productive because you are shipping, but you are validating a solution to a problem you never confirmed exists. This is exactly the path to the \"no market need\" failure that tops the CB Insights list.\n\nA related mistake is confusing the two while interviewing: pitching your solution during what should be a discovery conversation. The moment you start selling, customers shift into polite mode and you lose the unbiased problem signal discovery depends on. Keep the phases — and the conversations — distinct.\n\n<figure class=\"koji-figure\"><img src=\"https://sybpuenocntpoywqhgkf.supabase.co/storage/v1/object/public/blog-images/customer-discovery-vs-customer-validation/inline-reframe.webp\" alt=\"Customer discovery reframe: do not skip discovery to build and sell; confirm the real problem first, then validate the solution.\" title=\"Confirm the problem before you validate\" width=\"800\" height=\"555\" loading=\"lazy\" /><figcaption><span class=\"caption-text\">Confirm the real problem before validating the solution.</span><span class=\"caption-source\">Steve Blank, customer development model</span></figcaption></figure>\n\n## How to know you've moved from discovery to validation\n\nYou are ready to leave discovery and enter validation when:\n\n- You can state the problem and the most affected segment in one sentence, and customers consistently confirm it unprompted.\n- Multiple interviewees describe the same pain and the same inadequate workarounds.\n- You have heard enough that new interviews mostly repeat what you already know (you have reached saturation).\n- People ask \"can I use that when it's ready?\" before you pitch anything.\n\nIf you are still hearing wildly different problems, or you are the one convincing customers their pain is real, you are still in discovery. Stay there — it is far cheaper than a failed launch.\n\n<figure class=\"koji-figure\"><img src=\"https://sybpuenocntpoywqhgkf.supabase.co/storage/v1/object/public/blog-images/customer-discovery-vs-customer-validation/inline-numbered.webp\" alt=\"Four signs you have moved from customer discovery to validation: one-sentence problem, repeated pain, saturation, unprompted demand.\" title=\"4 signs you are ready to validate\" width=\"800\" height=\"805\" loading=\"lazy\" /><figcaption><span class=\"caption-text\">Four signals that you have reached validation-readiness.</span><span class=\"caption-source\"></span></figcaption></figure>\n\n## Map each phase to structured question types\n\nBoth phases get sharper when you mix open-ended depth with structured, comparable signal. Koji's **six structured question types** — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` — map naturally onto each:\n\n- **Discovery:** `open_ended` for problem stories with AI follow-up probing; `scale` to rate how painful a problem is; `ranking` to order which problems hurt most; `multiple_choice` to standardize current workarounds.\n- **Validation:** `yes_no` for clean commitment reads (\"Would you pay for this?\"); `scale` for purchase intent and willingness to pay; `single_choice` for preferred pricing tiers; `open_ended` to capture objections.\n\nSee the [structured questions guide](/docs/structured-questions-guide) for how each type aggregates across many respondents — turning a pile of interviews into a defensible, quantified case.\n\n## How Koji accelerates both phases\n\nThe classic constraint on customer development is throughput: discovery and validation each need dozens of conversations, and scheduling them one at a time stretches a search that should take weeks into a quarter. Koji removes that bottleneck. Its **AI interviewer runs problem and solution interviews by voice or text, asynchronously and in parallel**, so you can gather discovery signal from dozens of customers in days — and automatically probes for the specific past behavior that separates real pain from politeness.\n\nWhen you move to validation, the same engine tests solution reactions, purchase intent, and pricing, while **analyzing every transcript automatically** — clustering the problems customers actually raised, flagging where they diverge, and aggregating your scale and ranking answers into charts. Instead of a founder's memory of \"most people seemed to like it,\" you get an evidence-grounded readout of whether the problem is real (discovery) and whether the solution earns commitment (validation). That is the difference between guessing and knowing before you scale.\n\n## Discovery and validation are a loop, not a line\n\nIn practice you rarely move through these phases once. A failed validation — customers agreed the problem was real but would not pay for your solution — sends you back into discovery to learn why, sharpen the segment, or reframe the offer. Even a successful first pass usually loops: you validate with early adopters, then re-enter discovery to understand the mainstream customer whose needs differ. Treat the two as a cycle you spin quickly and cheaply, not a gate you pass through once. The faster and cheaper each loop, the more times you can afford to learn before the money runs out — which is the entire point of doing customer development *before* scaling.\n\n<figure class=\"koji-figure\"><img src=\"https://sybpuenocntpoywqhgkf.supabase.co/storage/v1/object/public/blog-images/customer-discovery-vs-customer-validation/inline-process-flow.webp\" alt=\"Customer development loop: discover a problem, hypothesize a solution, validate it, then loop back to discovery if validation fails.\" title=\"The customer discovery to validation loop\" width=\"800\" height=\"282\" loading=\"lazy\" /><figcaption><span class=\"caption-text\">Discovery and validation form a fast, repeatable loop.</span><span class=\"caption-source\">Steve Blank, customer development model</span></figcaption></figure>\n\n## A quick worked example\n\nImagine a team building scheduling software for clinics. In **discovery**, they interview 20 clinic managers and consistently hear that double-booking and no-shows are the real pain — not calendar syncing, which they assumed. That is a problem hypothesis confirmed and an assumption killed. In **validation**, they offer a paid pilot of a no-show-reduction feature; three of eight clinics sign and pay, and the others object to the price. The problem was validated in discovery; the solution and pricing were only partly validated. The right move is not to scale — it is a short loop back to discovery on pricing sensitivity before committing the roadmap. That disciplined sequencing is what separates teams that find fit from teams that guess.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and how each aggregates\n- [Customer Development Methodology](/docs/customer-development-methodology) — the full four-stage model\n- [Customer Discovery Call Guide](/docs/customer-discovery-call-guide) — how to run the discovery interview\n- [Startup Idea Validation Guide](/docs/startup-idea-validation-guide) — validating before you build\n- [Product-Market Fit Interviews](/docs/product-market-fit-interviews) — confirming fit in the validation phase\n- [Lean Startup Methodology](/docs/lean-startup-methodology) — the broader build-measure-learn frame","category":"Research Methods","lastModified":"2026-07-30T07:14:03.339017+00:00","metaTitle":"Customer Discovery vs. Customer Validation: Key Differences (2026)","metaDescription":"Customer discovery confirms a real problem; customer validation confirms people will buy your solution. Learn the differences, the order, and when to do each.","keywords":["customer discovery vs customer validation","customer validation","customer discovery","customer development model","steve blank","problem validation vs solution validation","startup validation phases"],"aiSummary":"Customer discovery and customer validation are the first two stages of Steve Blank customer development model. Discovery is the learning phase that confirms a real problem and who has it; validation is the confirming phase that tests whether your solution earns commitment and scales. This guide covers the model, what each phase is, a side-by-side comparison, the mistake of skipping discovery, how to know you have moved between phases, mapping each to Koji six structured question types, and how Koji accelerates both with asynchronous AI interviews.","aiPrerequisites":["A startup or product idea","Access to potential customers to interview"],"aiLearningOutcomes":["Distinguish customer discovery from customer validation","Run the right interviews for each phase","Recognize when you have passed discovery","Avoid the no-market-need failure mode","Accelerate both phases with AI interviews"],"aiDifficulty":"beginner","aiEstimatedTime":"11 min read"},{"type":"documentation","id":"5beb3eec-7a73-4127-a1ce-be713a45bf4a","slug":"insight-repository-methodology","title":"Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)","url":"https://www.koji.so/docs/insight-repository-methodology","summary":"A methodology-layer guide for research insight repositories that goes beyond tooling. Covers the four pillars (taxonomy, atomic insight structure, governance/decay, insight-to-action workflow), a 2-week setup plan, common failure modes, and how AI auto-tagging eliminates the librarian bottleneck. Cites NN/G's State of ResearchOps (39% have any repository, 8% have dedicated manager) and UXPA 2025 (60% faster time-to-insight with AI). Positions Koji's auto-extraction, auto-tagging, and insights chat as the operational layer that makes the methodology sustainable.","content":"# Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)\n\n**Bottom line up front:** A research insight repository is only useful if it's *queryable*, *fresh*, and *connected to decisions*. Most teams stop at storage — a Notion page or Airtable base full of past reports — and wonder why no one uses it. The methodology that separates a thriving repository from a digital graveyard has four pillars: a **stable taxonomy**, **atomic insight structure**, **governance with decay rules**, and an explicit **insight-to-action workflow**. Only **39% of organizations have a research repository at all**, and only **8% have a dedicated role to manage it** ([NN/G, State of ResearchOps](https://www.nngroup.com/articles/researchops-state-untapped/)) — which is why most repositories rot within 18 months. AI-native platforms like Koji eliminate the librarian bottleneck with automatic tagging, natural-language insight chat, and built-in decay tracking.\n\nThis guide is the methodology layer most repository how-to articles skip.\n\n---\n\n## Why most repositories fail\n\nThe repository tooling debate (Dovetail vs Notion vs Marvin vs Airtable) hides the real problem: **methodology, not tooling, kills repositories**. The same Notion base that works at Company A becomes a graveyard at Company B because:\n\n- **No taxonomy.** Insights get tagged \"user experience\" (meaningless) or \"from the Q3 study\" (unsearchable for anyone outside that study).\n- **No atomic structure.** Whole reports are filed, but no one can find a specific quote or finding without re-reading the report.\n- **No decay rules.** A 2022 insight about competitor pricing sits next to a 2026 one, with no signal which is current.\n- **No activation workflow.** Insights are stored but never linked to PRDs, OKRs, or product decisions. The repository becomes write-only.\n- **Librarian bottleneck.** Tagging falls to a single ResearchOps person; when they leave or get overloaded, the repo decays.\n\nThe NN/G State of ResearchOps survey confirms the scale: **only 39% of organizations have any insight repository**, **only 35% maintain a recording library**, and **only 24% have a participant-management system** ([NN/G](https://www.nngroup.com/articles/researchops-state-untapped/)). Among the 39% that *do* have a repository, NN/G's qualitative findings suggest fewer than half are actively used after the first year.\n\n> \"For an atomic system to function, it requires infrastructure for storing and retrieving these 'small nuggets' of insight. ResearchOps' efficient repositories ensure the necessary metadata standards are met so that each atomic insight is discoverable, accessible, and usable.\" — NN/G, Research Repositories for Tracking UX Research ([source](https://www.nngroup.com/articles/research-repositories/))\n\nThe fix is methodology. The four pillars below are what working repositories have in common, regardless of which tool they're built in.\n\n---\n\n## Pillar 1: The taxonomy (stable, hierarchical, evergreen)\n\nA taxonomy is the controlled vocabulary that makes the repository searchable in 18 months — not just this week. Without one, every researcher invents tags (\"onboarding pain,\" \"trial pain,\" \"first-run issue\") that mean the same thing but break filtering forever.\n\nA working taxonomy has 4 axes:\n\n1. **Theme** — the customer-level concept (e.g., \"pricing transparency,\" \"onboarding friction,\" \"data export\"). Aim for 30–60 themes total. Less is too coarse; more is unmanageable.\n2. **Segment** — who said it (plan tier, role, industry, tenure).\n3. **Source** — study type (churn interview, NPS follow-up, usability test, sales call).\n4. **Outcome area** — which business outcome this insight informs (activation, retention, expansion).\n\nThemes should be **stable** (rarely added or renamed), **mutually exclusive** where possible, and **collectively exhaustive** at the abstraction level you operate. The most common failure: themes are too granular (\"button placement on signup screen\") instead of conceptual (\"first-run cognitive load\"). Granular themes don't aggregate across studies.\n\nMaintain a written **taxonomy guide** with definitions and examples for each theme. Without it, you'll get tag drift within a quarter.\n\n---\n\n## Pillar 2: The atomic insight (the unit of the repository)\n\nThe atomic research nugget is the irreducible unit of storage. Whole reports are too big — no one searches a 14-page PDF for a single quote. The atomic structure (popularized by Daniel Pidcock and the [atomic research nuggets guide](/docs/atomic-research-nuggets-guide)) has four parts:\n\n```\n[Observation]     — What was said or observed\n[Evidence]        — The quote, timestamp, or artifact\n[Insight]         — The interpretation (what it means)\n[Tags]            — Theme + segment + source + outcome\n```\n\nExample:\n- **Observation:** Enterprise customers can't self-serve API token rotation\n- **Evidence:** *\"We had to open a ticket every quarter just to rotate keys — for security audits we need this in our own hands.\"* — Director of Security, 800-person fintech, Q1 2026 churn interview\n- **Insight:** Lack of self-serve key rotation is a compliance blocker for regulated-industry buyers; correlates with security-audit timing as a churn trigger\n- **Tags:** `theme: security-self-serve` `segment: enterprise-regulated` `source: churn-interview` `outcome: retention`\n\nEach insight is **immutable** — you don't edit it later, you add new ones. Immutability matters because insights can be *cited* in PRDs, OKRs, and dashboards; if they mutate, those citations break.\n\nA single 45-minute interview should produce **5–12 atomic insights**, not a single dumped transcript.\n\n---\n\n## Pillar 3: Governance and decay\n\nInsights age. A customer pain point about onboarding in 2024 may be solved by 2026 (or worse, still real but mis-attributed). Without governance, the repository becomes untrustworthy — a problem worse than not having one.\n\nThree governance rules:\n\n1. **Date-stamp every insight.** Filter by recency in every search.\n2. **Decay flags.** Any insight older than 12 months gets flagged \"needs revalidation.\" Either re-confirm with a new interview or archive.\n3. **Citation tracking.** When an insight is cited in a PRD, OKR, or decision, log the citation. High-cited insights deserve more validation; uncited insights deserve archival.\n\nA small ResearchOps team can run governance, but only **8% of organizations have a dedicated repository manager** ([NN/G](https://www.nngroup.com/articles/researchops-state-untapped/)) — which is why automation matters (see Koji section below).\n\n---\n\n## Pillar 4: The insight-to-action workflow\n\nThe repository's job is to influence decisions. If it doesn't, it dies. The workflow has three required hooks:\n\n- **PRDs reference repository insight IDs.** Every product requirement document cites the underlying atomic insights. No insight, no PRD section.\n- **Quarterly review of activation.** Which insights drove shipped work? Which sat unused? Which were contradicted by later data?\n- **Open-question backlog.** Insights that *raise* a question (not answer one) feed a research backlog. The repository becomes the source of next quarter's research plan.\n\nRead the dedicated [activating research insights](/docs/activating-research-insights) guide for the activation workflow in detail.\n\n---\n\n## How AI auto-tagging eliminates the librarian bottleneck\n\nManual tagging is the single biggest reason repositories fail. A researcher spending 2 hours after every interview tagging insights to a taxonomy can't sustain it past 20 studies. Modern AI-native platforms — Koji included — solve this by tagging atomically and automatically during analysis:\n\n- **Auto-extraction of atomic insights.** Each interview is parsed into 5–12 atomic nuggets — observation + evidence + insight — without manual coding.\n- **Auto-tagging against your taxonomy.** Themes, segments, and outcome areas are applied from a controlled vocabulary you maintain once.\n- **Quality scoring (1–5 scale).** Low-quality interviews (refused, off-topic) get flagged so they don't pollute the repo.\n- **Insights chat across the entire repository.** Ask natural-language questions: *\"Show me every insight where Enterprise customers mentioned security audits in the last 12 months.\"* The chat is the search interface a repository always needed but never had.\n- **Decay tracking built in.** Date-stamping is automatic. Filter by recency in every chat or report.\n- **Six structured question types** ([structured questions guide](/docs/structured-questions-guide)) — open_ended, scale, single_choice, multiple_choice, ranking, yes_no — let you store *both* the qualitative quote and the quantitative segment data on the same atomic insight, which is essential for cross-segment analysis.\n\nThe combined effect: a 3-person product team can maintain a repository that historically required a 2-person ResearchOps function — because the tagging, retrieval, and decay tracking are automated.\n\nTeams using AI-assisted insight platforms report **60% faster time-to-insight** ([UXPA, 2025](https://uxpa.org/ux-research-in-2025-from-insights-to-action/)) — most of that delta is in repository activation, not interview moderation.\n\n---\n\n## A 2-week setup plan\n\n**Week 1 — Foundation:**\n- Day 1: Draft taxonomy (30–60 themes, 4 axes). Use existing studies to validate coverage.\n- Day 2: Define the atomic insight template. Pick a storage tool.\n- Day 3: Back-tag the last 10 studies into atomic insights. This stress-tests the taxonomy.\n- Day 4: Write the taxonomy guide.\n- Day 5: Decide governance rules — date-stamp format, decay flag, citation log.\n\n**Week 2 — Activation:**\n- Day 6: Wire PRD template to require insight IDs.\n- Day 7: Schedule the quarterly review ritual.\n- Day 8: Set up auto-tagging (in Koji or your platform of choice).\n- Day 9: Run one new study end-to-end through the repository.\n- Day 10: Demo the chat-style query to the broader product org. This is the moment the repository becomes *used*, not just *built*.\n\nBy day 14, the repository is operational. By day 90, if governance is held, you'll have 100+ atomic insights, weekly citations in PRDs, and a research backlog driven by repository gaps.\n\n---\n\n## Common failure modes\n\n1. **Tool first, methodology second.** Buying Dovetail or building a Notion base before defining taxonomy and atomic structure guarantees a future migration.\n2. **Tags invented per-study.** Without a controlled vocabulary, the repo is unsearchable within 6 months.\n3. **Storing whole reports instead of atomic insights.** The unit of the repository is the insight, not the study.\n4. **No decay.** A 2022 insight presented as current undermines trust in the entire repo.\n5. **The repository is a write-only system.** If insights are never cited in PRDs, the activation workflow is broken.\n6. **The librarian bottleneck.** A single ResearchOps person can't manually tag at the rate a working product org generates insights. Automate tagging.\n\n---\n\n## Related Resources\n\n- [Research Repository Guide](/docs/research-repository-guide)\n- [Atomic Research Nuggets Guide](/docs/atomic-research-nuggets-guide)\n- [Activating Research Insights](/docs/activating-research-insights)\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide)\n- [Structured Questions Guide](/docs/structured-questions-guide)\n- [How to Prioritize Customer Feedback](/docs/how-to-prioritize-customer-feedback)\n- [Opportunity Solution Tree](/docs/opportunity-solution-tree)\n- [How to Conduct User Interviews](/docs/how-to-conduct-user-interviews)","category":"analysis","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Insight Repository Methodology: Taxonomy, Atomic Insights & Activation — Koji","metaDescription":"Build a research insight repository that gets used, not abandoned. Four-pillar methodology covering taxonomy design, atomic insight structure, governance with decay rules, and the insight-to-action workflow — plus AI auto-tagging that removes the librarian bottleneck.","keywords":["insight repository","research repository","atomic research","research operations","ResearchOps","insight taxonomy","knowledge management","insight activation","UX research repository","customer insights database"],"aiSummary":"A methodology-layer guide for research insight repositories that goes beyond tooling. Covers the four pillars (taxonomy, atomic insight structure, governance/decay, insight-to-action workflow), a 2-week setup plan, common failure modes, and how AI auto-tagging eliminates the librarian bottleneck. Cites NN/G's State of ResearchOps (39% have any repository, 8% have dedicated manager) and UXPA 2025 (60% faster time-to-insight with AI). Positions Koji's auto-extraction, auto-tagging, and insights chat as the operational layer that makes the methodology sustainable.","aiPrerequisites":["Awareness of UX research or product research practices","Familiarity with at least one repository tool (Dovetail, Notion, Airtable, Marvin)","Basic understanding of qualitative analysis"],"aiLearningOutcomes":["Identify why most insight repositories rot within 18 months","Design a stable, hierarchical research taxonomy with 30–60 themes","Structure findings as atomic insights (observation, evidence, insight, tags)","Apply governance rules including decay flags and citation tracking","Build an insight-to-action workflow that ties repository to PRDs and OKRs","Use AI auto-tagging to scale a repository without a dedicated librarian"],"aiDifficulty":"advanced","aiEstimatedTime":"15 min read"},{"type":"documentation","id":"d503cd31-4865-40ab-a83a-1c4600f2754f","slug":"presenting-research-findings","title":"Presenting Research Findings to Stakeholders","url":"https://www.koji.so/docs/presenting-research-findings","summary":"Effective presentation of qualitative research findings increases the likelihood of recommendations being implemented by 2.6x. Key techniques include tailoring format to audience, leading with participant stories, using the data sandwich structure (quantitative context, qualitative quote, implication), and always pairing findings with prioritized recommendations.","content":"The most rigorous research in the world is worthless if it sits in a document nobody reads. Presenting findings is not an afterthought — it is the mechanism through which research drives decisions. According to a study by Yoo and Kim (2019) published in the International Journal of Design, research teams that invested in structured presentation of findings were 2.6x more likely to see their recommendations implemented compared to teams that shared raw reports without a narrative structure.\n\nYour audience does not care about your methodology (much). They care about what you found and what they should do about it.\n\n## Know Your Audience\n\nBefore you design your presentation, understand who is receiving it and what they need:\n\n| Audience | What They Want | Format Preference |\n|----------|---------------|-------------------|\n| Executives / C-suite | Big picture: What did we learn? What should we do? What is the business impact? | Executive summary (1-2 pages), key metrics |\n| Product Managers | Specific pain points, user needs, feature implications, prioritization | Detailed findings with quotes, prioritization framework |\n| Designers | User mental models, workflow patterns, emotional moments, exact language | Journey maps, annotated quotes, video clips |\n| Engineers | Specific use cases, edge cases, technical constraints from users | Structured requirements, user stories derived from findings |\n| Sales / Marketing | Customer language, objections, value perception, competitive context | Quotable sound bites, persona summaries, competitive mentions |\n\nTailor your deliverable to your audience. A single \"research report\" rarely serves all of these stakeholders equally well. Consider creating a core report and audience-specific summaries.\n\n## The Three Report Formats\n\n### 1. Executive Summary\n\n**Length:** 1-2 pages\n**Purpose:** Communicate the headline findings and recommended actions\n**Structure:**\n\n1. **Study overview** (1 paragraph): What was the research question, who did you talk to, and why?\n2. **Top findings** (3-5 bullet points): The most important things you learned\n3. **Recommended actions** (3-5 bullet points): What should we do about it?\n4. **What is at stake** (1 paragraph): What happens if we do not act?\n\n**Example bullet:**\n> \"8 of 12 participants could not find the export feature without help. This means approximately two-thirds of users who need to share reports externally are either using workarounds or abandoning the task entirely.\"\n\n### 2. Detailed Findings Report\n\n**Length:** 5-15 pages depending on study scope\n**Purpose:** Provide the full picture for product teams and design\n**Structure:**\n\n1. **Study background**: Research questions, methodology, participant summary\n2. **Participant overview**: Demographics table, screening criteria, anonymized profiles\n3. **Theme-by-theme findings**: Each theme as a section with supporting evidence\n4. **Cross-cutting patterns**: Observations that span multiple themes\n5. **Prioritized recommendations**: Actionable next steps ranked by impact and effort\n6. **Appendix**: Full interview guide, participant matrix, methodology notes\n\n### 3. Quick-Reference Insight Cards\n\n**Length:** One card per insight (postcard-sized)\n**Purpose:** Shareable, digestible nuggets for team walls, Slack, or documentation\n**Structure per card:**\n\n- **Insight headline** (1 sentence)\n- **Supporting data** (frequency, 1-2 quotes)\n- **Recommended action** (1 sentence)\n- **Evidence strength** (strong / moderate / emerging)\n\nThese are particularly effective for keeping research top-of-mind between formal presentations. Pin them in your team's communication channel or print them for the office wall.\n\n## Storytelling With Data\n\n### Lead With the Story, Not the Method\n\nA common mistake is spending the first 10 minutes of a presentation explaining your methodology. Executives will tune out before you get to the findings.\n\nInstead, start with a participant story:\n\n*\"Let me tell you about Sarah. She's a product manager at a mid-size SaaS company. She signed up for our tool on a Monday, spent 45 minutes trying to set up her first project, and by Wednesday she had downgraded to a competitor's free tier. When I asked her why, she said: 'I could see it was powerful, but I couldn't figure out how to make it do the basic thing I needed.'\"*\n\nNow your audience is hooked. They want to know: How many Sarahs are there? And what can we do about it?\n\n### Use Participant Quotes Effectively\n\nQuotes are the most powerful tool in your presentation arsenal. They bring abstract findings to life with human voice and specificity.\n\n**Rules for effective quote usage:**\n\n1. **Select quotes that are vivid and specific**: \"I gave up after the third time the page loaded wrong\" is better than \"It was frustrating.\"\n2. **Keep quotes short**: 1-2 sentences maximum in a presentation. Longer quotes lose the audience.\n3. **Always attribute to an anonymized participant**: \"P7, Product Manager, Enterprise\" gives the quote context and credibility.\n4. **Use quotes to illustrate, not to prove**: A single quote supports a finding; it does not establish one. Always frame quotes within the broader pattern.\n\n### The Data Sandwich\n\nWhen presenting a finding, use this structure:\n\n1. **Quantitative context** (top bread): \"9 of 12 participants experienced this.\"\n2. **Qualitative depth** (filling): A specific quote or story that makes the number human.\n3. **Implication** (bottom bread): \"This suggests that our onboarding is losing the majority of new users before they experience core value.\"\n\nThis format satisfies both the data-driven and narrative-driven people in your audience simultaneously.\n\n## Visualizing Qualitative Data\n\nQualitative data can be visualized — it just requires different approaches than charts and graphs.\n\n### Theme Maps\n\nShow how themes relate to each other. A simple diagram with themes as nodes and connections as lines helps stakeholders see the system, not just individual findings.\n\n### Quote Walls\n\nA curated collection of participant quotes organized by theme. This works especially well in physical spaces or as a shared digital board.\n\n### Journey Maps\n\nIf your research followed a user journey, map the findings onto the stages of that journey. Annotate each stage with participant quotes, emotional states, and pain points.\n\n### Frequency Tables\n\nWhile qualitative research is not about counting, showing how many participants mentioned a theme provides useful signal:\n\n| Theme | Participants (of 15) | Intensity |\n|-------|---------------------|-----------|\n| Onboarding friction | 12 | High |\n| Pricing confusion | 9 | Medium |\n| Feature discovery gap | 8 | High |\n| Positive support experience | 6 | Medium |\n\nA study by Braun and Clarke (2006) in Qualitative Research in Psychology — the most cited paper on thematic analysis with over 100,000 citations — emphasizes that frequency should supplement, not replace, the researcher's judgment about theme importance.\n\n## Generating Reports With AI\n\nWhen you are working with a large volume of interviews, generating the initial report structure manually is time-consuming. AI-powered platforms like Koji can generate research reports that include theme identification, supporting quotes, and preliminary recommendations based on your interview data.\n\nThese auto-generated reports give you a strong starting draft: themes are surfaced with evidence, participant quotes are linked to findings, and patterns across interviews are highlighted. Your job is to validate the AI's interpretation, add your contextual knowledge, and tailor the narrative for your specific audience.\n\nFor details on how to generate and customize these reports, see [generating research reports](/docs/generating-research-reports) and [publishing and sharing reports](/docs/publishing-sharing-reports).\n\n## Presenting Live: Tips for the Room\n\n### Structure Your Presentation\n\n1. **Hook** (2 minutes): Start with a compelling participant story\n2. **Context** (3 minutes): Brief study overview — who, why, how many\n3. **Findings** (15-20 minutes): Theme by theme, using the data sandwich\n4. **Recommendations** (5 minutes): What to do, prioritized\n5. **Discussion** (10+ minutes): Open the floor for questions and debate\n\n### Handle Pushback Gracefully\n\nStakeholders may challenge your findings. This is healthy. Prepare for common objections:\n\n- *\"That's just 12 people\"*: \"You are right that this is qualitative data, not a statistically representative survey. What qualitative research tells us is *why* people behave a certain way. The consistency across 12 diverse participants gives us confidence in the direction of the finding.\"\n\n- *\"I talked to a customer who said the opposite\"*: \"That is valuable context. In our study, we found [X number] participants who felt this way. There may be a segment difference worth exploring. Can you share which customer that was so we can compare profiles?\"\n\n- *\"We already knew this\"*: \"If the organization already has this insight, that is great validation. The question is whether we have acted on it. Here are specific recommendations for what to do next.\"\n\n## Common Mistakes to Avoid\n\n1. **Burying the lead**: Do not save your most important finding for slide 47. Lead with impact.\n\n2. **Presenting every finding equally**: Not all themes are equally important. Focus 60% of your time on the top 2-3 findings and briefly acknowledge the rest.\n\n3. **Forgetting the \"so what\"**: Every finding needs a recommendation. Data without action is trivia.\n\n4. **Making it about you**: The presentation is about the participants and the team's next moves. Minimize references to your process and maximize focus on what was learned and what to do about it.\n\n5. **Not following up**: Schedule a follow-up meeting 2-4 weeks after the presentation to check whether findings have been incorporated into planning. Research that is never revisited is research that was never used.\n\n## Key Takeaways\n\n- Tailor your format to your audience: executives want headlines, product teams want detail, designers want mental models\n- Start with a participant story, not your methodology\n- Use the data sandwich: quantitative context, qualitative depth, implication\n- Short, vivid, attributed quotes are your most powerful presentation tool\n- AI-generated reports provide strong first drafts that need human refinement and audience tailoring\n- Always pair findings with specific, prioritized recommendations\n\nFor the analysis process that feeds into your presentation, see [turning interviews into insights](/docs/turning-interviews-into-insights). For details on auto-generating report drafts, explore [generating research reports](/docs/generating-research-reports).\n\n## Frequently Asked Questions\n\n**How long should a research presentation be?**\n\nFor a live presentation, 30-45 minutes including discussion time is ideal. For a written report, the executive summary should be 1-2 pages, and the detailed report should be 5-15 pages. Stakeholders rarely read reports longer than 15 pages end-to-end.\n\n**Should I include my interview guide in the report?**\n\nInclude it as an appendix for transparency and reproducibility, but do not expect anyone to read it. The people who want to see the methodology will appreciate having it available. Everyone else will skip to the findings.\n\n**How do I handle findings that contradict what stakeholders want to hear?**\n\nPresent them directly but empathetically. Lead with the strongest evidence, use participant quotes to make the finding human, and frame your recommendation constructively: \"The data suggests [finding], which creates an opportunity to [positive action].\"\n\n**When should I present research findings versus share a written report?**\n\nPresent live when findings are high-stakes, require discussion, or involve organizational change. Share written reports for lower-stakes updates, ongoing tracking studies, or when stakeholders are geographically distributed and scheduling is impractical.\n\n**How do I measure whether my research presentation was effective?**\n\nTrack whether your recommendations appear in sprint planning, roadmap discussions, or design briefs within 4 weeks of the presentation. If they do, the presentation worked. If they do not, follow up to understand what barrier exists between insight and action.\n\n## Further reading on the blog\n\n- [Agile User Research: How to Run Continuous Research in Sprint Cycles (2026)](/blog/agile-user-research-2026) — Most teams know they should do user research every sprint. Almost none actually do. Here's the practical playbook for integrating continuous\n- [AI Agents for User Research in 2026: How Autonomous Research Is Reshaping Customer Insight](/blog/ai-agents-user-research-2026) — AI agents are taking over user research in 2026 — moderating interviews, synthesizing themes, and producing insight reports in hours. The fu\n- [Best AI Market Research Tools in 2026: The Complete Buyer's Guide](/blog/ai-market-research-tools-2026) — AI has fundamentally changed market research. This guide compares the leading AI market research platforms—from AI-native interview tools li\n\n<!-- further-reading:blog -->\n","category":"Analysis & Synthesis","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Presenting Research Findings","metaDescription":"Present qualitative research effectively to stakeholders using storytelling, participant quotes, and structured report formats that drive action.","keywords":["research presentation","stakeholder reporting","qualitative findings","research storytelling","executive summary","research reports","data visualization"],"aiSummary":"Effective presentation of qualitative research findings increases the likelihood of recommendations being implemented by 2.6x. Key techniques include tailoring format to audience, leading with participant stories, using the data sandwich structure (quantitative context, qualitative quote, implication), and always pairing findings with prioritized recommendations.","aiPrerequisites":["turning-interviews-into-insights"],"aiLearningOutcomes":["Structure research reports for executives, product teams, and designers","Use storytelling techniques including the data sandwich and participant quotes","Create executive summaries, detailed findings reports, and insight cards","Handle stakeholder pushback on qualitative research findings"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min read"},{"type":"documentation","id":"1b16f433-4237-4eea-9496-b53a09b88482","slug":"linear-research-integration","title":"Send Koji Insights to Linear: Auto-File Engineering Tickets from Customer Interviews","url":"https://www.koji.so/docs/linear-research-integration","summary":"Connect Koji to Linear so interviews that surface real customer pain points auto-create tagged Linear issues — with verbatim quotes, theme tags, quality scores, and direct links to the live interview attached. The guide covers the critical design decision (which interviews should become tickets) using filters on quality score, theme, structured-answer thresholds, and sentiment; the full Zapier setup; severity-based team routing; deduplication via 'Find Issue' lookups; and the direct Linear GraphQL approach for teams that outgrow Zapier. The result replaces the Slack-thread-to-screenshot-to-ticket workflow that loses customer context with a closed-loop research-to-execution pipeline.","content":"**TL;DR:** Wire Koji to Linear so every customer interview that surfaces a real pain point auto-creates a tagged Linear issue with the verbatim quote, theme, study link, and quality score attached. Setup takes about 20 minutes via Zapier (or 2–3 hours via Linear's GraphQL API directly). Most engineering teams report this single integration replaces the \"Slack thread → screenshot → ticket title from memory\" workflow that everyone hates, and dramatically increases the percentage of customer-driven tickets that actually get prioritized.\n\n## Why teams sync Koji to Linear\n\nLinear is where engineering work actually happens. Tickets there get triaged, sprinted, and shipped — which means anything that wants to influence the roadmap needs to live there. Research findings that never reach Linear effectively never reach the roadmap.\n\nThe traditional flow looks like this: PM watches an interview recording → takes notes in Notion → writes a Linear ticket from memory → engineering reads the ticket without context → builds something slightly different than what the customer described. By the time the feature ships, the link back to \"this is what Dana said in the August interview\" is gone, and so is any way to evaluate whether the feature actually solved the problem.\n\nA Koji → Linear sync replaces that whole chain with a webhook. Every completed Koji interview that surfaces a pain point can auto-create a Linear issue containing:\n\n- Verbatim customer quote (not a paraphrase)\n- Theme tag from Koji's AI analysis\n- Quality score (so engineering can deprioritize low-confidence signal)\n- Direct link back to the live Koji interview, so anyone can play the voice clip\n- Participant segment and metadata\n- Study context and AI summary\n\nThis is the modern alternative to manual ticket creation. Platforms like Koji automate moderation, transcription, theme extraction, and quality scoring; Linear becomes the surface where engineering executes against that evidence. The two pieces fit together cleanly through Linear's webhook-friendly architecture.\n\n## Integration paths\n\nThere are two production-ready ways to send Koji interviews to Linear:\n\n1. **Zapier (recommended)** — point-and-click, ~20 minutes, supports filters, no code\n2. **Linear GraphQL API (advanced)** — write a webhook receiver that calls Linear's API directly\n\nMost teams should start with Zapier. The GraphQL path is worth it once you need rich formatting (markdown tables, attachments) or you're sending more than a few thousand tickets per month.\n\n## Prerequisites\n\n- A Koji study with at least one completed interview\n- A Linear workspace where you have permission to create issues in at least one team\n- A Zapier account (the Starter plan is needed for filters)\n- About 20 minutes\n\n## Step 1: Decide which interviews become tickets\n\nThis is the most important design decision and the one most teams skip. Do NOT auto-create a Linear ticket for every interview — your engineering backlog will explode and the signal-to-noise will tank within a week.\n\nThe patterns that work in production:\n\n- **Filter by quality score.** Only create tickets from interviews with quality score 4 or higher (out of 5). This drops gibberish and accidental submissions automatically.\n- **Filter by theme.** Only create tickets when Koji's AI detects a theme that maps to engineering work — e.g. \"bug-report\", \"performance-issue\", \"missing-feature\". Themes like \"general-praise\" or \"competitor-mention\" should never become tickets.\n- **Filter by structured-answer threshold.** If your study has a scale question like \"How easy was this?\", auto-create a ticket only when the participant rates it 2 or below — those are the real problems worth filing.\n- **Filter by sentiment.** Negative-sentiment interviews are usually the higher-leverage ones to file.\n- **Combine filters.** \"Quality score 4+ AND theme contains 'bug' OR scale rating <= 2\" is the kind of multi-condition filter Zapier handles natively.\n\nWrite down which filter you want before you start the integration. You can always change it later, but having a default rule prevents the first-day backlog explosion.\n\n## Step 2: Set up the Koji webhook\n\nIn Koji:\n\n1. Open the study you want to monitor\n2. Go to Settings → Webhooks (or the Webhooks tab in the study editor)\n3. Add a new webhook destination — leave the URL field empty for now\n\nYou'll paste the Zapier URL here in the next step.\n\n## Step 3: Create the Zap\n\nIn Zapier:\n\n1. New Zap → Trigger: **Webhooks by Zapier** → **Catch Hook**\n2. Zapier gives you a URL. Copy it.\n3. Paste it into the Koji webhook destination, save the Koji study, then fire a \"Test webhook\" from Koji\n4. Back in Zapier, click \"Test trigger\" — you should see a real Koji payload with study ID, interview ID, quality score, themes, structured answers, transcript text, and the public URL\n\n## Step 4: Add the filter step\n\nClick \"+ Add step\" → **Filter by Zapier**. Add your conditions. For the most common \"bug-or-pain\" filter:\n\n- `interview.quality_score` is greater than or equal to `4`\n- AND `interview.themes` text contains `bug` OR `pain` OR `frustration` OR `confusion`\n\nYou can chain ANDs and ORs natively. Zapier's filter UI makes this easier than writing it in code.\n\n## Step 5: Create the Linear issue\n\nAdd a new action: **Linear → Create Issue**. Connect your Linear account. Configure:\n\n- **Team:** the engineering team that owns this study's surface area (e.g. \"Onboarding\", \"Billing\", \"Search\")\n- **Title:** map to `interview.headline` if Koji generates one, otherwise build it from the theme + participant ID — e.g. `[Research] {theme}: {participant_name}`\n- **Description:** the most important field. Build it from these blocks (markdown supported in Linear):\n\n```\n**Quote:** \"{interview.headline_quote}\"\n\n**Theme:** {interview.themes}\n**Quality score:** {interview.quality_score}/5\n**Segment:** {interview.metadata.segment}\n**Study:** {study.title}\n\n**AI summary:** {interview.ai_summary}\n\n**Listen to the interview:** {interview.public_url}\n\n---\nAuto-filed from Koji. Edit or close if not actionable.\n```\n\n- **Labels:** map to `interview.themes` — Linear will create the labels on the fly if they don't exist (or you can pre-create the labels you care about)\n- **Priority:** map to a formula — if quality score >= 5 then \"High\", if 4 then \"Medium\", else \"Low\"\n- **Assignee:** leave empty so triage handles it (most teams) or set to a designated triage owner\n\nTest the action. The Linear issue should appear in your team's inbox within seconds.\n\n## Step 6: Turn the Zap on and tune\n\nToggle the Zap to \"On.\" Watch it for a week and tune the filter based on what you see. The most common adjustment: tightening the theme list because your initial filter was too permissive.\n\n## Patterns that pay off\n\n### Auto-link to existing issues\n\nAdd a Zapier \"Find Issue\" step before \"Create Issue.\" Search by theme tag. If an issue already exists for that theme, add a comment to the existing one instead of creating a duplicate. This is huge for keeping the backlog clean when the same pain shows up across multiple interviews.\n\n### Severity-based team routing\n\nUse a Zapier Paths step. Route interviews with the `bug` theme to the platform team, `pricing` to the growth team, and `onboarding` to the activation team. Each path can have its own filter and its own Linear team mapping.\n\n### Auto-close confirmation tickets\n\nWhen Koji's AI detects a `feature-praise` theme that mentions a recently shipped feature, you can auto-comment on the Linear issue that originally tracked the work and close it. This closes the customer-research loop without any human intervention.\n\n### Weekly digest issues\n\nInstead of one ticket per interview, run a Zapier Digest action: aggregate all interviews from a study for the past 7 days, then create one Linear ticket with all the quotes grouped by theme. This is the right pattern for steady-state insight ingestion vs. urgent bug surfacing.\n\n### Block deploys on research signal\n\nPair this with your release process: a Linear ticket tagged `block-release` from a Koji interview can fail your release checklist. This is overkill for most teams but useful for safety-critical products.\n\n## The Linear API approach (if you outgrow Zapier)\n\nLinear's GraphQL API is one of the cleanest issue-tracker APIs around. If you outgrow Zapier:\n\n1. Stand up a webhook receiver (Vercel, Cloudflare Workers, AWS Lambda)\n2. Verify Koji's webhook signature (HMAC)\n3. Transform the payload into a Linear `IssueCreateInput`\n4. Call `issueCreate` via Linear's GraphQL endpoint with a personal API key\n5. Optionally: call `issueLabelCreate` to ensure labels exist before tagging\n\nThis typically takes a senior engineer half a day. The advantage is markdown formatting, file attachments, custom views, and per-issue webhook callbacks back to Koji for closed-loop research.\n\n## Common pitfalls\n\n- **Label explosion.** If you map every Koji theme to a Linear label without filtering, your label list will balloon. Pre-create the 10–15 labels you care about and let Zapier skip the rest.\n- **Quote truncation.** Linear issue titles are limited to ~256 characters. Use a Zapier Formatter step to truncate the headline before mapping.\n- **Stale evidence.** Six months in, the original interview may be archived in Koji. Make sure your Linear ticket includes the participant ID and study ID — not just the URL — so the evidence can be re-fetched even if the URL changes.\n- **Triage burnout.** Even with good filters, expect 5–15 tickets per week from a moderately active research program. Assign a rotating triage role so no one person owns the whole queue.\n- **Customer privacy.** Verbatim quotes can contain identifying information. If your participants are under a strict NDA, anonymize quotes in the Zapier Formatter step before they hit Linear.\n\n## Why this changes how engineering teams ship\n\nThe biggest lift from a Koji + Linear integration is not the time saved on ticket creation — it's the change in what gets prioritized. Once every backlog ticket has a verbatim customer quote attached, the discussion in standup shifts from \"should we build this?\" to \"how do we solve what this person said?\" Quality scores let you weight signal honestly. Theme tags let you batch similar pains. And the link back to the live interview means anyone — engineer, designer, exec — can play the voice clip and form their own opinion in 30 seconds. That's a fundamentally different decision-making loop than the one most teams have today, and it's the kind of compounding advantage that the modern AI-native research stack makes possible.\n\n## Related resources\n\n- [Connect Koji to Zapier](/docs/zapier-research-automation) — the foundational webhook + Zapier guide\n- [Sync Koji Research Insights to Notion](/docs/notion-research-integration) — pair with Linear for full research-to-execution flow\n- [Sync Koji to HubSpot](/docs/hubspot-research-integration) — push insights to sales workflows\n- [Send Research Insights to Slack](/docs/slack-research-insights-integration) — real-time notifications\n- [Koji Webhook Setup](/docs/webhook-setup) — payload reference\n- [Koji Structured Questions Guide](/docs/structured-questions-guide) — the 6 question types Linear will receive\n- [Closing the Loop on Customer Feedback](/docs/closing-the-loop-customer-feedback) — turn shipped tickets into follow-up interviews\n- [User Research API](/docs/user-research-api-guide) — full API reference for direct integration\n","category":"API Reference","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Sync Koji to Linear: Auto-File Tickets from Customer Interviews","metaDescription":"Connect Koji to Linear via Zapier or GraphQL API to auto-create Linear issues from completed AI interviews. Filter by quality score and theme; attach verbatim quotes and study links.","keywords":["koji linear integration","linear customer feedback","auto-create linear tickets","customer interview to linear","research linear automation","linear webhook research","koji zapier linear","product feedback linear"],"aiSummary":"Connect Koji to Linear so interviews that surface real customer pain points auto-create tagged Linear issues — with verbatim quotes, theme tags, quality scores, and direct links to the live interview attached. The guide covers the critical design decision (which interviews should become tickets) using filters on quality score, theme, structured-answer thresholds, and sentiment; the full Zapier setup; severity-based team routing; deduplication via 'Find Issue' lookups; and the direct Linear GraphQL approach for teams that outgrow Zapier. The result replaces the Slack-thread-to-screenshot-to-ticket workflow that loses customer context with a closed-loop research-to-execution pipeline.","aiPrerequisites":["A Koji study with at least one completed interview","A Linear workspace with issue-creation permission","A Zapier account (Starter plan recommended for filters)","About 20 minutes"],"aiLearningOutcomes":["Choose the right filter rules so your Linear backlog doesn't explode","Set up a Zapier Catch Hook for Koji interview completions","Build a Linear issue description that includes verbatim quotes and metadata","Map quality scores to Linear priority automatically","Route different themes to different Linear teams","Deduplicate issues using Linear's Find Issue lookup","Migrate to direct Linear GraphQL API when volume justifies it","Anonymize quotes for privacy-sensitive participants"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"e2955d2d-eb5c-49c1-a997-616a75f3704a","slug":"jira-research-integration","title":"Jira + Koji: Auto-File Customer-Research-Backed Tickets and Close the Loop on Every Fix","url":"https://www.koji.so/docs/jira-research-integration","summary":"The Koji + Jira integration is bidirectional. From Koji into Jira: every interview theme that crosses a configurable severity/sentiment/quality threshold creates a Jira ticket via the REST API, pre-populated with the theme summary, 2-3 representative customer quotes, quality score, and a deep link to the source transcript. Subsequent interviews surfacing the same theme comment on the existing ticket (keyed on `koji-theme-{id}` label) rather than creating duplicates, so evidence accumulates. From Jira back into Koji: when a research-labeled ticket moves to Done, a Jira webhook fires back to Koji, which notifies the participants whose interviews originally surfaced the theme — closing the loop most research teams never execute. Three integration paths: Zapier (no code, 30 min), Koji → forwarder → Jira REST API (45 min), or a custom Atlassian app. Works on Jira Software, Service Management, and Product Discovery. Routing rules typically filter on sentiment (graduate only negative themes), quality (only 3+), and severity tag (bugs immediately, preferences weekly). Available on the Interviews plan (€79/mo) and Enterprise; Zapier path works on any plan.","content":"# Jira + Koji: Customer Research That Turns Into Shipped Tickets\n\n**Answer first:** Jira is where engineering work lives. Koji is where the customer-research signal that should drive that work lives. The integration pattern is bidirectional: (1) every Koji interview theme that crosses a configurable severity/sentiment threshold creates a Jira ticket pre-populated with the theme description, three representative customer quotes, sentiment, quality score, and a link back to the source transcript — and (2) when the Jira ticket is resolved, the participants whose interviews surfaced that theme can be notified automatically, closing the loop on the feedback they gave you. End-to-end this takes about 45 minutes to wire up via Koji webhooks + Jira REST API, or about 30 minutes with no code via Zapier. With tools like Koji, the gap between \\\"customers keep complaining about onboarding step 4\\\" and \\\"PROD-2812 is in the next sprint with five customer quotes attached\\\" closes inside the same tools your engineers already use.\n\nIf your engineering team won't look at a research report but does live in Jira, this is the integration that brings the research to them in the format they already prioritize from.\n\n## Why combine Jira and Koji\n\nMost product organizations have a fundamental impedance mismatch between research and engineering:\n\n- **Research outputs sit in slide decks and Notion pages.** They get presented in a meeting, then never re-opened. The ticket backlog goes on as if the research never happened.\n- **Engineers prioritize from Jira.** A ticket with a clear acceptance criterion, a screenshot, and a sprint label gets done. A reference to \\\"the Q2 research findings\\\" does not.\n- **The closing-the-loop step is missing.** Even when research-driven work ships, the customers who reported the issue almost never hear back. So the next round of research starts from zero trust.\n\nThe Koji + Jira integration fixes all three. Every theme that meets your threshold lands in Jira as a ticket the eng team can actually triage. The original transcripts are one click away. And when the ticket moves to Done, an outbound message goes back to the customers who flagged it — turning research into a visible feedback loop instead of a one-way report.\n\n## What flows in each direction\n\n### Koji → Jira (auto-file tickets from research themes)\n\n- An interview reaches `analysis_ready` in Koji.\n- The webhook payload includes themes (with severity and sentiment), customer quotes, quality score, and a transcript URL.\n- The forwarder checks each theme against your routing rules — for example: \\\"create a Jira ticket for any theme tagged as a bug with sentiment ≤ -0.4 and quality ≥ 3.\\\"\n- Matching themes become Jira issues via `POST /rest/api/3/issue` with: title (theme name), description (theme summary + 2-3 customer quotes + transcript link), labels (`koji-research`, theme tags), and (optionally) component, priority, and assignee.\n- Subsequent interviews that surface the same theme don't create new tickets — they comment on the existing one with the new quote, so the ticket's evidence accumulates rather than fragmenting.\n\n### Jira → Koji (close the loop when fixes ship)\n\nWhen a Jira ticket created from Koji moves to a configurable status (`Done`, `Released`, or a custom `Shipped to customers` status), a webhook fires back to your forwarder. The forwarder:\n\n- Looks up the participants whose interviews originally surfaced the theme.\n- Sends them a personalized thank-you message via Koji's [personalized interview links](/docs/personalized-interview-links) — \\\"You told us in your interview last month that X was broken. We shipped a fix this week. Here's what changed.\\\"\n- Optionally triggers a short Koji follow-up study asking the same participants whether the fix actually solved their problem.\n\nThe second loop is the rare research practice. It's also the one that turns participants into long-term advocates, because they see their feedback turn into a shipped change with their name attached.\n\n## Step 1 — Decide which themes graduate to Jira tickets\n\nNot every theme deserves a ticket. The most useful filters are:\n\n- **Severity.** Themes tagged `bug` or `blocker` create tickets immediately. Themes tagged `preference` or `nice-to-have` accumulate in Koji and get reviewed weekly.\n- **Sentiment.** Themes with negative sentiment below a threshold (e.g., -0.4 on a -1 to +1 scale) graduate to tickets. Positive themes go to a separate \\\"wins\\\" board, not Jira.\n- **Quality.** Only quality 3+ interviews ([understanding quality scores](/docs/understanding-quality-scores)) generate tickets. Low-quality conversations are excluded to keep ticket noise down.\n- **Volume.** A theme that appears in only one interview becomes a candidate ticket on a watch list. A theme that appears across three or more interviews creates the ticket immediately.\n\nThe [understanding themes and patterns](/docs/understanding-themes-patterns) doc explains how Koji generates themes; the [activating research insights](/docs/activating-research-insights) doc covers the broader \\\"insight → action\\\" pattern.\n\n## Step 2 — Build the forwarder (Koji → Jira)\n\nSubscribe to Koji's `interview.analysis_ready` event (full reference in [webhook setup](/docs/webhook-setup)). Your forwarder:\n\n1. **Verifies the Koji HMAC signature.**\n2. **For each theme in the payload, runs your routing rules.**\n3. **For each theme that passes, calls Jira:**\n   - First, query existing issues with label `koji-theme-{theme_id}` — if one exists, post a comment with the new quote and exit.\n   - If not, create a new issue via `POST /rest/api/3/issue` with:\n     - **Summary:** the theme name, prefixed with `[Research] ` so eng triage can spot it.\n     - **Description:** the theme summary, 2-3 representative quotes from the transcript, the quality score, and a deep link to the Koji study/report.\n     - **Labels:** `koji-research`, `koji-theme-{theme_id}`, plus any theme tags from Koji.\n     - **Priority:** mapped from sentiment (very negative → High; moderately negative → Medium; preference → Low).\n     - **Component or Epic:** routed from theme tags via your config (auth themes → Auth component, billing themes → Billing component, etc.).\n4. **Stores the Jira issue key in a local table** keyed on `koji_theme_id` so future interviews surfacing the same theme can find and comment on the existing issue.\n\nFor any path, the [user research API guide](/docs/user-research-api-guide) and [API authentication](/docs/api-authentication) doc cover the Koji side; Atlassian's Jira REST API docs cover the destination.\n\n### Path A: Zapier or Make (no code, ~30 minutes)\n\nFor teams that don't want to run a forwarder:\n\n1. **Trigger:** Koji → Interview Analysis Ready (via Koji's Zapier app).\n2. **Filter:** by sentiment and quality (Zapier filter step).\n3. **Search:** Jira → Find Issue by label `koji-theme-{id}` (Zapier action).\n4. **Action:** Jira → Create Issue (if none found) or Add Comment (if one exists).\n\nThis path skips the deduplication elegance of a custom forwarder, but it's a 30-minute setup and works for low-volume research programs.\n\n## Step 3 — Close the loop (Jira → Koji)\n\nIn Jira, configure a webhook that fires on issue status change. When an issue labeled `koji-research` moves to `Done` (or your custom `Shipped` status):\n\n1. **The Jira webhook calls your forwarder** with the issue key and new status.\n2. **The forwarder looks up the `koji_theme_id`** from the issue labels.\n3. **The forwarder queries Koji** for the list of respondents whose interviews surfaced that theme.\n4. **The forwarder sends each respondent** a thank-you message — either via your email tool or via Koji's [personalized interview links](/docs/personalized-interview-links) with a short follow-up study attached.\n\nThis is the step almost no research team executes today. It's also the highest-leverage one. A customer who flags a problem, then hears back when it's fixed, doubles their willingness to participate next time and becomes a vocal advocate. The CS team will tell you they've never had an easier renewal conversation than the one that includes \\\"you told us about X, we fixed it, here's the changelog.\\\"\n\n## What you can build in Jira once the data lands\n\n- **A research-backed backlog.** Filter Jira on label `koji-research` to see every ticket that came from customer interviews, sortable by sentiment, theme volume, or quality.\n- **Sprint review with customer quotes.** Every research-tagged ticket carries 2-3 customer quotes in its description, so engineers and PMs ship work without losing the \\\"why.\\\"\n- **Theme-to-epic mapping.** Group tickets by theme tag into epics that reflect customer-defined problem areas, not internal architecture.\n- **A \\\"voice of customer\\\" dashboard.** A Jira filter view of all `koji-research`-labelled issues by status answers \\\"how much customer-driven work are we actually shipping?\\\" — a question most product orgs cannot answer today.\n- **Re-prioritization on theme volume.** When the same theme accumulates 10+ quotes across two months, that's evidence to bump priority. The ticket already has the quotes attached.\n\n## Comparison: Koji → Jira vs. ad-hoc research-to-engineering handoff\n\n- **Manual transcription.** A PM reads a research report, writes a Jira ticket, and pastes one or two quotes. Most quotes are forgotten; tickets carry minimal qualitative weight. The Koji integration carries the *actual* customer language into the ticket automatically.\n- **\\\"Reference the Notion page.\\\"** A Jira ticket whose acceptance criterion is \\\"see Q2 research synthesis\\\" never gets implemented faithfully. Tickets need self-contained context. Koji puts that context in the ticket body.\n- **No closing the loop.** Without an integration, customers never hear back when their feedback ships. The Jira → Koji direction fixes this with one webhook.\n- **No deduplication.** A research-driven backlog written by hand fragments — \\\"onboarding step 4 is confusing\\\" becomes five different tickets. The Koji forwarder's theme-keyed deduplication keeps everything related to one root issue in one place.\n\n## Plan requirements and cost\n\nWebhooks and the headless API are included on the Interviews plan (€79/month, 79 credits) and Enterprise. The Insights plan (€29/month) doesn't include webhooks — for that tier, the Zapier path works on any plan. Text interviews cost 1 credit, voice interviews cost 3, and only conversations scoring 3 or higher on Koji's quality gate consume credits. See [plan comparison guide](/docs/plan-comparison-guide).\n\nJira pricing is independent — Standard, Premium, and Enterprise plans all support the REST API, webhooks, and custom labels needed here. The integration works equally on Jira Software, Jira Service Management, and Jira Product Discovery.\n\n## Identity, attribution, and the privacy default\n\nFor product research, the participant identity does *not* need to flow to Jira. The ticket carries the theme, the quotes, the quality score, and a link to the Koji study — but the participant's email never appears in the issue. This keeps Jira out of scope for any privacy regime that classifies participant identity as sensitive while still giving engineering everything they need to ship the fix.\n\nFor the closing-the-loop direction, identity lives only in your forwarder's lookup table (mapping theme → respondent list), not in Jira. Anonymous-mode studies cannot close the loop by definition — there's no respondent to notify — but the inbound Koji → Jira direction still works fully for anonymous research.\n\nFor regulated industries, see [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research) and [HIPAA-compliant AI user research](/docs/hipaa-compliant-ai-user-research).\n\n## A 45-minute first run\n\nThe fastest end-to-end test:\n\n1. In Koji, pick a recent study that surfaced a clear bug or friction theme.\n2. In Jira, create a label `koji-research` and a filter view that pins all such issues.\n3. Manually create a ticket from the theme: paste the theme summary, three quotes, the transcript link, and label `koji-research`.\n4. Run a sprint review with that ticket in the room. Watch how the customer quotes change the conversation about priority.\n5. If the eng team responds to the format, wire up the webhook → Jira forwarder so the next 20 themes graduate to tickets automatically.\n\nThe manual first run is the cheapest way to learn whether your eng team will actually engage with research-tagged tickets — almost always yes, but worth confirming before you invest in the automation.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types whose answers become evidence in every Jira ticket.\n- [Linear Research Integration](/docs/linear-research-integration) — sister guide for teams using Linear instead of Jira.\n- [Notion Research Integration](/docs/notion-research-integration) — build a self-updating research repository in Notion alongside the Jira backlog.\n- [Slack Research Insights Integration](/docs/slack-research-insights-integration) — pipe theme alerts to Slack channels in real time.\n- [Webhook Setup](/docs/webhook-setup) — full reference for Koji webhook events and HMAC signature verification.\n- [Research Automation Webhooks](/docs/research-automation-webhooks) — idempotency, retries, and replay patterns for production forwarders.\n- [Activating Research Insights](/docs/activating-research-insights) — the broader \\\"insight → action\\\" methodology this integration operationalizes.\n- [Understanding Themes and Patterns](/docs/understanding-themes-patterns) — how Koji generates the themes that become Jira tickets.","category":"API Reference","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Jira + Koji Integration: Auto-File Research-Backed Tickets and Close the Loop","metaDescription":"Auto-create Jira tickets from Koji AI interview themes — with customer quotes, sentiment, and quality scores attached. Close the loop by notifying participants when their feedback ships.","keywords":["jira integration","jira customer research","jira user research","jira webhook integration","research backlog jira","customer feedback to jira","jira ai survey","close the loop customer feedback"],"aiSummary":"The Koji + Jira integration is bidirectional. From Koji into Jira: every interview theme that crosses a configurable severity/sentiment/quality threshold creates a Jira ticket via the REST API, pre-populated with the theme summary, 2-3 representative customer quotes, quality score, and a deep link to the source transcript. Subsequent interviews surfacing the same theme comment on the existing ticket (keyed on `koji-theme-{id}` label) rather than creating duplicates, so evidence accumulates. From Jira back into Koji: when a research-labeled ticket moves to Done, a Jira webhook fires back to Koji, which notifies the participants whose interviews originally surfaced the theme — closing the loop most research teams never execute. Three integration paths: Zapier (no code, 30 min), Koji → forwarder → Jira REST API (45 min), or a custom Atlassian app. Works on Jira Software, Service Management, and Product Discovery. Routing rules typically filter on sentiment (graduate only negative themes), quality (only 3+), and severity tag (bugs immediately, preferences weekly). Available on the Interviews plan (€79/mo) and Enterprise; Zapier path works on any plan.","aiPrerequisites":["Active Jira workspace (Software, Service Management, or Product Discovery)","Koji account on Interviews plan or higher for webhook automation","Permission to create issues and configure webhooks in Jira"],"aiLearningOutcomes":["Auto-file Jira tickets from Koji AI interview themes with quotes attached","Set sentiment, severity, and quality thresholds that gate ticket creation","Deduplicate themes so accumulating evidence lives on a single ticket","Close the loop by notifying participants when their flagged issue ships","Build a research-backed Jira backlog that engineering will actually prioritize"],"aiDifficulty":"intermediate","aiEstimatedTime":"20 min read"},{"type":"documentation","id":"f928e394-eb87-4eb5-9dca-754e6e6552be","slug":"critical-incident-technique","title":"Critical Incident Technique: The Interview Method That Captures What Really Matters","url":"https://www.koji.so/docs/critical-incident-technique","summary":"The Critical Incident Technique (CIT) is a qualitative interview method developed by Flanagan (1954) that collects specific incidents — positive and negative — to identify the behaviours and moments that most impact user experience. Unlike general surveys, CIT anchors participants in real, memorable events.","content":"\n## What Is the Critical Incident Technique?\n\nThe Critical Incident Technique (CIT) is a qualitative research method that uncovers the moments that truly shape user experience. Rather than asking people how they *generally* feel about a product, CIT asks them to recall **specific events** — the moments where something went remarkably right or frustratingly wrong. This specificity is what makes CIT so powerful for product research.\n\nJohn Flanagan, a psychologist working with the US Air Force, developed CIT in 1954 to understand what distinguished effective pilots from ineffective ones. Standard surveys couldn't capture the nuance needed. His solution: collect specific, observable incidents and analyze them for patterns. The result was a landmark paper in *Psychological Bulletin* that defined CIT as \"a set of procedures for collecting direct observations of human behaviour in such a way as to facilitate their potential usefulness in solving practical problems and developing broad psychological principles.\"\n\nToday there are approximately 200 CIT studies in the marketing and consumer research literature alone. The technique has spread across healthcare, education, service quality research, and product UX — anywhere the difference between excellent and poor performance needs to be understood precisely.\n\n## The Core Problem CIT Solves\n\nAsk someone \"How satisfied are you with our onboarding?\" and you get a number shaped by their current mood, social desirability bias, and cognitive distortions. Ask them instead: \"Tell me about a specific time when you got completely stuck while setting up our product\" — and you unlock something far more actionable.\n\nThe Nielsen Norman Group notes that \"memory is fallible, and so details can often be lost, or critical incidents can be forgotten\" — which is precisely why CIT interviews should happen soon after the events in question, and why the method's structured prompting matters so much. By anchoring participants to a specific remembered incident, CIT bypasses the generalization problem that plagues most survey research.\n\n## The Three-Question Framework\n\nA CIT interview is built around three core questions:\n\n1. **What happened?** — Describe the specific incident\n2. **What led to it?** — Context and circumstances\n3. **What was the outcome?** — Result and impact\n\nThese questions cut through vague opinions and anchor responses in real, memorable events. The first question elicits the incident itself; the second surfaces context that helps you understand why it happened; the third reveals whether the incident actually mattered to the user's behaviour (did they churn? did they tell others? did they find a workaround?).\n\n## Origins: From Wartime Aviation to Product Research\n\nDuring World War II, the US Army Air Force needed to understand what made effective combat pilots. Flanagan's research team collected hundreds of specific incidents observed by supervisors — not ratings of general performance, but accounts of specific effective and ineffective behaviours in specific situations. Categories emerged from the incidents, not from the researchers' assumptions.\n\nThis inductive approach — letting the data define the categories — is CIT's most important methodological contribution. You don't ask \"was navigation confusing?\" You collect incidents, and if 60% of incidents involve users failing to find specific features, you know navigation is the problem. The insight emerges from the pattern, not from the question.\n\n## How to Structure a CIT Study\n\nA typical CIT interview lasts 30–60 minutes and covers 3–5 incidents. Here is the framework:\n\n### Phase 1: Situation Priming (5 minutes)\nOrient the participant. Explain you are looking for specific stories, not general opinions. \"I want to hear about actual experiences you have had — specific moments you remember clearly.\"\n\n### Phase 2: Incident Elicitation (15–20 minutes per incident)\nUse the core CIT prompt: \"Think back to a specific time when using [product] was [very helpful / very frustrating]. Can you walk me through exactly what happened?\"\n\nFollow up with:\n- \"What were you trying to accomplish?\"\n- \"What specifically triggered this moment?\"\n- \"What did you do next?\"\n- \"What was the outcome?\"\n\n### Phase 3: Impact Assessment (5 minutes per incident)\n- \"How significant was this moment for your overall experience?\"\n- \"Did this change how you use [product]?\"\n- \"Would you have done anything differently?\"\n\n### Phase 4: Pattern Reflection (10 minutes)\nAfter collecting 3–5 incidents, ask about patterns: \"Were these incidents typical of your experience? What would make these moments less likely to happen?\"\n\n## Positive vs. Negative Incidents: Collect Both\n\nOne of CIT's most underused features is its ability to collect **positive** critical incidents alongside negative ones. Most research over-indexes on problems. But understanding what delights users is equally important for identifying strengths to double down on, understanding the emotional peaks that drive word-of-mouth, and finding features that are working before inadvertently removing them.\n\nFor every \"tell me about a time it went wrong\" prompt, always ask: \"Tell me about a time it went especially well.\"\n\nNetflix uses a variant of CIT in their content research to understand the specific moments that cause viewers to abandon a show or become deeply engaged — the exact scenes that trigger \"just one more episode\" versus \"I'm done.\" The technique scales to any domain where understanding the *specific moment* of a behavioural shift matters more than measuring average sentiment.\n\n## Analysing CIT Data\n\nCIT produces rich qualitative data that requires systematic analysis. The standard approach:\n\n1. **Extract incidents** — Pull each distinct event from transcripts\n2. **Categorize by behaviour type** — Group incidents that share the same underlying pattern\n3. **Label categories** — Give each category a descriptive name (\"Failed to find feature\", \"Unexpectedly delighted by X\")\n4. **Quantify frequency** — Count how many incidents fall into each category\n5. **Prioritize by frequency and impact** — High-frequency, high-impact categories get addressed first\n\nModern AI analysis tools accelerate this dramatically. Koji's thematic analysis automatically clusters CIT responses by pattern, surfacing the most common incident types without hours of manual coding.\n\n## CIT with Koji's Structured Question Types\n\nCIT adapts naturally to Koji's structured question framework. A well-designed CIT study in Koji might combine:\n\n- **Open-ended questions** for incident elicitation: \"Describe a specific time when you got stuck in our product\"\n- **Scale questions** for impact measurement: \"On a scale of 1–10, how much did this incident affect your likelihood to continue using the product?\"\n- **Single-choice questions** for incident classification: \"Which area of the product was involved? (Onboarding / Core workflow / Settings / Reporting)\"\n- **Yes/No questions** for outcome tracking: \"Did this incident cause you to stop using the feature?\"\n- **Multiple-choice questions** for contributing factors: \"What factors contributed? (Select all that apply)\"\n- **Ranking questions** for solution prioritization: \"Rank these potential fixes in order of importance to you\"\n\nThis structured approach lets Koji aggregate CIT data across many respondents, converting qualitative incidents into quantifiable patterns — something traditional CIT could never do at scale.\n\n## Running CIT at Scale with AI\n\nTraditional CIT was limited to small samples (10–20 interviews) because each interview required a skilled moderator and hours of analysis. AI changes this fundamentally.\n\nWith Koji:\n1. **Create a CIT study** with incident-elicitation prompts\n2. **AI interviewer** probes naturally when participants give vague answers — \"Can you be more specific about what happened next?\"\n3. **Automated thematic analysis** clusters incidents by category across all respondents\n4. **Structured question widgets** capture quantitative ratings alongside qualitative incidents\n5. **Report generation** produces incident frequency distributions, not just themes\n\nThis makes CIT viable at 50–200 respondent scale — unlocking statistical significance while preserving the depth of incident-based research.\n\n## When to Use CIT (and When Not To)\n\n**Best for:**\n- Understanding the specific moments that cause churn\n- Identifying onboarding failure points\n- Discovering delight moments to amplify\n- Post-mortems: what went wrong in specific user journeys\n- Service quality research in customer experience teams\n\n**Not ideal for:**\n- Measuring overall satisfaction (use NPS or CSAT)\n- Quantifying feature preferences (use choice-based ranking)\n- Large-scale quantitative studies where breadth matters more than depth\n- Exploratory research without specific behaviours to investigate\n\n## CIT vs. Other Interview Methods\n\n| Method | Focus | Sample Size | Output |\n|--------|-------|-------------|--------|\n| CIT | Specific incidents | 10–20 | Incident categories |\n| Jobs to Be Done | Motivations & context | 10–30 | Job statements |\n| Usability Testing | Task performance | 5–8 | Usability issues |\n| NPS Follow-up | Overall sentiment | 100s | Themes |\n| Think Aloud | Real-time cognition | 5–8 | UI friction points |\n\nCIT occupies a unique niche: it bridges the gap between real behaviour and retrospective storytelling, giving you incidents you can act on without the artificiality of lab-based testing.\n\n## Related Resources\n\n- [How to Conduct User Interviews](/docs/how-to-conduct-user-interviews) — foundational interview skills and techniques\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) — how to analyze qualitative data from CIT sessions\n- [Structured Questions Guide](/docs/structured-questions-guide) — combining quantitative questions with CIT interviews\n- [Jobs to Be Done Framework](/docs/jobs-to-be-done-framework) — complementary technique for motivation research\n\n\n## Further reading on the blog\n\n- [AI-Moderated vs Human-Moderated Interviews: Which Should You Choose?](/blog/ai-moderated-vs-human-moderated-interviews) — AI-moderated and human-moderated interviews each have a time and a place. Here is the honest comparison to help you choose the right approac\n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n- [How to Analyze Customer Interview Data: A Complete Guide](/blog/how-to-analyze-customer-interview-data) — You ran the interviews. Now what? Here is a step-by-step process for turning raw transcripts into clear, actionable insights your team will \n\n<!-- further-reading:blog -->\n","category":"Interview Techniques","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Critical Incident Technique: CIT User Research Guide | Koji","metaDescription":"Master the Critical Incident Technique (CIT) for user research. Learn Flanagan’s interview framework to collect specific incidents that reveal what really matters to your users.","keywords":["critical incident technique","CIT user research","incident-based interviewing","Flanagan CIT","qualitative research methods","user experience research","UX research methods"],"aiSummary":"The Critical Incident Technique (CIT) is a qualitative interview method developed by Flanagan (1954) that collects specific incidents — positive and negative — to identify the behaviours and moments that most impact user experience. Unlike general surveys, CIT anchors participants in real, memorable events.","aiPrerequisites":["How to Conduct User Interviews","Basic qualitative research concepts"],"aiLearningOutcomes":["Design CIT interview protocols using Flanagan's three-question structure","Collect and categorize both positive and negative critical incidents","Analyze incident data to identify actionable patterns","Combine CIT with Koji's structured question types for scaled analysis"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"},{"type":"documentation","id":"e7d042a6-1597-40d7-8b75-63a4e8f1cb58","slug":"mcp-setup-cursor","title":"Connect Koji to Cursor: MCP Setup Guide for Product Engineers","url":"https://www.koji.so/docs/mcp-setup-cursor","summary":"Connect the Koji MCP server to Cursor (0.45+) by adding a single block to ~/.cursor/mcp.json with a bearer token. Once connected, Cursor can call all 15 Koji MCP tools — listing studies, pulling transcripts, fetching structured answers, generating reports, and even creating new studies — directly from any chat or Composer session. This brings live customer evidence into the editor for spec writing, ticket grooming, pre-deploy sentiment checks, and stakeholder report generation, with no copy-pasting and full quality-gate filtering.","content":"\nCursor's Model Context Protocol (MCP) support turns the editor into a research-aware coding environment. By connecting Koji's MCP server to Cursor, your AI assistant gains direct access to live customer interview transcripts, structured answers, themes, quality scores, and full study reports — without copy-pasting a single quote.\n\nThis guide walks through the full setup, the most useful Koji + Cursor workflows, and the gotchas to know about authentication and tool selection.\n\n## Why Connect Koji to Cursor\n\nMost teams already use Cursor to write code. Most teams also have customer feedback scattered across Slack threads, Notion docs, and a forgotten interview tool. The MCP integration closes that gap: Cursor can query your Koji workspace the same way it queries the filesystem.\n\nConcretely, this unlocks four high-value workflows for product engineers:\n\n- **Context-aware feature work.** Before adding a new onboarding step, ask Cursor \"what did the last 20 interview participants say about onboarding friction?\" and it pulls the actual quotes.\n- **PRDs grounded in evidence.** Generate a one-page spec that cites real interview transcripts instead of speculative bullet points.\n- **Triage with proof.** When grooming the backlog, Cursor can rank tickets against the themes appearing most frequently in published Koji reports.\n- **Post-launch sense-check.** After shipping, ask Cursor to summarise interviews collected since the deploy and flag regressions in sentiment or quality scores.\n\nThis is a wedge no traditional research tool can offer. SurveyMonkey, Typeform, and Qualtrics still expect you to log into a separate dashboard. Koji's MCP server brings the data to where engineers already work.\n\n## Prerequisites\n\nBefore you start, make sure you have:\n\n- A Koji workspace with at least one published study and a few completed interviews to query.\n- Cursor 0.45 or later. Earlier versions do not support MCP.\n- A Koji API key. Generate one in **Settings → API Keys** in your workspace. See [Managing API Keys](/docs/managing-api-keys) for guidance on scopes and rotation.\n- Familiarity with editing JSON config files. The setup is one short paste.\n\nThe Koji MCP server is included on every paid plan and on the Free plan for evaluation. Tool calls are not metered separately — they share the standard interview credit pool described in [Plan Comparison Guide](/docs/plan-comparison-guide).\n\n## Step 1: Open Your Cursor MCP Config\n\nCursor stores MCP server definitions in a JSON file. The location depends on your operating system:\n\n- macOS / Linux: `~/.cursor/mcp.json`\n- Windows: `%USERPROFILE%\\.cursor\\mcp.json`\n\nIf the file does not exist, create it. You can also open the same config from inside Cursor via **Cursor → Settings → MCP → Edit Config**.\n\n## Step 2: Add the Koji MCP Server Block\n\nPaste the following into the `mcpServers` object of your `mcp.json`. Replace `YOUR_KOJI_API_KEY` with the key you generated above.\n\n```json\n{\n  \"mcpServers\": {\n    \"koji\": {\n      \"url\": \"https://www.koji.so/api/mcp/mcp\",\n      \"headers\": {\n        \"Authorization\": \"Bearer YOUR_KOJI_API_KEY\"\n      }\n    }\n  }\n}\n```\n\nIf you already have other MCP servers configured (GitHub, Linear, filesystem, etc.), add the `koji` entry alongside them rather than replacing the file.\n\nSave the file and restart Cursor. The first time you open a chat after the restart, Cursor will list `koji` in its connected servers panel and surface its 15 tools.\n\n## Step 3: Verify the Connection\n\nOpen a new chat in Cursor (Cmd/Ctrl + L) and type:\n\n```\n@koji list my recent studies\n```\n\nCursor should call `koji_list_studies` and return your study titles, statuses, and interview counts. If it reports an authentication error, double-check the bearer token in `mcp.json` — a stray space or quote is the most common culprit.\n\nFor a deeper smoke test, try:\n\n```\n@koji summarise the most common themes from the most recent published study\n```\n\nThis chains `koji_list_studies`, `koji_get_study`, and `koji_get_interviews` automatically. The model picks the right tool for each step — you do not need to name them explicitly. See [MCP Tool Reference](/docs/mcp-tool-reference) for the full list.\n\n## The 15 Koji Tools You Now Have in Cursor\n\nOnce connected, Cursor can call any tool the Koji MCP server exposes:\n\n**Read tools** — `koji_list_studies`, `koji_get_study`, `koji_get_interviews`, `koji_get_transcript`, `koji_get_account`, `koji_get_study_data`, `koji_get_report`\n\n**Write and analysis tools** — `koji_create_study`, `koji_update_brief`, `koji_publish_study`, `koji_generate_report`, `koji_publish_report`, `koji_configure_study`, `koji_export_data`, `koji_import_respondents`\n\nThe read tools are safe to call freely. The write tools modify your workspace, so configure Cursor's tool-approval policy to require confirmation before they run. See [MCP Authentication and Security](/docs/mcp-authentication-security).\n\n## Recommended Cursor + Koji Workflows\n\nThe integration works best when you anchor it to a concrete coding task. Here are five patterns teams ship with on day one.\n\n### 1. Writing a feature spec from interview evidence\n\nOpen the spec doc and prompt Cursor:\n\n```\nDraft a one-page PRD for a new onboarding checklist. Cite at least\nfive direct quotes from the most recent onboarding study in Koji,\nand group user pain points by theme.\n```\n\nCursor will fetch transcripts and structured answers, then generate the spec with the quotes inline. The structured answers come from Koji's six question types (`open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no`) — see [Structured Questions in AI Interviews](/docs/structured-questions-guide) for how those become chartable data the model can reason over.\n\n### 2. Composer-driven ticket grooming\n\nOpen Composer (Cmd/Ctrl + I) on your `tickets/` directory and prompt:\n\n```\nFor each open ticket in this folder, score 1–5 how strongly the\nlast 30 days of Koji interviews support shipping it. Cite the\nthemes that match.\n```\n\nThe model walks each ticket, calls `koji_get_study_data` for theme aggregations, and writes the score plus a short rationale into the ticket file.\n\n### 3. Pre-deploy regression check\n\nBefore merging, ask Cursor to scan recent interview themes for any spike in negative sentiment. Useful when you ship onboarding, pricing, or auth changes — exactly the surfaces participants comment on most.\n\n### 4. Generating a stakeholder summary\n\nAfter a study concludes, prompt:\n\n```\nGenerate a stakeholder report for the latest study, then publish it.\n```\n\nThis chains `koji_generate_report` and `koji_publish_report`. The published report URL comes back inline in Cursor — paste into Slack and you are done. This is the same flow described in [Real-Time Research Insights](/docs/real-time-research-insights), but driven from your editor.\n\n### 5. Creating a quick discovery study from a code TODO\n\nHighlight a `// TODO: validate this with users` comment, open chat, and prompt:\n\n```\nCreate a 5-minute Koji study to validate this assumption with 20\nmobile users. Use one open_ended question and one scale question\nfor confidence.\n```\n\nCursor calls `koji_create_study` and returns a shareable interview URL. From idea to live study in under a minute.\n\n## Troubleshooting\n\n**\"Tool not found\" errors.** Restart Cursor — MCP servers are only registered on startup. If the issue persists, run `cat ~/.cursor/mcp.json | jq .` to confirm the JSON is valid.\n\n**Authentication failures.** Bearer tokens are sensitive; never paste them into a shared `.cursor` config. Use Cursor's environment-variable substitution (`\"Authorization\": \"Bearer ${env:KOJI_API_KEY}\"`) and load the key from your shell.\n\n**Tool selection drift.** If Cursor calls `koji_get_interviews` when you wanted `koji_get_study_data`, name the tool explicitly in your prompt. The model improves with corrections inside the same chat.\n\n**Rate limits.** The Koji MCP endpoint shares the standard API rate limits documented in [Rate Limits and CORS](/docs/rate-limits-and-cors). Heavy bulk reads should batch by `limit` and `cursor`.\n\n## How This Compares to Pasting Transcripts\n\nSome teams still copy interview snippets into Cursor by hand. That works for a single decision but breaks down at scale: stale data, missing quality scores, and no way to filter by theme or sentiment.\n\nThe MCP integration eliminates all three. Cursor sees the same authoritative data your researchers see, scoped by your API key, with quality scores that automatically suppress low-effort responses (only interviews scoring 3 or above are surfaced — see [How the Quality Gate Works](/docs/how-the-quality-gate-works)).\n\nThat is the difference between \"AI with vibes\" and \"AI with evidence.\"\n\n## Related Resources\n\n- [Koji MCP Integration Overview](/docs/mcp-overview)\n- [Connect Koji to Claude (Setup Guide)](/docs/mcp-setup-claude)\n- [MCP Tool Reference](/docs/mcp-tool-reference)\n- [MCP Authentication and Security](/docs/mcp-authentication-security)\n- [MCP Best Practices](/docs/mcp-best-practices)\n- [MCP Workflow Guide for Product Managers](/docs/mcp-workflow-product-managers)\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide)\n\n\n## Further reading on the blog\n\n- [B2B Customer Research: The Complete Guide for Product Teams (2026)](/blog/b2b-customer-research-guide-2026) — B2B customer research is harder than B2C — you are navigating buying groups of 10+ stakeholders, gatekeepers, and enterprise procurement cyc\n- [Best Product Discovery Tools in 2026: The Complete Buyer's Guide](/blog/best-product-discovery-tools-2026) — Product discovery is no longer a one-off pre-launch phase — it is a continuous loop. The best teams in 2026 are running weekly customer inte\n- [Concept Testing: The Complete Guide for Product Teams (2026)](/blog/concept-testing-guide-2026) — Concept testing validates whether your idea is worth building before you build it. This guide covers methods, question templates, analysis a\n\n<!-- further-reading:blog -->\n\n\n---\n\n<!-- crosslink:dev-hub -->\n## Related: other ways to connect\n\nThe [Developer Hub](/docs/building-with-ai) is the single starting point for every way to connect to Koji: the MCP server (this section), the REST Headless API, and the embed widget. For a code integration, see [API Authentication](/docs/api-authentication) and [Starting Interviews via API](/docs/starting-interviews-via-api).\n","category":"Claude & MCP Integration","lastModified":"2026-07-30T03:25:36.132045+00:00","metaTitle":"Connect Koji to Cursor: MCP Setup Guide (5 Minutes)","metaDescription":"Step-by-step guide to connect Koji to Cursor via MCP so engineers can pull live customer interview insights, transcripts, and themes directly into the editor.","keywords":["cursor mcp","koji mcp cursor","cursor user research","mcp ide research","cursor ai customer feedback","mcp setup cursor"],"aiSummary":"Connect the Koji MCP server to Cursor (0.45+) by adding a single block to ~/.cursor/mcp.json with a bearer token. Once connected, Cursor can call all 15 Koji MCP tools — listing studies, pulling transcripts, fetching structured answers, generating reports, and even creating new studies — directly from any chat or Composer session. This brings live customer evidence into the editor for spec writing, ticket grooming, pre-deploy sentiment checks, and stakeholder report generation, with no copy-pasting and full quality-gate filtering.","aiPrerequisites":["Koji workspace with at least one published study","Cursor 0.45 or later installed","A Koji API key generated from Settings"],"aiLearningOutcomes":["Add the Koji MCP server to Cursor in under 5 minutes","Verify the connection with a smoke-test query","Use all 15 Koji tools (read and write) inside Cursor chat and Composer","Run 5 production workflows: PRD writing, ticket grooming, regression checks, report generation, study creation","Troubleshoot auth, tool selection, and rate-limit issues"],"aiDifficulty":"beginner","aiEstimatedTime":"5 min read"}],"pagination":{"total":1268,"returned":100,"offset":0}}