{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-07-28T14:02:18.883Z"},"content":[{"type":"documentation","id":"bceabd4c-aede-4355-b9e8-e1e753cd8fde","slug":"research-data-retention-deletion","title":"Research Data Retention and Deletion: How Long Should You Keep Interview Data?","url":"https://www.koji.so/docs/research-data-retention-deletion","summary":"There is no universal legal retention period for research data - GDPR Article 5(1)(e) storage limitation requires keeping personal data only as long as necessary for the collection purpose, which means the absence of a documented schedule is itself the compliance failure. Retention should be decided per artifact, not per study: raw audio 30-90 days after transcription, identifiable transcripts 6-12 months, de-identified transcripts and quotes 24-36 months or indefinitely if genuinely anonymized, analysis outputs indefinitely, consent records as long as the data plus the limitation period. Genuinely anonymized data falls outside GDPR so the retention clock stops, while pseudonymized data remains personal data. Key regimes: COPPA prohibits indefinite retention of children data, Illinois BIPA requires a publicly available destruction schedule for biometric identifiers.","content":"## The short answer\n\n**There is no universal legal retention period for research data. The law requires you to decide, document, and enforce one — so the absence of a schedule is itself the compliance failure.**\n\nUnder GDPR's storage limitation principle (Article 5(1)(e)), personal data may be kept only as long as necessary for the purpose it was collected for. That is a *relative* standard: it obliges you to define the purpose, derive a period from it, and then actually delete on schedule.\n\nThe second key idea, which most teams miss:\n\n**Retention is a per-artifact decision, not a per-study decision.** A single 45-minute interview generates a recording, a transcript, a set of coded themes, a handful of verbatim quotes, a consent record, and a contact record for paying the incentive. Those six artifacts have wildly different risk profiles and wildly different useful lifespans. One blanket rule gets it wrong in both directions at once — holding raw audio far too long while deleting the analysis you actually needed.\n\n> Not legal advice. Retention periods interact with your jurisdiction, sector, and contracts; validate your schedule with counsel.\n\n## The recommended tiered schedule\n\nAdapt the periods, keep the structure. These are defensible defaults, not legal minimums.\n\n| Artifact | Suggested retention | Why |\n|---|---|---|\n| **Raw audio / video** | 30–90 days after transcription | Highest-risk, lowest marginal value once transcribed. The single biggest easy win. |\n| **Identifiable transcripts** | 6–12 months | Long enough to re-analyze and verify; short enough to limit exposure. |\n| **De-identified transcripts and quotes** | 24–36 months, or indefinitely if genuinely anonymized | Where nearly all analytical value lives. |\n| **Analysis outputs — themes, reports, aggregate scores** | Indefinite | Business records containing no personal data. |\n| **Consent records** | As long as you hold the data, plus your limitation period | This is your evidence of lawful basis. Do not delete it with the data. |\n| **Contact details and incentive/payment records** | Per finance and tax rules, in a separate system | Different purpose, different clock — never in the research corpus. |\n| **Voiceprints or other biometric identifiers** | Do not collect for research | If collected, a published destruction schedule is mandatory — see below. |\n\nThe recording tier deserves emphasis. Raw audio is the most sensitive artifact you hold and, once you have an accurate transcript, usually the least useful. Teams keep it out of vague anxiety about needing to \"go back to the tape.\" In practice they almost never do — and meanwhile the recording is the thing that turns a routine security incident into a serious one.\n\n## Anonymization is the lever that lets you keep insight forever\n\nThis is the most valuable mechanic in retention design, and it is widely misunderstood.\n\n- **Pseudonymized** data — real identifiers swapped for codes, with a key that still exists somewhere — is **still personal data**. The retention clock keeps running.\n- **Genuinely anonymized** data — where re-identification is no longer reasonably possible, and no key exists — falls **outside** GDPR's scope. The clock stops.\n\nSo the way to retain research value indefinitely without indefinite risk is not to argue for longer retention periods. It is to **de-identify at the point of synthesis**, so that your durable artifacts — reports, theme libraries, quote banks — never contain personal data in the first place.\n\nBe honest about the bar, though. Qualitative data resists anonymization more than survey data, because narrative detail identifies people. \"The VP of Engineering at a 40-person Berlin fintech who joined last March\" is identifiable no matter what you call them. Strip role-plus-company-plus-timeline combinations, not just names. See [anonymizing customer interview data](/docs/anonymizing-customer-interview-data) for the mechanics.\n\n## What specific regimes require\n\n**GDPR** — storage limitation (Art. 5(1)(e)) plus the right to erasure (Art. 17). You must be able to find and delete one participant's data on request. That is an architecture requirement, not a policy statement: if you cannot locate everything about one person, you cannot comply. See [GDPR-compliant AI user research](/docs/gdpr-compliant-ai-user-research).\n\n**COPPA** — the amended Rule explicitly prohibits retaining a child's personal information indefinitely, limits retention to what is reasonably necessary for the collection purpose, and expects a written, published retention policy. See [research with children and teens](/docs/user-research-with-children-teens).\n\n**Illinois BIPA** — requires a **publicly available** written retention schedule and destruction guidelines for biometric identifiers, with destruction when the initial purpose is satisfied or within three years of the individual's last interaction, whichever comes first. The simplest compliance posture for a research team is to not create voiceprints at all — see [interview recording consent laws](/docs/interview-recording-consent-laws).\n\n**US state privacy laws** — the comprehensive laws now in effect across roughly 20 states generally require disclosing retention practices and honoring deletion requests, with sensitive categories attracting opt-in consent.\n\n**Sector and contract terms** frequently override all of the above. Healthcare, financial services, and enterprise DPAs often specify periods directly; your customer's contract may be stricter than any statute.\n\n## Handling a deletion request without losing the finding\n\nThe hard case: a participant asks you to delete their data six months after their quote landed in a report that shaped your roadmap.\n\nWhat must go: the recording, the transcript, identifiers, the linkage between person and response.\n\nWhat generally survives: **aggregate findings and genuinely de-identified insights.** Deleting a source does not obligate you to unlearn a conclusion. If four of twelve participants struggled with the same step, the finding \"a third of participants struggled here\" remains valid and contains no personal data.\n\nWhere teams get stuck is verbatim quotes in published decks. A distinctive quote can identify its speaker. Two ways to avoid the problem entirely:\n\n1. **De-identify quotes at synthesis time**, so nothing downstream needs revisiting.\n2. **Keep a quote-to-source map in one place only** — the research platform — so a deletion request has a single point of execution rather than a scavenger hunt across Notion, Slack, decks, and someone's laptop.\n\nThat second point is the real argument for keeping research data in one system. Sprawl is what makes deletion requests unanswerable.\n\n## Writing the policy: six elements\n\n1. **Scope** — which artifacts, which systems, which studies\n2. **A period per artifact tier**, with the reasoning recorded\n3. **The trigger** — is the clock from collection date, study close, or last participant interaction?\n4. **The mechanism** — automatic deletion beats manual cleanup, which never happens\n5. **Exceptions** — legal hold, active dispute, contractual requirement, and who authorizes them\n6. **An owner and a review cadence** — annually, with a named accountable person\n\nThen do the thing almost nobody does: **run a deletion drill.** Pick one past participant and try to delete everything about them. Whatever you cannot find is your actual compliance gap, and it is invariably a spreadsheet or a slide deck rather than the research platform.\n\n## How this works with Koji\n\n- **The transcript is the durable artifact, and the recording does not have to be.** Because every session yields a complete verbatim transcript, deleting raw audio early costs you very little analytically — which makes the shortest, highest-value retention tier genuinely practical.\n- **Structured questions produce findings that survive deletion.** With the six question types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, `yes_no` — quantitative results aggregate into distributions and rankings that carry no personal data. When a participant exercises erasure, your scale distributions and choice frequencies remain intact because they were never personal data to begin with. That is a real structural advantage over a pile of recordings, and the [structured questions guide](/docs/structured-questions-guide) covers building studies that way.\n- **Reports are synthesized outputs**, so the insight layer you keep long-term is separable from the identifiable layer you delete on schedule.\n- **Collect less up front.** The cheapest retention policy is a study scoped so it never gathers identifiers — no full names, no employer, no free-text field inviting people to volunteer personal details.\n- **Text mode removes the audio tier entirely** for sensitive studies. See [voice vs text interviews](/docs/voice-vs-text-interviews).\n- **Export before deletion.** Take the de-identified analysis into your repository so the schedule never costs you institutional memory — see the [research repository guide](/docs/research-repository-guide).\n\nConfirm current storage locations, sub-processors, and configurable retention settings during procurement, and record the answers in your policy — see [enterprise security for AI research platforms](/docs/enterprise-security-ai-research-platforms) and [AI interview data privacy and security](/docs/ai-interview-data-privacy-security).\n\n## Common mistakes\n\n1. **No written schedule.** The most common failure, and the one that turns any incident or audit into an unnecessary problem.\n2. **One blanket period for everything.** Different artifacts, different risk, different value.\n3. **Keeping raw recordings forever** because deleting feels irreversible. Transcribe, verify, then delete.\n4. **Confusing pseudonymization with anonymization.** Only genuine anonymization stops the clock.\n5. **Deleting consent records along with the data.** You need them to show your lawful basis.\n6. **Manual cleanup.** If deletion is a calendar reminder, it will not happen. Automate it.\n7. **Data sprawl across decks, spreadsheets, and Slack**, making erasure requests impossible to fulfill honestly.\n8. **Mixing incentive and payment records into the research corpus.** Different purpose, different clock, different system.\n9. **Never testing the policy.** Run the deletion drill.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — build studies whose findings survive deletion\n- [Anonymizing Customer Interview Data](/docs/anonymizing-customer-interview-data) — the de-identification mechanics that stop the clock\n- [GDPR-Compliant AI User Research](/docs/gdpr-compliant-ai-user-research) — storage limitation and erasure in practice\n- [Interview Recording Consent Laws](/docs/interview-recording-consent-laws) — why biometric artifacts need a published schedule\n- [User Research With Children and Teens](/docs/user-research-with-children-teens) — the stricter COPPA retention rule\n- [Research Repository Guide](/docs/research-repository-guide) — keep the insight after the data is gone\n- [Enterprise Security for AI Research Platforms](/docs/enterprise-security-ai-research-platforms) — what to verify during procurement\n- [IRB Approval for User Research](/docs/irb-approval-user-research) — the data-management plan reviewers expect","category":"Research Operations","lastModified":"2026-07-27T03:19:42.141406+00:00","metaTitle":"Research Data Retention and Deletion: How Long to Keep Interview Data","metaDescription":"No universal legal retention period exists - so having no schedule is the compliance failure. A tiered per-artifact schedule for recordings, transcripts, quotes and reports, plus how to handle deletion requests without losing insights.","keywords":["research data retention policy","how long to keep interview data","research data deletion schedule","gdpr storage limitation research","interview recording retention","participant data deletion request","research data management plan","anonymization retention clock","ux research data policy","deletion request research data"],"aiSummary":"There is no universal legal retention period for research data - GDPR Article 5(1)(e) storage limitation requires keeping personal data only as long as necessary for the collection purpose, which means the absence of a documented schedule is itself the compliance failure. Retention should be decided per artifact, not per study: raw audio 30-90 days after transcription, identifiable transcripts 6-12 months, de-identified transcripts and quotes 24-36 months or indefinitely if genuinely anonymized, analysis outputs indefinitely, consent records as long as the data plus the limitation period. Genuinely anonymized data falls outside GDPR so the retention clock stops, while pseudonymized data remains personal data. Key regimes: COPPA prohibits indefinite retention of children data, Illinois BIPA requires a publicly available destruction schedule for biometric identifiers.","aiPrerequisites":["An existing research practice that collects interviews, recordings, or survey responses","Basic familiarity with GDPR or comparable privacy obligations"],"aiLearningOutcomes":["Build a tiered per-artifact retention schedule instead of one blanket period","Understand why genuine anonymization stops the GDPR retention clock and pseudonymization does not","Handle a participant deletion request without discarding valid aggregate findings","Write a retention policy with the six required elements and test it with a deletion drill","Scope studies to collect fewer identifiers so retention obligations shrink"],"aiDifficulty":"intermediate","aiEstimatedTime":"15 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}