Research Data Retention and Deletion: How Long Should You Keep Interview Data?
There is no universal legal number - which is exactly why having no retention schedule is itself the compliance failure. A tiered, per-artifact schedule for recordings, transcripts, quotes, and reports, plus how to handle deletion requests without losing your insights.
The short answer
There is no universal legal retention period for research data. The law requires you to decide, document, and enforce one — so the absence of a schedule is itself the compliance failure.
Under GDPR's storage limitation principle (Article 5(1)(e)), personal data may be kept only as long as necessary for the purpose it was collected for. That is a relative standard: it obliges you to define the purpose, derive a period from it, and then actually delete on schedule.
The second key idea, which most teams miss:
Retention is a per-artifact decision, not a per-study decision. A single 45-minute interview generates a recording, a transcript, a set of coded themes, a handful of verbatim quotes, a consent record, and a contact record for paying the incentive. Those six artifacts have wildly different risk profiles and wildly different useful lifespans. One blanket rule gets it wrong in both directions at once — holding raw audio far too long while deleting the analysis you actually needed.
Not legal advice. Retention periods interact with your jurisdiction, sector, and contracts; validate your schedule with counsel.
The recommended tiered schedule
Adapt the periods, keep the structure. These are defensible defaults, not legal minimums.
| Artifact | Suggested retention | Why |
|---|---|---|
| Raw audio / video | 30–90 days after transcription | Highest-risk, lowest marginal value once transcribed. The single biggest easy win. |
| Identifiable transcripts | 6–12 months | Long enough to re-analyze and verify; short enough to limit exposure. |
| De-identified transcripts and quotes | 24–36 months, or indefinitely if genuinely anonymized | Where nearly all analytical value lives. |
| Analysis outputs — themes, reports, aggregate scores | Indefinite | Business records containing no personal data. |
| Consent records | As long as you hold the data, plus your limitation period | This is your evidence of lawful basis. Do not delete it with the data. |
| Contact details and incentive/payment records | Per finance and tax rules, in a separate system | Different purpose, different clock — never in the research corpus. |
| Voiceprints or other biometric identifiers | Do not collect for research | If collected, a published destruction schedule is mandatory — see below. |
The recording tier deserves emphasis. Raw audio is the most sensitive artifact you hold and, once you have an accurate transcript, usually the least useful. Teams keep it out of vague anxiety about needing to "go back to the tape." In practice they almost never do — and meanwhile the recording is the thing that turns a routine security incident into a serious one.
Anonymization is the lever that lets you keep insight forever
This is the most valuable mechanic in retention design, and it is widely misunderstood.
- Pseudonymized data — real identifiers swapped for codes, with a key that still exists somewhere — is still personal data. The retention clock keeps running.
- Genuinely anonymized data — where re-identification is no longer reasonably possible, and no key exists — falls outside GDPR's scope. The clock stops.
So the way to retain research value indefinitely without indefinite risk is not to argue for longer retention periods. It is to de-identify at the point of synthesis, so that your durable artifacts — reports, theme libraries, quote banks — never contain personal data in the first place.
Be honest about the bar, though. Qualitative data resists anonymization more than survey data, because narrative detail identifies people. "The VP of Engineering at a 40-person Berlin fintech who joined last March" is identifiable no matter what you call them. Strip role-plus-company-plus-timeline combinations, not just names. See anonymizing customer interview data for the mechanics.
What specific regimes require
GDPR — storage limitation (Art. 5(1)(e)) plus the right to erasure (Art. 17). You must be able to find and delete one participant's data on request. That is an architecture requirement, not a policy statement: if you cannot locate everything about one person, you cannot comply. See GDPR-compliant AI user research.
COPPA — the amended Rule explicitly prohibits retaining a child's personal information indefinitely, limits retention to what is reasonably necessary for the collection purpose, and expects a written, published retention policy. See research with children and teens.
Illinois BIPA — requires a publicly available written retention schedule and destruction guidelines for biometric identifiers, with destruction when the initial purpose is satisfied or within three years of the individual's last interaction, whichever comes first. The simplest compliance posture for a research team is to not create voiceprints at all — see interview recording consent laws.
US state privacy laws — the comprehensive laws now in effect across roughly 20 states generally require disclosing retention practices and honoring deletion requests, with sensitive categories attracting opt-in consent.
Sector and contract terms frequently override all of the above. Healthcare, financial services, and enterprise DPAs often specify periods directly; your customer's contract may be stricter than any statute.
Handling a deletion request without losing the finding
The hard case: a participant asks you to delete their data six months after their quote landed in a report that shaped your roadmap.
What must go: the recording, the transcript, identifiers, the linkage between person and response.
What generally survives: aggregate findings and genuinely de-identified insights. Deleting a source does not obligate you to unlearn a conclusion. If four of twelve participants struggled with the same step, the finding "a third of participants struggled here" remains valid and contains no personal data.
Where teams get stuck is verbatim quotes in published decks. A distinctive quote can identify its speaker. Two ways to avoid the problem entirely:
- De-identify quotes at synthesis time, so nothing downstream needs revisiting.
- Keep a quote-to-source map in one place only — the research platform — so a deletion request has a single point of execution rather than a scavenger hunt across Notion, Slack, decks, and someone's laptop.
That second point is the real argument for keeping research data in one system. Sprawl is what makes deletion requests unanswerable.
Writing the policy: six elements
- Scope — which artifacts, which systems, which studies
- A period per artifact tier, with the reasoning recorded
- The trigger — is the clock from collection date, study close, or last participant interaction?
- The mechanism — automatic deletion beats manual cleanup, which never happens
- Exceptions — legal hold, active dispute, contractual requirement, and who authorizes them
- An owner and a review cadence — annually, with a named accountable person
Then do the thing almost nobody does: run a deletion drill. Pick one past participant and try to delete everything about them. Whatever you cannot find is your actual compliance gap, and it is invariably a spreadsheet or a slide deck rather than the research platform.
How this works with Koji
- The transcript is the durable artifact, and the recording does not have to be. Because every session yields a complete verbatim transcript, deleting raw audio early costs you very little analytically — which makes the shortest, highest-value retention tier genuinely practical.
- Structured questions produce findings that survive deletion. With the six question types —
open_ended,scale,single_choice,multiple_choice,ranking,yes_no— quantitative results aggregate into distributions and rankings that carry no personal data. When a participant exercises erasure, your scale distributions and choice frequencies remain intact because they were never personal data to begin with. That is a real structural advantage over a pile of recordings, and the structured questions guide covers building studies that way. - Reports are synthesized outputs, so the insight layer you keep long-term is separable from the identifiable layer you delete on schedule.
- Collect less up front. The cheapest retention policy is a study scoped so it never gathers identifiers — no full names, no employer, no free-text field inviting people to volunteer personal details.
- Text mode removes the audio tier entirely for sensitive studies. See voice vs text interviews.
- Export before deletion. Take the de-identified analysis into your repository so the schedule never costs you institutional memory — see the research repository guide.
Confirm current storage locations, sub-processors, and configurable retention settings during procurement, and record the answers in your policy — see enterprise security for AI research platforms and AI interview data privacy and security.
Common mistakes
- No written schedule. The most common failure, and the one that turns any incident or audit into an unnecessary problem.
- One blanket period for everything. Different artifacts, different risk, different value.
- Keeping raw recordings forever because deleting feels irreversible. Transcribe, verify, then delete.
- Confusing pseudonymization with anonymization. Only genuine anonymization stops the clock.
- Deleting consent records along with the data. You need them to show your lawful basis.
- Manual cleanup. If deletion is a calendar reminder, it will not happen. Automate it.
- Data sprawl across decks, spreadsheets, and Slack, making erasure requests impossible to fulfill honestly.
- Mixing incentive and payment records into the research corpus. Different purpose, different clock, different system.
- Never testing the policy. Run the deletion drill.
Related Resources
- Structured Questions Guide — build studies whose findings survive deletion
- Anonymizing Customer Interview Data — the de-identification mechanics that stop the clock
- GDPR-Compliant AI User Research — storage limitation and erasure in practice
- Interview Recording Consent Laws — why biometric artifacts need a published schedule
- User Research With Children and Teens — the stricter COPPA retention rule
- Research Repository Guide — keep the insight after the data is gone
- Enterprise Security for AI Research Platforms — what to verify during procurement
- IRB Approval for User Research — the data-management plan reviewers expect
Related Articles
AI Interview Data Privacy & Security: A Buyer's Evaluation Guide
How to evaluate the privacy and security of an AI customer research platform — the questions to ask about data handling, PII, retention, sub-processors, and compliance — plus how Koji approaches each one.
Anonymizing Customer Interview Data: A Practical Guide for Privacy-Safe Research
Five operational techniques for handling PII in AI customer interviews — from intake-time anonymization to stakeholder-safe quote sharing — without sacrificing research signal.
DPIA for User Research: When You Need One and How to Write It (2026)
A practical guide to Data Protection Impact Assessments for customer and user research: the Article 35 triggers, the WP29 nine criteria, what belongs in each section, and a worked example for AI-moderated interviews.
Enterprise Security for AI Customer Research Platforms: SOC 2, SSO, and Vendor Review
A procurement-ready guide to evaluating the security of an AI customer research platform — SOC 2, encryption, SSO/SAML, data residency, sub-processors, and the questions your security team should ask.
GDPR-Compliant AI User Research: A Practical Guide
How to run AI-moderated customer interviews under GDPR. Lawful basis, consent flows, data minimization, retention, sub-processors, and how Koji handles each requirement.
Interview Recording Consent Laws: One-Party, All-Party, and Biometric Rules (2026)
Federal law allows one-party consent recording, but roughly a dozen US states require all-party consent - and biometric laws like Illinois BIPA add a separate written-consent duty for voiceprints. Here is how research teams stay on the safe side of both.
Do You Need IRB Approval for User Research? A 2026 Decision Guide
Most commercial UX and product research does not require IRB approval - but four specific situations flip the answer to yes. Here is the actual regulatory test, the exempt categories, and how to prepare a submission that clears review fast.
Research Data Residency and International Transfers: Where Your Interview Data Actually Lives
Data residency, sovereignty, and localisation explained for research teams — GDPR Chapter V transfer mechanisms, transfer impact assessments, the DPF's current status, and the questions to ask any research vendor.
How to Build a UX Research Repository: The Complete Guide
A research repository transforms scattered insights into a searchable organizational asset. Learn how to build one that teams actually use.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
User Research With Children and Teens: COPPA, Parental Consent, and Assent
Researching under-13s triggers COPPA verifiable parental consent - including a separate consent before any child data trains AI. Here is the compliance path, the parent-mediated pattern most teams should use instead, and how to design sessions that actually work with young participants.