{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-21T09:59:07.817Z"},"content":[{"type":"documentation","id":"84ba2dd8-3076-4df6-9bf3-9dedb72ecce5","slug":"undocumented-deletion-research-repository","title":"Nobody Logged the Deletion: Why Your Repository Cannot Show You What It Lost","url":"https://www.koji.so/docs/undocumented-deletion-research-repository","summary":"Research repositories are the residue of removals that leave no trace: studies never written up, studies appraised away, files that quietly stopped opening, and deletions nobody recorded. Because a removed record takes the queryable object with it, an absence cannot be searched for and an empty result set is indistinguishable from uncovered ground. The remedy is a seven-column disposal register plus tombstone records that carry the subject of a study but never personal data. This is distinct from publication bias (records never created), survivorship bias (units never sampled), and access audit trails (who read what).","content":"**Short answer:** Your research repository looks like a record of what your company learned. It is actually the residue of a long series of removals that left no trace: studies never written up, studies appraised away, studies whose files quietly stopped opening, and studies someone deleted for a reason nobody wrote down. None of those events is visible from inside the repository, because a deleted record does not leave a gap you can search for. The fix is not to stop removing things. It is to make removal leave a mark: a disposal register, and a tombstone entry where the record used to be. Koji supports this by keeping the study-level record and its metadata even when raw data ages out, so the shape of what you no longer hold stays visible.\n\nThis is the closing article in a four-part series on treating a research repository as an archive. The first three each told you to make a decision: choose formats and migrate, appraise and discard, judge what may be reused. This one is about the property all three share. Every one of those decisions removes something, and none of them leaves evidence in the thing that remains.\n\n## The sliver, and the three things that made it one\n\nWriting in *Archivaria*, Rodney Carter summarises Verne Harris's assessment that archives preserve \"a sliver of a sliver of a sliver\" of everything that was created. What matters more than the phrase is Carter's enumeration of exactly [how the reduction happens](https://archivaria.ca/index.php/archivaria/article/view/12541/13687). What reaches the archive is small, he writes,\n\n> due to the active and passive destruction by records creators, the appraisal by the archivist of what does manage to come to them, and through the physical (and even more alarming, the electronic or virtual) records' inevitable self-destruction.\n\nThree filters. They map onto research with uncomfortable precision, and each has an article of its own.\n\n| Carter's filter | The research equivalent | Where it is covered |\n| --- | --- | --- |\n| Destruction by records creators | The study that was run and never written up, or written up and never filed | [Publication bias and the file drawer](/docs/publication-bias-product-research) |\n| Appraisal by the archivist | The deliberate decision not to keep something | [Most of your research should not be kept](/docs/archival-appraisal-research-what-to-keep) |\n| Records' inevitable self-destruction | Format obsolescence and context loss | [Will you be able to open your 2026 data in 2036](/docs/research-data-preservation-format-obsolescence) |\n\nEach of those is a known, manageable problem. The capstone problem is what they have in common: **after any of them operates, the repository looks exactly the same as it would have looked if the material had never existed.**\n\n## You cannot query an absence\n\nThis is the structural point, and it is why the usual defences do not work.\n\nCarter poses the question that has no easy answer: \"how can one 'prove the absence of an archive?' Where does one begin to look? How do we begin to look for absences?\" And he quotes Augustine on the one thin thread available: \"we do not entirely forget what we remember that we have forgotten. If we had completely forgotten it, we should not even be able to look for what was lost.\"\n\nTranslate that into repository terms and it is brutally practical. Every search you run returns results from what is present. There is no query that returns the studies that were removed, because the removal took the queryable object with it. A researcher joining your team next year will search the repository, find nothing on a topic, and draw the only available conclusion: nobody has studied it. That conclusion will sometimes be wrong in the most expensive possible way, and nothing in the system will contradict it.\n\nCarter's formulation of the general case is the sentence to keep: \"What is present in the archives is defined by what is not.\"\n\n## This is a different problem from the two you already know about\n\nTwo well-documented biases sit next to this one, and it is worth being precise about the boundary, because the remedies differ.\n\n**Publication bias** is about records that were never created. The study ran, the result was inconvenient or merely unexciting, and no write-up exists. The remedy is at the authoring stage: make write-up automatic.\n\n**Survivorship bias** is about units that were never sampled. The churned customers did not answer, so the archive over-represents the satisfied. The remedy is at the recruitment stage.\n\nThe gap this article covers is neither. The record *was* created, it *did* enter the repository, and then a custodial act removed it or made it unusable. The remedy is not at authoring or recruitment. It is at the moment of removal, and it is the only one of the three that a repository owner controls directly.\n\nNote also that this is not covered by your access audit trail. Access logging answers who read, exported, and shared data, which matters enormously for privacy and is covered in [research data access controls and audit trails](/docs/research-data-access-controls-audit-trail). It is designed to detect improper access, not to preserve the intellectual shape of the collection. A perfect access log can coexist with a repository that cannot tell you a single thing about what it used to hold.\n\n## The four silent removals\n\nIn practice, material leaves a research repository four ways, in roughly ascending order of how invisible they are.\n\n**1. Explicit deletion.** Someone deletes a study. Usually legitimate, occasionally required, and in most tools it leaves nothing behind but a slightly shorter list.\n\n**2. Retention expiry.** Data ages out automatically under policy. This is the best-behaved case, because at least the policy is written down somewhere, but the individual expiries are typically not recorded as events.\n\n**3. Appraisal without a record.** Someone decides not to accession a study, or quietly stops maintaining one. Nothing is deleted; the material simply never becomes findable. This is the most common and the least visible.\n\n**4. Silent unreadability.** The file survives, the format or the context does not. There was no decision at all, and consequently nothing to log. Only a fixity and format review will surface it, which is one of the arguments for having one.\n\nThere is a fifth case worth naming because it is genuinely ambiguous rather than merely undocumented. A checksum mismatch cannot distinguish corruption from an authorised transformation. Both change the fingerprint. Only a written record of the authorised change tells them apart, which means an archive that migrates files without logging migrations eventually cannot certify its own contents.\n\n## The disposal register\n\nThe archival answer is not to keep everything. It is to record the removal.\n\nPublic records regimes treat this as a first-order duty. The United States National Archives maintains a formal reporting channel for [unauthorized dispositions](https://www.archives.gov/records-mgmt/appraisal) of federal records, which tells you that the profession regards an undocumented removal as a distinct category of failure, separate from the question of whether the removal was justified.\n\nA research disposal register is a single table. Seven columns, and it can live in whatever system you already use.\n\n| Column | Why it earns its place |\n| --- | --- |\n| **What** | Study title, original slug or identifier, date range of fieldwork |\n| **Scope** | Roughly how many sessions, and with which population |\n| **What it was about** | One sentence. This is the field that makes the register searchable and is the only reason the whole thing works |\n| **When removed** | Date of the removal |\n| **Why** | Retention expiry, appraisal decision, participant deletion request, legal instruction, corruption |\n| **Under what authority** | The policy, the request, or the named decision-maker |\n| **What survives** | Whether a summary, a finding, or an anonymised subset was retained, and where |\n\nThe third column is the one to fight for. A register of identifiers is nearly useless; a register with one sentence of subject matter per row is a genuine finding aid for the collection's own absences. It turns Augustine's thin thread into an actual mechanism: your future colleague searching for a topic can find the record that a study on it once existed, even though the study is gone.\n\n## Tombstones, and the rule that makes them safe\n\nA **tombstone record** is a stub that remains in place after the content is removed: the title, the date, the subject sentence, the reason, and no participant data.\n\nThis is where the practice has to be handled carefully, because two obligations pull against each other. If material was deleted because a participant exercised a right to erasure, the tombstone must not defeat that erasure. The rule that resolves it:\n\n> A tombstone records the existence and subject of a study. It never contains personal data, quotes, or anything identifying a participant.\n\nUnder that constraint, tombstones for retention expiry and appraisal decisions are straightforwardly good practice. For deletions driven by a participant request, keep the disposal register entry at study level for your own accountability, note that a request was actioned, and record no detail that could reconstruct what was erased. Our guides to [research data retention and deletion](/docs/research-data-retention-deletion) and [DSARs for research data](/docs/dsar-research-data) cover the underlying obligations.\n\n## What this costs and what it buys\n\nThe cost is roughly two minutes per removal, and one recurring hour per quarter to review the register.\n\nWhat it buys is specific, and it shows up in four places.\n\n- **Search results stop lying.** A search that returns a tombstone tells a researcher that the ground has been covered, which is exactly the information they need and exactly what an empty result set fails to convey.\n- **Re-research becomes visible.** The most expensive consequence of an undocumented removal is paying again for an answer you already bought. That cost is measurable, and the method is in [the re-research audit](/docs/re-research-audit-duplicate-studies).\n- **Appraisal becomes reviewable.** A decision that is recorded can be argued with. Recording who decided and why is also the guard against the failure mode where the small, unglamorous studies get quietly discarded because nobody senior asked for them.\n- **Your archive can describe its own shape.** This is the real prize. A repository that knows what it no longer holds can tell you where the evidence base is thin, which is the beginning of a research strategy rather than the end of a cleanup.\n\n## How Koji supports this\n\n- **Study-level records persist independently of raw data.** Because Koji separates the study, its brief, and its report from the underlying transcripts, the descriptive layer can survive when the interview data ages out. That is a tombstone by construction rather than by discipline.\n- **The brief is the subject sentence.** The research brief already contains the problem context and research question, so the register column that teams find hardest to maintain is populated from something that already exists.\n- **Structured questions record what was asked.** The six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - mean a tombstone can carry the question set even after answers are gone. Knowing that a study asked a particular scale question of a particular segment is often enough to prevent a duplicate.\n- **Export before disposal is a routine, not a scramble.** With `koji_export_data` through the API or MCP, exporting the brief, respondents, transcripts, and report summary as JSON is a scheduled job. Anything you export before removal is a removal you can reverse.\n- **Anonymisation as an alternative to deletion.** Frequently the right disposal is not deletion at all. [Anonymising the data](/docs/anonymizing-customer-interview-data) removes the liability that motivated the deletion while keeping the informational value, which is the outcome the disposal register will show you wanting more often than you expect.\n\n## Start here\n\nDo not reconstruct history; you cannot, and that is the point of the article. Start from today.\n\n1. Create the seven-column register.\n2. Add one rule to your repository process: nothing is removed without a row.\n3. Set a quarterly hour to review the register alongside your appraisal decisions.\n4. Once a year, read the register on its own. The pattern of what you have been discarding is a report about your research programme that no other document produces.\n\n## Frequently asked questions\n\n### Is this not just an audit log?\n\nNo, and the distinction is practical. An audit log records access events - who viewed, exported, or shared data - and exists to detect improper access. A disposal register records removals and exists to preserve the intellectual shape of the collection. They serve different readers: the audit log serves a security reviewer, the register serves a researcher two years from now wondering whether a question has already been answered. Most teams have some form of the first and none of the second.\n\n### Does keeping tombstones conflict with the right to erasure?\n\nNot if the tombstone is scoped correctly. Erasure obligations attach to personal data. A record stating that a study on a given topic existed, ran on given dates, and was removed under a given authority contains no personal data and does not undo the erasure. The rule to hold is absolute: no quotes, no participant identifiers, no reconstructable detail. Where a deletion arises from a participant request, keep the entry minimal and record that a request was actioned rather than what it contained.\n\n### How is this different from publication bias?\n\nPublication bias concerns records that were never created - the study ran and no write-up exists. This concerns records that were created, entered the repository, and were later removed or became unusable. The remedies sit at different points in the lifecycle: publication bias is fixed at authoring, this is fixed at removal. Both produce the same symptom, which is an evidence base that looks more complete than it is.\n\n### Who should own the register?\n\nThe same person or team that owns the repository and the appraisal policy, because the register is mostly a byproduct of appraisal decisions they are already making. It should not be owned by legal or security, both of whom have adjacent registers with different purposes and will reasonably scope this one to their own needs.\n\n### What if our tool deletes things without telling us?\n\nThen that is a finding worth knowing, and the way to discover it is a periodic reconciliation: compare a list of study identifiers captured at one point against the same list a quarter later, and investigate anything that vanished without a register row. This is also a strong argument for holding your own exported copy, since a copy you control cannot be removed by a system you do not.\n\n### Is there a version of this that is worth doing in ten minutes?\n\nYes. Open a spreadsheet, add the seven columns, and populate it going forward only. A register that starts today and is honestly maintained is worth far more than a reconstruction project that stalls at the halfway point. The absences before today are already unrecoverable, which is the whole argument for starting now.\n\n## Related Resources\n\n- [Structured Questions: The Complete Guide](/docs/structured-questions-guide) - the six question types a tombstone can preserve after the answers are gone.\n- [Research Data Retention and Deletion](/docs/research-data-retention-deletion) - the policy that drives most scheduled removals.\n- [Research Data Access Controls and Audit Trails](/docs/research-data-access-controls-audit-trail) - the adjacent log, and what it does not cover.\n- [Publication Bias and the File-Drawer Problem](/docs/publication-bias-product-research) - the records that were never created.\n- [The Re-Research Audit](/docs/re-research-audit-duplicate-studies) - measuring what undocumented removal actually costs.\n- [Research Provenance and Authenticity](/docs/research-artifact-provenance-authenticity) - proving that what remains is genuine.\n","category":"Research Operations","lastModified":"2026-08-21T03:25:48.491759+00:00","metaTitle":"Undocumented Deletion: Why a Research Repository Cannot Show What It Lost (2026)","metaDescription":"Deleted studies leave no searchable gap. How a disposal register and tombstone records keep the shape of your research archive visible after removal.","keywords":["research repository deletion","disposal register","tombstone record archive","research data destruction log","repository governance deletion","archival silence research"],"aiSummary":"Research repositories are the residue of removals that leave no trace: studies never written up, studies appraised away, files that quietly stopped opening, and deletions nobody recorded. Because a removed record takes the queryable object with it, an absence cannot be searched for and an empty result set is indistinguishable from uncovered ground. The remedy is a seven-column disposal register plus tombstone records that carry the subject of a study but never personal data. This is distinct from publication bias (records never created), survivorship bias (units never sampled), and access audit trails (who read what)."}],"pagination":{"total":1,"returned":1,"offset":0}}