Back to docs
Research Operations

Nobody Logged the Deletion: Why Your Repository Cannot Show You What It Lost

A repository is the residue of removals that left no trace. Why you cannot query an absence, and the disposal register that makes removal visible.

Short answer: Your research repository looks like a record of what your company learned. It is actually the residue of a long series of removals that left no trace: studies never written up, studies appraised away, studies whose files quietly stopped opening, and studies someone deleted for a reason nobody wrote down. None of those events is visible from inside the repository, because a deleted record does not leave a gap you can search for. The fix is not to stop removing things. It is to make removal leave a mark: a disposal register, and a tombstone entry where the record used to be. Koji supports this by keeping the study-level record and its metadata even when raw data ages out, so the shape of what you no longer hold stays visible.

This is the closing article in a four-part series on treating a research repository as an archive. The first three each told you to make a decision: choose formats and migrate, appraise and discard, judge what may be reused. This one is about the property all three share. Every one of those decisions removes something, and none of them leaves evidence in the thing that remains.

The sliver, and the three things that made it one

Writing in Archivaria, Rodney Carter summarises Verne Harris's assessment that archives preserve "a sliver of a sliver of a sliver" of everything that was created. What matters more than the phrase is Carter's enumeration of exactly how the reduction happens. What reaches the archive is small, he writes,

due to the active and passive destruction by records creators, the appraisal by the archivist of what does manage to come to them, and through the physical (and even more alarming, the electronic or virtual) records' inevitable self-destruction.

Three filters. They map onto research with uncomfortable precision, and each has an article of its own.

Carter's filterThe research equivalentWhere it is covered
Destruction by records creatorsThe study that was run and never written up, or written up and never filedPublication bias and the file drawer
Appraisal by the archivistThe deliberate decision not to keep somethingMost of your research should not be kept
Records' inevitable self-destructionFormat obsolescence and context lossWill you be able to open your 2026 data in 2036

Each of those is a known, manageable problem. The capstone problem is what they have in common: after any of them operates, the repository looks exactly the same as it would have looked if the material had never existed.

You cannot query an absence

This is the structural point, and it is why the usual defences do not work.

Carter poses the question that has no easy answer: "how can one 'prove the absence of an archive?' Where does one begin to look? How do we begin to look for absences?" And he quotes Augustine on the one thin thread available: "we do not entirely forget what we remember that we have forgotten. If we had completely forgotten it, we should not even be able to look for what was lost."

Translate that into repository terms and it is brutally practical. Every search you run returns results from what is present. There is no query that returns the studies that were removed, because the removal took the queryable object with it. A researcher joining your team next year will search the repository, find nothing on a topic, and draw the only available conclusion: nobody has studied it. That conclusion will sometimes be wrong in the most expensive possible way, and nothing in the system will contradict it.

Carter's formulation of the general case is the sentence to keep: "What is present in the archives is defined by what is not."

This is a different problem from the two you already know about

Two well-documented biases sit next to this one, and it is worth being precise about the boundary, because the remedies differ.

Publication bias is about records that were never created. The study ran, the result was inconvenient or merely unexciting, and no write-up exists. The remedy is at the authoring stage: make write-up automatic.

Survivorship bias is about units that were never sampled. The churned customers did not answer, so the archive over-represents the satisfied. The remedy is at the recruitment stage.

The gap this article covers is neither. The record was created, it did enter the repository, and then a custodial act removed it or made it unusable. The remedy is not at authoring or recruitment. It is at the moment of removal, and it is the only one of the three that a repository owner controls directly.

Note also that this is not covered by your access audit trail. Access logging answers who read, exported, and shared data, which matters enormously for privacy and is covered in research data access controls and audit trails. It is designed to detect improper access, not to preserve the intellectual shape of the collection. A perfect access log can coexist with a repository that cannot tell you a single thing about what it used to hold.

The four silent removals

In practice, material leaves a research repository four ways, in roughly ascending order of how invisible they are.

1. Explicit deletion. Someone deletes a study. Usually legitimate, occasionally required, and in most tools it leaves nothing behind but a slightly shorter list.

2. Retention expiry. Data ages out automatically under policy. This is the best-behaved case, because at least the policy is written down somewhere, but the individual expiries are typically not recorded as events.

3. Appraisal without a record. Someone decides not to accession a study, or quietly stops maintaining one. Nothing is deleted; the material simply never becomes findable. This is the most common and the least visible.

4. Silent unreadability. The file survives, the format or the context does not. There was no decision at all, and consequently nothing to log. Only a fixity and format review will surface it, which is one of the arguments for having one.

There is a fifth case worth naming because it is genuinely ambiguous rather than merely undocumented. A checksum mismatch cannot distinguish corruption from an authorised transformation. Both change the fingerprint. Only a written record of the authorised change tells them apart, which means an archive that migrates files without logging migrations eventually cannot certify its own contents.

The disposal register

The archival answer is not to keep everything. It is to record the removal.

Public records regimes treat this as a first-order duty. The United States National Archives maintains a formal reporting channel for unauthorized dispositions of federal records, which tells you that the profession regards an undocumented removal as a distinct category of failure, separate from the question of whether the removal was justified.

A research disposal register is a single table. Seven columns, and it can live in whatever system you already use.

ColumnWhy it earns its place
WhatStudy title, original slug or identifier, date range of fieldwork
ScopeRoughly how many sessions, and with which population
What it was aboutOne sentence. This is the field that makes the register searchable and is the only reason the whole thing works
When removedDate of the removal
WhyRetention expiry, appraisal decision, participant deletion request, legal instruction, corruption
Under what authorityThe policy, the request, or the named decision-maker
What survivesWhether a summary, a finding, or an anonymised subset was retained, and where

The third column is the one to fight for. A register of identifiers is nearly useless; a register with one sentence of subject matter per row is a genuine finding aid for the collection's own absences. It turns Augustine's thin thread into an actual mechanism: your future colleague searching for a topic can find the record that a study on it once existed, even though the study is gone.

Tombstones, and the rule that makes them safe

A tombstone record is a stub that remains in place after the content is removed: the title, the date, the subject sentence, the reason, and no participant data.

This is where the practice has to be handled carefully, because two obligations pull against each other. If material was deleted because a participant exercised a right to erasure, the tombstone must not defeat that erasure. The rule that resolves it:

A tombstone records the existence and subject of a study. It never contains personal data, quotes, or anything identifying a participant.

Under that constraint, tombstones for retention expiry and appraisal decisions are straightforwardly good practice. For deletions driven by a participant request, keep the disposal register entry at study level for your own accountability, note that a request was actioned, and record no detail that could reconstruct what was erased. Our guides to research data retention and deletion and DSARs for research data cover the underlying obligations.

What this costs and what it buys

The cost is roughly two minutes per removal, and one recurring hour per quarter to review the register.

What it buys is specific, and it shows up in four places.

  • Search results stop lying. A search that returns a tombstone tells a researcher that the ground has been covered, which is exactly the information they need and exactly what an empty result set fails to convey.
  • Re-research becomes visible. The most expensive consequence of an undocumented removal is paying again for an answer you already bought. That cost is measurable, and the method is in the re-research audit.
  • Appraisal becomes reviewable. A decision that is recorded can be argued with. Recording who decided and why is also the guard against the failure mode where the small, unglamorous studies get quietly discarded because nobody senior asked for them.
  • Your archive can describe its own shape. This is the real prize. A repository that knows what it no longer holds can tell you where the evidence base is thin, which is the beginning of a research strategy rather than the end of a cleanup.

How Koji supports this

  • Study-level records persist independently of raw data. Because Koji separates the study, its brief, and its report from the underlying transcripts, the descriptive layer can survive when the interview data ages out. That is a tombstone by construction rather than by discipline.
  • The brief is the subject sentence. The research brief already contains the problem context and research question, so the register column that teams find hardest to maintain is populated from something that already exists.
  • Structured questions record what was asked. The six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - mean a tombstone can carry the question set even after answers are gone. Knowing that a study asked a particular scale question of a particular segment is often enough to prevent a duplicate.
  • Export before disposal is a routine, not a scramble. With koji_export_data through the API or MCP, exporting the brief, respondents, transcripts, and report summary as JSON is a scheduled job. Anything you export before removal is a removal you can reverse.
  • Anonymisation as an alternative to deletion. Frequently the right disposal is not deletion at all. Anonymising the data removes the liability that motivated the deletion while keeping the informational value, which is the outcome the disposal register will show you wanting more often than you expect.

Start here

Do not reconstruct history; you cannot, and that is the point of the article. Start from today.

  1. Create the seven-column register.
  2. Add one rule to your repository process: nothing is removed without a row.
  3. Set a quarterly hour to review the register alongside your appraisal decisions.
  4. Once a year, read the register on its own. The pattern of what you have been discarding is a report about your research programme that no other document produces.

Frequently asked questions

Is this not just an audit log?

No, and the distinction is practical. An audit log records access events - who viewed, exported, or shared data - and exists to detect improper access. A disposal register records removals and exists to preserve the intellectual shape of the collection. They serve different readers: the audit log serves a security reviewer, the register serves a researcher two years from now wondering whether a question has already been answered. Most teams have some form of the first and none of the second.

Does keeping tombstones conflict with the right to erasure?

Not if the tombstone is scoped correctly. Erasure obligations attach to personal data. A record stating that a study on a given topic existed, ran on given dates, and was removed under a given authority contains no personal data and does not undo the erasure. The rule to hold is absolute: no quotes, no participant identifiers, no reconstructable detail. Where a deletion arises from a participant request, keep the entry minimal and record that a request was actioned rather than what it contained.

How is this different from publication bias?

Publication bias concerns records that were never created - the study ran and no write-up exists. This concerns records that were created, entered the repository, and were later removed or became unusable. The remedies sit at different points in the lifecycle: publication bias is fixed at authoring, this is fixed at removal. Both produce the same symptom, which is an evidence base that looks more complete than it is.

Who should own the register?

The same person or team that owns the repository and the appraisal policy, because the register is mostly a byproduct of appraisal decisions they are already making. It should not be owned by legal or security, both of whom have adjacent registers with different purposes and will reasonably scope this one to their own needs.

What if our tool deletes things without telling us?

Then that is a finding worth knowing, and the way to discover it is a periodic reconciliation: compare a list of study identifiers captured at one point against the same list a quarter later, and investigate anything that vanished without a register row. This is also a strong argument for holding your own exported copy, since a copy you control cannot be removed by a system you do not.

Is there a version of this that is worth doing in ten minutes?

Yes. Open a spreadsheet, add the seven columns, and populate it going forward only. A register that starts today and is honestly maintained is worth far more than a reconstruction project that stalls at the halfway point. The absences before today are already unrecoverable, which is the whole argument for starting now.

Related Resources

Related Articles

Most of Your Research Should Not Be Kept: Archival Appraisal for Research Repositories

Keeping everything is a decision with victims. How archivists decide what has enduring value, and how to run the same appraisal on your research repository.

Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)

Publication bias is not an academic curiosity. In product research it is worse, because nobody rejects your null study - you simply never write it up. Learn how big the file drawer is, what it does to your confidence, and how to build a study register that closes it.

The Re-Research Audit: How Much of Your Budget Buys an Answer You Already Own

Count how many of your last twenty studies answered a question you already owned. The protocol, the four causes, and where the duty belongs.

Research Provenance: How to Prove an Interview, a Quote, or a Report Is Genuine (2026)

When any text can be generated, a customer quote proves nothing on its own. Content Credentials, the EU AI Act marking rules, and the hash-anchored capture method that actually works for research text.

Research Data Access Controls and Audit Trails: Who Can See Your Interview Data

Your vendor's SOC 2 report proves the vendor is secure. It says nothing about which colleague opened a raw transcript last Tuesday. Here is how to build access tiers, audit trails and access reviews for research data — and what an auditor will actually ask you for.

Research Data Retention and Deletion: How Long Should You Keep Interview Data?

There is no universal legal number - which is exactly why having no retention schedule is itself the compliance failure. A tiered, per-artifact schedule for recordings, transcripts, quotes, and reports, plus how to handle deletion requests without losing your insights.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.