Most of Your Research Should Not Be Kept: Archival Appraisal for Research Repositories
Keeping everything is a decision with victims. How archivists decide what has enduring value, and how to run the same appraisal on your research repository.
Short answer: Keeping everything is not a neutral, safe default. It is a decision that quietly denies real care to the small share of your research that deserves it, because preservation work costs the same per item whether the item matters or not. Archivists solved this a century ago with appraisal: the deliberate judgement of which records have enduring value, made explicitly and recorded. Less than five percent of records created by United States federal agencies are ever appraised as permanent. Your research repository is not more valuable than the federal record. Deciding what to keep is the skill, and Koji makes it easier by attaching the brief, the question set, and the quality signal to every study, so value is visible without reopening the data.
This article is the counterweight to the previous one. Choosing durable formats, checking fixity, and capturing context are all necessary, and every one of them has a per-item price. Apply them to everything and the budget runs out before the work starts. That is why teams with a preservation plan and no appraisal policy end up preserving nothing well.
Keeping everything is a choice, and it has victims
The instinct is understandable. Storage is cheap, deletion feels irreversible, and nobody was ever fired for keeping a transcript. So the repository accumulates: every study, every pilot, every abandoned screener, every duplicate export someone made for a deck.
Three costs follow, and none of them appear on a storage bill.
Retrieval cost. Every low-value item is noise in every future search. The signal-to-noise ratio of a repository is set by what you declined to add.
Care cost. Fixity checks, format migration, and context capture consume attention that is finite. Spread evenly across a pile that is ninety percent disposable, each genuinely important study gets a tenth of the attention it needed.
Trust cost. A repository where most results are stale, duplicated, or unattributable trains people not to search it. Once that habit sets in, the good material inside it might as well not exist.
The archival profession settled this argument long ago. Hilary Jenkinson held that archivists ought not be in the business of destroying records at all. The Society of American Archivists Dictionary records the modern verdict, quoting Eastwood: that view "has received, as might be expected, almost universal condemnation by archivists who routinely conduct appraisal." The same entry quotes Brichford's 1977 formulation: "The most significant archival function is the appraisal or evaluation of the mass of source material and the selection of that portion that will be kept."
The most significant function. Not storage, not description, not search. Selection.
Appraisal is not your retention schedule
These get conflated constantly, and the confusion is expensive because it lets a compliance document stand in for a judgement nobody has made.
The SAA dictionary is explicit that appraisal "is distinguished from evaluation, which is typically used by records managers to indicate a preliminary assessment of value based on existing retention schedules."
- Retention answers: how long are we permitted and required to hold this? It is a legal floor and ceiling, driven by privacy law, consent terms, and contractual commitments. Our guide to research data retention and deletion covers it properly.
- Appraisal answers: of the material we are permitted to keep, which deserves it? That is a value judgement about future usefulness, and no regulation will make it for you.
A retention schedule that says "keep interview data for three years" tells you nothing about whether a given study should be curated, indexed, and carried forward, or simply allowed to age out. Most teams have the first document and have never written the second.
The two values every study has
The United States National Archives operates a clean two-perspective model in its guidance on record values, and it transfers to research almost without translation.
Business use value is the value to the unit that created the record. NARA splits it three ways: administrative (usefulness in conducting the work), fiscal (documenting financial obligations), and legal (protecting rights and interests). For a research study, this is the decision it was commissioned for. It is high on the day the readout happens and it decays fast, often to near zero once the feature ships.
Archival value is "the potential for ongoing use of the records" after the original business use ends. NARA divides it into evidential value, which documents the organisation's own functions and significant activities, and informational value, which documents the persons, places, and matters the organisation dealt with.
The translation is direct and useful:
- Evidential value in research is what the study proves about how your company decided things. Why did we build it this way? What did customers tell us before the pivot? This is what a new head of product, an acquirer's diligence team, or a regulator asks for.
- Informational value in research is what the study says about the customers themselves, independent of the decision it informed. Segment vocabulary, workflow descriptions, purchase criteria, the way people actually describe the problem in their own words. This is the reusable material.
Note where the mismatch lies. Teams appraise on business use value, because that is what they felt when the study landed. Almost all durable worth is in the other two.
The five analyses: a real appraisal kit
Gerald Ham's formulation, also recorded in the SAA dictionary, names five analyses archivists use to identify records of enduring value. Run them as five questions per study and appraisal stops being vague.
- Functional. Who made this and for what purpose? A study commissioned to justify a decision already taken has different value from one commissioned to make it.
- Informational. How significant and how good is the content? Ten thin sessions with a mis-specified screener are not equivalent to ten strong ones, and quality is a legitimate appraisal criterion.
- Contextual. How does this sit against parallel or related sources? Ham's test asks whether the information exists elsewhere. Three studies covering the same ground do not each carry full value; the best one carries most of it.
- Use. What use is likely to be made of this, and what physical, legal, or intellectual limits apply to access? A study whose consent terms forbid reuse has low archival value regardless of content quality.
- Cost. What does preserving this cost, weighed against the benefit of retaining the information? Ham puts the economics inside the framework rather than treating it as an afterthought.
Question four is the one research teams skip and then regret. Consent language, participant anonymity commitments, and NDA terms all constrain future use, and they are far cheaper to check at appraisal time than at the moment someone wants to reuse the data. Where reuse matters, anonymising the interview data is often what converts a study from a three-year liability into a permanent asset.
Two levels, and the one most teams get wrong
Bass distinguishes macro-level appraisal, which judges significance against a collecting policy, from micro-level appraisal, which separates the wheat from the chaff inside a collection.
Research teams almost always do micro and never macro. They argue about whether a particular transcript is worth keeping while never having written down what the repository is for. Macro first: a one-paragraph collecting policy stating which research this repository exists to hold. Everything else becomes tractable once that exists, because appraisal is a comparison against a standard and without the policy there is no standard.
There is a documented failure mode here worth naming. The dictionary quotes Duranti's warning that top-down appraisal, driven by the importance of the creator's mandate, "has proved to be unsatisfactory because it excludes the 'powerless transactions,' which might throw light on the broader social context." The research translation is precise: appraise only by seniority of the requesting stakeholder and you will systematically discard the small, unglamorous studies about workarounds, edge cases, and the customers nobody was building for. Those are frequently the studies that turn out to matter. Guard against it by appraising on evidence quality and informational value, not on who asked.
A practical appraisal policy
Four tiers, reviewed quarterly, and it takes an hour once the collecting policy exists.
| Tier | What qualifies | Treatment |
|---|---|---|
| Permanent | Studies that anchor a strategic decision, define a segment, or establish a baseline you will measure against later | Full curation: study-level description, preserved raw data, fixity, indefinite keeping subject to consent |
| Long-term | Solid studies with clear informational value about customers, still likely to inform work | Indexed and searchable, raw data kept, reviewed at three years |
| Short-term | Studies tied to a shipped decision with little reuse potential; routine validation | Keep the summary and the finding, let the raw data age out on the retention schedule |
| Not accessioned | Pilots, aborted studies, duplicate exports, sessions that failed quality screening | Never enters the repository; disposed under policy and recorded |
The fourth tier is the one that pays. Most repository bloat is material that was never worth accessioning in the first place, and the cheapest possible moment to appraise a study is before it is added, not three years later when nobody remembers it.
Reappraisal is legitimate. Archives revisit earlier decisions and deaccession material that no longer merits keeping. So should you, particularly for studies about products that no longer exist. The one requirement is that the decision is recorded, which is the subject of the companion article on undocumented deletion in this series.
Where Koji makes appraisal cheap
Appraisal is expensive when judging a study means reopening it. Koji attaches most of the evidence you need to the study itself.
- The brief is the functional analysis. Because every Koji study is generated from a research brief with the problem context, the methodology, and the question set, Ham's first question is answered without opening a transcript. You can see who asked and why.
- Quality scoring is the informational analysis. Koji scores conversation quality, and only conversations meeting the quality bar consume credits. A study whose sessions scored poorly is visible as low-value at a glance, rather than after an hour of reading.
- Structured questions make content comparable. The six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - mean a study's quantitative spine can be compared against a neighbouring study directly, which is what Ham's contextual analysis requires. Two studies that asked the same scale question of different segments are visibly complementary; two that asked it of the same segment are visibly redundant.
- Re-running beats hoarding. This is the part traditional research economics got wrong. When a study takes six weeks and a moderator's calendar, keeping everything forever is rational insurance. When AI-moderated interviews return a fresh study in days without a moderator, the calculus flips: it is often cheaper to re-run a current study than to preserve, migrate, and continually revalidate an old one. Appraisal is easier to do honestly when discarding is not irreversible in practice.
That last point is the real argument. Traditional tools made research so expensive to produce that teams kept everything out of fear. Platforms like Koji make the marginal study cheap enough that you can afford to be selective about the archive, which is precisely what makes the archive good.
Frequently asked questions
Is not storage cheap enough that we can just keep everything?
Storage is cheap; attention is not. The costs of keeping everything are paid in retrieval noise, in preservation care spread too thin, and in the loss of trust that follows when people search the repository and find mostly stale results. NARA keeps under five percent of federal records permanently, and it is not short of disk. The constraint has never been the bytes.
How is this different from our data retention policy?
Retention sets the legal floor and ceiling: how long you are required and permitted to hold personal data. Appraisal operates inside those limits and asks which material deserves to be kept, curated, and carried forward. The SAA dictionary draws the distinction explicitly, treating assessment based on retention schedules as a different activity from appraisal proper. You need both documents; most teams have only the first.
Who should do appraisal?
Whoever owns the repository, working from a written collecting policy, with a quarterly review rather than a continuous one. The one arrangement to avoid is letting each study's requesting stakeholder appraise their own study, because everyone rates their own work as permanent. Note also Duranti's warning about top-down appraisal: judging by the seniority of who asked will systematically discard the small studies about edge cases, which are often the ones with the most informational value.
What if we discard something and later need it?
That is a real risk and the reason appraisal is a judgement rather than a rule. Three things reduce it: appraise conservatively at the permanent tier, keep the finding and the study-level description even when the raw data ages out, and record every disposal so a future colleague can at least tell that something existed. The last point matters more than it sounds and is covered separately in this series.
Does appraisal apply to individual interviews or only whole studies?
Both, and the SAA note confirms appraisal may be done "at the collection, creator, series, file, or item level." In practice, appraise at study level for keeping decisions and at session level only for quality exclusions. Sessions that failed screening or quality checks should not enter the repository at all, which is a data quality decision that happens to have appraisal consequences.
How do we start if we already have thousands of items?
Do not start with the backlog. Write the collecting policy, apply the four tiers to everything new from today, and set a recurring quarterly hour to appraise the previous quarter. Then work backwards only through the most recent two years, since older material is usually resolved by the retention schedule anyway. A policy applied going forward beats a heroic cleanup that never finishes.
Related Resources
- Structured Questions: The Complete Guide - the six question types that make studies comparable at appraisal time.
- Research Data Retention and Deletion - the legal floor and ceiling appraisal operates inside.
- The Study-Level Description - the description a permanent-tier study needs.
- Insight Repository Methodology - taxonomy, governance, and activation for the repository.
- The Re-Research Audit - how much budget buys an answer you already own.
- Anonymizing Customer Interview Data - the step that converts a liability into a keepable asset.
Related Articles
Anonymizing Customer Interview Data: A Practical Guide for Privacy-Safe Research
Five operational techniques for handling PII in AI customer interviews — from intake-time anonymization to stakeholder-safe quote sharing — without sacrificing research signal.
Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)
The methodology layer most repository guides skip — taxonomy design, atomic insight structure, governance, freshness/decay rules, and the insight-to-action workflow that turns a static archive into a decision engine. Includes a 2-week setup plan and how AI auto-tagging from Koji eliminates the librarian bottleneck.
The Re-Research Audit: How Much of Your Budget Buys an Answer You Already Own
Count how many of your last twenty studies answered a question you already owned. The protocol, the four causes, and where the duty belongs.
Will You Be Able to Open Your 2026 Research Data in 2036? A Preservation Plan for Interview Archives
Backups protect the bits. Preservation protects the meaning. A practical plan for keeping interview transcripts, audio, and study context usable for a decade.
Research Data Retention and Deletion: How Long Should You Keep Interview Data?
There is no universal legal number - which is exactly why having no retention schedule is itself the compliance failure. A tiered, per-artifact schedule for recordings, transcripts, quotes, and reports, plus how to handle deletion requests without losing your insights.
How to Build a UX Research Repository: The Complete Guide
A research repository transforms scattered insights into a searchable organizational asset. Learn how to build one that teams actually use.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Study-Level Description: How to Make Research Findable Without Reading It
Repositories are tagged at the item level and described at no level. A nine-field study record, adapted from the archival description standard.