Why a Tagged Quote Is Not Evidence: The Archival Bond in Research Repositories
A quote filed under a theme is information; the interview it came from is evidence. The five bonds repositories delete at intake, and how to keep them.
Short answer: a customer quote filed under a theme is information. The interview it came from, in the order it happened, is evidence. The difference is a set of relationships that most research repositories delete at intake and cannot reconstruct afterwards. Archivists have studied this failure for eighty years and named the thing you lose: the archival bond, defined by Duranti (1988) as the interrelationships between a record and the other records produced by the same activity. A quote without its bond can still be true, still be quotable, and still be completely unusable as proof — because nobody can tell what question produced it, what was said immediately before it, or what the respondent had already been told.
This is not an argument against tagging, atomising, or building a repository. It is an argument about what has to travel with the tag.
What archivists learned about re-filing by subject
Archival practice rests on two principles that sound bureaucratic and are actually epistemic.
Provenance says records are kept grouped by who created them, never merged with records from another creator. Original order — described in the Society of American Archivists dictionary as the organisation and sequence established by the creator of the records — says you keep the sequence the creator used, even when a subject arrangement would be more convenient.
Both rules exist because archivists discovered, repeatedly, that re-arranging records by topic destroys information that is not recorded anywhere in the records themselves. A memo filed between two others tells you what the writer was responding to. Pull it out, file it under "pricing", and that context is gone — not hidden, gone, because it was never a property of the memo. It was a property of the memo's position.
The Italian archivist Giorgio Cencetti formalised this in 1939; Duranti developed it as the archival bond. The modern formulation from Trace (2021) is the network of relationships each record has with the other records in the same aggregation. The claim is strong and worth stating plainly: a document removed from its aggregation is no longer a record. It is just a document.
Now look at what a research repository does at intake. It takes a 40-minute interview, extracts twelve quotes, attaches three tags to each, and files them under themes alongside quotes from other people in other studies asked other questions. That is precisely the subject re-arrangement archivists spent a century learning not to do.
The five bonds a quote loses when you atomise it
| Bond | What it tells you | What goes wrong when it is missing |
|---|---|---|
| The question | What was actually asked to produce this | A quote answering "what frustrates you about billing?" reads as unprompted volunteering of a billing complaint |
| Position in sequence | What the respondent had already discussed | A price objection after fifteen minutes of feature discussion means something different from one in minute two |
| The preceding turn | What the interviewer said immediately before | Agreement with a leading follow-up gets stored as an independent opinion |
| The interview | Everything else this person said | The one person who contradicted themselves four times looks identical to a consistent respondent |
| The study and its brief | Who was recruited, why, and what the study was for | A quote from a churned-customer study gets cited as evidence about active users |
Every one of these failures is a real thing that happens in product organisations, and none of them are detectable from the quote itself. That is the whole problem: decontextualisation is silent. A stripped quote does not look damaged. It looks clean.
The re-attachment test
Here is a five-minute audit you can run on any repository today.
Pick five quotes at random from your repository — ideally five that have been cited in a deck or a PRD. For each one, try to answer these five questions using only what the repository stores:
- What question was the respondent answering?
- What did the interviewer say immediately before this?
- Where in the interview did it occur — first third, middle, last third?
- Who was this person, in terms of the study's own recruitment criteria?
- What study was this, and what was that study trying to decide?
Score one point per question you can answer without opening a raw transcript or asking a colleague. A repository scoring below 4 out of 5 on average is storing anecdotes, not evidence. Most score 1 or 2 — usually the study name and nothing else.
The follow-up question is the uncomfortable one: how many decisions have been made on quotes that would score 1?
What this does not mean
This is a correction to atomic research, not a rejection of it. Our own atomic research guide makes a strong case for breaking findings into searchable units, and that case survives intact. Searchability is real value; so is cross-study synthesis; so is letting non-researchers self-serve.
The distinction that matters:
- Copying a quote into a themed view while the original stays linked and intact = arrangement. This is what a good archive does — multiple access points over one preserved order.
- Extracting a quote so that the themed view becomes the only place it lives = re-filing. This is what destroys the bond.
The first is a view. The second is a migration. Repositories that are built on manual copy-paste almost always become the second, because the transcript is a file in a folder somewhere and the nugget is what people actually open.
The rule: never store a quote without its bond
Any quote entering your repository should carry five fields, populated automatically, that cannot be edited by the person doing the tagging:
- Source interview ID and a link that opens the transcript at the quote's position
- The verbatim question text that preceded it, not a paraphrase of the topic
- Turn index or timestamp within the interview
- Study ID plus the study's recruitment criteria
- Speaker role — which side of the conversation said it
Two design rules follow, and they are the ones teams get wrong:
- Populated automatically. A field a human fills in during tagging is a field that will be blank on the busiest week of the quarter. If your intake ritual depends on discipline, you are measuring discipline, not context.
- Not editable in the themed view. The moment someone can "clean up" the question text on a nugget, the bond is a summary rather than a record.
How Koji preserves the bond by default
The reason this problem is so widespread is structural: in a tool-chain where interviewing, transcription, and the repository are three separate products, the bond has to be re-created by hand at every boundary — and it never is.
Koji collapses that chain, which means the bond is a property of the data rather than a discipline:
- Quotes are extracted from interviews the platform itself conducted, so the question that produced each quote is known exactly, not inferred from a transcript. See customer quotes for how candidates surface.
- Structured questions are pre-bonded by design. All six types —
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no— store the answer against the question that asked it, so ascaleresponse can never drift away from its stem the way a copy-pasted number can. - Follow-ups are recorded as follow-ups. When the AI probes, the probe is part of the record, so you can tell an unprompted statement from an answer to a pointed question — the single most commonly lost distinction in product research.
- The transcript is never the second copy. Themed views, cross-study search, and reports all read from the interview record rather than replacing it, so re-attachment is always one click, not an archaeology project.
- Study context travels with the artifact. The brief, recruitment criteria, and question set stay attached to every derived object, which is what makes a quote usable eighteen months later by someone who was not there.
A worked example
A PM pastes this into a roadmap doc: "I would pay double if it just worked with our SSO."
Without the bond, this reads as a pricing insight with a feature dependency. It gets cited in a pricing discussion three months later.
With the bond, the record shows: the question was "if we solved your integration problems, how would that change what this is worth to you?" — a hypothetical, framed by the interviewer, asked in minute 34 of a churn-risk study whose recruitment criteria selected accounts that had already filed an SSO support ticket. The same words, now correctly readable as a prompted willingness-to-pay hypothetical from a pre-selected frustrated segment. Still useful. Completely different weight.
Nothing about the quote changed. What changed is that somebody could see what it was.
Frequently asked questions
Is this an argument against research repositories?
No — it is an argument about what a repository stores. A repository that keeps quotes and their bonds is more valuable than raw transcripts, because it adds retrieval without subtracting context. A repository that keeps quotes instead of their bonds trades evidence for convenience and usually does not know it has made the trade. The test is whether re-attachment is one click or a research project.
What is the archival bond, in one sentence?
It is the set of relationships between a record and the other records created by the same activity — the interrelationships, in Duranti's 1988 formulation, that give a record meaning beyond its own content. In research terms: the question, the sequence, the interview, and the study are not metadata about a quote. They are part of what the quote is.
Our tags are AI-generated, so isn't the context captured?
Auto-tagging solves the labour problem, not the bond problem. A tag describes what a quote is about; a bond describes where it came from. If your AI tags a quote "pricing sensitivity" but the stored record still cannot tell you that the question was leading, you have made a decontextualised quote easier to find, which is arguably worse than leaving it hard to find.
How is this different from provenance and authenticity?
Research provenance answers "is this quote genuine, and can I prove it was not fabricated or altered?" The archival bond answers "what did this quote mean in the setting that produced it?" A quote can be perfectly authentic and completely misleading. You need both properties, and they are established by different mechanisms.
Does this apply to survey verbatims too?
Yes, and more severely, because open-text survey responses arrive already stripped of sequence and follow-up. The mitigation is the same: store the exact question stem with every verbatim, keep the respondent's other answers linked, and never present a verbatim in a deck without the question that produced it visible on the same slide.
What is the minimum viable version of this if we cannot rebuild our repository?
Add one field to your quote template: the verbatim question text. It is the single highest-value bond, it is the one that most often reverses a reading, and it can be backfilled for recent studies faster than any of the others. Then add a working link back to the transcript position. Those two get most repositories from a score of 1 to a score of 3 on the re-attachment test.
Related Resources
- Atomic Research and Research Nuggets — the practice this article qualifies rather than rejects
- Customer Quotes Guide — extracting quotes without severing them from their source
- Research Provenance and Authenticity — proving a quote is genuine, a separate property from context
- Insight Repository Methodology — taxonomy and governance for the layer above the bond
- Search Across Interview Transcripts — retrieval that reads from the record instead of replacing it
- Structured Questions Guide — the six question types that bond an answer to its stem automatically
Related Articles
Atomic Research: The Complete Guide to Research Nuggets and Insight Repositories
Learn the atomic research framework developed by Daniel Pidcock. Break research findings into reusable nuggets — observations, evidence, and tags — that prevent insight rot and make your repository searchable across teams.
Customer Quotes: How to Extract, Tag, and Use the Voice of Your Customer
Customer quotes are the most persuasive evidence in product, marketing, and research. This guide covers how to extract them, what makes a quote useful, and how Koji surfaces them automatically.
Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)
The methodology layer most repository guides skip — taxonomy design, atomic insight structure, governance, freshness/decay rules, and the insight-to-action workflow that turns a static archive into a decision engine. Includes a 2-week setup plan and how AI auto-tagging from Koji eliminates the librarian bottleneck.
Research Provenance: How to Prove an Interview, a Quote, or a Report Is Genuine (2026)
When any text can be generated, a customer quote proves nothing on its own. Content Credentials, the EU AI Act marking rules, and the hash-anchored capture method that actually works for research text.
How to Search Across All Customer Interview Transcripts (Semantic + Keyword)
Find the exact moment a customer said the thing across every study in your research repository — semantic search, keyword search, theme filters, and jump-to-quote deep links in Koji.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.