Terminus Post Quem: How Old Is Your Feedback Pile, Really? (2026)
A mixed research corpus gives you a lower bound on its own age, not a point estimate. The archaeological rule for dating a layer, plus the two contaminations that break it.
A pile of customer feedback is not as old as its average item and not as fresh as its newest one: the only date you can defend is the latest date any item in it establishes with certainty. Archaeologists formalized this rule a century ago and gave it a name, terminus post quem, and they use it every time they date a layer of soil. Product teams handle the same problem constantly, usually by quoting whichever research is most convenient, and almost never with a stated rule.
The short version: a mixed corpus of interviews, tickets, and survey responses gives you a lower bound on its own age, not a point estimate. Dating it correctly changes which conclusions you are allowed to draw from it, and two specific contamination patterns will break the dating entirely if you do not check for them.
The rule, and the coin that explains it
Terminus post quem, abbreviated TPQ, is the limit after which. It marks the earliest date the event may have happened or the item was in existence. Its mirror image, terminus ante quem or TAQ, is the limit before which: the latest possible date.
The canonical worked example is a burial. Consider an archaeological find of a burial that contains coins dating to 1588, 1595, and others less securely dated to 1590 to 1625. The terminus post quem for the burial would be the latest date established with certainty: in this case, 1595.
Two things in that example matter more than the arithmetic. First, the answer is set by the latest secure item, not the earliest, not the average, and not the most numerous. Second, the loosely-dated coins contribute nothing to the bound until they are pinned down: a secure dating of an older coin to an earlier date would not shift the terminus post quem, while securing the later date of 1625 would make that date the terminus post quem.
Now translate. You have 40 interviews about onboarding. Thirty-two are from a study run fourteen months ago, six are from a refresh five months ago, and two came in last week through an always-on study. What is the date of that corpus? Not fourteen months. Not the weighted average of eleven months. The defensible statement is that the corpus describes a product state no earlier than last week, which sounds absurd until you notice it is exactly the claim people make when they cite it: "our research says onboarding is confusing," present tense, as though the whole pile were current.
The archaeological formulation of the trap is unimprovable: the date of artifacts in a context does not represent the date of the context, but just the earliest date the context could be.
Superposition, and why your repository is a layer cake
Archaeological stratigraphy rests on a principle worth quoting in full, because the precision is the point: within a series of layers and interfacial features, as originally created, the upper units of stratification are younger and the lower are older, for each must have been deposited on, or created by the removal of, a pre-existing mass of archaeological stratification.
A research repository is a stratified deposit whether or not anyone treats it as one. Studies are laid down in sequence, each on top of the accumulated pile, and the sequence carries real information: a theme that appears in the 2024 layer, vanishes in the 2025 layer, and returns in the 2026 layer is telling you something a flat search across all three layers will actively hide.
The discipline that makes this legible is the Harris matrix, a two-dimensional representation of a site's formation in space and time. The product equivalent is cheap: every insight carries the study it came from and the date that study closed, and no insight is ever cited without both. Most repositories store the first and drop the second somewhere between the tag and the slide.
The rule for dating a layer follows directly, and it is a rule about discipline rather than technique: it is crucial that dating a context is based on the latest dating evidence drawn from the context.
The two contaminations that break the dating
This is where the archaeology earns its keep, because it names both failure modes precisely, and both have exact analogues that quietly corrupt research corpora.
Residual finds. When a later feature cuts through earlier layers, old material gets carried upward into the new deposit. Artifacts from layers 9 and 10 may be redeposited higher up the sequence in the context representing the backfill of a construction cut, and these artifacts are referred to as residual or residual finds. They are genuinely old, genuinely present, and they make the new layer look older than it is.
In a repository, a residual find is a stale quote in a current synthesis. Someone builds a 2026 insight page and pulls in a vivid 2024 verbatim because it expresses the theme so well. The quote is real. The person really said it. And the page now claims a product state that has not existed for two years, with no marker that one of its load-bearing pieces of evidence was redeposited from a lower layer.
Intrusive finds. The reverse direction is worse, because it produces confident wrong answers. Non-residual artifacts from later contexts can contaminate the excavation of earlier contexts and give false dating information; these artifacts may be termed intrusive finds.
The research version: you go back to last year's study to check what customers thought before a redesign, and the corpus has been quietly amended since. A follow-up interview was filed into the old study. A tag was updated. A summary was rewritten after the redesign shipped, using post-redesign language. The old layer now contains material that could not have existed when it was deposited, and it will tell you customers anticipated a problem they were actually reacting to.
Both patterns are failures of the same property, and archaeology names that too. A context is trustworthy in proportion to how sealed it is: the date of a context which is totally sealed between two datable layers will fall between the dates of the two layers sealing it. A sealed study is one that is closed, dated, and never edited again. An open one is not a context; it is a mixing bowl.
The practical procedure
Five steps, and none of them require new tooling.
- Seal studies on completion. When a study closes, freeze it: no new interviews filed in, no retroactive tag edits, no rewritten summaries. Later thinking goes in a new layer that cites the old one.
- Date every corpus by its latest secure item, and say so out loud. Write "no earlier than March 2026" rather than "2025 research". The awkwardness of the phrasing is the useful part, because it forces the question of whether the pile should be split.
- Split mixed piles before analyzing them. If 32 interviews are fourteen months old and 8 are recent, that is two contexts, not one. Analyze them separately and compare, which is the only way the sequence yields information instead of noise.
- Flag redeposited quotes. Any verbatim older than the synthesis that contains it gets its original date shown inline. If a 2024 quote is the best evidence for a 2026 claim, that is worth seeing rather than smoothing over.
- Check for intrusion before trusting an old layer. Before citing a historical study, confirm nothing has been added or edited since it closed. If you cannot confirm that, the layer is unsealed and its date is unknown.
This is a different question from how long a finding stays true, which is covered in how long user research is valid. Insight decay asks when to re-run a study. Terminus post quem asks what date a pile of existing evidence can legitimately claim, which you need to answer first, because a decay calculation applied to a misdated corpus is arithmetic on a fiction.
How Koji handles this
Koji's data model treats a study as a dated context rather than a folder, which is what makes the stratigraphic discipline enforceable rather than aspirational.
- Every interview carries its own completion timestamp. Koji records when each conversation actually happened, so the latest secure item in a corpus is a queryable fact rather than something you reconstruct from memory.
- Reports are bound to their study. A Koji report is generated from a specific study's interviews, so a synthesis cannot silently acquire evidence from a different layer. Refreshing a report is an explicit, credited action, which leaves a mark rather than quietly re-dating the conclusion.
- Transcripts stay under every theme. Because each theme in a Koji report links back to the transcript segment behind it, a redeposited quote is traceable to its original interview and date instead of floating free in a synthesis.
- Always-on studies are separate contexts by design. Running continuous discovery through Koji produces a stream of dated interviews rather than an ever-edited document, so the sequence stays readable and the newest layer is obvious.
- Programmatic access for the dating check. Koji's MCP tools, including koji_list_studies, koji_get_study_data, and koji_get_interviews, let you pull study and interview dates directly, so "what is the latest secure item in this corpus" is a query rather than a meeting.
- Structured questions make the comparison possible. Koji's six question types, open_ended, scale, single_choice, multiple_choice, ranking, and yes_no, produce identically-shaped quantitative answers across studies, so two sealed contexts can be compared on the same scale instead of by vibes. See structured questions in AI interviews.
Frequently asked questions
What is terminus post quem in plain language?
It is the earliest date something could possibly be, established by the latest securely-dated item found in it. If a deposit contains a coin minted in 1595, the deposit cannot be older than 1595, regardless of how many older coins sit beside it. Applied to research, a corpus containing an interview from last week cannot be described as a picture of last year.
Should I date a feedback corpus by its newest item or its average age?
By its newest secure item, for claims about what the corpus could describe, and by its actual distribution for everything else. The average is the one number that is always wrong: it describes a moment when nothing was collected. Better practice is to split a mixed pile into separate dated contexts and analyze each.
What is the difference between a residual and an intrusive find?
A residual find is genuinely old material redeposited in a newer layer, which makes the new layer look older. An intrusive find is newer material that has worked its way into an older layer, which makes the old layer look like it knew things it could not have known. In research terms: a stale quote in a current synthesis versus a post-hoc edit filed into a closed study.
How does this differ from insight decay?
Insight decay asks how long a finding remains true and when to re-run the study, which is the subject of how long user research is valid. Terminus post quem asks a prior question: what date can this specific pile of evidence legitimately claim? You need the date before the decay calculation means anything.
Does sealing a study mean I can never correct a mistake in it?
No. It means corrections are deposited as a new, dated layer that references the old one, rather than being written over the original. The audit trail is the whole value: a reader can see both what was concluded at the time and what was revised later, which is impossible once the original has been edited in place.
Can I reconstruct dates for a repository that was never maintained this way?
Partially, and it is worth doing. Recover the closing date of each study and the timestamp of each underlying interview, then re-date every synthesis by the latest secure item it actually rests on. Anything whose evidence you cannot date becomes an unsealed context: still readable, but not citable as evidence about a particular moment. Research provenance covers the mechanics of establishing that chain.
Related Resources
- Structured Questions in AI Interviews - the six question types, and how consistent structure makes two dated contexts comparable
- How Long Is User Research Valid? - insight decay and re-run timing, the question that comes after dating
- Research Provenance - proving an interview, quote, or report is genuine and unaltered
- The Taphonomy of Customer Feedback - which evidence classes are destroyed before they ever reach the deposit
- Why a Tagged Quote Is Not Evidence - why a verbatim needs its context to mean anything
- Insight Repository Methodology - building a repository that keeps its layers legible
Related Articles
Why a Tagged Quote Is Not Evidence: The Archival Bond in Research Repositories
A quote filed under a theme is information; the interview it came from is evidence. The five bonds repositories delete at intake, and how to keep them.
Insight Repository Methodology: How to Build, Tag, and Activate a Research Insight Library (Beyond Just Storage)
The methodology layer most repository guides skip — taxonomy design, atomic insight structure, governance, freshness/decay rules, and the insight-to-action workflow that turns a static archive into a decision engine. Includes a 2-week setup plan and how AI auto-tagging from Koji eliminates the librarian bottleneck.
Research Provenance: How to Prove an Interview, a Quote, or a Report Is Genuine (2026)
When any text can be generated, a customer quote proves nothing on its own. Content Credentials, the EU AI Act marking rules, and the hash-anchored capture method that actually works for research text.
How Long Is User Research Valid? Insight Decay and When to Re-Run a Study
Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Taphonomy of Customer Feedback: Which Complaints Survive to Reach You (2026)
Most customer feedback is destroyed before it reaches you, and the filter has a predictable shape. A taphonomic method for naming the evidence classes your channels systematically lose.