{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-21T09:55:59.972Z"},"content":[{"type":"documentation","id":"f575f709-6064-49a5-887a-e29e051c8b52","slug":"reusing-interview-data-new-question","title":"Reusing Interview Data for a Question It Was Not Collected to Answer","url":"https://www.koji.so/docs/reusing-interview-data-new-question","summary":"Qualitative secondary analysis re-analyses existing primary interview data to answer a new question. Heaton identifies three access modes, of which auto-data (a team returning to its own data) is the most common, meaning most product teams already do this without safeguards. Validity depends on data fit, tested against sampling frame, question wording, depth, denominator, and analytic frame. Context and tacit field knowledge cannot be recovered, and consent limits are answered before methodology. Koji keeps the brief, screener, exact question wording, and full transcripts attached to every study, which is the context reuse normally lacks.","content":"**Short answer:** You can reuse existing interview data to answer a new question, and it is often far cheaper than fielding a fresh study. But the reuse is only valid if the old data actually fits the new question, and fit is not about topic overlap. It is about whether the original sampling plan, screener, and question wording can support the claim you now want to make. The research literature calls this **qualitative secondary analysis**, and the discipline that governs it is a short list of checks you run before analysing anything. Koji makes reuse unusually tractable because the brief, the screener, the exact question wording, and the full transcript stay attached to every study, which is precisely the context that reuse normally lacks.\n\nThere is a specific moment this matters. Someone asks a question, you remember that a study eighteen months ago talked to roughly the right people, and the temptation is to search the transcripts and answer from what comes back. Sometimes that is excellent practice. Sometimes it produces a confident answer that the data never supported. The difference is knowable in advance.\n\n## First, this is not desk research\n\nThe terminology collides badly, so let us separate it once.\n\n**Secondary research** in the common product sense means desk research: reading published reports, analyst notes, competitor material, and public data. That is a different activity with a different guide, [secondary research](/docs/secondary-research-guide), and it is about sources someone else published.\n\n**Qualitative secondary analysis** means re-analysing *primary data that already exists* - real transcripts from real interviews - to answer a question different from the one they were collected for. The data is still primary. Only the question is new.\n\nThe rest of this article is about the second thing.\n\n## Three ways you get the data, and one of them is yours already\n\nHeaton's typology, summarised in [Tluczek and colleagues' case exemplar](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5911239/), identifies three modes of access.\n\n- **Formal.** Analysing data deposited in an archive by researchers you have no relationship with. Common in academia, rare in product work.\n- **Informal.** Getting direct access to another investigator's data through a personal request or network.\n- **Auto-data.** The primary research team returning to its own dataset to explore further. The paper notes this is **the most common mode**.\n\nAuto-data is what a product team does every time it searches its own repository for a question the study was not designed to answer. Which means most teams are already doing secondary analysis continuously, without calling it that and therefore without applying any of its safeguards.\n\nThat is the actual risk. Not that reuse is illegitimate, but that it happens invisibly and is never held to a standard.\n\nThe same paper flags the specific hazard of auto-data: because qualitative work is iterative, \"it may be difficult to determine where the original study questions end and discrete, distinct analysis begins.\" When the same team reuses its own data, the new question tends to drift into the old findings and inherit their authority. Writing the new question down before you search is the cheapest available control.\n\n## Data fit: the check that decides everything\n\nThe literature's central concept is **fit**, and the test is not whether the old interviews mention your new topic. It is whether the study's design can carry the weight of your new claim.\n\nTluczek and colleagues put it directly: \"The underlying assumptions, sampling plan, research questions, and conceptual framework selected to answer the original study question may not fit the question posed\" in the reuse. And crucially: \"The researchers of the primary study may have selectively sampled participants and analyzed the resulting data in a manner that produced a narrow or uneven scope of data.\"\n\nRun five checks before you analyse anything. Any single failure does not necessarily stop you, but it does change what you are allowed to conclude.\n\n**1. Sampling frame.** Who was eligible for the original study? If the screener recruited existing customers, the transcripts cannot tell you why non-customers stay away, no matter how many participants volunteered an opinion about it. This is the most common fatal mismatch.\n\n**2. Question wording.** Was your new topic asked about directly, or did it surface spontaneously? These support entirely different claims. Direct questioning supports statements about prevalence within the sample. Spontaneous mention supports only \"this was salient enough that some people raised it unprompted,\" which is interesting but is not a base rate.\n\n**3. Depth.** Did the original protocol probe the topic or pass over it? A single unfollowed mention is an anecdote. Koji's AI follow-up questions help here, because open-ended answers are probed automatically rather than left wherever the participant stopped, which means old transcripts often contain more depth on incidental topics than a human-moderated study would have captured.\n\n**4. Denominator.** How many participants were asked, as distinct from how many answered? If your new topic came up with eleven of forty participants, you cannot report \"28 percent\" unless all forty were actually asked. If they were not, your denominator is unknown and no percentage is defensible.\n\n**5. Analytic frame.** Was the original coded against a codebook that presupposed the old question? Existing tags will pull your new analysis toward the old conclusions. Where the new question is genuinely different, code from the raw transcripts rather than from the existing themes.\n\n## What you cannot recover, and how to handle it\n\nSome things are simply gone, and the honest move is to name them rather than to pretend the transcripts are complete.\n\nThe context critique is the strongest objection to qualitative reuse in the methodological literature. Tluczek and colleagues summarise it: \"Tacit understandings developed in the field may be difficult or impossible to reconstruct,\" and \"because the context in which the data were originally produced cannot be recovered, the ability of the researcher to react to the lived experience may be curtailed.\" Fieldworkers redirect their questioning based on knowledge of the setting, and they filter what counts as data in ways \"that may not be apparent in either the written or spoken records.\"\n\nThe archival standard sets the bar this should be measured against. The [OAIS reference model](https://public.ccsds.org/Pubs/650x0m2.pdf) requires preserved information to be **Independently Understandable**, meaning interpretable by its intended community \"without having to resort to special resources not widely available, including named individuals.\" Data you can only interpret by asking the researcher who collected it has not met that standard.\n\nThere is a genuine advantage for AI-moderated research here, and it is worth stating plainly rather than as a sales point. A great deal of the irrecoverable context in traditional qualitative work is irrecoverable because it lived in one moderator's head: their read of the room, the reason they dropped a question, the participant who seemed evasive. AI-moderated interviews in Koji externalise more of that into the record. The protocol is the brief, the follow-ups are in the transcript verbatim, and moderator variation between sessions is not a hidden variable. The tacit layer is thinner because less of it was ever tacit.\n\n**Passage of time is the second irrecoverable.** Data can carry historical bias: the product changed, the market changed, the participants' own perspectives changed. If your new question is about attitudes rather than about experiences that already happened, old data is weak evidence. See [how long research stays valid](/docs/research-refresh-cadence) for finding-type decay rates.\n\n## What reuse is legitimately for\n\nDrawing on Hinds and colleagues, the paper lists four defensible purposes, and they are worth using as a menu because they set expectations about what the output can claim.\n\n1. **Additional analysis of the original dataset** - going deeper on something the original noticed but did not pursue.\n2. **Analysis of a subset** - examining one segment more closely than the original report did.\n3. **A new perspective or focus** - applying a different lens to the same material.\n4. **Validating or expanding the original findings** - checking whether the conclusion holds under a different reading.\n\nNotice what is absent: producing a fresh prevalence estimate for a population the study did not sample. That is the use case reuse cannot serve, and it is the one people most often attempt.\n\n## Holding a reuse study to a standard\n\nApply the four established rigor criteria, adapted for reuse.\n\n- **Trustworthiness** - a logical relationship between the data and your claims. State which transcripts support each claim.\n- **Fit** - the context in which findings apply. Name the original sampling frame in the write-up, not just in your notes.\n- **Transferability** - how far the claims generalise. Reuse narrows this rather than widening it.\n- **Auditability** - transparency of procedural steps. For reuse this means recording which studies you searched, which you excluded and why, and when the new question was written relative to the search.\n\nTriangulation is the recommended remedy for reduced access to the original participants. The literature suggests treating the original researchers as sources of confirmation, and bringing in new informants or related datasets as genuine triangulation. In practice, the strongest pattern is a hybrid: reuse the archive to develop the hypothesis, then run a small fresh study to test it. Our guide to [triangulation](/docs/triangulation-in-research-guide) covers the general form.\n\n## Consent is a hard limit, not a methodological one\n\nThe methodological checks above tell you whether reuse is *valid*. Consent tells you whether it is *permitted*, and that question is answered first.\n\nIf participants consented to a specific study purpose, reusing their data for a materially different purpose may exceed what they agreed to, and under GDPR the compatibility of a new purpose with the original is a legal test rather than a matter of judgement. Anonymised data escapes much of this, which is a strong practical argument for [anonymising interview data](/docs/anonymizing-customer-interview-data) at the point a study is archived rather than at the point someone wants to reuse it. Write reuse into the consent form up front where you can. It is a sentence, and it converts your archive from a single-use asset into a compounding one.\n\n## How Koji makes reuse actually work\n\nReuse fails in most stacks for a mundane reason: the information needed to assess fit is scattered across a survey tool, a scheduling tool, a recording, and a slide deck, so nobody can reconstruct the sampling frame two years later. Koji keeps it in one object.\n\n- **The brief states the original question.** Every study is generated from a research brief holding the problem context, methodology, and question set, so check one of the fit test is answerable in seconds.\n- **The screener is preserved.** The recruitment criteria that define your sampling frame stay attached to the study rather than living in a recruiter's inbox.\n- **Exact question wording survives.** With Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - you can see precisely what each participant was asked and which options they were offered. That settles checks two and four, including the denominator question, because it is unambiguous who was actually asked.\n- **Full transcripts, not summaries.** Every answer stays bound to its question, and AI follow-ups appear verbatim, so you can code from raw material rather than from someone's prior interpretation.\n- **Search across the whole archive.** Semantic and keyword search over every transcript makes finding candidate studies practical; see [searching interview transcripts](/docs/search-interview-transcripts).\n- **And when reuse fails the fit test, you are not stuck.** This is the quiet advantage. In a traditional stack, discovering that old data does not fit means either six weeks of fresh fieldwork or forcing the claim. With AI-moderated interviews you can field a properly designed study in days, which means the fit test is safe to fail honestly.\n\n## Frequently asked questions\n\n### Is reusing old interview data legitimate research?\n\nYes. Qualitative secondary analysis is an established method with a substantial methodological literature, and reuse of existing data is standard practice across the social sciences. It is legitimate when the new question fits the original design and when you report the limits that reuse imposes. It becomes illegitimate when it is used to manufacture prevalence claims about populations the original study never sampled.\n\n### How is this different from secondary research or desk research?\n\nDesk research analyses sources someone else already published: reports, articles, market data. Qualitative secondary analysis re-analyses primary data that you or a colleague collected, to answer a question it was not collected for. The data is primary in both senses that matter; only the research question is secondary. See our [secondary research guide](/docs/secondary-research-guide) for the desk research discipline.\n\n### What is the single most common mistake?\n\nReporting a percentage without a valid denominator. A topic mentioned by eleven participants out of forty is only 28 percent if all forty were asked about it. If it arose spontaneously, the correct statement is that eleven participants raised it unprompted, which is a claim about salience, not prevalence. Structured question types make this unambiguous because you can see exactly who was asked.\n\n### Can we reuse data if the original consent did not mention reuse?\n\nTreat that as a legal question rather than a methodological one, and answer it before the analysis. Under GDPR, whether a new purpose is compatible with the original is a defined test, and a materially different purpose may require fresh grounds. Anonymised data is substantially less constrained, which is why anonymising at archive time is good practice. The cheapest fix is prospective: add a reuse clause to your consent form now.\n\n### How old is too old?\n\nIt depends entirely on what you are asking about, not on a fixed interval. Historical experiences that already happened do not expire. Attitudes, preferences, and comparisons to competitors decay quickly, and product-specific findings become invalid the moment the product changes. If the new question is about the current state of anything, old data generates hypotheses rather than answers.\n\n### Should we reuse or just run a new study?\n\nWith AI-moderated research the economics have shifted enough that this is a real question rather than a rhetorical one. Reuse first when you want to develop a hypothesis, understand vocabulary, or check whether something was already asked. Run fresh when you need a defensible prevalence estimate, when the population differs from the original sampling frame, or when the topic is time-sensitive. The strongest pattern is both: mine the archive to design a sharper study, then field it.\n\n## Related Resources\n\n- [Structured Questions: The Complete Guide](/docs/structured-questions-guide) - the six question types that make original wording and denominators recoverable.\n- [How to Search Across All Interview Transcripts](/docs/search-interview-transcripts) - semantic and keyword search for finding candidate studies.\n- [Secondary Research: The Complete Guide](/docs/secondary-research-guide) - desk research, the other thing called secondary.\n- [The Re-Research Audit](/docs/re-research-audit-duplicate-studies) - finding out how often you buy an answer you already own.\n- [Why a Tagged Quote Is Not Evidence](/docs/archival-bond-research-quotes-context) - why quotes need their surrounding context to stay evidential.\n- [Triangulation in Research](/docs/triangulation-in-research-guide) - combining reuse with fresh evidence.\n","category":"Analysis & Synthesis","lastModified":"2026-08-21T03:25:48.491759+00:00","metaTitle":"Qualitative Secondary Analysis: Reusing Interview Data for a New Question (2026)","metaDescription":"How to reuse existing interview transcripts for a new research question: the five data-fit checks, what context cannot be recovered, and the denominator trap.","keywords":["qualitative secondary analysis","reusing interview data","secondary analysis qualitative data","research data reuse","data fit secondary analysis","reanalyzing interview transcripts"],"aiSummary":"Qualitative secondary analysis re-analyses existing primary interview data to answer a new question. Heaton identifies three access modes, of which auto-data (a team returning to its own data) is the most common, meaning most product teams already do this without safeguards. Validity depends on data fit, tested against sampling frame, question wording, depth, denominator, and analytic frame. Context and tacit field knowledge cannot be recovered, and consent limits are answered before methodology. Koji keeps the brief, screener, exact question wording, and full transcripts attached to every study, which is the context reuse normally lacks."}],"pagination":{"total":1,"returned":1,"offset":0}}