{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-25T08:47:32.901Z"},"content":[{"type":"documentation","id":"ccec3578-d0f0-4384-8cdd-8dff6c55b0cc","slug":"research-repository-retrieval-metrics-known-item-test","title":"Is Your Research Repository Working? The Retrieval Metrics Nobody Collects","url":"https://www.koji.so/docs/research-repository-retrieval-metrics-known-item-test","summary":"Repository health should be measured by retrieval rather than volume. Zero-result rate, known-item recall from a 40-probe sample, time-to-first-relevant and the recall ceiling implied by list length together answer whether colleagues can reach existing research. The largest failure is the query nobody thought to formulate, which only push-based distribution addresses.","content":"Most teams evaluate their research repository by how much is in it. That is a measure of the warehouse, not of the service. The question that matters is whether a colleague who needs a finding can reach it, and almost nobody instruments that. Four metrics answer it, all cheap: zero-result rate, known-item recall, time-to-first-relevant, and the retrieval ceiling implied by your list length. The most important of the four costs an afternoon and will probably tell you your repository is operating somewhere near 45 percent.\n\nUnderneath all four sits a failure mode worth naming up front, because it is the one your governance process cannot see. The research was done. It was done well. It is correct, it is still true, and it is sitting in your repository right now. And the decision it should have informed was made without it, because the person making that decision never knew to look.\n\n## Metric 1: zero-result rate — the only failure that announces itself\n\nA search that returns nothing is the single visible retrieval failure you get. Log every query, count the proportion that return no results, and segment by who searched.\n\nIt is genuinely useful. It is also the smallest part of the problem, and it is worth being precise about why: a zero-result search tells you someone tried and failed. It says nothing about the far larger category of searches that returned four plausible results while missing eleven, and nothing at all about the questions nobody typed.\n\nTreat a rising zero-result rate as a vocabulary signal rather than a content signal. In most repositories the material exists and the words do not match. Read the actual failed query strings — they are the cheapest research you will ever do, and each one names a specific gap.\n\n## Metric 2: known-item recall — the number that actually matters\n\nYou cannot measure true recall, because verifying it means judging your whole corpus against every question. What you can do is what NIST does: estimate it by sampling rather than measuring it exhaustively. In the 2007 TREC Legal Track, NIST abandoned exhaustive judgment in favour of a statistical sampling method precisely because complete assessment of a large collection is not feasible.\n\nThe repository version is a known-item test, and it takes an afternoon:\n\n1. **Select 40 findings you know exist.** Draw them from studies spanning at least two years. Have someone who did not run those studies do the selecting, so you are not unconsciously picking memorable ones.\n2. **Write the question each finding answers** in the words a product manager or designer would actually use — not the words in the report title.\n3. **Give each question to a colleague who did not run that study.** Two-minute limit per question, using only the repository.\n4. **Record found or not found, and capture every query they tried.**\n5. **Compute the proportion found.** That is your estimated recall.\n\nOn sample size: 40 probes gives you roughly ±15 percentage points at 95 percent confidence in the worst case. If you want ±10 points you need about 97 probes, and ±5 points needs about 385 — which is why 40 is the right starting number. It distinguishes a repository running at 45 percent from one running at 80 percent, and that is the decision in front of you.\n\nA worked example. Forty probes, eighteen found. That is 45 percent, with a 95 percent confidence interval of roughly 30 to 60 percent. Notice what that interval does and does not let you say: it does not let you claim a precise recall figure, and it comfortably rules out the *our repository works fine* hypothesis. That is the whole job.\n\nThen do the part everyone skips: **read the 22 misses, not the 18 hits.** Each miss is a named, fixable gap — a synonym that was never attached, a study whose description never mentioned the segment, a finding buried in an open_ended answer that should have been a structured field. Aggregate recall tells you the size of the problem. The misses tell you where it is.\n\n## Metric 3: time-to-first-relevant\n\nRecall assumes the searcher keeps going. Real people do not. They scan a few results, and if nothing looks promising they conclude the repository has nothing and move on — which converts a retrieval failure into a confident, wrong statement of fact.\n\nTime-to-first-relevant is measured during the known-item test at no extra cost: how long until the searcher opens something they judge useful? Anything past about ninety seconds is effectively a miss, because that is roughly where people stop.\n\nThis is why Marcia Bates's model of searching matters more here than the classical retrieval model does. In her 1989 work on browsing and berrypicking, Bates argued that real information seeking does not work as one query producing one result set; instead, as she puts it, \"the query is satisfied not by a single final retrieved set, but by a series of selections of individual references and bits of information at each stage of the ever-modifying search.\" Each thing you find changes what you are looking for.\n\nThe practical consequence: a repository that returns one good result quickly beats one that returns a comprehensive set slowly, because the first result reshapes the query and the second search is better. Optimising purely for a complete result set optimises for a behaviour nobody exhibits.\n\n## Metric 4: the retrieval ceiling you cannot search your way past\n\nThis one is arithmetic, and it is the reason repository performance degrades even when nothing is broken.\n\nYour search results list shows some fixed number of items — call it ten, since that is what people actually read. If a question has R genuinely relevant findings in the repository, then recall from that one screen cannot exceed 10/R, no matter how good the ranking is:\n\n| Relevant findings that exist (R) | Ceiling on recall from a 10-item list |\n|---|---|\n| 4 | 100% |\n| 8 | 100% |\n| 20 | 50% |\n| 40 | 25% |\n\nNow add growth. Every study you run increases R for the questions it touches. The list length does not grow. So **the ceiling falls as your repository succeeds** — double the corpus, halve the ceiling. A two-year-old repository with excellent search will show worse recall-per-screen than a six-month-old repository with mediocre search, and nothing has gone wrong in between.\n\nNIST's evaluation puts a floor under how much retrieving deeper can rescue this. Even allowing systems to return 25,000 documents, the best run in the 2007 Legal Track reached 47 percent estimated recall. Depth helps. Depth does not deliver completeness.\n\nTwo responses actually work. Consolidate — replace twelve findings that say the same thing with one that says it well and links to its evidence, which lowers R without losing anything. And make more of your evidence non-textual, so it is aggregated rather than retrieved.\n\n## The failure your governance process cannot see\n\nEvery metric above assumes someone ran a search. The most expensive retrieval failures happen before that.\n\nNielsen Norman Group's Raluca Budiu makes the point directly in her analysis of site search: \"In order to formulate a good search query users need to know fairly well what they are searching for.\" Bates's model says the same thing from the other side — you refine a query by encountering things, so before the first encounter you have very little to go on.\n\nPut those together and you get the failure mode that no audit catches. To retrieve a finding you must suspect it exists. A product manager scoping a pricing change does not search for onboarding research, because there is no reason to think onboarding research would say anything about pricing. The finding is present, correct, indexed, and perfectly retrievable by a query nobody had a reason to type.\n\nThis is a different failure from the ones teams usually chase. It is not bad data — the study was fine. It is not a bad metric definition. It is not a duplicate study, where at least someone knew to ask the question twice. Here the evidence was in the building, it was right, it was yours, and the decision went ahead without it. And the only defences are ones that do not require the searcher to have the idea first: pushing findings to people rather than waiting for pull, and lowering the cost of asking so far that \"just ask again\" stops being a defeat.\n\n## How Koji instruments this\n\n**Search that does not require the exact words.** Koji retrieves on meaning, so a question phrased in a product manager's language reaches a transcript phrased in a participant's language. That directly raises the known-item numbers above rather than explaining them away.\n\n**Automatic thematic analysis across every interview.** Themes get applied consistently rather than drifting with whoever wrote the summary, which is what makes cross-study aggregation possible instead of aspirational. It also supports consolidation, because you can see the twelve findings that are really one.\n\n**Koji real-time reporting that pushes rather than waits.** The unformulated-query problem is not solved by better search — it is solved by findings arriving in front of people who did not ask. Koji live reports and the MCP and API surfaces put research where decisions are being made, which is the only defence against a question nobody thought to ask.\n\n**Structured questions, which have no retrieval problem at all.** Koji's six question types are open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. The five non-open types produce values, so they are filtered and aggregated rather than searched — recall is 100 percent by construction, and they lower R for everything else by taking countable questions out of the prose entirely. Legacy tools like SurveyMonkey can give you the structured half; what they cannot do is analyse the open_ended half well enough for it to stay findable.\n\n**And the economics.** With Koji AI-moderated interviews running in parallel and voice interviews collecting depth without scheduling calls, re-asking a question costs hours instead of weeks. That does not repair a bad repository. It changes what a retrieval miss costs, which is what makes it safe to publish an honest recall number instead of defending an assumed one.\n\n## The scorecard\n\nRun this once a quarter against your own repository, whether or not it runs on Koji. It is half a day.\n\n| Metric | How to get it | A useful threshold |\n|---|---|---|\n| Zero-result rate | Query logs | Under 10%, and read every failed query |\n| Known-item recall | 40-probe test | Under 60% means fix retrieval before adding content |\n| Time-to-first-relevant | Timed during the same test | Median under 90 seconds |\n| Retrieval ceiling | Count relevant items for 5 common questions | If R regularly exceeds 20, consolidate |\n\nIf you only ever do one of them, do the known-item test. It is the only one that answers the question you actually care about, and the number it returns is usually about half what the team expected.\n\n## Frequently asked questions\n\n### What is a known-item test?\n\nA known-item test measures retrieval by starting from findings you already know exist. You select a sample of them, phrase the question each one answers, and ask colleagues who did not run those studies to find them within a time limit. The proportion found is an unbiased estimate of your repository's recall, obtained by sampling rather than by judging the whole corpus.\n\n### How many probes do I need?\n\nForty probes gives roughly ±15 percentage points at 95 percent confidence in the worst case, which is enough to tell a repository at 45 percent from one at 80 percent. Tightening to ±10 points requires about 97 probes and ±5 points about 385, so 40 is the sensible starting point and you only go bigger if the first result lands in an ambiguous range.\n\n### Isn't a low zero-result rate a sign the repository is healthy?\n\nNo. A zero-result search is the only retrieval failure that reports itself, so a low rate mostly means people are getting some results. It says nothing about searches that returned four items while missing eleven, and nothing about questions that were never typed. Use it as a vocabulary signal and rely on known-item recall for the health judgment.\n\n### Why does repository performance get worse as we add research?\n\nBecause the number of relevant findings per question grows while the results list stays the same length. If a question has 20 relevant findings and the searcher reads ten results, recall from that screen cannot exceed 50 percent regardless of ranking quality. Doubling the corpus roughly halves that ceiling, which is why consolidation matters as much as collection.\n\n### What is the difference between this and auditing for duplicate studies?\n\nA duplicate-study audit counts how often the same question was commissioned twice and assigns ownership for prevention. This is instrumentation of the search system itself. They catch different things: an audit finds cases where someone knew to ask and asked anyway, while these metrics find cases where the search was run and quietly failed, or was never run at all.\n\n### How do I protect against questions nobody thinks to ask?\n\nYou cannot fix that with better search, because retrieval requires the searcher to suspect the finding exists. The defences are push rather than pull — routing findings to the teams whose decisions they touch, surfacing research inside the tools where decisions are made, and lowering the cost of asking a fresh question so far that re-asking is cheaper than an undetected miss.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types and why structured answers never need retrieving\n- [Research Repository Guide](/docs/research-repository-guide) — setting up and maintaining a repository\n- [Re-Research Audit: Duplicate Studies](/docs/re-research-audit-duplicate-studies) — the governance counterpart to these retrieval metrics\n- [Activating Research Insights](/docs/activating-research-insights) — pushing findings to decisions instead of waiting for search\n- [Research Refresh Cadence](/docs/research-refresh-cadence) — deciding when a finding has expired\n- [Research Ops Guide](/docs/research-ops-guide) — the operating layer these metrics belong to","category":"Research Operations","lastModified":"2026-08-25T03:26:57.609936+00:00","metaTitle":"Research Repository Retrieval Metrics: Zero-Result Rate and Known-Item Recall","metaDescription":"Four cheap metrics that show whether your repository actually works: zero-result rate, known-item recall, time-to-first-relevant and the retrieval ceiling.","keywords":["research repository metrics","known-item test","zero result rate","repository health","retrieval metrics","search abandonment","repository recall"],"aiSummary":"Repository health should be measured by retrieval rather than volume. Zero-result rate, known-item recall from a 40-probe sample, time-to-first-relevant and the recall ceiling implied by list length together answer whether colleagues can reach existing research. The largest failure is the query nobody thought to formulate, which only push-based distribution addresses.","aiPrerequisites":["An existing research repository with some search history"],"aiLearningOutcomes":["Instrument four retrieval metrics for a research repository","Run a 40-probe known-item test and interpret its confidence interval","Explain why repository recall degrades as the corpus grows","Identify retrieval failures that governance audits cannot detect"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}