Back to docs
Research Methods

Research Pipeline Yield: Why Every Stage Passes and the Finding Still Arrives Wrong

Research stages sit in series, so their pass rates multiply rather than average. Seven stages at 95 percent deliver a correct finding 69.8 percent of the time. How to run a yield audit and fund the lowest stage.

Bottom line up front: A research finding has to survive seven handoffs before it changes a decision, and each handoff is judged on its own. Reliability engineering says that is the wrong operation. Stages in series multiply, they do not average. Seven stages that each pass at 95 percent deliver a correct finding 69.8 percent of the time; at 90 percent each, they deliver it 47.8 percent of the time. Nobody's stage is failing. The chain is. Measure first-pass yield per stage, multiply, and fund the lowest number.

The operation almost every research team gets wrong

Ask a research team how good its process is and you will get a list. Recruiting is solid. The discussion guide was reviewed. Transcription is "pretty accurate now." Two people code. The readout was well received. Every item on that list is a local judgement, and every one of them can be true while the output is unreliable, because the list is being read as an average when the mathematics is a product.

The National Institute of Standards and Technology states the rule in one line in its Engineering Statistics Handbook: "Add failure rates and multiply reliabilities in the Series Model." The series model itself is defined there as the case used "to go from individual components to the entire system, assuming the system fails when the first component fails and all components fail or survive independently of one another."

That is exactly a research pipeline. A finding that survives recruiting but dies in coding is dead. A finding that survives coding but is misstated in the readout is dead. There is no parallel path. One bad stage sinks the finding regardless of how good the others were, which is why the arithmetic is multiplication.

Yield at every stage5 stages7 stages9 stages
99%95.1%93.2%91.4%
97%85.9%80.8%76.0%
95%77.4%69.8%63.0%
90%59.0%47.8%38.7%

Computed directly from the series model as yield raised to the power of the number of stages.

Read the 95 percent row again. A team where every single stage is doing an A-grade job, by its own honest assessment, ships a correct finding about seven times in ten. That is the good case. The 90 percent row, which is a fair description of most real pipelines, is a coin flip.

The seven stages, and what "pass" means at each

The stage list is not arbitrary. It is the set of points where the finding changes hands or changes medium, because those are the points where information is lost.

#StageWhat a pass meansWhat a failure looks like
1Frame the questionThe question, as written, can be answered by the people you can reachThe study answers a question nobody asked
2Recruit and screenThe participant is who the screener says they areA professional respondent who learned the right answers
3ElicitThe participant said what they actually thinkSatisficing, acquiescence, a moderator steering the answer
4CaptureThe transcript matches the audioWords dropped, especially the unusual ones
5CodeTwo competent analysts assign the same codeCodes that track the coder rather than the content
6SynthesizeThe claim is entailed by the evidence behind itA theme with three quotes and no denominator
7Transmit and decideThe decision-maker acts on what the study foundThe deck says "customers want X"; the study said "eleven of forty said X"

Every one of these is a real gate with a real pass rate. Three of them have published pass rates that are far worse than practitioners assume.

Three stages where the published numbers are brutal

Stage 2, screening. Andrew Bell and Thomas Gift hired a well-known commercial market research firm and fielded a survey of United States Army members and veterans across two rounds in 2021. Their finding, published in the Journal of Experimental Political Science 10(1):148-153, 2023, is that over 81 percent of respondents appeared to misrepresent their credentials to get into the study: 43.3 percent failed a basic Army knowledge question, a further 35.5 percent passed the knowledge screen but supplied non-viable service information, and 3.0 percent reported improbable Army backgrounds. That is a stage yield below 20 percent on a targeted business-to-business style population.

Stage 2, again, and why the channel matters more than the vendor. Joshua Gordon and colleagues, writing in Health Expectations 27(3), 2024, ran a survey open to a small target population and classified responses for suspected fraud. During the first twelve days the study was open, they classified 12 of 69 responses, or 17.4 percent, as suspected fraud. After the study was posted to social media, they classified 1,475 of 1,774 responses, or 83.1 percent, as suspected fraud. The same instrument, the same screener, the same team. The yield of a stage is not a property of your process alone; it is a property of your process crossed with the population flowing through it.

Stage 4, capture. Adam Miner and colleagues at Stanford measured automatic speech recognition against psychotherapy session audio and reported in npj Digital Medicine 3:82, 2020 that their HIPAA-compliant system "demonstrated a transcription word error rate of 25%." The number that should worry a research team is the next one: "For clinician-identified harm-related sentences, the word error rate was 34%." The error rate was higher on exactly the sentences that mattered most. Unusual vocabulary, emotional speech, and domain terms are harder to transcribe, and those are the utterances a researcher is hunting for. A stage does not lose information uniformly. It loses the tails first, and the tails are where the finding lives.

First-pass yield, not final yield

The number teams instinctively report is final yield: of the studies we ran, how many produced something usable. That number is flattering because it counts rework as success. Quality engineering calls the rework you do not count the hidden factory: the re-recruits after a bad screener batch, the second coding pass after the first one disagreed, the analyst who quietly re-listened to six recordings because the transcript was garbled.

First-pass yield is the fraction that clears the stage without rework. It is the only version of the number that tells you where the money is going, because rework is not free even when it is invisible; it is paid in researcher weeks and in elapsed time before the decision.

The practical test is one question per stage: in the last ten studies, how many times did this stage have to be done twice? Three out of ten is a 70 percent first-pass yield, whatever the final quality looked like.

Only the lowest number is worth funding

Here is a worked chain for a mid-sized product research team. The numbers are illustrative, but they are the shape of what a yield audit typically finds.

StageFirst-pass yield
Frame the question95%
Recruit and screen90%
Elicit97%
Capture (transcription)75%
Code85%
Synthesize95%
Transmit and decide90%

Multiplied through, the chain delivers 45.2 percent. Now consider two investments. Push the best stage, elicitation, from 97 to 99 percent, and the chain goes to 46.1 percent. Push the worst stage, capture, from 75 to 95 percent, and the chain goes to 57.2 percent. The second move is worth more than twelve times the first, and it costs less, because fixing a 75 percent stage is easy and fixing a 97 percent stage is hard.

This is the whole practical payoff of the series model. It does not tell you to improve everything. It tells you that improving anything other than the minimum is close to wasted, and it gives you the ratio that proves it.

This is a different question from total survey error. Total survey error asks how far a single estimate is from the truth and how to split a budget across the sources of that distance; see total survey error for that framework, which is the right one when the output is a number. Yield asks a blunter question: what fraction of the findings that enter the pipeline come out the other end intact? The two are complements, and most teams have neither.

Running a yield audit in an afternoon

  1. List your stages by handoff, not by activity. If the same person does two things without the artefact changing hands or changing medium, it is one stage. Handoffs are where yield is lost.
  2. Define pass at each stage as a binary. Not a score. "Did this stage produce something the next stage could use without going back?"
  3. Score the last ten studies retrospectively. Ten is enough to find the minimum, which is all you need. Do not build a dashboard.
  4. Multiply. Say the number out loud to the team. It is usually the first time anyone has.
  5. Fund the minimum, then re-audit. The minimum moves. When capture goes from 75 to 95 percent, coding at 85 percent becomes your constraint.
  6. Separate first-pass from final. Count the rework. That gap is your hidden factory and it is where the elapsed time goes.

A word of warning on step 2: the reflex after a low audit is to add a review step at each gate. That is the single most common response and it does not work for reasons that are exactly quantifiable. A review step is itself a stage in the chain, and its own detection rate is far lower than people assume. Spot-check inspection and the all-or-none rule covers why, and what to do instead.

How Koji changes the chain rather than patching it

The reason research pipelines have low yield is that they have too many stages, each run by a different person, each with a handoff. AI-native research does not improve stage 4 from 75 to 95 percent. It deletes stages 4 and 5 as separate steps.

  • Stages 3, 4 and 5 collapse into one. A Koji AI-moderated interview produces the transcript and the coded output as a by-product of the conversation itself. There is no separate capture step to lose the tails and no separate coding step to disagree with itself, because the analysis is generated from the same record, not from a re-typing of it.
  • Stage 2 gets a real gate. Every interview is scored 1 to 5 for quality against the study's research goals as it completes, so a fraudulent or satisficing participant is visible immediately rather than at synthesis, when the batch is already closed.
  • Stage 3 loses its variance. A human moderator's yield depends on who they are and what time it is. An AI moderator asks the same question the same way to participant 1 and participant 240. See interviewer bias for the evidence on how large that variance is.
  • Stage 1 gets structure. Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - force the framing decision to be made once, up front, and carry a stable question ID from interview plan through analysis to report. A question with an ID cannot quietly become a different question by the time it reaches the readout, which is the classic stage 7 failure.
  • Stage 7 shortens. When the report is generated from the interviews rather than transcribed from them by an analyst under deadline, the distance between what the study found and what the deck says collapses.

Traditional tooling optimises stages. A survey platform makes stage 2 cheaper. A transcription vendor makes stage 4 faster. A repository makes stage 6 tidier. None of them changes the exponent, and the exponent is what is killing the number.

Honest objections

"Our stages are not independent, so the multiplication is wrong." Correct, and it usually makes things worse rather than better. The series model assumes components "fail or survive independently." In a research pipeline the failures are positively correlated: the study with the rushed screener is also the study with the rushed guide. Positive correlation means the true joint yield is lower than the product at the low end of the distribution. Treat the multiplied number as an optimistic bound.

"Ninety-five percent per stage is pessimistic for a good team." Then use your own numbers. The point of the audit is that the exponent is doing the damage, not the base. Even 99 percent per stage over nine stages is 91.4 percent, which means roughly one finding in eleven is wrong for purely procedural reasons before anyone has argued about interpretation.

"This treats research as manufacturing." It treats the handling of research as manufacturing, which it is. Framing the question and interpreting the finding are craft. Getting a true sentence from a participant's mouth into a decision-maker's head without corrupting it is logistics, and logistics has laws.

"We cannot measure yield at the decide stage." You can measure a proxy: how often the claim in the final artefact is traceable to the evidence behind it. The roadmap evidence audit is a working procedure for exactly that check.

Frequently asked questions

What is rolled throughput yield in a research context?

Rolled throughput yield is the probability that a unit passes every stage of a process without rework. For research, the unit is a finding and the stages are the handoffs between framing, recruiting, eliciting, capturing, coding, synthesizing and deciding. You compute it by multiplying the first-pass yield of each stage. Seven stages at 95 percent give 69.8 percent; seven at 90 percent give 47.8 percent.

Why multiply the stage yields instead of averaging them?

Because the stages are in series: the finding must survive all of them, and any one failure kills it. The NIST Engineering Statistics Handbook states the rule as "Add failure rates and multiply reliabilities in the Series Model." Averaging assumes a good stage can compensate for a bad one, and in a series chain it cannot. A 99 percent stage does not rescue a 60 percent stage.

What is the difference between first-pass yield and final yield?

First-pass yield counts only the units that clear a stage with no rework. Final yield counts everything that eventually clears, including work that had to be redone. Final yield is always higher and always more flattering. The gap between them is the hidden factory: the second coding pass, the re-recruit, the re-listen. That gap is where research calendars disappear.

How accurate is automatic transcription in practice?

The best public measurement on conversational clinical audio is Miner and colleagues in npj Digital Medicine 3:82, 2020, which reported a 25 percent word error rate overall and 34 percent on clinician-identified harm-related sentences. The important pattern is that the error rate was higher on the most consequential utterances, because unusual and emotionally loaded language is harder to recognise. Assume your capture stage loses the tails first.

How many participants misrepresent themselves to get into a study?

It depends enormously on the recruitment channel, not just the vendor. Bell and Gift found that over 81 percent of respondents in a commercially recruited military sample appeared to misrepresent their credentials. Gordon and colleagues found 17.4 percent suspected fraud from their initial channel and 83.1 percent after the study was posted to social media. Treat your screening yield as a property of the channel and re-measure it whenever the channel changes.

Does adding a quality review step raise the pipeline yield?

Usually much less than expected, and sometimes not at all. A review step is itself a stage in the series, it has its own detection rate, and a sample review of a handful of items detects a moderate defect rate far less reliably than intuition suggests. The all-or-none inspection rule gives the exact condition under which reviewing is worth doing at all, and it is a corner solution: inspect nothing, or inspect everything.

Related Resources

Related Articles

Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement

Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.

Interviewer Bias: How Moderators Distort Research (and How AI Removes the Variance)

Interviewer bias is the distortion caused by a moderator's wording, reactions, expectations, and characteristics. Learn the types, the evidence, mitigation techniques, and why an AI interviewer eliminates interviewer variance.

The Re-Research Audit: How Much of Your Budget Buys an Answer You Already Own

Count how many of your last twenty studies answered a question you already owned. The protocol, the four causes, and where the duty belongs.

You Cannot Spot-Check Your Way to Data Quality: The All-or-None Rule for Research QA

A ten-item spot check accepts a 5 percent defective batch 59.9 percent of the time. Deming's all-or-none rule says inspect nothing or inspect everything, and sampling is optimal essentially never.

When No Study Was Wrong: Why Research Programs Fail Without a Defective Study

Some of the worst research-driven decisions trace to no bad study at all. Every study was true; the loss came from the interactions between them. The safety-engineering framework for losses with no component failure, applied to research.

Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)

Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.