Back to docs
Research Methods

Why Your First Three Interviews Always Go Worst: The Failure-Rate Curve of an Interview Guide

Interview guides fail on a curve, not at a constant rate. Learn the three regions - early-life, useful life, and wear-out - and the opposite fixes each one needs.

Your interview guide does not fail at a constant rate. It fails on a curve with three distinct regions, and the defects in each region have different causes, different costs, and opposite fixes. Reliability engineers have described this shape for decades. Most research teams have never applied it to their own instrument, so they keep treating a wear-out failure in session 22 as if it were the same problem as an ambiguity failure in session 2.

It is not. Early failures live in the guide. Late failures live in the interviewer. Confusing the two is why teams rewrite a guide that was never broken.

The three regions of a failure-rate curve

Reliability engineering calls the shape the bathtub curve: high failure rate at the start, a long flat middle, a rising tail. The NIST/SEMATECH e-Handbook of Statistical Methods describes the first region as one in which "The initial region that begins at time zero when a customer first begins to use the product is characterized by a high but rapidly decreasing failure rate."

The middle region is the one you plan for. NIST: "Next, the failure rate levels off and remains roughly constant for (hopefully) the majority of the useful life of the product."

And the tail: "Finally, if units from the population remain in use long enough, the failure rate begins to increase as materials wear out and degradation failures occur at an ever increasing rate."

Swap "product" for "interview guide" and "customer" for "interviewer" and the curve describes a research study almost exactly.

Region 1: early-life failure, and why it is always steeper than you expect

The first sessions of any study fail at an elevated rate because a guide is a piece of untested engineering until a real human has pushed on it. A question that reads cleanly on the page turns out to have two readings. A branch you wrote for an edge case catches half your participants. An opening question that seemed warm turns out to prime everything after it.

These are not interviewer mistakes. They are defects in the artifact, present from the moment it was written, and they surface on first contact with a real respondent. That is exactly the signature NIST describes: high, but rapidly decreasing. Each one you find is a defect removed permanently.

The five defects that only a live session reveals

  • Ambiguity that survived the draft. Two participants answer the same question about two different things, and neither is wrong.
  • Dead-end branches. A follow-up that assumes a "yes" and has nothing to say to a "no".
  • Order effects you introduced yourself. An early question that frames the vocabulary for every answer after it.
  • Unanswerable specificity. A question that requires a participant to recall something nobody remembers.
  • Length miscalibration. A guide written for 30 minutes that takes 52, so the last third is always rushed.

The last one is the most expensive because it is silent. Nothing looks broken; you just never get to the questions that mattered most.

Region 2: the flat middle you are actually paying for

Once the artifact defects are out, the failure rate levels off. What remains is irreducible participant variance: some people are more articulate, some had a bad day, some genuinely do not have the experience you recruited them for. This is the region where your sample size math is valid, and it is the only region where a usable-rate assumption is stable enough to plan against.

The planning error most teams make is assuming the whole study lives here.

Region 3: wear-out, and the tail nobody pilots for

Here is the part that the pilot-study literature does not cover, because a pilot by definition cannot reach it. At the end of a study, the failure rate climbs again - and the cause has moved.

The guide has not changed. The interviewer has. Late-study degradation shows up as:

  • Shortened probes. You already know what they are going to say, so you stop asking the third follow-up.
  • Anticipatory coding. You hear the answer you have heard eleven times and write it down before the participant finishes qualifying it.
  • Vocabulary drift. Your phrasing in session 23 is not the phrasing in session 3, because you have absorbed the participants' words.
  • Attention decay. The genuinely novel answer arrives late and gets filed under an existing theme.

Vocabulary drift is the worst of these, because it silently converts a standardized instrument into an unstandardized one partway through, and your analysis will treat all 24 sessions as if they were the same study.

The asymmetry: early failures live in the artifact, late failures live in the operator

This is the whole point of borrowing the curve, and it is the thing a generic "run a pilot" instruction misses. The two tails of the curve need opposite interventions:

  • Early-life failure is fixed by changing the guide. The interviewer is fine.
  • Wear-out failure is fixed by changing or resting the interviewer. The guide is fine.

Teams that only know about piloting apply the first fix to both problems. They reach session 22, see quality drop, and rewrite the guide - which destroys comparability across the study and fixes nothing, because the defect was never in the guide.

A worked illustration: what the curve costs on a 24-session study

The figures below are planning illustrations, not measured constants - your own rates will differ, and the point is the shape, not the numbers.

SessionsRegionUsable rateDominant defectWhere the defect lives
1-3Early-life60%Ambiguity, dead ends, bad orderThe guide
4-20Useful life90%Irreducible participant varianceNeither
21-24Wear-out70%Shortened probes, drift, anticipationThe interviewer

Total usable: (3 x 0.60) + (17 x 0.90) + (4 x 0.70) = 1.8 + 15.3 + 2.8 = 19.9 sessions.

If both tails were flattened to the useful-life rate, you would get 21.6. The two tails together cost 1.7 sessions - about 7% of a study you have already paid for in full. On a recruited B2B panel at a few hundred per participant, that is a real line item, and it is invisible on every dashboard because no single session is marked "failed".

How to shorten the early-life region

  • Pilot against the hardest recruit, not the easiest. Accelerated life testing stresses a unit deliberately to surface failures faster; the research equivalent is piloting with the participant most likely to break the guide, not the friendly colleague who will nod along.
  • Pilot the branches, not just the happy path. Walk each "no" explicitly.
  • Time the pilot end to end and cut to fit before session 1, not after session 6.
  • Treat every pilot defect as permanent removal. Write the fix into the guide immediately, or you will rediscover it in session 4.

How to detect wear-out before it eats your last sessions

  • Compare average probe count in your first five and last five sessions. A drop is drift, not saturation.
  • Diff your own phrasing. Pull the exact wording of one core question from session 3 and session 23 and read them side by side.
  • Watch for novel answers arriving late and being coded into old themes.
  • Do not schedule your highest-value participants last. Most teams do exactly this, saving the "best" interview for the end, which is precisely when the operator is least reliable.

How Koji handles this

The curve is a human-operator model, and the interesting consequence is what happens when the operator is not human.

  • Koji's AI interviewer has no wear-out region. Session 200 is conducted with the same probing depth, the same phrasing, and the same patience as session 1. Fatigue, anticipation, and vocabulary drift are operator properties, and an AI interviewer does not have them.
  • Koji's quality scoring makes the early-life region visible while you can still act on it. Every conversation is scored 1-5 across relevance, depth, and coverage, so a cluster of low scores in your first sessions shows up as a pattern instead of a vague feeling.
  • Koji's quality gate means low-scoring conversations do not consume credits, so early-life failures are not billed as if they were usable data.
  • Koji's structured questions remove a whole class of early-life defects by construction. The six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - cannot drift in wording between session 3 and session 23, because the instrument, not the interviewer, holds the phrasing.
  • Koji's AI follow-up questions keep probe depth constant. The third follow-up still gets asked in session 23, which is exactly the behaviour that decays first in a tired human moderator.

The honest framing: an AI interviewer does not eliminate the early-life region, because a badly written guide is still a badly written guide. It eliminates the wear-out region, which is the half of the curve that traditional research has no answer for at all.

Common mistakes

  • Rewriting the guide when late-session quality drops. The defect is in the operator by then; rewriting mid-study destroys comparability and fixes nothing.
  • Reading the early-life dip as "bad participants". Three weak sessions in a row at the start is an instrument signal, not a recruiting signal.
  • Piloting only the happy path, which leaves every branch defect to be discovered by a paying participant.
  • Assuming a flat usable rate across the whole study when you size the sample. Koji users who plan against the flat middle alone consistently under-recruit.
  • Scheduling the most important participant last.

Frequently asked questions

Is the early-life failure region just another name for a pilot study?

No. A pilot is one way to move the early-life region out of your paid sample, but the region exists whether or not you pilot. If you skip piloting, the region simply happens inside your real study and you pay for it with participants. The distinction matters because it tells you what a pilot is actually buying: not "practice", but the relocation of a predictable block of failures to cheaper sessions.

How many sessions does the early-life region usually last?

There is no universal constant, and anyone quoting one is guessing. What is reliable is the shape: the rate is highest at first contact and drops quickly as each artifact defect is permanently removed. Track it directly rather than assuming a number. Compare the usable rate of your first three sessions against sessions four onward, and you will see your own curve within a single study.

Why does the failure rate rise again at the end of a study?

Because the failure source moves from the artifact to the operator. The guide is stable by then, but the interviewer has absorbed the participants' vocabulary, can anticipate common answers, and starts shortening probes. The result is a genuine drop in data quality that looks like saturation and is not. Saturation means you are learning nothing new; wear-out means you have stopped asking properly.

Does an AI interviewer really have no wear-out region?

On the operator side, yes - and this is the cleanest structural advantage AI-moderated research has. Koji runs session 200 with identical probing depth and phrasing to session 1, so drift and fatigue do not accumulate. The early-life region still applies, because a guide with an ambiguous question is ambiguous to everyone. The curve does not disappear; its rising tail does.

How do I tell wear-out apart from genuine saturation?

Saturation is a property of the findings, wear-out is a property of the process. Check the process directly: count follow-up probes per session and compare your first five with your last five. If probe count fell, you have wear-out. If probe count held steady and the answers stopped surprising you, that is saturation. Teams routinely declare saturation when what actually happened is that they got tired.

Can I just run the whole study twice to average out the curve?

That doubles the cost and reproduces both tails a second time, so it is the most expensive possible fix. The cheap interventions are asymmetric by design: shorten the early-life region by stress-piloting the guide against your hardest recruit, and remove the wear-out region by not relying on a human operator for session 23. Running more of the same study does neither.

Related Resources