Back to docs
Research Methods

Blanks and Spikes: Two Control Samples That Audit Your Research Instrument (2026)

Analytical labs run two different controls because there are two different ways to be wrong. A spike measures what your process misses; a blank measures what it invents. Here is how to run both on a research pipeline.

Answer first: Your research process can fail in two independent directions: it can miss a theme that was genuinely there, and it can report a theme that was not. Analytical laboratories audit these with two separate control samples, because neither one detects the other's failure. A spike is material with a known quantity added, and it measures recovery. A blank is material known to be clean, run through the entire procedure, and it measures contamination introduced by the process itself. You can run both on an interview study this week, and the blank is the one almost nobody runs. In Koji both are ordinary studies rather than special projects.

Two failure directions, two controls

Ask a chemist how they know their method works and you will not get one answer. You will get a list, because a method can be wrong in ways that do not overlap.

The EPA's quality control guidance for its SW-846 methods names the division precisely. Of the laboratory control sample, a clean reference matrix carried through the whole method, it says: "The primary purpose of the laboratory control sample (LCS) is to demonstrate that the laboratory can perform the overall analytical approach in a matrix free of interferences (e.g., in reagent water, clean sand, or another suitable reference matrix) and its analytical system is in control."

Then, on why that is not sufficient on its own, it makes the structural point: "Therefore, the LCS results should be used in conjunction with MS/MSD results to separate issues of laboratory performance and 'matrix effects.'" The matrix spike and its duplicate, it notes, "are an important measure of the performance of the method relative to the specific sample matrix of interest."

Two controls, used together, to separate two causes. That is the idea worth importing, and it is more useful than either control alone.

The blank

LCGC describes the concept: "The concept of blanks-samples lacking the analyte of interest used to determine or track the source of contamination or sample degradation taken through the analytical process-is somewhat straightforward." The inference it licenses is clean: "Any analytical signal emanating from a blank sample that is absent in a blank solvent can be attributed to contamination."

Crucially, blanks are not a separate cheap test. They must travel the whole route: "Blank samples are collected, stored, treated, and analyzed in a manner as close to that for authentic samples as possible to account for contaminants and other potential interferents or potential sample degradation." And different blanks localise different stages, because "Method blanks are used to determine background contamination or interferences in the analytical system," whereas "Field blanks can determine contaminants or analytical errors or bias, stemming from sample collection and analysis."

That distinction is the whole trick, and it transfers exactly.

Translating both controls to research

The spike: does your process find what is there?

Plant something you know is present and check whether the process reports it.

In practice: take a set of participants you have independently confirmed experience a specific problem, perhaps from support tickets, session data or a prior verified interview. Run them through your normal study, normal guide, normal analysis. Count how many come out the other end with that theme attached.

What you get is a recovery rate. If eight participants definitely have the problem and your analysis surfaces it for seven, recovery is 87.5 percent. That is a real, defensible number about your pipeline's sensitivity, and most teams have never computed one.

The blank: does your process invent what is not there?

Now the control almost nobody runs. Take participants you have independently confirmed do not have the problem, and run them through the identical study and analysis. Count how many come out with the theme attached anyway.

Any theme that appears here was manufactured somewhere between your question and your report. It is the research equivalent of a signal in a blank sample, and by the same logic it can be attributed to contamination rather than to the customer.

Why you need both, in one table

Take 12 participants: 8 confirmed to have the issue, 4 confirmed not to (the blanks).

MeasureResultWhat it tells you
Recovery (spike)7 of 8 = 87.5 percentThe process finds most real instances
Blank contamination2 of 4 = 50 percentThe process invents the theme half the time
Reported instances7 + 2 = 9What lands in the report
Precision7 of 9 = 77.8 percentNearly a quarter of reported instances are artefacts

Look at the recovery figure on its own: 87.5 percent is excellent, and a team that measured only recovery would conclude the pipeline is healthy. The blank says otherwise. Two of the four clean participants produced a theme out of nothing, and one reported instance in five is an artefact. No amount of spiking would have revealed that, because a spike can only tell you about material that genuinely contains the analyte.

This is the same structural blindness described in When the Human Baseline Is Wrong: a validation design can be incapable of seeing an entire error class. The blank exists specifically to close that gap.

Localising the contamination

When a blank comes back dirty, you know something manufactured the theme. The EPA's partition logic tells you how to find out what. Run the blank at two different stages.

Stage one, the elicitation blank. A confirmed-clean participant goes through the full interview. If the theme appears in the transcript, the conversation produced it. The guide planted it, the probe suggested it, or the participant inferred what you wanted to hear. That is a question-design defect, and the existing literature on it is good: see How to Avoid Leading Questions, Demand Characteristics and The Priming Effect.

Stage two, the analysis blank. Take a transcript you have read and verified contains no mention of the theme, and put it through the analysis step alone. If the theme appears in the output, the analysis produced it, not the interview. That is a coding or model defect rather than a guide defect.

Those two blanks split a single symptom into two distinct fixes, and they are what makes this a diagnostic rather than a warning. Because Koji keeps the transcript and the analysis as separate artefacts, each blank can be run against the exact stage it is meant to test. The bias articles above tell you to avoid contaminating your instrument. The blank tells you how much your instrument is contaminating right now, which is a different and more actionable thing.

What makes a valid blank

The requirement that a blank travel the identical route is where most attempts fail, so it is worth being explicit.

  • Same guide, same probes, same length. A blank participant who gets a shortened interview has only blanked part of the chain.
  • Same analysis path. If your real studies get AI analysis plus a human review pass, the blank gets both.
  • Blind to everyone involved. Whoever moderates or reviews must not know which participants are blanks, or the control measures their vigilance instead of your process.
  • Confirmed clean on independent evidence. Do not use self-report to establish a blank, because self-report is part of the instrument you are auditing. Use behavioural or account data.
  • Enough of them to mean something. Four blanks give you a very coarse estimate. The detection arithmetic in Nobody Mentioned It applies here too: a small number of blanks can only detect a high contamination rate.

Common mistakes

  • Running only spikes. Recovery looks reassuring and says nothing about false positives. This is the single most common version of the error.
  • Using self-report to define the blank. Circular. The instrument under test cannot certify its own control.
  • Telling the reviewer which cases are controls. You then measure attention, not process.
  • Treating one dirty blank as a verdict. It is a signal to investigate and localise, not a reason to discard a study.
  • Auditing the model and not the guide. Most manufactured themes in interview research originate in the conversation, not the analysis. Blank both stages, as Research Process Controls vs Output Checks argues more generally.

How Koji makes both controls practical

The reason almost no research team runs blanks is not that the idea is unknown. It is that under a human-moderated model, control samples are pure overhead: four extra interviews that produce no findings, cost four hours of moderation, and exist only to audit your own process. That is the first thing cut from a timeline.

Koji changes that calculation in a few concrete ways.

  • Control interviews cost no moderator time. Because the AI interviewer runs sessions asynchronously, adding four blanks and eight spikes to a study costs recruitment and credits, not a researcher's week. Overhead that was prohibitive becomes routine.
  • Blinding is structural rather than procedural. There is no human moderator who could know which participants are controls, so the most fragile requirement of a good blank is satisfied by the architecture instead of by discipline.
  • The instrument is genuinely identical across participants. A human moderator cannot ask twelve people the same question the same way; Koji's AI interviewer follows the same guide and the same probing logic every time, which is what makes a control sample comparable to a real one at all.
  • Structured questions give you unambiguous scoring. With six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), a yes_no or single_choice question about the target issue produces a typed value per participant, so recovery and contamination are counted rather than judged.
  • Analysis blanks are easy to run. Because transcripts and the analysis step are separable, you can push a verified-clean transcript through analysis on its own and see whether a theme appears, which is the stage-two test above.
  • Themes carry their supporting quotes. Koji's analysis attaches the participant's verbatim words to each theme, so a suspected artefact can be checked against what was actually said instead of argued about.

None of this makes your instrument clean. It makes the cleanliness measurable, which is the precondition for improving it.

Frequently asked questions

Is a blank the same as a control group?

No, and conflating them causes confusion. A control group in an experiment does not receive the intervention, and you compare outcomes between groups to estimate an effect. A blank receives the full measurement procedure on material known to lack what you are measuring, and you examine the blank's own reading to detect contamination introduced by the procedure. One estimates an effect; the other audits an instrument.

How many blanks do I need?

Enough to detect a contamination rate you would care about, which follows the same detection arithmetic as any other sampling question. Four blanks can only reliably reveal fairly high contamination. If you want to detect a 10 percent false-positive rate with reasonable confidence you need on the order of twenty, so a sensible pattern is a small number of blanks in every study and a larger audit periodically.

What if I cannot confirm that a participant is clean?

Then you cannot construct a blank for that theme, and you should not pretend otherwise. Blanks work for issues with independent behavioural evidence: a feature never used, an error never logged, a workflow never triggered. For purely subjective themes such as a feeling about a brand there is no clean material available, and the honest response is to rely on elicitation-design safeguards instead.

Does the AI analysis need its own blank if the interview was clean?

Yes, and they are not interchangeable. A clean elicitation blank tells you the conversation did not plant the theme. It says nothing about whether the analysis step invents themes from a transcript that does not contain them. Those are separate failure points with separate fixes, which is exactly why the EPA guidance pairs two controls to separate two causes.

Will a dirty blank invalidate my study?

Usually not, and the useful response is quantitative rather than binary. A blank gives you a contamination rate, and that rate lets you discount your reported frequencies rather than discard them. If 50 percent of blanks produced a theme, the reported prevalence of that theme is substantially inflated and you should say so in the report, treating the affected finding as directional while you fix the cause.

Is this worth it for a small study?

The spike usually is not; the blank often is, because it audits your guide rather than your sample. One or two blank participants in a study of a dozen is a small cost, and a manufactured theme discovered in a blank is usually a defect in the interview guide, which means finding it once improves every study that reuses that guide.

Related Resources

Related Articles

How to Avoid Leading Questions in Surveys and Interviews

Leading questions quietly bias your research data. Learn how to spot and rewrite leading, loaded, and double-barreled questions — and how Koji's AI writes neutral questions and probes without steering respondents.

Demand Characteristics: When Participants Tell You What They Think You Want

Demand characteristics are the cues in a study that let participants guess your hypothesis and change their behavior to fit it. Learn where they come from, how they differ from social desirability and the Hawthorne effect, and how to design research that captures honest behavior.

When the Human Baseline Is Wrong: Validating AI Analysis Against an Imperfect Gold Standard (2026)

Checking AI coding against one senior researcher does not measure accuracy - it measures agreement with that person, errors included. Here is how imperfect reference standards bias the number, which direction, and what to do instead.

Process Controls vs Output Checks: How to Earn the Right to Read Fewer Transcripts

Evidence that your research process worked substitutes for evidence about each individual output. The trade auditors formalized, why existence is not operation, and how reperformance proves a control actually ran.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.