{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-30T18:50:41.011Z"},"content":[{"type":"documentation","id":"c395c644-9bb3-4409-85a3-e3a17e2463a5","slug":"research-blank-spike-control-samples","title":"Blanks and Spikes: Two Control Samples That Audit Your Research Instrument (2026)","url":"https://www.koji.so/docs/research-blank-spike-control-samples","summary":"Research pipelines fail in two independent directions: missing real themes and manufacturing absent ones. Analytical labs audit these with separate controls, and EPA SW-846 guidance pairs a laboratory control sample with matrix spikes to separate causes. A spike (participants confirmed to have an issue) measures recovery; a blank (participants confirmed clean, run through the identical procedure) measures contamination. Worked example: 87.5 percent recovery with 50 percent blank contamination yields only 77.8 percent precision.","content":"**Answer first:** Your research process can fail in two independent directions: it can miss a theme that was genuinely there, and it can report a theme that was not. Analytical laboratories audit these with two separate control samples, because neither one detects the other's failure. A spike is material with a known quantity added, and it measures recovery. A blank is material known to be clean, run through the entire procedure, and it measures contamination introduced by the process itself. You can run both on an interview study this week, and the blank is the one almost nobody runs. In Koji both are ordinary studies rather than special projects.\n\n## Two failure directions, two controls\n\nAsk a chemist how they know their method works and you will not get one answer. You will get a list, because a method can be wrong in ways that do not overlap.\n\nThe EPA's quality control guidance for its SW-846 methods names the division precisely. Of the laboratory control sample, a clean reference matrix carried through the whole method, it says: \"The primary purpose of the laboratory control sample (LCS) is to demonstrate that the laboratory can perform the overall analytical approach in a matrix free of interferences (e.g., in reagent water, clean sand, or another suitable reference matrix) and its analytical system is in control.\"\n\nThen, on why that is not sufficient on its own, it makes the structural point: \"Therefore, the LCS results should be used in conjunction with MS/MSD results to separate issues of laboratory performance and 'matrix effects.'\" The matrix spike and its duplicate, it notes, \"are an important measure of the performance of the method relative to the specific sample matrix of interest.\"\n\n**Two controls, used together, to separate two causes.** That is the idea worth importing, and it is more useful than either control alone.\n\n### The blank\n\n*LCGC* describes the concept: \"The concept of blanks-samples lacking the analyte of interest used to determine or track the source of contamination or sample degradation taken through the analytical process-is somewhat straightforward.\" The inference it licenses is clean: \"Any analytical signal emanating from a blank sample that is absent in a blank solvent can be attributed to contamination.\"\n\nCrucially, blanks are not a separate cheap test. They must travel the whole route: \"Blank samples are collected, stored, treated, and analyzed in a manner as close to that for authentic samples as possible to account for contaminants and other potential interferents or potential sample degradation.\" And different blanks localise different stages, because \"Method blanks are used to determine background contamination or interferences in the analytical system,\" whereas \"Field blanks can determine contaminants or analytical errors or bias, stemming from sample collection and analysis.\"\n\nThat distinction is the whole trick, and it transfers exactly.\n\n## Translating both controls to research\n\n### The spike: does your process find what is there?\n\nPlant something you know is present and check whether the process reports it.\n\nIn practice: take a set of participants you have independently confirmed experience a specific problem, perhaps from support tickets, session data or a prior verified interview. Run them through your normal study, normal guide, normal analysis. Count how many come out the other end with that theme attached.\n\nWhat you get is a **recovery rate**. If eight participants definitely have the problem and your analysis surfaces it for seven, recovery is 87.5 percent. That is a real, defensible number about your pipeline's sensitivity, and most teams have never computed one.\n\n### The blank: does your process invent what is not there?\n\nNow the control almost nobody runs. Take participants you have independently confirmed do *not* have the problem, and run them through the identical study and analysis. Count how many come out with the theme attached anyway.\n\nAny theme that appears here was manufactured somewhere between your question and your report. It is the research equivalent of a signal in a blank sample, and by the same logic it can be attributed to contamination rather than to the customer.\n\n### Why you need both, in one table\n\nTake 12 participants: 8 confirmed to have the issue, 4 confirmed not to (the blanks).\n\n| Measure | Result | What it tells you |\n| --- | --- | --- |\n| Recovery (spike) | 7 of 8 = 87.5 percent | The process finds most real instances |\n| Blank contamination | 2 of 4 = 50 percent | The process invents the theme half the time |\n| Reported instances | 7 + 2 = 9 | What lands in the report |\n| Precision | 7 of 9 = 77.8 percent | Nearly a quarter of reported instances are artefacts |\n\nLook at the recovery figure on its own: 87.5 percent is excellent, and a team that measured only recovery would conclude the pipeline is healthy. The blank says otherwise. **Two of the four clean participants produced a theme out of nothing, and one reported instance in five is an artefact.** No amount of spiking would have revealed that, because a spike can only tell you about material that genuinely contains the analyte.\n\nThis is the same structural blindness described in [When the Human Baseline Is Wrong](/docs/imperfect-gold-standard-ai-analysis-validation): a validation design can be incapable of seeing an entire error class. The blank exists specifically to close that gap.\n\n## Localising the contamination\n\nWhen a blank comes back dirty, you know something manufactured the theme. The EPA's partition logic tells you how to find out what. Run the blank at two different stages.\n\n**Stage one, the elicitation blank.** A confirmed-clean participant goes through the full interview. If the theme appears in the transcript, the *conversation* produced it. The guide planted it, the probe suggested it, or the participant inferred what you wanted to hear. That is a question-design defect, and the existing literature on it is good: see [How to Avoid Leading Questions](/docs/avoiding-leading-questions), [Demand Characteristics](/docs/demand-characteristics) and [The Priming Effect](/docs/priming-effect-research).\n\n**Stage two, the analysis blank.** Take a transcript you have read and verified contains no mention of the theme, and put it through the analysis step alone. If the theme appears in the output, the *analysis* produced it, not the interview. That is a coding or model defect rather than a guide defect.\n\nThose two blanks split a single symptom into two distinct fixes, and they are what makes this a diagnostic rather than a warning. Because Koji keeps the transcript and the analysis as separate artefacts, each blank can be run against the exact stage it is meant to test. The bias articles above tell you to avoid contaminating your instrument. The blank tells you how much your instrument is contaminating *right now*, which is a different and more actionable thing.\n\n## What makes a valid blank\n\nThe requirement that a blank travel the identical route is where most attempts fail, so it is worth being explicit.\n\n- **Same guide, same probes, same length.** A blank participant who gets a shortened interview has only blanked part of the chain.\n- **Same analysis path.** If your real studies get AI analysis plus a human review pass, the blank gets both.\n- **Blind to everyone involved.** Whoever moderates or reviews must not know which participants are blanks, or the control measures their vigilance instead of your process.\n- **Confirmed clean on independent evidence.** Do not use self-report to establish a blank, because self-report is part of the instrument you are auditing. Use behavioural or account data.\n- **Enough of them to mean something.** Four blanks give you a very coarse estimate. The detection arithmetic in [Nobody Mentioned It](/docs/detection-floor-nobody-mentioned-interviews) applies here too: a small number of blanks can only detect a high contamination rate.\n\n## Common mistakes\n\n- **Running only spikes.** Recovery looks reassuring and says nothing about false positives. This is the single most common version of the error.\n- **Using self-report to define the blank.** Circular. The instrument under test cannot certify its own control.\n- **Telling the reviewer which cases are controls.** You then measure attention, not process.\n- **Treating one dirty blank as a verdict.** It is a signal to investigate and localise, not a reason to discard a study.\n- **Auditing the model and not the guide.** Most manufactured themes in interview research originate in the conversation, not the analysis. Blank both stages, as [Research Process Controls vs Output Checks](/docs/research-process-controls-vs-output-checks) argues more generally.\n\n## How Koji makes both controls practical\n\nThe reason almost no research team runs blanks is not that the idea is unknown. It is that under a human-moderated model, control samples are pure overhead: four extra interviews that produce no findings, cost four hours of moderation, and exist only to audit your own process. That is the first thing cut from a timeline.\n\nKoji changes that calculation in a few concrete ways.\n\n- **Control interviews cost no moderator time.** Because the AI interviewer runs sessions asynchronously, adding four blanks and eight spikes to a study costs recruitment and credits, not a researcher's week. Overhead that was prohibitive becomes routine.\n- **Blinding is structural rather than procedural.** There is no human moderator who could know which participants are controls, so the most fragile requirement of a good blank is satisfied by the architecture instead of by discipline.\n- **The instrument is genuinely identical across participants.** A human moderator cannot ask twelve people the same question the same way; Koji's AI interviewer follows the same guide and the same probing logic every time, which is what makes a control sample comparable to a real one at all.\n- **Structured questions give you unambiguous scoring.** With six question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), a yes_no or single_choice question about the target issue produces a typed value per participant, so recovery and contamination are counted rather than judged.\n- **Analysis blanks are easy to run.** Because transcripts and the analysis step are separable, you can push a verified-clean transcript through analysis on its own and see whether a theme appears, which is the stage-two test above.\n- **Themes carry their supporting quotes.** Koji's analysis attaches the participant's verbatim words to each theme, so a suspected artefact can be checked against what was actually said instead of argued about.\n\nNone of this makes your instrument clean. It makes the cleanliness measurable, which is the precondition for improving it.\n\n## Frequently asked questions\n\n### Is a blank the same as a control group?\n\nNo, and conflating them causes confusion. A control group in an experiment does not receive the intervention, and you compare outcomes between groups to estimate an effect. A blank receives the full measurement procedure on material known to lack what you are measuring, and you examine the blank's own reading to detect contamination introduced by the procedure. One estimates an effect; the other audits an instrument.\n\n### How many blanks do I need?\n\nEnough to detect a contamination rate you would care about, which follows the same detection arithmetic as any other sampling question. Four blanks can only reliably reveal fairly high contamination. If you want to detect a 10 percent false-positive rate with reasonable confidence you need on the order of twenty, so a sensible pattern is a small number of blanks in every study and a larger audit periodically.\n\n### What if I cannot confirm that a participant is clean?\n\nThen you cannot construct a blank for that theme, and you should not pretend otherwise. Blanks work for issues with independent behavioural evidence: a feature never used, an error never logged, a workflow never triggered. For purely subjective themes such as a feeling about a brand there is no clean material available, and the honest response is to rely on elicitation-design safeguards instead.\n\n### Does the AI analysis need its own blank if the interview was clean?\n\nYes, and they are not interchangeable. A clean elicitation blank tells you the conversation did not plant the theme. It says nothing about whether the analysis step invents themes from a transcript that does not contain them. Those are separate failure points with separate fixes, which is exactly why the EPA guidance pairs two controls to separate two causes.\n\n### Will a dirty blank invalidate my study?\n\nUsually not, and the useful response is quantitative rather than binary. A blank gives you a contamination rate, and that rate lets you discount your reported frequencies rather than discard them. If 50 percent of blanks produced a theme, the reported prevalence of that theme is substantially inflated and you should say so in the report, treating the affected finding as directional while you fix the cause.\n\n### Is this worth it for a small study?\n\nThe spike usually is not; the blank often is, because it audits your guide rather than your sample. One or two blank participants in a study of a dozen is a small cost, and a manufactured theme discovered in a blank is usually a defect in the interview guide, which means finding it once improves every study that reuses that guide.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that make recovery and contamination countable\n- [When the Human Baseline Is Wrong](/docs/imperfect-gold-standard-ai-analysis-validation) - validating analysis when the reference standard is itself imperfect\n- [How to Avoid Leading Questions](/docs/avoiding-leading-questions) - the most common source of a dirty elicitation blank\n- [Demand Characteristics](/docs/demand-characteristics) - when participants tell you what they think you want\n- [Research Process Controls vs Output Checks](/docs/research-process-controls-vs-output-checks) - where instrument audits sit in a quality system\n- [Nobody Mentioned It](/docs/detection-floor-nobody-mentioned-interviews) - the detection arithmetic that also governs how many blanks you need\n","category":"Research Methods","lastModified":"2026-09-29T03:36:26.140063+00:00","metaTitle":"Blanks and Spikes: Two Control Samples That Audit Your Research Instrument (2026)","metaDescription":"Two control samples audit your research instrument: a spike measures what it misses, a blank measures what it invents. You need both.","keywords":["research quality control","validate interview guide","false positive themes research","spike recovery blank sample","audit research process"],"aiSummary":"Research pipelines fail in two independent directions: missing real themes and manufacturing absent ones. Analytical labs audit these with separate controls, and EPA SW-846 guidance pairs a laboratory control sample with matrix spikes to separate causes. A spike (participants confirmed to have an issue) measures recovery; a blank (participants confirmed clean, run through the identical procedure) measures contamination. Worked example: 87.5 percent recovery with 50 percent blank contamination yields only 77.8 percent precision.","aiPrerequisites":["Familiarity with interview study workflow","Basic understanding of false positives and false negatives"],"aiLearningOutcomes":["Distinguish recovery failure from contamination failure","Design a spike and a blank for an interview study","Localise a manufactured theme to elicitation or to analysis","Convert a blank result into a discount on reported frequencies"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}