Process Controls vs Output Checks: How to Earn the Right to Read Fewer Transcripts
Evidence that your research process worked substitutes for evidence about each individual output. The trade auditors formalized, why existence is not operation, and how reperformance proves a control actually ran.
Answer first: most research quality advice tells you to check more outputs - read more transcripts, re-code more responses, sanity-check more numbers. Auditing takes the opposite route and makes it explicit. If you can produce evidence that the process which generated the outputs was working, you are entitled to examine fewer outputs. The catch, and it is the whole discipline, is that the evidence has to be about the control operating, not about the control existing.
The short answer
Auditing splits every procedure into exactly two categories. PCAOB AS 2301.10: "The audit procedures performed in response to the assessed risks of material misstatement can be classified into two categories: (1) tests of controls and (2) substantive procedures."
A test of controls asks: did the machine that produces these outputs work? A substantive procedure asks: is this particular output right?
They trade against each other. AS 2301.16 explains that obtaining evidence that controls operate effectively allows the auditor to assess control risk at less than the maximum by relying on those controls, which in turn changes the extent of substantive work required. That is the inversion worth importing: more evidence about the process means less evidence needed about each output.
Almost every research quality practice is substantive. You read the transcripts. You spot-check the coding. You eyeball the open-ended responses for nonsense. All of that is checking outputs one at a time, and its cost scales linearly with the size of the study while its coverage does not.
What a control looks like in research
A control is a step in your process whose specific job is to prevent or detect a particular error. Naming them is most of the work, because most teams have controls they have never described as controls.
| Control | The error it exists to stop | How you would test that it operated |
|---|---|---|
| Screener logic with verification | Wrong people enter the study | Re-run the screener criteria against the final participant list |
| Piloted interview guide | Ambiguous or leading questions | Check the pilot happened and that questions changed as a result |
| Identical brief across comparison segments | Unequal probing depth between groups | Compare the brief and follow-up configuration used per segment |
| Randomized option order | Order effects in choice questions | Verify randomization was on and inspect the realized order distribution |
| Codebook plus double-coding on a slice | Analyst-invented themes | Independently re-code a slice and compare agreement |
| Consistency and attention checks | Low-effort or fraudulent responses | Confirm the checks were present and read the flag rate |
Once these exist and you can evidence that they ran, the argument for reading every transcript weakens considerably - not because reading transcripts is bad, but because you have already tested the thing that would have produced the error.
The catch: existence is not operation
This is where the import earns its keep, and it is the part research teams get wrong.
AS 2301.21 requires the auditor to test whether "the control is operating as designed and whether the person performing the control possesses the necessary authority and competence." Two separate conditions, and neither is satisfied by a document.
A written interview guide that nobody followed is a failed control. A codebook that one analyst used and the other ignored is a failed control. A screener whose criteria were overridden three times by a recruiter under deadline pressure is a failed control. In each case the control exists, is documented, and would pass any inspection that consists of asking whether it exists.
The distinction has a formal name in auditing - design effectiveness versus operating effectiveness - and the standards insist on testing the second. A research team that only ever confirms its templates are in place has tested design and called it quality.
Reperformance: the procedure that settles it
There is one procedure that cleanly distinguishes "the control exists" from "the control worked": do a slice of the work again, independently, and compare.
Reperformance is not review. Review reads what somebody produced and looks for problems, which inherits the original analyst's frame - if they missed a theme entirely, a reviewer reading their output will usually miss it too. Reperformance ignores the output until the second pass is finished, then compares.
Practical reperformance in research:
- Coding. A second analyst codes ten transcripts from raw, without seeing the first codebook. Then compare codebooks, not just code assignments. Divergent themes are the finding; divergent labels for the same theme are noise.
- Screening. Re-apply the screener criteria to the final participant roster and count how many participants would not have qualified.
- Analysis. Recompute the two or three headline numbers from the raw response data rather than from the summary table.
- Interview execution. Take five transcripts and check, against the brief, whether every required question was actually covered and whether follow-up depth was comparable across segments.
That last one is the highest-yield check in most qualitative studies, and it almost never gets run.
How much substantive work can you actually skip?
Auditing is disciplined about not letting controls testing become a blank cheque, and the rules translate well.
You cannot rely on a control you have not tested recently. ISA 330 permits using evidence about a control's operating effectiveness from a previous audit, but only after establishing that the control has not changed - and even then, the control must be tested at least once in every third audit. The standard also bars relying on prior-period evidence for controls that mitigate a significant risk; those must be tested in the current period.
The research translation is direct. If your screener and interview guide have not changed and you tested them last quarter, you can lean on that for routine studies, on a rotation, so that some controls get re-tested every cycle. But for the study answering your highest-risk question, test the controls now. Never coast on a control that guards the thing you most need to be right about.
Substantive work never goes to zero. AS 2301.17 requires testing controls where "substantive procedures alone cannot provide sufficient appropriate audit evidence" - but the reverse duty stands too. Controls testing tells you the process was sound; it cannot tell you what people said. You still have to read the material. What changes is how much of it you have to read defensively rather than curiously.
Risk sets the floor. AS 2301.09a requires the auditor to "Obtain more persuasive audit evidence the higher the auditor's assessment of risk," and AS 2301.37 adds that as assessed risk increases, "the evidence from substantive procedures that the auditor should obtain also increases." Strong controls lower the required substantive extent; they do not lower it below what the question's risk demands. Assessing that risk is covered in the research risk model.
How this differs from the two neighbours you already have
This is a different move from two practices your team may already run, and the distinctions are worth stating.
Pre-launch peer review examines one study design before it fields and catches broken questions, missing segments and unusable scales. It is a control, and a good one. See the research peer review QA gate. But reviewing a design is testing design effectiveness. This article is about testing whether the control operated across the study, after fielding.
Response-level data quality work - detecting straightlining, speeders, contradictory answers, duplicate identities - is substantive testing of outputs, applied to respondents. See survey data quality and survey fraud and respondent quality. It is necessary and it is not a substitute. Those techniques tell you which responses to discard; they say nothing about whether the instrument that collected them was administered consistently.
The three fit together: peer review checks the design, controls testing checks the administration, and data quality checks the output. Most teams run the first and third and skip the middle one entirely.
A control matrix you can start from
Auditing classifies controls by when they act, which is a useful lens for deciding what to build next.
| Type | When it acts | Research example | Failure signature |
|---|---|---|---|
| Preventive | Before the error can occur | Screener logic; piloted wording; locked option order | Errors that simply do not appear in your data |
| Detective | After the error, before the decision | Attention checks; double-coding; reperformance of headline numbers | A flag rate you can read and act on |
| Compensating | Covers a known gap in another control | Post-hoc weighting when a quota underfilled | A stated adjustment in the report |
The practical guidance: preventive controls are cheaper per study but invisible, which is exactly why teams under-invest in them and over-invest in end-of-study cleanup. A detective control that fires often is telling you a preventive control is missing. Track the flag rate over time and let it drive template changes rather than one-off fixes.
Where an AI moderator changes the control risk itself
Most research controls are asked to constrain a human under time pressure, which is why they degrade. A moderator running ten enterprise and ten SMB interviews across two weeks will not probe both segments identically, no matter what the guide says - and that drift is invisible in the output, because a richer transcript reads like a richer segment rather than like a harder-probed one.
This is where an AI-moderated platform changes the arithmetic rather than just the cost. In Koji, the brief and the follow-up configuration are the control, and they operate identically on interview one and interview two hundred, in voice or in text, at 2am or at 2pm. Control risk on consistency of administration is genuinely lower, and it is lower structurally rather than through discipline.
Koji's structured questions supply the other half - a detection-independent baseline. Closed question types mean the same thing regardless of how long the conversation ran: scale for magnitude, single_choice and multiple_choice for selection, ranking for relative order, and yes_no for a binary. Those aggregate deterministically across every participant. open_ended questions carry the depth, with AI follow-up probing and coded themes tied back to verbatim quotes. The useful test falls out of the pairing: if open-ended themes diverge sharply between two segments but the closed data does not, you are probably looking at a probing artifact rather than a real difference. That is a control you can actually run. See the structured questions guide.
Honest limits: an AI moderator makes administration consistent, not correct. A leading question asked identically two hundred times is a consistently administered bad instrument, which is precisely why the pre-launch design review remains a separate and necessary control.
Common mistakes
Documenting controls and calling it testing. A folder of templates is design evidence. You need evidence of operation.
Reviewing instead of reperforming. Reading someone's coding inherits their frame. Re-code from raw, then compare.
Testing every control every time. Rotate the stable ones; always test the controls guarding your highest-risk question.
Letting controls testing replace reading the research. It changes how much you must read defensively. It does not replace the part where you learn something.
Treating a low flag rate as good news. A detective control that never fires may be working, or may not be running. Confirm it fired at least once on data you seeded.
Frequently asked questions
What is the difference between a test of controls and a substantive check?
A test of controls asks whether the process that produced your data was working - was the screener applied, was the guide followed, did double-coding happen. A substantive check asks whether a specific output is correct. Auditing standards treat them as the two categories of procedure, and evidence about the first reduces how much of the second you need.
Why is it not enough that a control exists?
Because standards require evidence that a control is operating as designed and that the person performing it has the necessary authority and competence. A documented interview guide that nobody followed is a failed control that would pass any check consisting of confirming the guide exists.
What is reperformance in a research context?
Doing a slice of the work again independently and comparing. A second analyst codes ten transcripts from raw without seeing the first codebook; the screener criteria are re-applied to the final roster; headline numbers are recomputed from raw data. It differs from review, which reads the existing output and therefore inherits its frame.
How often should research controls be re-tested?
Rotate them. Controls that have not changed can be re-tested periodically rather than every study, on the model of ISA 330, which permits reliance on prior evidence for unchanged controls but requires testing at least once in every third audit. Controls guarding your highest-risk question should be tested every time.
Can strong controls let me skip reading transcripts entirely?
No. Controls testing tells you the process was administered soundly; it cannot tell you what people said. It reduces how much material you must read defensively looking for errors, which frees time for reading it curiously looking for insight.
Does this replace our pre-launch research review?
No. Pre-launch review tests design effectiveness before fielding. This tests operating effectiveness after fielding. Together with response-level data quality checks they cover design, administration and output, and most teams run only the first and last.
Related Resources
- The research risk model - how much assurance a question needs in the first place
- The research peer review QA gate - the pre-launch design control this sits alongside
- Choosing what to review - how to pick the slice you reperform
- Survey data quality - substantive checking at the response level
- The structured questions guide - the six question types and the detection-independent baseline
- Data annotation quality - agreement metrics and gold tasks for coded data
Related Articles
Data Annotation Quality: Guidelines, Agreement Metrics, and Gold Tasks That Actually Work (2026)
A practical guide to running an annotation and data-labeling operation: writing guidelines that resolve edge cases, choosing between Cohen's kappa and Krippendorff's alpha, setting gold-task and honeypot rates, and adjudicating disagreement instead of averaging it away.
Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support
Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
How to Choose What to Review: Sampling Rules for Tickets, Recordings, and Transcripts You Already Have
Every sampling guide covers who to recruit. This one covers what to read from evidence you already own - and why selecting the interesting items means you can describe but never estimate.
The Research Risk Model: How to Decide Which Questions Deserve a Rigorous Study
Rigor is a residual, not an input. Borrow the audit risk model to allocate research effort by what is left over after inherent risk and your existing controls, instead of by how important the question feels.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Data Quality: How to Detect and Prevent Bad Responses (2026)
The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.
Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)
Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.