Test the Process or Test the Output: When Research Controls Let You Review Less Data
Auditing splits its work into tests of controls and substantive procedures. Learn how testing your research process lets you justifiably review less raw data, where the floor is, and which seven research controls to test first.
Answer first: auditing splits its work into two procedure families that answer two different questions. Substantive procedures look at the output and try to find the error directly. Tests of controls look at the process that produced the output and ask whether it reliably prevents the error. The rule that follows is the one research teams never apply: if you have tested your controls and they work, you are entitled to look at less raw evidence. If you have not, reading more transcripts buys you far less than you think.
The short answer
Most research advice pushes in one direction: more interviews, more responses, more coverage, more confidence. That advice is right when your process is untrusted, and it is expensive and partly wasted when your process is sound.
Auditing formalized the alternative decades ago. ISA 330 defines the two families with unusual precision:
| Procedure family | ISA 330 definition | What it looks at |
|---|---|---|
| Substantive procedure | "An audit procedure designed to detect material misstatements at the assertion level" | The output: the numbers, the transcripts, the responses |
| Test of controls | "An audit procedure designed to evaluate the operating effectiveness of controls in preventing, or detecting and correcting, material misstatements at the assertion level" | The process that produced the output |
The translation to research is direct. Reading forty transcripts to check whether your headline holds is a substantive procedure. Checking that your screener actually excluded the people it was supposed to exclude is a test of controls. Almost every research team does the first constantly and the second never.
The rule that inverts the usual advice
ISA 330.8 says the auditor shall test controls if the risk assessment "includes an expectation that the controls are operating effectively (that is, the auditor intends to rely on the operating effectiveness of controls in determining the nature, timing and extent of substantive procedures)."
Read the parenthesis carefully, because it contains the whole trade. Reliance on controls determines the nature, timing and extent of the direct evidence you then have to gather. Demonstrate that the process is sound and you may legitimately gather less. That is the inversion: in a controlled process, confidence goes up as the volume of direct review goes down, because the confidence is coming from somewhere else.
There is a price attached, and ISA 330.9 states it plainly: the auditor "shall obtain more persuasive audit evidence the greater the reliance the auditor places on the effectiveness of a control." You do not get the reduction for free. You buy it once, with evidence about the control, and then it pays out across every study that runs through that control.
The floor you can never go below
This is where the borrowed idea has to be constrained, and auditing constrains it explicitly. ISA 330.18: "Irrespective of the assessed risks of material misstatement, the auditor shall design and perform substantive procedures for each material class of transactions, account balance, and disclosure."
Irrespective. No matter how good the controls are, you never stop looking at the output entirely. A profession that spent decades building the case for process reliance still refused to let anyone rely on it alone.
The research version of that floor: however clean your recruiting pipeline and however consistent your moderation, somebody still reads the actual transcripts for every study that informs a real decision. Controls reduce how much you read and let you read it faster. They never replace reading it.
When testing the process is not optional
ISA 330.8(b) adds a second trigger, and it is the one that most resembles modern research at scale. You must test controls when "substantive procedures alone cannot provide sufficient appropriate audit evidence at the assertion level."
ISA 330.A24 explains when that happens: it "may occur when an entity conducts its business using IT and no documentation of transactions is produced or maintained, other than through the IT system."
That is an exact description of a large unmoderated study. If you field two hundred conversations, nobody is going to review two hundred conversations. There is no realistic substantive procedure that covers the population. The evidence that the study was well-run cannot come from reading the output, because you will only ever read a fraction of it. It has to come from the process. At scale, testing your controls stops being an efficiency and becomes the only available source of assurance.
The controls a research process actually has
Most teams have controls. Very few have ever tested one. A control is only a control if it is specified tightly enough that you could check whether it operated.
| Control | What it prevents | How you test it |
|---|---|---|
| Screener logic | Wrong-population data entering the sample | Re-run the screener answers of everyone admitted and count how many should have been excluded |
| Guide fidelity | Different participants being asked different things | Sample sessions and check every required question was actually asked |
| Probe symmetry | Deeper probing of the answers that confirm the hypothesis | Compare average follow-up depth on confirming versus disconfirming answers |
| Disconfirming coverage | A sample that structurally cannot contradict you | Check the recruit included churned, declined and non-adopting segments |
| Response quality gate | Low-effort and fraudulent responses reaching analysis | Count what the gate rejected and hand-check a sample of what it passed |
| Analysis independence | The person with the hypothesis grading the hypothesis | Check who coded, and whether a second coder saw the first coder's themes |
| Cut-off discipline | Findings that were overtaken between fieldwork and readout | Check the readout carries a fieldwork close date |
Notice that every test in the right column is cheap and none of them involve reading more transcripts. That is the point. Control tests are small, repeatable and reusable, and substantive review is large, bespoke and consumed by a single study.
Having a control is not evidence that it operated
ISA 330.A21 draws a line that research teams reliably blur: "Testing the operating effectiveness of controls is different from obtaining an understanding of and evaluating the design and implementation of controls."
A written discussion guide is a control design. Evidence that the guide was followed in this study is operating effectiveness. Teams routinely present the first as if it were the second: we have a standard guide, so the sessions were consistent. That claim has never been tested in most organizations, and when it is tested it usually fails, because human moderators adapt, skip, reorder and run out of time.
The same trap applies to screeners. Having a screener is not evidence that the wrong people were excluded. It is common to find that a screener admitted people it should have blocked, because the logic had a gap nobody re-read after the study was cloned from a previous one.
Rotation: how to keep control evidence fresh without retesting everything
ISA 330.14(b) contains a scheduling rule worth stealing verbatim. Where controls have not changed, the auditor "shall test the controls at least once in every third audit, and shall test some controls each audit to avoid the possibility of testing all the controls on which the auditor intends to rely in a single audit period with no testing of controls in the subsequent two audit" periods.
Two ideas in one sentence. First, control evidence has a shelf life, and three cycles is the outer limit. Second, and more subtle, you must stagger the tests so that you are never in a period with no control testing at all. Batching all your process checks into one quarter and then coasting is specifically prohibited.
The research version: pick one control per month and test it. Seven controls, tested on rotation, means every control has current evidence within a quarter and no month passes without a process check. This is perhaps two hours of work a month, and it is the highest-leverage two hours in a research operation.
ISA 330.13 also lists what shortens that interval, and one item transfers cleanly: whether "there have been personnel changes that significantly affect the application of the control." A new moderator, a new analyst or a new recruiting vendor resets your evidence about guide fidelity and probe symmetry. Test after the change, not on the old schedule.
When a control test fails
ISA 330.17 handles the failure case: if deviations are detected, the auditor determines whether the tests still support reliance, whether more control testing is needed, or whether "the potential risks of misstatement need to be addressed using substantive procedures."
That third option is the honest one and the one to plan for. A failed control test does not mean the study is void. It means the discount is withdrawn and you are back to reading the output. If your screener let through eight people who should not have been there, you do not throw away the study, you go and read those eight transcripts and decide what they did to the finding.
Auditing also offers a shortcut worth knowing. ISA 330.A23 describes the dual-purpose test: a single procedure designed to test a control and produce substantive evidence at the same time, "designed and evaluated by considering each purpose of the test separately." In research, reading five transcripts end to end is a dual-purpose test. You learn what participants said, and you learn whether the guide was followed. Score them separately or you will let a good finding excuse a broken process.
What this changes about allocating rigor
Our guide to the research risk model sets out the arithmetic for deciding which questions deserve a heavy study, and it treats control risk as something set by your org design rather than something you buy down by studying harder. Within a single study, that is exactly right. You cannot fix a weak recruiting pipeline in the two weeks you have for this question.
Across a year, it is the opposite. Control risk is the one component of that model you can permanently lower, and testing controls is how auditing does it. The distinction worth holding onto:
- Buying detection with a bigger, more rigorous study pays out once, for one question.
- Lowering control risk by fixing and testing a control pays out on every study that runs through it, forever.
A team that only ever buys detection is paying full price for every question. This is the strongest argument available for spending research-operations time on process rather than on more fieldwork, and it is an argument from efficiency rather than from tidiness.
How Koji makes controls testable rather than aspirational
The reason most research controls are never tested is that testing them requires reconstructing what a human moderator did across dozens of sessions, from recordings nobody has time to watch.
An AI-moderated platform changes the economics because the control is enforced at the point of execution rather than audited afterwards:
- Guide fidelity is structural. Koji's AI interviewer works from the study brief, so every participant is asked the required questions. Fidelity stops being a discipline problem and becomes a property of the system. See AI interview best practices.
- Probe symmetry is enforced by construction. The same follow-up logic runs for every participant, which means the depth of probing does not vary with whether the answer pleased the researcher. That is the single hardest control to maintain with human moderators, as interviewer bias sets out.
- The quality gate is an automated control test on every conversation. Each conversation is scored 1 to 5, and anything below 3 does not consume a credit. The score is visible per conversation, which means the deviation rate is a number you can read rather than a thing you assume. See how the quality gate works and understanding quality scores.
- Structured questions give you a detection-independent baseline. Koji supports six question types alongside conversational probing: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Closed types mean the same thing regardless of how long a conversation ran, so if your open-ended themes diverge across two segments but the closed answers do not, you are looking at a probing artifact rather than a real difference. That comparison is a control test you can run in minutes, and it is described in the structured questions guide.
The honest framing is not that the platform removes the need for control testing. It is that it makes the tests cheap enough to actually run, and it removes the two controls that human processes fail most often.
Common mistakes
- Presenting control design as control evidence. We have a standard guide is not a finding. We checked twelve sessions and the guide was followed in eleven is a finding.
- Relying on controls you have never tested. Untested reliance is not a discount, it is an assumption. ISA 330.8 only permits reduction where an expectation of effectiveness has been established by testing.
- Going to zero on substantive review. ISA 330.18 exists for a reason. Someone reads the raw material on every consequential study.
- Testing everything once and then coasting. The rotation rule specifically prohibits leaving subsequent periods with no control testing at all.
- Treating a failed control test as a failed study. It is a withdrawn discount. Go back and read the output.
- Testing only the controls that are easy to test. The two that matter most, probe symmetry and disconfirming coverage, are the two nobody checks.
One further limit is worth stating. Controls reduce the evidence you need; they do not decide when you have enough. That judgment is a separate discipline, covered in professional skepticism. And when a control failure is discovered too late to repair, the finding needs a formal flag rather than a softer verb, which is the subject of the graded hedge.
Frequently asked questions
What is the difference between a test of controls and a substantive procedure in research terms?
A substantive procedure examines the research output directly to find errors in it, such as reading transcripts to check whether the headline claim is supported. A test of controls examines the process that produced the output, such as verifying that your screener excluded the people it was designed to exclude. ISA 330 defines the second as a procedure "designed to evaluate the operating effectiveness of controls." Research teams do the first constantly and the second almost never.
Can strong controls really justify reviewing fewer transcripts?
Yes, within limits. ISA 330.8 explicitly permits reliance on effective controls to reduce the nature, timing and extent of direct evidence gathering. But two constraints apply: you must have actually tested the controls rather than assumed them, and ISA 330.18 requires substantive procedures regardless of how good the controls are. Controls reduce how much output you review. They never take it to zero.
Which research controls should I test first?
Start with probe symmetry and disconfirming coverage, because they are the two that most often fail and the two whose failure most directly corrupts a finding. Screener integrity is a close third because it is cheap to test: re-run the screener answers of everyone you admitted and count how many should have been excluded. Most teams find at least one admission error the first time they do this.
How often do research controls need retesting?
ISA 330.14 sets the outer limit at once every third cycle for unchanged controls, with some controls tested every cycle so no period passes untested. A practical research version is one control per month on rotation. Retest immediately after any personnel or vendor change, since ISA 330.13 identifies personnel changes as a specific reason that prior control evidence stops being reliable.
Is this the same as the research risk model?
No, and the difference matters. The research risk model decides how much rigor a given question deserves, treating your control environment as a fixed input. This article is about changing that input. Buying more rigor pays out for one question; lowering control risk pays out on every study that runs through the control.
Does an AI moderator remove the need for control testing?
It removes the need to test some controls and makes others cheap, but it does not remove the principle. Guide fidelity and probe symmetry become structural properties rather than discipline problems, and the quality score gives you a deviation rate you can read directly. Screener logic, disconfirming coverage and analysis independence still need testing, because those are decisions a human made when designing the study.
Related Resources
- The Research Risk Model: How to Decide Which Questions Deserve a Rigorous Study
- Structured Questions in AI Interviews
- Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
- Research Independence: Why the Team That Built the Feature Should Not Grade It
- Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support
- How to Choose What to Review: Sampling Rules for Tickets, Recordings, and Transcripts You Already Have
- Professional Skepticism in Research: The Duty to Doubt an Answer That Sounds Right
- The Graded Hedge: How to Flag a Finding You Cannot Fully Support
Related Articles
Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support
Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.
Research Independence: Why the Team That Built the Feature Should Not Grade It
The five threats to independence from professional ethics codes, applied to product research - and why structural independence is a different problem from cognitive bias.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
How to Choose What to Review: Sampling Rules for Tickets, Recordings, and Transcripts You Already Have
Every sampling guide covers who to recruit. This one covers what to read from evidence you already own - and why selecting the interesting items means you can describe but never estimate.
The Research Risk Model: How to Decide Which Questions Deserve a Rigorous Study
Rigor is a residual, not an input. Borrow the audit risk model to allocate research effort by what is left over after inherent risk and your existing controls, instead of by how important the question feels.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.