Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support
Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.
Answer first: the audit profession defines three distinct service levels - reasonable assurance, limited assurance, and agreed-upon procedures - and each one has its own required wording, its own evidence burden, and its own explicit statement of what it does not cover. Product research has one register. A readout from six interviews and a readout from six hundred responses arrive in the same font, with the same confident headline, and the reader has no way to tell them apart. Declaring the assurance level before you field costs nothing and prevents the single most common failure in research communication: a small exploratory study being quoted six months later as if it settled the question.
This is not an argument for doing more research. It is an argument for saying, in advance and in writing, what kind of conclusion the study you are about to run will be entitled to state.
The three levels a profession invented when it had to say how sure it was
Accountants had a problem product teams will recognise. Clients wanted a signature on a number. Some clients could afford an exhaustive examination and some could not, and some only wanted a few specific checks. Rather than let one word cover all three situations, the International Auditing and Assurance Standards Board built three separate engagement types with three separate report formats.
| Level | Standard | Form of conclusion | Typical wording | Research equivalent |
|---|---|---|---|---|
| Reasonable assurance | ISA 200 | Positive opinion | "In our opinion, the financial statements present fairly..." | A powered study with a pre-registered primary outcome |
| Limited assurance | ISRE 2400 (Revised) | Negative conclusion | "Based on our review, nothing has come to our attention that causes us to believe..." | A moderate sample checked against a specific hypothesis |
| No assurance | ISRS 4400 (Revised) | Findings only, no conclusion | "We performed the procedures below and report the following findings" | Five discovery interviews, a read-out of what was said |
The interesting part is not the hierarchy. It is that the lowest level has a report format at all. A profession that lives on its opinions built a formal, standard-governed way to hand over facts and explicitly decline to draw a conclusion from them. Research has no such format, so every study output drifts upward into the language of the top row.
Reasonable assurance is high, and it is still not absolute
The top level is worth understanding precisely, because its name is a piece of deliberate honesty. Reasonable assurance is defined as a high level of assurance, obtained when the auditor has gathered sufficient appropriate evidence to reduce audit risk to an acceptably low level. ISA 500 puts the mechanism plainly: reasonable assurance is obtained when the auditor has obtained sufficient appropriate audit evidence to reduce audit risk - that is, the risk that the auditor expresses an inappropriate opinion when the financial statements are materially misstated - to an acceptably low level.
Acceptably low. Not zero. The assurance framework is explicit that engagement risk is never reduced to nil and there can therefore never be absolute assurance. The strongest thing the profession will say about a set of accounts it has spent months examining is a bounded claim with a named residual error rate.
Compare that with the standard product research headline. "Users want simpler onboarding." No sample, no residual risk, no scope, no statement of what would have falsified it. The claim is grammatically stronger than anything an auditor is permitted to write after a full-year engagement.
Limited assurance: the most useful level, and the one research never uses
The middle row is the one worth stealing. In a review engagement under ISRE 2400 (Revised), the practitioner gathers less evidence than in an audit and is therefore not permitted to state a positive opinion. Instead the standard prescribes the exact phrase for an unmodified conclusion: "Based on our review, nothing has come to our attention that causes us to believe that the financial statements do not present fairly, in all material respects..."
Read that construction carefully, because it is doing something research writing almost never does. It reports the absence of a contrary finding rather than the presence of a positive one. It is a true statement about what the work turned up, and it stays true even if the underlying reality is different, because it only ever claimed to describe what surfaced.
The standard also refuses to let this become a loophole. The evidence gathered must be at least sufficient for the practitioner to obtain a meaningful level of assurance, and the standard defines meaningful: to be meaningful, the level of assurance obtained is likely to enhance the intended users confidence about the financial statements. You cannot do two hours of work and shelter behind the negative form. The floor is that the work has to move somebody.
Almost every mid-sized research study is honestly a limited assurance engagement. Twenty interviews on a specific question do not license "users prefer X". They license "nothing in twenty interviews suggested users would reject X, and three named risks did surface". That second sentence survives contact with the next six months. The first one does not.
Agreed-upon procedures: reporting facts and declining to conclude
The bottom row is the most disciplined of the three and the most useful for early discovery work. Under ISRS 4400 (Revised), findings are defined as the factual results of the agreed-upon procedures performed, and findings are capable of being objectively verified. The standard then requires the report to exclude opinions or conclusions in any form, as well as recommendations.
The required statements in the report are worth reproducing, because each has a direct research analogue:
- A statement that the practitioner makes no representation regarding the appropriateness of the agreed-upon procedures. In research: we ran the study you asked for, and we are not vouching that it was the right study to answer your question.
- A statement that the engagement is not an assurance engagement and accordingly the practitioner does not express an opinion or an assurance conclusion.
- A statement that, had the practitioner performed additional procedures, other matters might have come to the practitioner's attention that would have been reported.
That last one is the single most useful sentence in the entire assurance canon for a research team. It is a standing, pre-agreed acknowledgement that the study saw what it was pointed at and nothing else, written into the report by default rather than negotiated defensively after a stakeholder is disappointed. On the effect of that disappointment when it is not pre-empted, see our guide to the research expectation gap.
Choosing the level before you field
The choice has to be made at brief time, because it determines the design. Deciding afterwards is just picking whichever verb the results can bear, which is exactly the researcher degree of freedom described in p-hacking and researcher degrees of freedom.
| Question to settle at brief time | Reasonable | Limited | Findings only |
|---|---|---|---|
| Is there a pre-specified primary outcome? | Yes, one | Yes, one | No |
| Is the sample sized against a target effect? | Yes | Approximately | No |
| Is the population defined and is coverage checked? | Yes | Partially | No |
| Can the conclusion be stated positively? | Yes | No, negative form only | No conclusion at all |
| What does a null result mean? | Evidence of absence, within power | Nothing surfaced | Nothing surfaced |
Sizing the top row honestly is a separate exercise covered in statistical power and minimum detectable effect; deciding when you have enough qualitative depth is covered in how many interviews are enough and data saturation.
Two practical rules follow. First, most studies should be in the middle or bottom row, and saying so out loud is a mark of seniority rather than weakness. Second, a study cannot be promoted after the fact. If you fielded a findings-only study and the results look decisive, you have a decisive-looking findings-only study.
This is not evidence grading, and the distinction matters
Grading a body of accumulated evidence after the fact is a different exercise, and one the corpus already covers - see evidence synthesis for the confidence ledger approach, which assigns confidence to a claim once several studies exist. Assurance level is upstream of that. It is a contract about what a single engagement will be able to say, agreed before any data exists, in the same spirit as a pre-registered analysis plan. One is a verdict on evidence you have. The other is a limit on the verdict you will be allowed to reach.
What changes when the level is written down
Three things, all of them cheap.
Reports stop overclaiming by default. When the top line of the readout says "limited assurance, negative form", the writer physically cannot open with "users want". The format does the discipline that willpower does not.
Stakeholders learn to ask for the level rather than for more confidence. "Can you make this more conclusive?" becomes "do we want to upgrade this to a reasonable-assurance study, and can we fund the sample?" That is a budget conversation with a real answer, instead of a pressure conversation with a rhetorical one.
The repository stops lying over time. Six months on, a claim carries the level it was born with, which is the only durable defence against the memory filter described in publication bias in product research.
How Koji makes the level visible rather than aspirational
Declaring a level only works if the artefacts of the study can support the declaration. Three parts of the platform do that work automatically.
Structured questions make the primary outcome unambiguous. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and a study that names a scale or yes_no question as its primary outcome before fielding has done the single hardest part of qualifying for a higher assurance level. An open_ended question with AI follow-up probing gives you the depth; the structured types give you the countable outcome the level is claimed against. The structured questions guide covers how the six types map to report visualisations.
Quality scores make the evidence base inspectable. Every interview is scored 1 to 5 with a breakdown across relevance, depth, and coverage, so "sufficient appropriate evidence" stops being a feeling. A limited-assurance claim resting on twenty interviews with a median quality score of 4 is a materially different object from the same claim resting on twenty scored 2, and the difference is visible in the report rather than buried in the transcripts. See understanding quality scores.
Real-time results make the level a live decision. Because interviews complete continuously and themes surface as they land, a team can see mid-field that a study is not going to reach the bar for the level it was scoped at, and either extend the sample or downgrade the claim on purpose. Doing that mid-study has rules of its own, covered in changing a study mid-field and interim analysis and sequential testing. A traditional agency engagement surfaces the same problem in the debrief, when both options have expired.
The contrast with survey-first tools is structural rather than rhetorical. A form builder collects responses; it has no concept of the evidentiary weight of a response, so every export arrives at the same implied confidence. AI-moderated interviews that probe, score, and report continuously produce evidence with a grain to it, and a grain is what an assurance level is claimed against.
Honest objections
This is bureaucracy for a five-person team. The whole mechanism is one line in the brief and one line at the top of the readout. If that is too heavy, the study did not have a brief, which is a different problem.
Stakeholders will read "limited assurance" as "ignore this". They will for about two studies, and then they will notice that the limited-assurance claims keep surviving and the old confident ones did not. Credibility comes from claims that hold, and the fastest way to make claims hold is to make them smaller than the evidence.
Auditing has statutory backing and research does not. True, and it cuts the other way. Auditors were forced into this precision by liability. Research teams get to adopt the useful half voluntarily, without the liability, which is a bargain.
Frequently asked questions
What are the three levels of assurance?
Reasonable assurance (ISA 200) supports a positive opinion and requires sufficient appropriate evidence to reduce engagement risk to an acceptably low level. Limited assurance (ISRE 2400 Revised) supports only a negative conclusion, phrased as "nothing has come to our attention that causes us to believe...". Agreed-upon procedures (ISRS 4400 Revised) support no conclusion at all - the practitioner reports factual findings and explicitly states that no assurance is expressed.
Is reasonable assurance the same as certainty?
No, and the standards are explicit about it. Engagement risk is never reduced to nil, so absolute assurance is not attainable. Reasonable assurance is a high level of assurance with a residual error rate that is accepted rather than eliminated. Any research claim phrased more strongly than that is claiming more than a full audit does.
Which level fits a typical 15 to 25 interview study?
Almost always limited assurance. You have enough evidence to say that nothing surfaced which contradicts the proposition, and to name the risks that did surface, but not enough to state a positive population-level claim. Write the conclusion in the negative form and the study will age well.
How is this different from grading evidence with GRADE or a confidence ledger?
Evidence grading is retrospective and applies to an accumulated body of findings across studies. An assurance level is prospective and applies to a single engagement - it is agreed before fielding and constrains what the report is permitted to conclude. Teams benefit from both, and they operate at different moments.
Can a study be upgraded to a higher level after the data comes in?
No. The level is a property of the design, not of how good the results look. Upgrading after seeing results is a researcher degree of freedom, and it reintroduces exactly the bias the pre-specification was there to prevent. If the results justify a stronger claim, run a confirmatory study scoped at the higher level.
Does declaring a low assurance level make research look weak to executives?
The opposite, over a horizon of a few months. Teams that label levels build a track record where the confident claims hold up, because they are only made when the evidence supports them. Teams that state everything at full confidence spend that credibility once and then argue about why the last three certainties did not replicate.
Related Resources
- Structured Questions Guide - the six question types that give a study a countable primary outcome
- The Research Expectation Gap - what happens when the level is never stated
- Corroboration in Research - why adding sources does not automatically raise the level
- Evidence Synthesis - grading a body of evidence once several studies exist
- Statistical Power and Minimum Detectable Effect - sizing a study that can carry a positive claim
- Research Peer Review: The Pre-Launch QA Gate - where the assurance level should be checked
Related Articles
Corroboration in Research: Why Three Sources Saying the Same Thing Can Be One Source
Evidence from multiple sources only multiplies confidence when the sources are independent. Four ways research sources secretly share an origin, and a ten-minute test for catching it.
Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)
Most teams have dozens of studies and no way to say what they collectively know. Evidence synthesis is the discipline of pooling findings across studies into a single rated conclusion - adapted from GRADE and systematic review practice for product research.
How Many Interviews Are Enough? A Guide to Sample Size
Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.
The Research Expectation Gap: Why Stakeholders Are Disappointed by Studies That Were Done Right
Auditing measured its own credibility gap and found only 16% was sub-standard work. Half was scope. Here is how that decomposition changes what research teams should fix.
Research Independence: Why the Team That Built the Feature Should Not Grade It
The five threats to independence from professional ethics codes, applied to product research - and why structural independence is a different problem from cognitive bias.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.