Back to docs
Analysis & Synthesis

The Graded Hedge: How to Flag a Finding You Cannot Fully Support

Auditing built a five-rung ladder for conclusions it could not fully support, with fixed wording and a two-by-two rule for choosing the rung. Research has one register, so every impairment gets rounded to clean. Here is the ladder, translated.

Answer first: sometimes you reach the readout with a problem you cannot fix. The segment refused, the log data was missing, fieldwork was cut short. Auditing's response to this was not to work harder. It was to build a five-rung ladder of graded hedges with fixed wording and a formal two-by-two for choosing the rung, so a reader can tell a clean conclusion from a flagged one without reading the methodology. Research has exactly one register, so every impairment gets silently rounded to clean.

The short answer

The other articles in this series are all about prevention and repair. Test your controls so the process is sound. Hold a skeptical stance so a plausible answer does not slip through. Check for subsequent events so a finding is not overtaken before you present it.

All three are covered here: testing research controls, professional skepticism, and subsequent events. All three share an assumption: that the problem can be fixed if you handle it properly. This article is about what to do when it cannot.

Auditing's answer is a graded vocabulary. Not a softer verb chosen by feel, but a defined set of rungs, each with required wording and a rule that determines which one you are on. ISA 705 establishes three modified opinions - qualified, adverse, and disclaimer - and ISA 706 adds a way to flag something important without modifying the conclusion at all.

The two-by-two that decides the rung

This is the reusable object, and it is worth reproducing exactly. ISA 705.A1 sets out how the type of modification is determined:

Nature of matter giving rise to the modificationMaterial but not pervasiveMaterial and pervasive
The conclusion is materially misstatedQualified opinionAdverse opinion
Unable to obtain sufficient appropriate evidenceQualified opinionDisclaimer of opinion

Two axes, and research collapses both.

The first axis is the one that matters most, and almost nobody separates it. There is a categorical difference between what I found is wrong and I could not find out. These are opposite failures. One is a defect in the evidence you have; the other is an absence of evidence you needed. In a research readout they receive identical treatment: a hedging adverb. We think, it appears, directionally, early signal suggests. The reader cannot tell which failure they are being warned about, and the two imply completely different next steps. A misstatement means the claim needs correcting. An inability means the claim needs evidence that does not yet exist.

The second axis is pervasiveness, and ISA 705 defines it as whether the effects are confined to specific elements or represent a substantial proportion of the whole. Translated: is the problem in one part of the finding, or does it run through everything? A study that could not reach the enterprise segment has a confined problem if the headline is about SMB behavior, and a pervasive one if the headline is about all customers.

Put together, the grid answers a question research teams argue about endlessly and never resolve: how bad is bad enough to change what we say? The answer is not a feeling. It is two questions with four outcomes.

The five rungs, translated

RungAudit formWhen to use it in researchWhat the readout says
1UnmodifiedScope was met, evidence was sufficientThe finding, stated plainly
2Unmodified with Emphasis of MatterThe finding holds, but one disclosed fact is fundamental to reading it correctlyThe finding, plus a prominent flag
3Qualified - misstatementOne element is not supported; the rest is"Except for X, the finding holds"
4Qualified - scopeA material part could not be covered; the rest stands"Except for the segment we could not reach..."
5Adverse or disclaimerThe headline is wrong, or you cannot conclude at all"We cannot conclude"

Rung 2 is the most immediately useful and the most underused. ISA 706 defines an Emphasis of Matter paragraph as one referring to a matter "appropriately presented or disclosed" that is "of such importance that it is fundamental to users' understanding." The critical feature: it does not modify the conclusion. You are not hedging. You are saying this finding is sound and you will misread it if you do not know this one thing.

Research examples that belong on rung 2: all twelve sessions ran during the week of a major outage; the sample skews to customers who opted into a beta; fieldwork closed before a competitor launch. The finding is not weakened. The context is load-bearing.

Rung 3 gives you the single most useful borrowed phrase in this entire lane: "except for." It is a precision instrument. It lets you carve out the unsupported part and keep the rest at full strength, instead of the usual move of applying a vague hedge to everything and thereby degrading the well-supported claims along with the weak one. Most research reports either overclaim across the board or underclaim across the board. "Except for" is how you avoid both.

Rung 5 is the one research never issues. ISA 705 requires a disclaimer when the auditor cannot obtain sufficient evidence and the possible effects are pervasive: the practitioner formally declines to conclude. Consider how rare that is in product research. Studies that failed to recruit, lost half their sessions, or fielded a broken instrument get written up anyway, with hedged language, because a deck that concludes nothing feels like a wasted quarter. The profession that invented this ladder concluded the opposite: a formal refusal to conclude is more valuable than a conclusion nobody should rely on, because the refusal tells the organization it still does not know, and the hedged version tells it that it does.

What this is not

Two distinctions, both of which matter because the neighbouring ideas are already covered here and it would be easy to conflate them.

This is not an assurance level. Our guide to levels of assurance covers declaring, before you field, how much confidence the design can support. That is prospective: it describes what the study was built to deliver. A modification is retrospective and involuntary. It is forced on you at report time by something that went wrong during execution. The two are independent, which is the crisp version: a limited-assurance study can still be qualified. Choosing a modest level at brief time does not immunize you against a scope limitation later.

This is not a scope preamble. The research expectation gap article proposes four lines at the top of every readout covering what was and was not in scope. That declares the boundary you intended. A modification reports the boundary you actually hit. A study can meet its declared scope exactly and still need a qualification because the evidence within that scope turned out to be insufficient, and a study can miss its declared scope and need one for that reason. The preamble sets expectations; the ladder grades delivery against them.

Why one register is worse than no register

The reason this matters more than it sounds is that hedged language is not a weaker signal than plain language. It is an unreadable one.

When every impairment is expressed as an adverb, the reader has to infer severity from tone. And tone is exactly what varies with the author's confidence, seniority, and relationship to the audience rather than with the evidence. A cautious researcher's solid finding and a bold researcher's broken one can arrive with identical hedging. The reader's only real cue is the presenter's manner.

Fixed wording solves this in the same way it solved the audit expectation gap: it removes the author's discretion over how bad things sound. If a scope limitation always produces the phrase "except for," then a reader who sees that phrase knows precisely what happened, and a reader who does not see it knows nothing was carved out. The signal becomes readable because it is not being generated fresh each time.

There is a second effect that is worth more than the first. A ladder makes it socially possible to flag a problem. The reason researchers bury limitations at the end is that raising one in the headline feels like an admission of failure, and there is no established, non-catastrophic way to do it. Rung 2 gives you one. Flagging a matter without modifying the conclusion is a standard, expected, entirely normal act. Once the middle rungs exist, teams stop choosing between overclaiming and self-flagellation.

Making the call: three questions

  1. Which failure is it? Is the evidence I have wrong, or is the evidence I need missing? If you cannot answer this, you are not ready to write the hedge.
  2. Is it confined or does it run through the whole thing? Name the specific claims affected. If you can list them, it is confined and rung 3 or 4 applies. If you cannot, it is pervasive.
  3. Would a reasonable reader change their decision if they knew? If yes, it belongs at the top of the readout in fixed wording. If no, it is a methodology note.

A practical rule for the borderline: if you find yourself writing a sentence that hedges the headline itself, you are on rung 3 or above and should say so explicitly. Hedging the headline while presenting it as a finding is the failure mode this whole ladder exists to prevent.

Where the cheap-study argument actually applies

It would be easy to claim that a fast research platform means you never need to hedge. That is not true, and the honest version is more limited and more useful.

Modifications come from two causes, and only one of them is fixable by cheaper fieldwork:

  • Scope limitations are often a cost problem. The segment you could not reach, the market you could not cover, the sample too small to split - these are usually resource constraints rather than facts of nature. When an additional cell can be fielded in a day, a rung 4 qualification frequently converts back to rung 1. This is the same argument that appears in the research expectation gap: a large share of stakeholder disappointment is a coverage problem, and coverage is a cost problem.
  • Misstatements are not. If your instrument was leading, your sample was structurally unable to disconfirm you, or your analysis was done by the person who wanted the answer, running it again cheaply reproduces the same defect faster. Those are control problems, addressed by testing your research controls, not by volume.

Two features do make the ladder easier to apply honestly:

  • Coverage is visible rather than argued. Because Koji reports response counts per question, a scope limitation is a number you can point at. The debate about whether the enterprise segment was adequately covered ends when the count is on the page.
  • Structured questions make partial coverage explicit. Koji supports six question types: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Closed types produce per-question response counts, so if 40 people answered the scale item and 9 answered a follow-up, the pervasiveness question answers itself. Guidance on designing them is in the structured questions guide.

Common mistakes

  • Using one hedge for both failure types. "We think" covers both a wrong finding and a missing one. They need opposite responses.
  • Hedging the whole report to cover one weak claim. That is what "except for" is for. Degrading your good findings to protect one bad one helps nobody.
  • Putting the flag at the end. A limitations section after the recommendations is decoration. Rung 2 exists because some matters belong at the top.
  • Never issuing a disclaimer. If no study in your history concluded nothing, some of them concluded anyway.
  • Confusing a modification with an assurance level. One is chosen in advance; the other is forced on you. A limited-assurance study can still be qualified.
  • Letting the hedge be chosen by tone. The point of fixed wording is to remove the author's discretion over how serious it sounds.

Frequently asked questions

What is a modified conclusion in research?

It is a finding that carries a formal flag because something prevented you from supporting it fully. Borrowing from ISA 705, there are three grades - qualified, adverse, and a disclaimer where you decline to conclude - plus an emphasis of matter under ISA 706 that highlights something important without modifying the conclusion at all. The point is a fixed vocabulary rather than an ad hoc hedge.

How do I decide how strongly to flag a finding?

Use the ISA 705 two-by-two. First, is the problem that your evidence is wrong, or that you could not obtain evidence you needed? Second, is the effect confined to specific claims, or does it run through the whole finding? Confined problems produce a qualified conclusion in either case. Pervasive problems produce an adverse conclusion if the evidence is wrong, or a disclaimer if the evidence is missing.

What is an emphasis of matter and when should I use one?

It is a prominent flag on a finding that is sound. ISA 706 defines it as referring to a matter properly disclosed that is "fundamental to users' understanding." Use it when the conclusion holds but a reader will misinterpret it without one specific fact, such as fieldwork running during an outage week or a sample skewed to beta opt-ins. Crucially it does not weaken the finding, which is why it is the most usable rung.

What does "except for" actually do?

It carves out the unsupported part of a finding so the rest keeps its full strength. Without it, teams tend to apply a vague hedge across an entire report to cover one weak claim, which degrades every well-supported conclusion alongside it. "Except for the enterprise segment, which we could not recruit, the finding holds" is far more useful to a decision-maker than "these are early directional signals."

Is this the same as declaring an assurance level?

No. An assurance level is prospective and voluntary: you choose before fielding how much confidence the design can support. A modification is retrospective and involuntary: something went wrong during execution and the conclusion must carry a flag. They are independent, so a study that honestly declared limited assurance can still end up qualified.

Should a failed study still be written up?

Yes, but as a formal refusal to conclude rather than a hedged conclusion. That is what a disclaimer is for. A readout that states plainly that the question remains unanswered, and why, leaves the organization correctly informed that it still does not know. A hedged write-up of the same study leaves it believing it does, which is the more expensive outcome.

Related Resources

Related Articles

Corroboration in Research: Why Three Sources Saying the Same Thing Can Be One Source

Evidence from multiple sources only multiplies confidence when the sources are independent. Four ways research sources secretly share an origin, and a ten-minute test for catching it.

Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)

Most teams have dozens of studies and no way to say what they collectively know. Evidence synthesis is the discipline of pooling findings across studies into a single rated conclusion - adapted from GRADE and systematic review practice for product research.

Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support

Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.

The Research Expectation Gap: Why Stakeholders Are Disappointed by Studies That Were Done Right

Auditing measured its own credibility gap and found only 16% was sub-standard work. Half was scope. Here is how that decomposition changes what research teams should fix.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

How to Create Effective UX Research Reports (+ Free Template)

A complete guide to writing UX research reports that drive decisions — with a reusable template, best practices, and how AI tools like Koji auto-generate research reports in minutes.