{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-15T20:20:20.222Z"},"content":[{"type":"documentation","id":"6fa35b75-665e-40d9-a1ab-7b423a3f019d","slug":"research-modified-conclusions","title":"The Graded Hedge: How to Flag a Finding You Cannot Fully Support","url":"https://www.koji.so/docs/research-modified-conclusions","summary":"When a study reaches the readout with an impairment that cannot be fixed, auditing's response was a graded vocabulary rather than more work. ISA 705 establishes three modified opinions (qualified, adverse, disclaimer) chosen by a two-by-two: the nature of the matter (the conclusion is misstated versus evidence could not be obtained) crossed with pervasiveness (confined to specific elements versus running through the whole). ISA 706 adds an emphasis of matter that flags something fundamental without modifying the conclusion. Research collapses both axes into a single hedging adverb, making severity unreadable because it varies with the author's tone rather than the evidence. The article translates five rungs, argues that except for is the most useful borrowed construction, and notes that a formal refusal to conclude is more valuable than a conclusion nobody should rely on. Modifications differ from assurance levels, which are prospective and voluntary.","content":"**Answer first: sometimes you reach the readout with a problem you cannot fix. The segment refused, the log data was missing, fieldwork was cut short. Auditing's response to this was not to work harder. It was to build a five-rung ladder of graded hedges with fixed wording and a formal two-by-two for choosing the rung, so a reader can tell a clean conclusion from a flagged one without reading the methodology. Research has exactly one register, so every impairment gets silently rounded to clean.**\n\n## The short answer\n\nThe other articles in this series are all about prevention and repair. Test your controls so the process is sound. Hold a skeptical stance so a plausible answer does not slip through. Check for subsequent events so a finding is not overtaken before you present it.\n\nAll three are covered here: [testing research controls](/docs/research-controls-vs-evidence), [professional skepticism](/docs/professional-skepticism-research), and [subsequent events](/docs/research-subsequent-events). All three share an assumption: that the problem can be fixed if you handle it properly. This article is about what to do when it cannot.\n\nAuditing's answer is a graded vocabulary. Not a softer verb chosen by feel, but a defined set of rungs, each with required wording and a rule that determines which one you are on. ISA 705 establishes three modified opinions - qualified, adverse, and disclaimer - and ISA 706 adds a way to flag something important without modifying the conclusion at all.\n\n## The two-by-two that decides the rung\n\nThis is the reusable object, and it is worth reproducing exactly. ISA 705.A1 sets out how the type of modification is determined:\n\n| Nature of matter giving rise to the modification | Material but not pervasive | Material and pervasive |\n| --- | --- | --- |\n| The conclusion is materially misstated | Qualified opinion | Adverse opinion |\n| Unable to obtain sufficient appropriate evidence | Qualified opinion | Disclaimer of opinion |\n\nTwo axes, and research collapses both.\n\n**The first axis is the one that matters most, and almost nobody separates it.** There is a categorical difference between *what I found is wrong* and *I could not find out*. These are opposite failures. One is a defect in the evidence you have; the other is an absence of evidence you needed. In a research readout they receive identical treatment: a hedging adverb. We think, it appears, directionally, early signal suggests. The reader cannot tell which failure they are being warned about, and the two imply completely different next steps. A misstatement means the claim needs correcting. An inability means the claim needs evidence that does not yet exist.\n\n**The second axis is pervasiveness**, and ISA 705 defines it as whether the effects are confined to specific elements or represent a substantial proportion of the whole. Translated: is the problem in one part of the finding, or does it run through everything? A study that could not reach the enterprise segment has a confined problem if the headline is about SMB behavior, and a pervasive one if the headline is about all customers.\n\nPut together, the grid answers a question research teams argue about endlessly and never resolve: how bad is bad enough to change what we say? The answer is not a feeling. It is two questions with four outcomes.\n\n## The five rungs, translated\n\n| Rung | Audit form | When to use it in research | What the readout says |\n| --- | --- | --- | --- |\n| 1 | Unmodified | Scope was met, evidence was sufficient | The finding, stated plainly |\n| 2 | Unmodified with Emphasis of Matter | The finding holds, but one disclosed fact is fundamental to reading it correctly | The finding, plus a prominent flag |\n| 3 | Qualified - misstatement | One element is not supported; the rest is | \"Except for X, the finding holds\" |\n| 4 | Qualified - scope | A material part could not be covered; the rest stands | \"Except for the segment we could not reach...\" |\n| 5 | Adverse or disclaimer | The headline is wrong, or you cannot conclude at all | \"We cannot conclude\" |\n\nRung 2 is the most immediately useful and the most underused. ISA 706 defines an Emphasis of Matter paragraph as one referring to a matter \"appropriately presented or disclosed\" that is \"of such importance that it is fundamental to users' understanding.\" The critical feature: **it does not modify the conclusion.** You are not hedging. You are saying this finding is sound and you will misread it if you do not know this one thing.\n\nResearch examples that belong on rung 2: all twelve sessions ran during the week of a major outage; the sample skews to customers who opted into a beta; fieldwork closed before a competitor launch. The finding is not weakened. The context is load-bearing.\n\nRung 3 gives you the single most useful borrowed phrase in this entire lane: **\"except for.\"** It is a precision instrument. It lets you carve out the unsupported part and keep the rest at full strength, instead of the usual move of applying a vague hedge to everything and thereby degrading the well-supported claims along with the weak one. Most research reports either overclaim across the board or underclaim across the board. \"Except for\" is how you avoid both.\n\nRung 5 is the one research never issues. ISA 705 requires a disclaimer when the auditor cannot obtain sufficient evidence and the possible effects are pervasive: the practitioner formally declines to conclude. Consider how rare that is in product research. Studies that failed to recruit, lost half their sessions, or fielded a broken instrument get written up anyway, with hedged language, because a deck that concludes nothing feels like a wasted quarter. The profession that invented this ladder concluded the opposite: **a formal refusal to conclude is more valuable than a conclusion nobody should rely on**, because the refusal tells the organization it still does not know, and the hedged version tells it that it does.\n\n## What this is not\n\nTwo distinctions, both of which matter because the neighbouring ideas are already covered here and it would be easy to conflate them.\n\n**This is not an assurance level.** Our guide to [levels of assurance](/docs/research-assurance-levels) covers declaring, before you field, how much confidence the design can support. That is prospective: it describes what the study was built to deliver. A modification is retrospective and involuntary. It is forced on you at report time by something that went wrong during execution. The two are independent, which is the crisp version: **a limited-assurance study can still be qualified.** Choosing a modest level at brief time does not immunize you against a scope limitation later.\n\n**This is not a scope preamble.** The [research expectation gap](/docs/research-expectation-gap) article proposes four lines at the top of every readout covering what was and was not in scope. That declares the boundary you intended. A modification reports the boundary you actually hit. A study can meet its declared scope exactly and still need a qualification because the evidence within that scope turned out to be insufficient, and a study can miss its declared scope and need one for that reason. The preamble sets expectations; the ladder grades delivery against them.\n\n## Why one register is worse than no register\n\nThe reason this matters more than it sounds is that hedged language is not a weaker signal than plain language. It is an unreadable one.\n\nWhen every impairment is expressed as an adverb, the reader has to infer severity from tone. And tone is exactly what varies with the author's confidence, seniority, and relationship to the audience rather than with the evidence. A cautious researcher's solid finding and a bold researcher's broken one can arrive with identical hedging. The reader's only real cue is the presenter's manner.\n\nFixed wording solves this in the same way it solved the audit expectation gap: it removes the author's discretion over how bad things sound. If a scope limitation always produces the phrase \"except for,\" then a reader who sees that phrase knows precisely what happened, and a reader who does not see it knows nothing was carved out. The signal becomes readable because it is not being generated fresh each time.\n\nThere is a second effect that is worth more than the first. **A ladder makes it socially possible to flag a problem.** The reason researchers bury limitations at the end is that raising one in the headline feels like an admission of failure, and there is no established, non-catastrophic way to do it. Rung 2 gives you one. Flagging a matter without modifying the conclusion is a standard, expected, entirely normal act. Once the middle rungs exist, teams stop choosing between overclaiming and self-flagellation.\n\n## Making the call: three questions\n\n1. **Which failure is it?** Is the evidence I have wrong, or is the evidence I need missing? If you cannot answer this, you are not ready to write the hedge.\n2. **Is it confined or does it run through the whole thing?** Name the specific claims affected. If you can list them, it is confined and rung 3 or 4 applies. If you cannot, it is pervasive.\n3. **Would a reasonable reader change their decision if they knew?** If yes, it belongs at the top of the readout in fixed wording. If no, it is a methodology note.\n\nA practical rule for the borderline: if you find yourself writing a sentence that hedges the headline itself, you are on rung 3 or above and should say so explicitly. Hedging the headline while presenting it as a finding is the failure mode this whole ladder exists to prevent.\n\n## Where the cheap-study argument actually applies\n\nIt would be easy to claim that a fast research platform means you never need to hedge. That is not true, and the honest version is more limited and more useful.\n\nModifications come from two causes, and only one of them is fixable by cheaper fieldwork:\n\n- **Scope limitations are often a cost problem.** The segment you could not reach, the market you could not cover, the sample too small to split - these are usually resource constraints rather than facts of nature. When an additional cell can be fielded in a day, a rung 4 qualification frequently converts back to rung 1. This is the same argument that appears in [the research expectation gap](/docs/research-expectation-gap): a large share of stakeholder disappointment is a coverage problem, and coverage is a cost problem.\n- **Misstatements are not.** If your instrument was leading, your sample was structurally unable to disconfirm you, or your analysis was done by the person who wanted the answer, running it again cheaply reproduces the same defect faster. Those are control problems, addressed by [testing your research controls](/docs/research-controls-vs-evidence), not by volume.\n\nTwo features do make the ladder easier to apply honestly:\n\n- **Coverage is visible rather than argued.** Because Koji reports response counts per question, a scope limitation is a number you can point at. The debate about whether the enterprise segment was adequately covered ends when the count is on the page.\n- **Structured questions make partial coverage explicit.** Koji supports six question types: open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. Closed types produce per-question response counts, so if 40 people answered the scale item and 9 answered a follow-up, the pervasiveness question answers itself. Guidance on designing them is in the [structured questions guide](/docs/structured-questions-guide).\n\n## Common mistakes\n\n- **Using one hedge for both failure types.** \"We think\" covers both a wrong finding and a missing one. They need opposite responses.\n- **Hedging the whole report to cover one weak claim.** That is what \"except for\" is for. Degrading your good findings to protect one bad one helps nobody.\n- **Putting the flag at the end.** A limitations section after the recommendations is decoration. Rung 2 exists because some matters belong at the top.\n- **Never issuing a disclaimer.** If no study in your history concluded nothing, some of them concluded anyway.\n- **Confusing a modification with an assurance level.** One is chosen in advance; the other is forced on you. A limited-assurance study can still be qualified.\n- **Letting the hedge be chosen by tone.** The point of fixed wording is to remove the author's discretion over how serious it sounds.\n\n## Frequently asked questions\n\n### What is a modified conclusion in research?\n\nIt is a finding that carries a formal flag because something prevented you from supporting it fully. Borrowing from ISA 705, there are three grades - qualified, adverse, and a disclaimer where you decline to conclude - plus an emphasis of matter under ISA 706 that highlights something important without modifying the conclusion at all. The point is a fixed vocabulary rather than an ad hoc hedge.\n\n### How do I decide how strongly to flag a finding?\n\nUse the ISA 705 two-by-two. First, is the problem that your evidence is wrong, or that you could not obtain evidence you needed? Second, is the effect confined to specific claims, or does it run through the whole finding? Confined problems produce a qualified conclusion in either case. Pervasive problems produce an adverse conclusion if the evidence is wrong, or a disclaimer if the evidence is missing.\n\n### What is an emphasis of matter and when should I use one?\n\nIt is a prominent flag on a finding that is sound. ISA 706 defines it as referring to a matter properly disclosed that is \"fundamental to users' understanding.\" Use it when the conclusion holds but a reader will misinterpret it without one specific fact, such as fieldwork running during an outage week or a sample skewed to beta opt-ins. Crucially it does not weaken the finding, which is why it is the most usable rung.\n\n### What does \"except for\" actually do?\n\nIt carves out the unsupported part of a finding so the rest keeps its full strength. Without it, teams tend to apply a vague hedge across an entire report to cover one weak claim, which degrades every well-supported conclusion alongside it. \"Except for the enterprise segment, which we could not recruit, the finding holds\" is far more useful to a decision-maker than \"these are early directional signals.\"\n\n### Is this the same as declaring an assurance level?\n\nNo. An [assurance level](/docs/research-assurance-levels) is prospective and voluntary: you choose before fielding how much confidence the design can support. A modification is retrospective and involuntary: something went wrong during execution and the conclusion must carry a flag. They are independent, so a study that honestly declared limited assurance can still end up qualified.\n\n### Should a failed study still be written up?\n\nYes, but as a formal refusal to conclude rather than a hedged conclusion. That is what a disclaimer is for. A readout that states plainly that the question remains unanswered, and why, leaves the organization correctly informed that it still does not know. A hedged write-up of the same study leaves it believing it does, which is the more expensive outcome.\n\n## Related Resources\n\n- [Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support](/docs/research-assurance-levels)\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide)\n- [The Research Expectation Gap: Why Stakeholders Are Disappointed by Studies That Were Done Right](/docs/research-expectation-gap)\n- [How to Create Effective UX Research Reports (+ Free Template)](/docs/ux-research-report-template)\n- [Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)](/docs/evidence-synthesis-research-findings)\n- [Corroboration in Research: Why Three Sources Saying the Same Thing Can Be One Source](/docs/evidence-corroboration-independent-sources)\n- [Subsequent Events in Research: What You Owe a Decision After Fieldwork Closes](/docs/research-subsequent-events)\n- [Professional Skepticism in Research: The Duty to Doubt an Answer That Sounds Right](/docs/professional-skepticism-research)\n","category":"Analysis & Synthesis","lastModified":"2026-08-15T03:23:05.617281+00:00","metaTitle":"How to Flag a Research Finding You Cannot Fully Support (2026)","metaDescription":"A five-rung ladder of graded hedges borrowed from ISA 705 and 706. Learn the two-by-two that decides the rung, why except for is the most useful phrase in reporting, and when to formally decline to conclude.","keywords":["research limitations reporting","qualified research findings","research caveats","flagging research findings","scope limitation research","research report hedging"],"aiSummary":"When a study reaches the readout with an impairment that cannot be fixed, auditing's response was a graded vocabulary rather than more work. ISA 705 establishes three modified opinions (qualified, adverse, disclaimer) chosen by a two-by-two: the nature of the matter (the conclusion is misstated versus evidence could not be obtained) crossed with pervasiveness (confined to specific elements versus running through the whole). ISA 706 adds an emphasis of matter that flags something fundamental without modifying the conclusion. Research collapses both axes into a single hedging adverb, making severity unreadable because it varies with the author's tone rather than the evidence. The article translates five rungs, argues that except for is the most useful borrowed construction, and notes that a formal refusal to conclude is more valuable than a conclusion nobody should rely on. Modifications differ from assurance levels, which are prospective and voluntary.","aiPrerequisites":["Experience writing research reports and readouts","Familiarity with research limitations and scope"],"aiLearningOutcomes":["Apply the ISA 705 two-by-two to choose how strongly to flag a finding","Distinguish a misstatement from an inability to obtain evidence","Use an emphasis of matter to flag context without weakening a finding","Use except for to carve out an unsupported element without degrading the rest","Recognize when to formally decline to conclude rather than hedge"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}