{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T14:43:12.577Z"},"content":[{"type":"documentation","id":"9188f984-36f5-4858-8102-6958e757bfbd","slug":"answer-extraction-confidence-flags","title":"Answer Confidence in Koji Reports: What High, Medium and Low Actually Mean","url":"https://www.koji.so/docs/answer-extraction-confidence-flags","summary":"Koji attaches a high, medium or low confidence flag to every structured answer, describing only how certain the analysis is that it mapped the correct transcript span to the correct question. It is distinct from the participant's own certainty and from the study's level of assurance. Because language models are documented to be overconfident when verbalizing confidence, the flag should be read asymmetrically: a low flag is a strong signal to open the transcript, while a high flag means only that no problem was detected. Every answer stores its transcript message indices, making a low flag cheap to resolve.","content":"Every structured answer in a Koji report carries a confidence flag: high, medium, or low. That flag describes exactly one thing - how certain the analysis is that it matched the right part of the transcript to the right question. It is not a claim about whether the participant told you the truth, and it is not the interview quality score. The practical rule is asymmetric: treat a low flag as an instruction to open the transcript, and treat a high flag as weak evidence that nothing needs checking.\n\nThat asymmetry is the whole point of this article, and it is the opposite of how most people read a confidence number.\n\n## Three different things get called confidence\n\nMost confusion about this field comes from the word itself. Three separate measures in research all get called confidence, and they answer different questions.\n\n### The machine's extraction confidence\n\nThis is the field this article is about. When Koji analyses a completed interview, it maps the conversation back onto the questions in your study and produces a structured answer for each one. Alongside each answer it records whether that mapping was high, medium, or low confidence. The question it answers is *did I read this transcript correctly?*\n\n### The participant's certainty\n\nEntirely different. A participant can be completely certain and completely wrong, and interviewer behaviour can inflate that certainty. That is a measurement problem, not an extraction problem, and it is covered separately in [Participant Confidence Is Not Accuracy](/docs/participant-confidence-vs-accuracy).\n\n### The study's level of assurance\n\nDifferent again. This is the question of how much weight the whole study can bear - sample, method, and design. See [Levels of Assurance in Research](/docs/research-assurance-levels). An extraction can be high confidence inside a study that supports very little assurance overall.\n\nKeeping these apart matters because the remedies are different. Low extraction confidence is fixed by reading the transcript. Low assurance is fixed by running better research.\n\n## How Koji assigns a confidence level\n\nThe analysis pass reads the full conversation and, for each question in your interview plan, records the structured value, the qualitative answer in the participant's own words, the indices of the transcript messages the answer came from, and the confidence level.\n\nBecause the answer carries its message indices, every extraction is traceable. You are never asked to trust the flag on its own - you can go straight to the exchange it came from. That traceability is what makes a low flag cheap to resolve rather than alarming.\n\n### What pushes an answer toward low confidence\n\nIn practice, a few patterns account for most low flags:\n\n- **The question was never really asked.** In exploratory interviews the conversation sometimes runs out of time before covering everything.\n- **The participant answered a different question.** They responded to what they thought you meant, or answered two questions at once.\n- **The answer was hedged or self-contradictory.** *It depends, I guess sometimes but not really* does not map cleanly onto a single choice.\n- **The answer arrived indirectly.** The participant conveyed a rating through a story rather than a number.\n- **A closed question was answered in prose.** Someone talks around a yes/no instead of answering it.\n\nNotice that most of these are facts about your questionnaire or the conversation, not defects in the analysis. A cluster of low flags on one question is telling you that the question is not working. That is the same diagnostic logic as [paradata signals](/docs/paradata-interview-process-signals): concentration points at the instrument, not the person.\n\n## Why a low flag is worth more than a high flag\n\nThere is a good reason to weight these levels unevenly rather than treating confidence as a symmetric scale.\n\nLanguage models are systematically overconfident when they state confidence in words. In *Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs* (Zhao, He, Zheng, Zhang and Chen, arXiv, submitted 1 April 2026), the authors open with the problem directly: \"Large language models are often not just wrong, but confidently wrong: when they produce factually incorrect answers, they tend to verbalize overly high confidence rather than signal uncertainty.\" They add that such verbalized overconfidence \"can mislead users and weaken confidence scores as a reliable uncertainty signal.\"\n\nTwo honest implications follow.\n\nFirst, a high flag is the default output, so it carries little information. It means no specific problem was detected, which is not the same as verified.\n\nSecond, a low flag is informative precisely because it is the harder admission to make. When a system with a documented bias toward overconfidence tells you it is unsure, that is a signal worth acting on every time.\n\nOne important scoping note, because the distinction is easy to blur. The overconfidence literature above concerns open-domain factual claims, where a model asserts things about the world. Koji's confidence field is scoped much more narrowly - it asks only whether a span of a transcript sitting in front of you was mapped to the right question. That is a genuinely easier and better-grounded task, and crucially it is one you can audit in seconds because the transcript is right there. The lesson to carry over is the direction of the bias, not the error rate.\n\n## What to do at each level\n\n### High confidence: spot-check, do not verify\n\nRead these normally. Sample a handful per study to confirm the mapping looks sane, particularly early in a new study when your questions are untested. Do not treat the flag as verification.\n\n### Medium confidence: read before you quote\n\nMedium usually means the answer is present but arrived awkwardly. Fine for counting, risky for quoting. If a medium-confidence answer is about to appear in a slide, open the transcript and read the exchange first.\n\n### Low confidence: open the transcript\n\nAlways. The resolution takes seconds because the answer carries its message indices. You will usually find one of three situations: the answer is actually there and fine, the answer is genuinely ambiguous and should be treated as missing for that question, or the question failed and needs rewriting before your next study.\n\nWhat you should not do is filter low-confidence rows out of the report before looking at them. They are the most information-dense rows you have.\n\n## Confidence behaves differently across the six question types\n\nKoji's structured questions come in six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and the flag means something slightly different depending on which you are looking at. The full model is in the [structured questions guide](/docs/structured-questions-guide).\n\n### Closed types: the flag is about mapping\n\nFor scale, single_choice, multiple_choice, ranking and yes_no, there is a definite value the answer should resolve to. Low confidence here means the conversation did not pin that value down. In text mode these questions can be answered through an interactive widget, which captures the value directly and leaves very little to extract - so a low flag on a widget-answered question is unusual and worth a look.\n\n### Open-ended: the flag travels with themes and quotes\n\nFor open_ended questions there is no single value. Koji instead produces coded themes, each grounded in specific messages and carrying a verbatim supporting quote in the participant's original words. Low confidence typically means the answer was too thin to code rather than that the coding was wrong. Read it alongside [Understanding Themes and Patterns](/docs/understanding-themes-patterns).\n\n## How Koji handles this\n\n- Every structured answer is stored with its confidence level, so the flag is data you can filter and count, not a transient UI hint.\n- Each answer records the transcript message indices it came from, so Koji can take you from a flag to the exact exchange without searching.\n- Open-ended themes each carry a verbatim supporting quote preserved in the participant's original language, so a coded label can always be checked against what was actually said.\n- Koji separates extraction confidence from the interview quality score, so a clean extraction of a weak interview is never mistaken for a good interview.\n- The quality gate operates independently: only conversations scoring 3 or above consume a credit, so genuinely unusable interviews do not quietly become line items. See [How the Quality Gate Works](/docs/how-the-quality-gate-works).\n\n## Common mistakes\n\n### Filtering out low confidence before reading it\n\nThe single most expensive mistake here. Teams build a habit of excluding low-confidence rows to *clean* the data, and in doing so delete exactly the rows that would have told them a question was broken. Read first, exclude second, and record what you excluded.\n\n### Treating a high flag as a fact-check\n\nA high flag says the mapping looks right. It says nothing about whether the participant was accurate, honest, or representative. Those are separate questions with separate remedies.\n\n### Reporting a confidence average\n\nAveraging the flags across a study produces a number that looks meaningful and is not. Confidence is a per-answer routing signal telling you where to look. Count the low flags per question instead - that number is actionable.\n\n## Frequently asked questions\n\n### What does a low confidence flag actually mean?\n\nIt means the analysis was not sure it mapped the right part of the transcript to that question. It does not mean the participant was unhelpful or the answer is wrong. Open the transcript at the recorded message indices and you will usually resolve it in seconds.\n\n### Should I delete low confidence answers before I report?\n\nNo. Read them first. Low-confidence answers cluster on questions that are not working, so they are your best free diagnostic. If an answer is genuinely ambiguous after reading it, treat it as missing for that one question and say so in your write-up rather than silently dropping the row.\n\n### Is confidence the same as the interview quality score?\n\nNo. Confidence is per answer and describes extraction accuracy. The quality score is per interview, runs 1 to 5, and feeds the billing quality gate where only interviews scoring 3 or above consume a credit. A weak interview can produce high-confidence extractions, and a strong interview can contain one badly worded question that produces a low-confidence answer.\n\n### Why would a yes/no question come back low confidence?\n\nUsually because the participant talked around it. Asked whether they would recommend the product, someone may explain their reasoning at length without ever landing on yes or no. The conversation is informative but there is no definite value to record, so the extraction is flagged low.\n\n### Can I see which part of the transcript an answer came from?\n\nYes. Each structured answer records the indices of the transcript messages it was drawn from, and open-ended themes additionally carry a verbatim supporting quote. This is why resolving a low flag is quick. See [Viewing Interview Transcripts](/docs/viewing-interview-transcripts) for how to jump straight to the exchange.\n\n### Does a high confidence flag mean the answer is true?\n\nNo, and this is the most important limit to understand. High confidence means the answer was extracted correctly from what the participant said. Whether what they said was accurate is a separate question, and one that published work on model overconfidence gives you good reason not to outsource to a confidence score.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and how each one is captured and reported\n- [Understanding Quality Scores](/docs/understanding-quality-scores) - the per-interview score, which is not the same as extraction confidence\n- [Viewing Interview Transcripts](/docs/viewing-interview-transcripts) - how to jump from an extracted answer to the exchange behind it\n- [Levels of Assurance in Research](/docs/research-assurance-levels) - how much weight a study can honestly bear\n- [Participant Confidence Is Not Accuracy](/docs/participant-confidence-vs-accuracy) - the respondent-side sense of confidence\n- [Partial Interviews and Breakoff](/docs/partial-interviews-breakoff-analysis) - what to do when an interview stops early\n","category":"Reports & Analysis","lastModified":"2026-09-27T03:39:08.406309+00:00","metaTitle":"Answer Confidence in Koji Reports: High, Medium and Low Explained","metaDescription":"Koji flags every extracted answer high, medium or low confidence. What each level means, and why a low flag matters more than a high one.","keywords":["answer extraction confidence","AI extraction accuracy","low confidence answer","interview analysis confidence","structured answer confidence","AI research report confidence"],"aiSummary":"Koji attaches a high, medium or low confidence flag to every structured answer, describing only how certain the analysis is that it mapped the correct transcript span to the correct question. It is distinct from the participant's own certainty and from the study's level of assurance. Because language models are documented to be overconfident when verbalizing confidence, the flag should be read asymmetrically: a low flag is a strong signal to open the transcript, while a high flag means only that no problem was detected. Every answer stores its transcript message indices, making a low flag cheap to resolve.","aiPrerequisites":["Familiarity with creating Koji studies","Understanding of the six structured question types"],"aiLearningOutcomes":["Distinguish extraction confidence from participant certainty and study assurance","Explain why a low confidence flag carries more information than a high one","Resolve a low confidence answer using its transcript message indices","Diagnose a broken question from a cluster of low confidence flags","Avoid the three common errors: pre-filtering, treating high as verified, and averaging confidence"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}