Back to docs
Reports & Analysis

Answer Confidence in Koji Reports: What High, Medium and Low Actually Mean

Every structured answer in a Koji report carries a high, medium or low confidence flag describing how certain the analysis is that it mapped the right transcript span to the right question. Here is what each level means and what to do about it.

Every structured answer in a Koji report carries a confidence flag: high, medium, or low. That flag describes exactly one thing - how certain the analysis is that it matched the right part of the transcript to the right question. It is not a claim about whether the participant told you the truth, and it is not the interview quality score. The practical rule is asymmetric: treat a low flag as an instruction to open the transcript, and treat a high flag as weak evidence that nothing needs checking.

That asymmetry is the whole point of this article, and it is the opposite of how most people read a confidence number.

Three different things get called confidence

Most confusion about this field comes from the word itself. Three separate measures in research all get called confidence, and they answer different questions.

The machine's extraction confidence

This is the field this article is about. When Koji analyses a completed interview, it maps the conversation back onto the questions in your study and produces a structured answer for each one. Alongside each answer it records whether that mapping was high, medium, or low confidence. The question it answers is did I read this transcript correctly?

The participant's certainty

Entirely different. A participant can be completely certain and completely wrong, and interviewer behaviour can inflate that certainty. That is a measurement problem, not an extraction problem, and it is covered separately in Participant Confidence Is Not Accuracy.

The study's level of assurance

Different again. This is the question of how much weight the whole study can bear - sample, method, and design. See Levels of Assurance in Research. An extraction can be high confidence inside a study that supports very little assurance overall.

Keeping these apart matters because the remedies are different. Low extraction confidence is fixed by reading the transcript. Low assurance is fixed by running better research.

How Koji assigns a confidence level

The analysis pass reads the full conversation and, for each question in your interview plan, records the structured value, the qualitative answer in the participant's own words, the indices of the transcript messages the answer came from, and the confidence level.

Because the answer carries its message indices, every extraction is traceable. You are never asked to trust the flag on its own - you can go straight to the exchange it came from. That traceability is what makes a low flag cheap to resolve rather than alarming.

What pushes an answer toward low confidence

In practice, a few patterns account for most low flags:

  • The question was never really asked. In exploratory interviews the conversation sometimes runs out of time before covering everything.
  • The participant answered a different question. They responded to what they thought you meant, or answered two questions at once.
  • The answer was hedged or self-contradictory. It depends, I guess sometimes but not really does not map cleanly onto a single choice.
  • The answer arrived indirectly. The participant conveyed a rating through a story rather than a number.
  • A closed question was answered in prose. Someone talks around a yes/no instead of answering it.

Notice that most of these are facts about your questionnaire or the conversation, not defects in the analysis. A cluster of low flags on one question is telling you that the question is not working. That is the same diagnostic logic as paradata signals: concentration points at the instrument, not the person.

Why a low flag is worth more than a high flag

There is a good reason to weight these levels unevenly rather than treating confidence as a symmetric scale.

Language models are systematically overconfident when they state confidence in words. In Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs (Zhao, He, Zheng, Zhang and Chen, arXiv, submitted 1 April 2026), the authors open with the problem directly: "Large language models are often not just wrong, but confidently wrong: when they produce factually incorrect answers, they tend to verbalize overly high confidence rather than signal uncertainty." They add that such verbalized overconfidence "can mislead users and weaken confidence scores as a reliable uncertainty signal."

Two honest implications follow.

First, a high flag is the default output, so it carries little information. It means no specific problem was detected, which is not the same as verified.

Second, a low flag is informative precisely because it is the harder admission to make. When a system with a documented bias toward overconfidence tells you it is unsure, that is a signal worth acting on every time.

One important scoping note, because the distinction is easy to blur. The overconfidence literature above concerns open-domain factual claims, where a model asserts things about the world. Koji's confidence field is scoped much more narrowly - it asks only whether a span of a transcript sitting in front of you was mapped to the right question. That is a genuinely easier and better-grounded task, and crucially it is one you can audit in seconds because the transcript is right there. The lesson to carry over is the direction of the bias, not the error rate.

What to do at each level

High confidence: spot-check, do not verify

Read these normally. Sample a handful per study to confirm the mapping looks sane, particularly early in a new study when your questions are untested. Do not treat the flag as verification.

Medium confidence: read before you quote

Medium usually means the answer is present but arrived awkwardly. Fine for counting, risky for quoting. If a medium-confidence answer is about to appear in a slide, open the transcript and read the exchange first.

Low confidence: open the transcript

Always. The resolution takes seconds because the answer carries its message indices. You will usually find one of three situations: the answer is actually there and fine, the answer is genuinely ambiguous and should be treated as missing for that question, or the question failed and needs rewriting before your next study.

What you should not do is filter low-confidence rows out of the report before looking at them. They are the most information-dense rows you have.

Confidence behaves differently across the six question types

Koji's structured questions come in six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and the flag means something slightly different depending on which you are looking at. The full model is in the structured questions guide.

Closed types: the flag is about mapping

For scale, single_choice, multiple_choice, ranking and yes_no, there is a definite value the answer should resolve to. Low confidence here means the conversation did not pin that value down. In text mode these questions can be answered through an interactive widget, which captures the value directly and leaves very little to extract - so a low flag on a widget-answered question is unusual and worth a look.

Open-ended: the flag travels with themes and quotes

For open_ended questions there is no single value. Koji instead produces coded themes, each grounded in specific messages and carrying a verbatim supporting quote in the participant's original words. Low confidence typically means the answer was too thin to code rather than that the coding was wrong. Read it alongside Understanding Themes and Patterns.

How Koji handles this

  • Every structured answer is stored with its confidence level, so the flag is data you can filter and count, not a transient UI hint.
  • Each answer records the transcript message indices it came from, so Koji can take you from a flag to the exact exchange without searching.
  • Open-ended themes each carry a verbatim supporting quote preserved in the participant's original language, so a coded label can always be checked against what was actually said.
  • Koji separates extraction confidence from the interview quality score, so a clean extraction of a weak interview is never mistaken for a good interview.
  • The quality gate operates independently: only conversations scoring 3 or above consume a credit, so genuinely unusable interviews do not quietly become line items. See How the Quality Gate Works.

Common mistakes

Filtering out low confidence before reading it

The single most expensive mistake here. Teams build a habit of excluding low-confidence rows to clean the data, and in doing so delete exactly the rows that would have told them a question was broken. Read first, exclude second, and record what you excluded.

Treating a high flag as a fact-check

A high flag says the mapping looks right. It says nothing about whether the participant was accurate, honest, or representative. Those are separate questions with separate remedies.

Reporting a confidence average

Averaging the flags across a study produces a number that looks meaningful and is not. Confidence is a per-answer routing signal telling you where to look. Count the low flags per question instead - that number is actionable.

Frequently asked questions

What does a low confidence flag actually mean?

It means the analysis was not sure it mapped the right part of the transcript to that question. It does not mean the participant was unhelpful or the answer is wrong. Open the transcript at the recorded message indices and you will usually resolve it in seconds.

Should I delete low confidence answers before I report?

No. Read them first. Low-confidence answers cluster on questions that are not working, so they are your best free diagnostic. If an answer is genuinely ambiguous after reading it, treat it as missing for that one question and say so in your write-up rather than silently dropping the row.

Is confidence the same as the interview quality score?

No. Confidence is per answer and describes extraction accuracy. The quality score is per interview, runs 1 to 5, and feeds the billing quality gate where only interviews scoring 3 or above consume a credit. A weak interview can produce high-confidence extractions, and a strong interview can contain one badly worded question that produces a low-confidence answer.

Why would a yes/no question come back low confidence?

Usually because the participant talked around it. Asked whether they would recommend the product, someone may explain their reasoning at length without ever landing on yes or no. The conversation is informative but there is no definite value to record, so the extraction is flagged low.

Can I see which part of the transcript an answer came from?

Yes. Each structured answer records the indices of the transcript messages it was drawn from, and open-ended themes additionally carry a verbatim supporting quote. This is why resolving a low flag is quick. See Viewing Interview Transcripts for how to jump straight to the exchange.

Does a high confidence flag mean the answer is true?

No, and this is the most important limit to understand. High confidence means the answer was extracted correctly from what the participant said. Whether what they said was accurate is a separate question, and one that published work on model overconfidence gives you good reason not to outsource to a confidence score.

Related Resources

Related Articles

Paradata: What Response Time, Hesitation and Drop-Off Tell You About Your Questions

Every interview produces a record of how the answers were produced. Most teams read it to judge respondents. Read it to judge your questions instead, and you get the cheapest instrument improvement available.

Partial Interviews: Should You Analyse Someone Who Answered Half Your Questions?

A partial interview is breakoff - a third category that is neither unit nonresponse nor item nonresponse. How Koji flags partials, why they usually cost you nothing, and when to include them.

Participant Confidence Is Not Accuracy: How Interviewer Feedback Inflates Certainty (2026)

A single "good, that is helpful" can inflate how certain a participant says they were, and how well they say they saw. The post-identification feedback research, why first-telling confidence is still informative, and how to capture it before you contaminate it.

Levels of Assurance in Research: How Much Confidence a Study Can Honestly Support

Auditing defines three levels of assurance - reasonable, limited, and none. Research reports use one voice for all three. Here is how to pick and state the level before you field a study.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Understanding Quality Scores

Learn how Koji evaluates interview quality on a 0-5 scale and why it matters for your research and billing.

Understanding Themes & Patterns

Learn how Koji identifies recurring themes across interviews and how to use them for decision-making.

Viewing Interview Transcripts

How to read, navigate, and get value from your interview transcripts in Koji.