Back to docs
Reports & Analysis

When Participants Go Off-Script: How AI Interviews Stay On-Brief

Participants wander off topic and occasionally test the interviewer with instructions of their own. How Koji separates useful drift from derailment using scope bounds, interview modes and quality scoring.

Two completely different things get called going off-script in an AI interview, and conflating them leads teams to the wrong fix. The first is ordinary topic drift - a participant answers a question you did not ask, tells a story, or wanders somewhere interesting. That is common, usually valuable, and something you want your interviewer to handle gracefully rather than suppress. The second is deliberate redirection - someone typing instructions at the AI to see what happens. That is rare in research, but it belongs to a real and well-documented class of vulnerability, and it deserves a different answer.

Koji treats them differently, and so should you.

Ordinary drift is often your best data

Start by noticing that a participant leaving your script is frequently a finding rather than a failure.

When someone answers why did you cancel by talking for two minutes about an unrelated billing surprise, the script did not fail - your model of the problem did. The classic interviewing methodologies are built around this. The Mom Test principles that ship with Koji's methodology frameworks explicitly favour talking about the participant's life over your idea, asking about the past rather than hypothetical futures, and digging for specifics when someone says always or never. Every one of those invites the participant somewhere you did not plan to go.

So the goal is not a participant who never deviates. The goal is an interview that can follow a productive digression and still come back and cover what you needed. That is a coverage problem, not a control problem.

When drift is a problem

Drift stops being useful in three situations: when it consumes the interview so your required questions never get asked, when it walks into territory you have deliberately excluded for legal or ethical reasons, and when it is not really drift but avoidance of a question the participant does not want to answer.

The first two are what the controls below are for. The third is a question-design problem - see Question Specificity and the Length Contract.

What actually bounds the conversation

Three mechanisms do the work, and they are worth understanding separately because they fail in different ways.

The brief's out-of-scope field

A Koji research brief carries an explicit out-of-scope field alongside the problem statement, the decision the research will inform, and the success criteria. This is where you write down what the interview must not pursue - we are not researching pricing sensitivity, or do not discuss the pending litigation.

This field is underused and it is the cheapest control you have. An interviewer that knows your boundaries can decline a topic gracefully and move on, which is a far better participant experience than one that either follows the participant anywhere or refuses without explanation. If a topic is genuinely off limits, write it in the brief rather than hoping it does not come up.

Structured, exploratory and hybrid modes

Koji studies run in one of three interview modes, and this is the main dial for how much wandering you want.

Structured follows your key questions closely. Use it for validation, for anything you intend to compare across participants, and for regulated topics.

Exploratory follows interesting threads. Use it for discovery, when you do not yet know what the real problem is.

Hybrid starts structured and goes exploratory on interesting topics - which is what most discovery work actually wants.

Choosing exploratory and then complaining that participants went off-topic is the most common self-inflicted version of this problem. The mode is a commitment about what you value.

Required questions and coverage

Each structured question in your plan is marked required or not, and coverage - how well the key questions and topics were actually covered - is one of the dimensions the analysis scores. Together these are what stop a pleasant, rambling conversation from quietly producing nothing.

A digression that still ends with every required question answered is a good interview. One that does not is visible in the score rather than hidden, which is the point. See Understanding Quality Scores and the AI interviewer tuning guide.

The adversarial case

Now the other half, which is smaller but should not be hand-waved.

What prompt injection is

Some participants will try to talk to the system rather than answer the question. Most are just curious. The underlying risk class is well defined. The OWASP Gen AI Security Project lists it as LLM01:2025 Prompt Injection in its LLM Top 10 for 2025, and describes it plainly: "A Prompt Injection Vulnerability occurs when user prompts alter the LLM's behavior or output in unintended ways."

OWASP splits it in two. Direct prompt injections "occur when a user's prompt input directly alters the behavior of the model in unintended or unexpected ways." Indirect prompt injections "occur when an LLM accepts input from external sources, such as websites or files. The content may have in the external content data that when interpreted by the model, alters the behavior of the model in unintended or unexpected ways."

An interview participant typing instructions is the direct case. It is the one you will actually see.

Why interview text is untrusted input

Any AI interview ingests free text from strangers, which makes that text untrusted input by definition. The reason to take the category seriously even when the attempts you see are playful is that the wider research on injection shows the attack surface is cheap to exploit and that obvious defences are not sufficient.

In Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems (Chang, Bao, Luo and Yu, arXiv, submitted 11 January 2026), the authors demonstrate end-to-end exploits costing "as little as $0.21 per target user query on OpenAI's embedding models", achieving "near-100% retrieval across 11 benchmarks and 8 embedding models". In one scenario a single poisoned email was enough to coerce a model into exfiltrating SSH keys "with over 80% success in a multi-agent workflow". They conclude that they "evaluate several defenses and find that they are insufficient to prevent the retrieval of malicious text".

Scope that honestly: their setting is retrieval corpora and agentic systems, not research interviews, and an interview transcript is not a retrieval corpus. The transferable lessons are narrower but real - untrusted text reaching a model is a demonstrated attack surface, cost is not a barrier, and a single input can be enough. None of that is a reason for alarm about interview data. It is a reason to keep the interviewer's job narrow and to treat what a participant types as content to be recorded rather than instructions to be obeyed.

What this means in practice

The structural defence is scope. An AI interviewer whose job is to ask the questions in a brief, record answers, and probe within bounds has very little surface worth attacking - it is not holding credentials, not calling tools on the participant's behalf, and not retrieving from a corpus a stranger can write to. Keeping the role narrow is worth more than any single filter.

What a derailed interview looks like in your report

The practical reassurance is that you do not have to detect this manually.

An interview that went badly off the rails scores poorly on coverage, because the required questions were not answered. Answers extracted from a conversation that never really addressed a question come back with low confidence flags, which is your signal to read the transcript - see Answer Confidence Flags. And because Koji's quality gate means only conversations scoring 3 or above consume a credit, a genuinely wrecked interview generally does not cost you anything.

Someone testing the interviewer is usually obvious the moment you open the transcript, and reading transcripts takes seconds once a flag has pointed you at one.

How Koji handles this

  • Koji briefs carry an explicit out-of-scope field, so boundaries are configuration rather than hope.
  • Three interview modes - structured, exploratory, and hybrid - let you choose how much productive wandering you want rather than accepting a fixed behaviour.
  • Required questions plus a coverage dimension in the interview score mean a rambling conversation that missed the point is visible instead of silently counted.
  • Koji's AI probes within the bounds of your brief, so follow-ups pursue depth on your questions rather than following the participant indefinitely.
  • All six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - give the conversation a spine to return to, and in text mode the closed types are answered through widgets that are not free-text at all.
  • Low-confidence extractions and low coverage scores surface a derailed interview automatically, and the quality gate means it usually does not consume a credit.

Common mistakes

Choosing exploratory mode and then wanting control

If comparability across participants matters, choose structured. Mode is the decision, not something to litigate afterwards.

Leaving out-of-scope empty

If there is a topic the interview genuinely must not pursue, writing it in the brief is the whole mechanism. An empty field is not a boundary.

Treating curiosity as an attack

Most participants who poke at the AI are just interested. That is not a security incident and it usually does not invalidate their answers - read the transcript and judge the interview on coverage.

Discarding a whole interview because one answer went sideways

Judge it per question. An interview where one answer is unusable and nine are excellent is a good interview. See the per-question denominator argument in Partial Interviews and Breakoff.

Frequently asked questions

What happens if a participant ignores the question and rambles?

Usually something useful. Koji's AI probes within the bounds of your brief and works to cover your required questions, so a digression that ends with everything answered is simply a good interview. If the rambling crowded out your questions, that shows up as a low coverage score rather than being hidden.

Can a participant manipulate the AI interviewer with instructions?

Someone can certainly try, and this is the direct case of what OWASP calls prompt injection - user input that alters the model's behaviour in unintended ways. The structural defence is scope: an interviewer whose job is to ask your questions, record answers, and probe within bounds is not holding credentials or calling tools on a participant's behalf, so there is little worth attacking. Attempts are also easy to spot in the transcript.

Does going off-topic ruin the interview data?

Rarely. Judge the interview per question rather than as a whole. The answers to questions that were properly covered remain valid, and answers extracted from a conversation that never really addressed a question come back flagged low confidence so you know which ones to check.

How does Koji know what is out of scope?

Because you tell it. The research brief has an explicit out-of-scope field alongside the problem statement and success criteria. Writing we are not researching pricing there lets the interviewer decline that topic gracefully and move on. If you leave the field empty, there is no boundary to enforce.

Should I use structured or exploratory mode to stay on track?

Structured, if staying on track is the priority - it follows your key questions closely and is the right choice for validation and for anything you plan to compare across participants. Exploratory deliberately follows interesting threads. Hybrid starts structured and opens up on interesting topics, which suits most discovery work.

What if a participant asks the AI a question?

This is normal and usually harmless - people ask what the study is for or whether their answers are anonymous. Those are reasonable questions and are better handled in your intake and consent step, where the answers are yours rather than improvised. See Intake Forms and Consent.

Related Resources

Related Articles

Can You Trust AI Interviewers? How Koji Prevents Hallucinations and Bias in Customer Research

A practical guide to how modern AI research platforms prevent hallucinations, model bias, and leading questions during auto-moderated customer interviews — with the verification techniques Koji uses to keep AI-generated insights faithful to the actual transcript.

AI Interviewer Tuning: How to Get Research-Grade Voice Interviews

A complete playbook for tuning Koji's AI interviewer — company context, probing depth, structured questions, and interview mode — to deliver interviews indistinguishable from a human researcher.

How Koji's AI Follow-Up Probing Works: Going Deeper Than Any Survey

Understand how Koji's AI interviewer automatically asks follow-up questions to go deeper on every answer — and how to configure probing depth, custom instructions, and anchor behavior for scale questions.

Answer Confidence in Koji Reports: What High, Medium and Low Actually Mean

Every structured answer in a Koji report carries a high, medium or low confidence flag describing how certain the analysis is that it mapped the right transcript span to the right question. Here is what each level means and what to do about it.

Intake Forms and Consent

Collect participant information and consent before interviews begin with customizable form fields.

Partial Interviews: Should You Analyse Someone Who Answered Half Your Questions?

A partial interview is breakoff - a third category that is neither unit nonresponse nor item nonresponse. How Koji flags partials, why they usually cost you nothing, and when to include them.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Understanding Quality Scores

Learn how Koji evaluates interview quality on a 0-5 scale and why it matters for your research and billing.