The Order You Learn Things In: Context Management for Research Analysis (2026)
An analyst who knows the hypothesis before reading the transcripts reaches a different conclusion from the same data. Here is the forensic procedure that fixes it.
You cannot un-see a hypothesis. Once an analyst knows which theory the team is hoping for, that knowledge shapes the analysis, and the distortion does not appear as a misread quote. It appears as a different set of questions asked, a different slice taken, a different ambiguous sentence resolved one way rather than the other. Forensic science has a procedure for this called linear sequential unmasking (LSU): treat the order in which information reaches the analyst, and the relevance of each piece, as explicit design decisions rather than accidents of who sent which Slack message first.
The short answer
Decide, before analysis starts, what the analyst sees and in what sequence. Give them the raw evidence first and the team's theory last, or not at all. Record the analyst's reading at each step, so that when context does arrive you can see exactly what it changed. The goal is not an analyst with no context, which is usually impossible and sometimes undesirable. The goal is an analyst whose revisions are visible instead of invisible.
The evidence: the same expert, the same data, a different answer
The foundational study is uncomfortable to read. Dror, Charlton and Peron (Forensic Science International, 2006; volume 156, pages 74-8) took fingerprints that latent print experts had already examined and used to make positive identifications. They showed the same prints to the same experts again, this time with a context suggesting the prints were not a match. The paper reports that "Within this new context, most of the fingerprint experts made different judgements, thus contradicting their own previous identification decisions."
Read that carefully, because the design removes every comfortable explanation. The evidence was identical. The examiners were the same people, not a less experienced comparison group. Expertise did not protect them. Only the surrounding story changed, and the conclusions moved with it.
The same research group later tested whether this was confined to the interpretive disciplines. Hamnett and Dror (Forensic Science International: Synergy, 2020; volume 2, pages 339-348) ran it in forensic toxicology, which they describe as one of the more objective domains. In their first experiment 58 participants were affected by irrelevant case information while analysing data from an immunoassay test for opiate-type drugs. A measurement that produces a number was still read differently depending on the story attached to it.
The failure you cannot detect afterwards
The second Hamnett and Dror experiment is the one that should change how you run analysis. With 53 participants, context biased not the reading of a result but the choice of which tests to run. The authors report that the age of the deceased influenced testing strategy: for older people, medicinal drugs were commonly chosen, whereas for younger people drugs of abuse were selected.
This is a categorically worse failure than misreading evidence, and it is the one that maps most directly onto product research. A misread transcript leaves a trace. Someone can re-read it and disagree with you. A question you never asked, a segment you never cut, a counter-hypothesis you never tested leaves no trace at all. There is no artifact to audit. Every quote in your report verifies, every number reconciles, and the analysis is still steered, because the steering happened in the choice of what to look at.
Your version of choosing which test to run looks like this:
- You filter the transcripts to the segment the team already suspects, and never run the same query on the others.
- You search the corpus for the feature name in the hypothesis, and not for the three features nobody mentioned in the kickoff.
- You stop coding once the expected theme is clearly supported, because the finding feels established.
- You pull the quotes that illustrate the agreed story, which is a reasonable thing to do for a readout and a disastrous thing to do before the conclusion is settled.
What the procedure actually prescribes
LSU was proposed by Dror and colleagues as a context management toolbox (Journal of Forensic Sciences, 2015; volume 60, pages 1111-2), and later expanded. The practical version, LSU-Expanded, is laid out by Quigley-McBride, Dror, Roy, Garrett and Kukucka (Forensic Science International: Synergy, 2022; volume 4, pages 100216), who note that "the order in which task-relevant information is received impacts human cognition and decision-making." The framework scores each piece of information on three parameters - its objectivity, its relevance to the task, and its biasing power - and uses those scores to sequence what the analyst sees. The stated aim is to increase the repeatability, reproducibility and transparency of decisions, not merely to suppress bias.
Translated into a research workflow, the sequence rule is: most objective and most task-relevant first, least objective and most biasing last. And critically, you commit your reading in writing before each new layer arrives.
| Information | Task-relevant? | Biasing power | When to reveal |
|---|---|---|---|
| The recording and the verbatim transcript | Primary evidence | Low | First, always |
| The research questions, phrased as questions | Yes | Low to medium | First, but never as expectations |
| Participant segment, plan tier, tenure | Sometimes | Medium | Only when an analysis step needs it |
| The existing codebook or last quarter top theme | Yes, but | High | After one open coding pass |
| Which feature the team believes is at fault | No | High | After coding is locked |
| The outcome a stakeholder is hoping for | No | Very high | Never, if you can manage it |
A worked example: how the same 40 interviews produce opposite roadmaps
Suppose you run 40 churn interviews. When you code them, the evidence falls into three piles: 12 transcripts contain an explicit, unambiguous pricing complaint, 9 contain explicit onboarding friction, and 19 contain some version of "it just was not worth it for us," which is genuinely compatible with either explanation.
Analyst A is told up front that the leadership theory is pricing. Faced with 19 ambiguous transcripts, A resolves 15 of them toward pricing. The result: pricing 27, onboarding 13. Pricing outranks onboarding by a factor of 2.1, and the readout recommends a pricing change.
Analyst B is told up front that the theory is onboarding. The same 19 ambiguous transcripts, 15 resolved toward onboarding. The result: onboarding 24, pricing 16. Onboarding leads by 1.5x, and the readout recommends an onboarding rebuild.
Same 40 interviews. Opposite decisions. Neither analyst invented a quote, and both can defend every individual classification as reasonable, because each one was reasonable. Note also what neither report mentions: that 19 of 40 transcripts, 48 percent of the corpus, did not determine the answer on their own. The ambiguity was absorbed rather than reported, which is the real loss.
The LSU fix is cheap. Code the 40 transcripts before anyone states a theory, with an explicit "ambiguous" bucket that an analyst is allowed to use. Now the honest output is 12 pricing, 9 onboarding, 19 undetermined, and the next step is obvious: the study cannot yet separate the two explanations, and the correct action is a targeted follow-up on those 19, not a roadmap commitment.
Why this is not the same thing as blind analysis
Blind analysis hides which group is which until the analysis is locked. It is an excellent default and it solves a different problem. Blinding works when there is a label you can remove without damaging the work: a condition, a variant, an arm.
Context management handles the information you cannot remove. An analyst usually needs to know what the product does, who the participant is, and what the study was for. You cannot strip all of that and still get competent coding. LSU accepts that some context is necessary, then asks a sharper question: of the context this analyst needs, what is genuinely task-relevant, and what is merely available? Task-irrelevant information gets withheld. Task-relevant information gets sequenced. Revisions get documented. Use blinding where you can; use sequencing for everything that is left.
How Koji handles this
Koji removes a large part of this problem by changing when the probing happens. In a traditional study the follow-up questions are asked by a human moderator who already knows the team's theory, and the deepest probes land wherever that moderator was most curious. In Koji the AI interviewer generates follow-ups live, inside each conversation, before any dataset-level hypothesis exists. Each question in a Koji study carries a required flag and a configurable probing depth, so coverage is enforced by the study design rather than by whatever the moderator found interesting on the day. That directly removes the Hamnett and Dror failure mode: the test still gets run even when nobody expects it to be informative.
Koji's structured questions help for the same reason. The six types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - commit you to an unambiguous item before the data arrives, so a scale answer or a ranking cannot be quietly re-resolved later by whoever is holding the hypothesis. The ambiguous middle that Analyst A and Analyst B fought over mostly disappears when the question was structured in the first place.
For the parts that stay qualitative, Koji preserves the chain back to the evidence. Every extracted answer carries a confidence flag of high, medium or low, plus the indices of the exact transcript messages it came from, and each theme Koji assigns stores the participant's verbatim supporting quote alongside the code label. That means a second analyst can re-run the interpretation from the grounded original rather than from someone else's summary, which is the practical precondition for ever detecting that context moved a conclusion. An analysis you cannot re-run from source is an analysis whose context effects are permanent.
Common mistakes
- Treating this as a character problem. The fingerprint examiners were experts acting in good faith. Asking an analyst to try harder to be objective is not a control. Sequencing is a control.
- Revealing the hypothesis in the kickoff and calling it alignment. Shared context is valuable for deciding what to study and poisonous for deciding what the data says. Separate those two meetings.
- Letting the codebook arrive first. A mature codebook is task-relevant and strongly biasing at the same time. Run one open pass before applying it.
- Forgetting to record the intermediate reading. If you do not write down the pre-context conclusion, you lose the only evidence of what the context changed, and the procedure degrades into a vibe.
- Blinding the things that were harmless and revealing the things that were not. Rank by biasing power, not by what happens to be easy to hide.
- Allowing no ambiguous category. If an analyst has nowhere to put an undetermined transcript, every undetermined transcript becomes evidence for whichever theory is in the room.
Frequently asked questions
What is linear sequential unmasking?
Linear sequential unmasking is a context management procedure from forensic science that controls the order in which information reaches an analyst. Information is assessed for objectivity, relevance and biasing power, then released in sequence from most objective and relevant to most biasing, with the analyst recording their conclusion before each new layer arrives.
How is this different from blind analysis?
Blind analysis hides a label, such as which group a participant was in, until the analysis is locked. Context management deals with information you cannot remove without making the work impossible, such as what the product does. Blinding removes; sequencing orders what remains and documents every revision.
Does an analyst need the research questions before they start?
Yes. The research questions are task-relevant and low in biasing power as long as they are phrased as questions rather than as expectations. "Why did these accounts churn" is fine. "Confirm that pricing drove churn" is the thing to withhold.
Can I just tell my team to be objective instead?
No, and the evidence is specific on this point. Dror, Charlton and Peron showed the same experts reversing their own prior conclusions on identical fingerprints when the surrounding context changed. Expertise and good intent did not prevent it, so an instruction to be objective will not either.
What is the cheapest version of this I can adopt this week?
Code your transcripts before you circulate the hypothesis, allow an explicit ambiguous bucket, and write down the counts for all three categories. The ambiguity you surface is the finding that tells you whether your study can actually separate the competing explanations.
Does using AI for analysis solve the problem by itself?
Only partly, and only if the evidence chain survives. An AI analysis prompted with the team's preferred conclusion inherits exactly the same bias. What helps is enforced question coverage, structured answers that cannot be silently reinterpreted, and per-answer traceability back to the original transcript so any later analyst can re-derive the finding from source.
Related Resources
- Blind Analysis - how to hide the labels you can remove, and where that technique stops working
- Confirmation Bias in User Research - the underlying bias this procedure is built to contain
- Analysis of Competing Hypotheses - what to do with the ambiguous transcripts once you have counted them
- Two Coders Cannot Rescue Categories That Were Too Close Together - why agreement between coders does not detect this failure
- How to Build a Qualitative Research Codebook - the artifact that is task-relevant and strongly biasing at once
- Structured Questions Guide - the six question types, and why an unambiguous item cannot be re-resolved later
Related Articles
Analysis of Competing Hypotheses: How to Test What Your Research Actually Supports
Most evidence that supports your favorite explanation also supports the ones you never wrote down. ACH is the matrix method that finds the evidence which actually discriminates.
Blind Analysis: How to Analyze Research Before You Know the Answer
Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.
Two Coders Cannot Rescue Categories That Were Too Close Together (2026)
Coding agreement is capped by the distance between your two closest categories - a number you fixed when you wrote them. No amount of coder training raises it.
Confirmation Bias in User Research: How to Recognize and Eliminate It
Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.
How to Build a Qualitative Research Codebook (With Examples and Templates)
A qualitative codebook is the rulebook for how you code your data — code names, definitions, inclusion criteria, examples, and exceptions. Done well, it makes coding consistent across analysts. Done badly, it produces findings nobody can defend.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.