Back to docs
Research Methods

Model Version Drift: What Happens to Your Research When the AI Changes Mid-Study (2026)

When the model behind your AI moderator or analyst is upgraded, your measuring instrument changed. The evidence, the three layers of drift, the bridge sample method, and how to make model version part of your method section.

Short answer: if the model behind your AI moderator or your AI analysis changes between wave 1 and wave 2, you have not measured a change in your customers - you may have measured a change in your instrument. This is not a new problem in research methodology. It is Campbell and Stanley instrumentation threat, first named in 1963, arriving in a new form. What is new is that the instrument can now change without anyone on the research team being told, between two studies fielded a month apart. The published evidence shows that the behaviour of a nominally identical model service can swing by tens of percentage points in a few months. This guide covers what the evidence shows, the three layers at which drift reaches your research, which of them are recoverable, the bridge sample method borrowed from survey mode-change practice, and how to make model version a first-class part of your method section.

The oldest threat in research design, wearing new clothes

In Experimental and Quasi-Experimental Designs for Research (1963), Donald Campbell and Julian Stanley enumerated the threats to internal validity that every research methods course still teaches: history, maturation, testing, instrumentation, statistical regression, selection and mortality.

Instrumentation is the one that covers changes in the measure - calibration, malfunction - and changes in the scorers, meaning the measurement procedures themselves - that may impact measurements. Classically it describes a scale that goes out of true, or a second rater who scores more harshly than the first, or a questionnaire that was revised between waves. The defining property is that the observed difference is produced by the measuring apparatus rather than by the thing being measured.

An AI moderator is a measuring apparatus. An AI thematic analyser is a scorer. When either is silently swapped for a different version, you have satisfied Campbell and Stanley definition exactly. The only thing that has changed since 1963 is who controls the swap and how much warning you get.

Research teams already have strong instincts about this in every other context. No competent researcher would change the wording of a tracking question between waves and then report the movement as a trend. No one would swap the response scale from 1-5 to 1-7 mid-study. But teams routinely let the underlying model change between waves without noting it, because the model does not feel like part of the instrument. It is.

What the evidence shows

The definitive study on this is How Is ChatGPT Behavior Changing over Time? by Lingjiao Chen, Matei Zaharia and James Zou (Harvard Data Science Review, vol. 6, no. 2, 2024). The authors evaluated the March 2023 and June 2023 versions of two commercial models on a fixed battery of tasks - maths problems, sensitive questions, opinion surveys, multi-hop knowledge questions, code generation, medical licensing exam questions and visual reasoning - and compared them.

The results were not small.

TaskModelMarch 2023June 2023
Prime vs composite identificationGPT-484.0%51.1%
Prime vs composite identificationGPT-3.549.6%76.2%
Directly executable generated codeGPT-452.0%10.0%
Directly executable generated codeGPT-3.522.0%2.0%
Answer rate on sensitive questionsGPT-421.0%5.0%
Answer rate on sensitive questionsGPT-3.52.0%8.0%

Three things in that table matter for research operations.

Drift is not directional. GPT-4 got much worse at prime identification while GPT-3.5 got much better at the same task in the same window. You cannot assume newer means better, and you cannot assume a change in one direction on one task predicts anything about another.

Behaviour changed as much as accuracy. GPT-4 average response length on the maths task fell from 638.3 characters to 3.9 - the model stopped showing its reasoning. Relatedly, the benefit of chain-of-thought prompting for GPT-4 fell from 24.4% to -0.1% while for GPT-3.5 it rose from -0.9% to 15.8%. A prompting technique that was load-bearing in March was inert in June. If your moderator prompt depends on a technique, that dependency is a silent version dependency.

The refusal surface moved. A fourfold change in the answer rate on sensitive questions is, for a research moderator, a change in what topics your instrument is capable of collecting data on. A study on a sensitive subject fielded either side of that change would produce different coverage for reasons that have nothing to do with participants.

The authors conclusion is the operational one to carry into your research ops: the behavior of the same LLM service can change substantially in a relatively short amount of time, highlighting the need for continuous monitoring of LLMs.

The three layers of drift, ranked by how badly they hurt

Model drift can reach a research programme at three distinct layers, and the crucial practical insight is that they differ enormously in how recoverable they are. Most teams spend their attention on the wrong one.

LayerWhat the model doesWhat drift corruptsRecoverable?
CollectionModerates the interview, decides follow-upsThe raw data itselfNo. The conversation happened once
AnalysisCodes transcripts, extracts themes, scores qualityThe coding layerYes. Re-run over the same transcripts
SynthesisWrites the report narrative, summarises findingsThe narrativeYes, cheaply. Regenerate

Collection drift is irreversible and should get essentially all of your pinning budget. If the moderator changes between waves, the probes were different, so the answers are different, so the raw transcripts are not comparable - and no amount of reanalysis fixes it, because you cannot go back and ask the question that was not asked. The participant is gone. That data is what it is, forever.

Analysis drift is annoying and recoverable. If your thematic coder changes and the themes shift, you re-run the new coder over the entire historical transcript corpus and get a consistent coding layer back. It costs compute, not fieldwork. The rule is simple and almost never followed: when the analysis model changes, re-code the whole corpus, not just the new wave. A trend line where waves 1-3 were coded by one model and wave 4 by another is measuring the coder, not the customer.

Synthesis drift barely matters. Report prose changing style between quarters is cosmetic. Regenerate it if you care.

This ranking gives you a clean allocation rule. Version-pin the moderator hard and treat any change to it as a study boundary. Version-track the analyser and re-run on change. Ignore the writer.

Silent drift, announced drift, and forced migration

Not all version changes arrive the same way, and the response differs.

Silent drift is a provider updating a model behind a stable endpoint name. You are not told, there is no changelog entry you will see, and the first symptom is usually an unexplained movement in a tracking metric. This is the one Chen, Zaharia and Zou were documenting, and it is the reason the continuous monitoring recommendation exists.

Announced drift is a new model version released alongside the old one, with you choosing when to move. This is the good case and the one you should insist on from any research platform: the change is an event with a date you control.

Forced migration is a deprecation. The version you pinned is retired on a published date and you must move. This is the case teams handle worst, because it arrives as an engineering ticket rather than a research decision, and the migration typically happens in whatever sprint has capacity - which is to say, in the middle of a fielding window.

The operational rule that follows: model deprecation dates belong on the research calendar, not just the engineering one. If your longitudinal study has waves in March, June, September and December and your moderator model retires in July, you now have a study design decision to make, and July is when to make it - not September, when the numbers look strange.

The bridge sample: what to do when the model has to change

Survey methodology solved this problem decades ago for a structurally identical situation: what to do when a tracking study has to change data collection mode, say from telephone to online. The answer is not to hope it does not matter, and it is not to quietly restate history. It is to measure the instrument effect directly and report it.

Adapted for model version changes, the method is:

Step 1. Field an overlap wave. For one wave, run a subset of your sample on the old model and a matched subset on the new one, at the same time, with the same guide. Everything except the model version is held constant. A few dozen interviews per arm is usually enough to see a material effect.

Step 2. Compute the delta on your headline metrics. For each tracked metric, the difference between arms is your instrument effect. It is not noise and it is not a finding about customers. It is the size of the step your trend line is about to take for purely technical reasons.

Step 3. Decide explicitly, and write the decision down. There are exactly three defensible choices, and the wrong move is making one implicitly:

  • Adjust. If the effect is stable and well-estimated, apply the offset to bring the new series onto the old basis, and footnote it on every chart that spans the change.
  • Break the series. Draw a visible discontinuity in the trend line at the change date and start a new baseline. This is honest, cheap and usually correct.
  • Accept it as immaterial. Only if the measured delta is small relative to your decision threshold - and you can only say that because you measured it.

Step 4. Record the model version in the wave metadata, so the next researcher can see exactly where the boundary sits.

The bridge sample costs one extra arm on one wave. The alternative is discovering, two years later, that the interesting inflection point in your brand tracker was a model upgrade.

The model of record

Here is the discipline that makes all of the above enforceable, and it costs nothing.

Every research method section already names the instrument: the questionnaire version, the sample source, the fielding dates, the mode. Add the model version. Not as an engineering detail buried in a config file, but as a line in the study record, next to the sample size.

Fielded 3-17 March 2026. n = 84 completed interviews. AI-moderated, text and voice. Moderator model: [version]. Analysis model: [version]. Guide version 2.1.

If you cannot fill in those two lines, you cannot state your method, and you should treat that as the same category of gap as not knowing your sample source. The test is simple and worth running today: pick your most recent study and try to write that paragraph. Most teams cannot, and discovering that is the single highest-value thing this guide can do for you.

Three things follow automatically once model version is in the study record:

  1. Comparability becomes checkable. Two studies are comparable on instrument grounds if the versions match. Right now that question is usually unanswerable.
  2. Drift becomes diagnosable. When a metric moves unexpectedly, "did the instrument change?" is a query rather than an investigation.
  3. Reproducibility claims become honest. A finding you cannot reproduce because the model changed is a different kind of finding from one that failed to replicate, and the distinction matters to whoever is making a decision on it.

Detecting drift you were not told about

If you have no version metadata for historical studies - which is the normal starting position - you can still detect drift retrospectively from the artefacts you already have.

Watch the shape of the output, not just its content. Chen, Zaharia and Zou found response length collapsing from 638.3 characters to 3.9. Structural properties of your AI outputs - mean moderator turn length, follow-ups per question, themes extracted per transcript, quality score distribution - are cheap to compute across your archive and change abruptly when a model changes. A step change in a distribution with no corresponding change in your sample or your guide is a version change until proven otherwise.

Keep a small frozen probe set. Ten to twenty fixed inputs with known-good outputs, run monthly through your analysis pipeline, and diffed. This is the continuous monitoring the paper recommends, at a scale a research team can actually sustain. It will not tell you what changed, but it will tell you when, and the date is what lets you bound the affected studies.

Diff the quality score distribution. If your platform scores interviews, the distribution of those scores is a sensitive drift detector. Scores are produced by a model, so a shift in the distribution with a stable sample is a strong signal that the scorer moved.

Re-run a historical study through the current pipeline. Take a completed corpus from six months ago, re-run the analysis, and compare the themes against what you originally shipped. The disagreement rate is a direct measurement of analysis drift over that window, computed from data you already own, at the cost of one batch job.

How Koji helps

Koji is designed so that the parts of your instrument that must stay stable are structural rather than emergent.

Structured questions are version-independent. This is the core defence, and it is not a marketing claim - it is a property of where the question definition lives. A conversational probe generated on the fly is a model output, and it changes when the model changes. A defined question with a defined answer space is data, and it does not.

TypeWhat it capturesDrift exposure
open_endedFree-form qualitative answer with AI follow-up probingHighest - the probe is model-generated
scaleNumeric rating such as 1-10 satisfaction or NPSLow - the scale is fixed, the number is comparable across versions
single_choiceOne option from a listLow - options are defined, not generated
multiple_choiceOne or more options from a listLow
rankingItems ordered by preferenceLow - fully specified task
yes_noBinary answerLowest - maximally comparable across any model change

The practical consequence for longitudinal work: anything you intend to trend should be carried by a structured question, and anything you intend to explore can be open-ended. A brand tracker whose headline numbers ride on scale and choice questions survives a model change intact, with the qualitative layer providing the colour that explains the movement. A tracker whose headline numbers are extracted by a model from free conversation does not.

Stable question IDs make re-analysis possible. Every study question carries an ID that follows it from interview plan through the AI interviewer to analysis and report aggregation. When your analysis model changes and you need to re-code the historical corpus, the questions line up across waves automatically - which is what turns "re-code everything" from a project into a job.

Full transcript export means you own the raw layer. Koji exports complete transcripts as CSV and JSON. This is what makes analysis drift recoverable in practice rather than in principle: if the coder changes, you re-run it over transcripts you hold, rather than over an API you do not control. A research platform that gives you only summaries and not raw transcripts has made analysis drift permanent for you, which is a question worth asking any vendor before you build a longitudinal programme on them.

Quality scores are per interview and inspectable. Each interview gets a 1-5 score with a written rationale, so a shift in scoring behaviour is visible in the distribution rather than hidden inside an aggregate.

The contrast with legacy tooling is instructive but cuts both ways, and it is worth being honest about it. A traditional survey platform has no drift problem, because it has no model - the trade is that it cannot probe, cannot follow a thread, and cannot ask why. Human moderators drift too, in well-documented ways, and unlike models they do it continuously and without version numbers. The AI-native answer is neither to pretend the problem does not exist nor to retreat from conversation: it is to build the trended layer out of things that do not drift, keep the exploratory layer conversational, record the version, and measure the step when it changes.

A drift policy you can adopt this quarter

  1. Record moderator and analysis model versions in every study record. Start with the next study. Backfill what you can.
  2. Treat a moderator version change as a study boundary. New version, new baseline, unless you bridged.
  3. Bridge when you must change mid-programme. One overlap wave, delta computed, decision documented.
  4. Re-code the whole corpus when the analysis model changes, never just the new wave.
  5. Put deprecation dates on the research calendar alongside fielding dates.
  6. Keep a frozen probe set and run it monthly.
  7. Carry every trended metric on a structured question, not on model-extracted prose.

Frequently asked questions

Is model version drift really a research validity problem, or just an engineering concern?

It is a textbook validity problem. Campbell and Stanley named instrumentation as a threat to internal validity in 1963, defined as changes in the measure or in the scorers that affect the measurement. An AI moderator is the measure and an AI analyser is the scorer. When either changes between waves, any observed movement is confounded with the instrument. The fact that the change is executed by an engineer does not make it an engineering matter any more than a printing error in a questionnaire would be.

How much can a model actually change between versions?

More than most teams assume. Chen, Zaharia and Zou found the same commercial model service dropping from 84.0% to 51.1% accuracy on prime identification over three months, directly executable generated code falling from 52.0% to 10.0%, and response length on one task collapsing from 638.3 characters to 3.9. Changes moved in different directions for different models on the same task, so newer does not reliably mean better and one benchmark does not predict another.

If the model gets better, is drift still a problem for trending?

Yes, and this is the counterintuitive part. For a trend line, consistency matters more than quality. A moderator that is uniformly good across all four waves gives you a usable trend; a moderator that improves between waves 2 and 3 gives you an artefact at that boundary that is indistinguishable from a real change in customers. Improve the instrument deliberately, at a boundary you choose, with a bridge sample to size the step - not silently, mid-programme.

What is a bridge sample and how big does it need to be?

A bridge sample runs a subset of one wave on both the old and new model simultaneously, holding the guide and sample frame constant, so the difference between arms isolates the instrument effect. A few dozen interviews per arm is usually enough to detect an effect large enough to matter for a decision. The point is not a precise estimate but a defensible answer to whether the step is big relative to your decision threshold.

Which is worse, drift in the moderator or drift in the analysis model?

Moderator drift, decisively, because it is irreversible. If the interviewer changed, different questions were asked and different answers were given, and no amount of reanalysis recovers the conversation that did not happen. Analysis drift corrupts only the coding layer, which you can rebuild by re-running the new analyser over your full transcript archive. Spend your version-pinning effort on collection and your re-processing effort on analysis.

How do I detect drift in studies where I never recorded the model version?

Use structural signals from the artefacts you already hold. Compute mean moderator turn length, follow-ups per question, themes per transcript and the quality score distribution across your archive, and look for step changes with no corresponding change in sample or guide. Then re-run a six-month-old corpus through your current analysis pipeline and measure the disagreement rate against what you originally shipped - that number is your analysis drift over that window.

Does version-pinning the model solve this permanently?

No, it buys you a controlled window rather than a permanent fix. Pinned versions eventually reach a published deprecation date and you are forced to migrate, which is why deprecation dates belong on the research calendar. Pinning converts silent drift into announced drift - a change you schedule, bridge and document - and that conversion is the whole benefit. It is a large one, but it is not permanence.

Related Resources


Run your first study free. Koji gives you 10 free credits when you sign up - enough to field a real study, export the raw transcripts, and start your study record with the model version written down from day one.

Related Articles

Conversation Memory and Long-Session Degradation: Why AI Interviews Get Worse After Turn 20 (2026)

AI moderators lose the thread in long conversations. The evidence, the four degradation symptoms, the session budget framework, and how to test your own moderator before it costs you a study.

Evaluation Datasets for AI Products: How to Build a Golden Set from Real User Research (2026)

How to construct and maintain the golden dataset your AI evals run against — sizing and confidence intervals, the four-bucket structure, label-error rates in published benchmarks, sourcing acceptance criteria from real users, and versioning against overfitting.

AI Model Cards and User Disclosure: Documenting Intended Use, Limitations, and What You Tell People (2026)

A practical guide to model cards, system cards, and user-facing AI disclosure — what belongs in each section, what the EU AI Act's Article 50 has required since 2 August 2026, and how to source the Limitations section from real user research instead of guesswork.

Brand Tracking Studies: How to Measure Brand Health Over Time (2026)

A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.

Exporting Research Data from Koji: CSV, JSON, and Transcript Access

A complete guide to every way you can get your interview data out of Koji — from one-click CSV downloads to real-time webhook pipelines.

Longitudinal Research: How to Track User Behavior and Attitudes Over Time

Longitudinal research captures how users change over time — not just a snapshot. This guide explains panel studies, cohort studies, and how AI-moderated interviews make multi-wave research feasible for any team.

Reliability vs. Validity in Research: What They Mean and How to Get Both

A clear guide to reliability versus validity in research: precise definitions, the dartboard analogy, the types of each, how to improve them, and how AI-moderated interviews deliver consistent, accurate insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.