The Zero Point Moved: Adaptation and Response Shift in Long-Running Research (2026)
Your satisfaction tracker is flat while the product got better. That is not a measurement failure: the internal scale users rate against re-zeroes itself every time you ship.
Your satisfaction tracker has been flat for a year while the product got substantially better, and that is not a measurement failure — it is what success looks like when the measuring instrument re-zeroes itself. Every rating question asks a person to compare their experience against an internal reference. That reference is not fixed. It adapts to whatever you most recently gave them. Ship an improvement, and the improvement becomes the new baseline against which the next rating is made.
This is the last article in a sequence about perception. The first three treat the user as an instrument with a known threshold: find the just-noticeable difference, separate detection from willingness to answer, measure the threshold efficiently. All three assume the instrument keeps its calibration between measurements. Over any horizon longer than a session, it does not.
The failure in one sentence
A longitudinal score is a difference between two measurements. If the scale used at time 2 is not the scale used at time 1, the difference is not a change in the quantity — it is a change in the quantity plus a change in the ruler, and nothing in the data separates them.
This is a different defect from every other one in this corpus. It is not that your denominator is wrong, or your sample is unrepresentative, or the question is ambiguous. The question is clear, the sample is fine, the respondent is honest and accurate at both time points. The two honest, accurate answers are still not comparable, because they were produced against different internal standards.
What the literature calls it
In health outcomes research — which has studied this harder than any other field, because treatments visibly work while quality-of-life scores refuse to move — the phenomenon has a name. Sprangers and Schwartz, in the foundational paper (Social Science and Medicine, 1999, volume 48, pages 1507-1515), define response shift as "changes in the meaning of one's self-evaluation of QOL resulting from changes in internal standards, values, or conceptualization."
Note the three separate mechanisms in that definition, because they map exactly onto product research:
- Internal standards. The user recalibrates. A two-second load time was fine last year and is unacceptable now, because everything else got faster. The scale stretched; the experience did not change.
- Values. What matters shifts. A user who once rated your product mainly on features now rates it mainly on reliability, because they have started depending on it. Two ratings, two different weightings, one number.
- Conceptualization. The construct itself is redefined. "Easy to use" meant "few clicks" when they were new and means "does not interrupt me" now that they are expert.
Response shift, in this framing, is not noise or bias in the pejorative sense. It is an adaptation process. Sprangers and Schwartz describe it as "an important mediator" of how people accommodate to change, driven by a catalyst — a change in circumstances — running through cognitive and behavioural mechanisms. In product terms: you are the catalyst. The shipping is what moves the reference.
The size of the effect
The best available demonstration of how large this gets comes from McPhail and Haines (Health and Quality of Life Outcomes, 2010, volume 8, article 65), who followed 103 hospitalised older adults, of whom 101 completed both assessments, over a median stay of 38 days. They measured change three ways: conventionally (score at discharge minus score at admission), by asking patients how much they thought they had changed, and by adjusting that perceived change for recall bias.
The agreement between the conventional change score and the patient's own perceived change was poor: intraclass correlations of 0.34 for the EQ-5D utility score and 0.40 for the visual analogue scale. The discrepancy between the two was "considered clinically meaningful for 84 (83.2%) of participants."
Read that again in product terms. For more than four in five people, the change the instrument recorded and the change the person believed they had experienced disagreed by an amount that would alter a decision. Not a rounding difference — a decision-altering one.
The same study also shows the repair is possible. After adjusting perceived change for recall bias, the intraclass correlations rose to 0.98 and 0.90, and the proportion with a clinically meaningful discrepancy fell from 83.2% to 7.9%. The disagreement was not irreducible. It was structured, and structure can be modelled.
The measurement that actually works: the then-test
If the problem is that the time-1 rating was made on a different scale, the fix is to obtain a time-1 rating on the time-2 scale. That is the then-test, also called the retrospective pretest: at the end of the period, you ask the participant to rate not only how things are now, but how things were at the start, judged from where they stand today.
- Conventional change = (rating now) minus (rating recorded then). Contaminated by response shift.
- Then-test change = (rating now) minus (rating of the past, given now). Both on today's scale, so response shift cancels.
The gap between those two change scores is an estimate of the response shift itself, which is often the most interesting number in the study. A large gap means your users recalibrated hard, which is usually evidence that something changed enough to reset expectations.
Be honest about the trade-off, because the literature is: the then-test replaces one problem with another. A retrospective judgement is subject to recall bias — people do not remember accurately, and they tend to reconstruct the past in ways that make the present coherent. McPhail and Haines built their whole design around measuring and adjusting for exactly that. The correct posture is to run both change scores, treat the difference as informative, and never present either one alone as the truth.
This is not the same as three things it resembles
It is not panel conditioning. Panel conditioning is caused by being measured — repeated participation teaches people how to answer, so their responses drift. Response shift is caused by living through the change. A user who has never been surveyed before, answering for the first time, has still recalibrated. If you swap in a fresh cohort you eliminate conditioning entirely and response shift is untouched, because the fresh cohort has also been using a product that changed.
It is not the "no true value" problem described in split-ballot experiments. There, the number depends on the question wording and no version is privileged. Here, each rating is a perfectly valid measurement of a real state — just measured in units that were redefined between the two readings. The quantity exists at both times. The metre stick changed length.
It is not a benchmarking question. Internal benchmarks and percentile norms answer whether a 4.1 is good relative to your own history. Response shift is the reason that history is not a stable comparison set: your 2024 norm bank was built on a population whose reference point no longer exists.
Detecting it without a then-test
The then-test is the direct instrument. Three indirect signals are cheaper and worth watching continuously:
The flat line with rising behaviour. Attitudinal scores flat, behavioural metrics improving. Retention up, task completion up, support contacts down, satisfaction unchanged. That combination is the signature.
Complaints migrating down the severity ladder. When the top complaint in your open-ended responses shifts from "it loses my data" to "the export is two clicks too many", the reference point has moved. Your users are no longer grading you on the same curve, and the numeric score can stay identical throughout that migration.
A fresh-cohort divergence. Ask new users and long-tenured users the same rating question. Newcomers judge against the market; veterans judge against your last release. A widening gap between the two, with no change in the product, is a moving veteran reference.
Anchor items with external referents. Include a small number of questions whose reference cannot drift — "how long does the report take to generate, in seconds", "how many times this week did you have to retry" — alongside the rating scales. Behavioural anchors do not recalibrate. Their stability while the ratings move, or their movement while the ratings hold, tells you which side the change is on.
Running this in Koji
A response-shift design is a longitudinal study with an extra retrospective question and a conversational probe. Both are awkward in survey tools and native to AI-moderated interviews.
The structured question types carry the quantitative frame:
- scale for the current rating and, asked separately, the retrospective rating of the earlier period — the two halves of the then-test.
- yes_no for a direct check on whether the participant believes their standards changed.
- single_choice for which of several reference points they used when rating — the market, their previous experience of your product, or a specific competitor.
- multiple_choice for which aspects they now weigh most heavily, which surfaces the values component of the shift.
- ranking for ordering what matters, repeated wave over wave, so a reordering is visible directly rather than inferred.
- open_ended for the reconstruction, where Koji's AI interviewer asks its own follow-up questions.
That final piece is where the method stops being theoretical. When a participant rates the product 7 today and rates last year as 7 as well, a survey stops. Koji's AI interviewer asks what a 7 meant to them last year and what it means now — and the answers routinely reveal that the same number is describing two very different experiences. That is the response shift, in the participant's own words, captured automatically in every session with no moderator scheduled and no discussion guide improvised on the fly.
Because sessions are AI-moderated, wave 2 asks in exactly the words wave 1 used, twelve months apart, which no human moderator can guarantee and no rotating research team ever does. Voice and text both work, so participants can answer the way they prefer across waves. Analysis runs automatically and reports arrive with the wave-over-wave distributions and the supporting transcript quotes together — the two things you need to tell a real change from a recalibrated scale.
What to do with all of this
Three habits follow, and they cost very little.
Report two change scores, always. Conventional and then-test, side by side, with the gap labelled as the estimated response shift. A stakeholder who sees a flat conventional score and a strongly positive then-test score learns something a single number can never convey.
Keep behavioural anchors in every wave. They are the fixed points that let you tell scale drift from real movement.
Stop treating a flat tracker as a failure by default. In a product that is genuinely improving, expectations rise to meet it. A flat satisfaction line alongside improving behaviour and complaints that have migrated to smaller problems is a portrait of a product that is winning and a scale that has moved underneath it. The number to defend in that situation is not the rating. It is the evidence that the reference point changed — and the only way to have that evidence is to have asked.
Frequently asked questions
What is response shift in research?
Response shift is a change in the meaning of a person's self-evaluation over time, resulting from changes in their internal standards, their values, or how they conceptualise the thing being rated. It means two honest ratings from the same person at two points in time can be measured against different internal scales, so the difference between them is not purely a change in the underlying experience.
Why is my satisfaction score flat when the product clearly improved?
Because the improvement moved the reference point that people rate against. Users adapt to whatever you most recently shipped, and the new capability becomes the baseline for the next judgement rather than an addition to the old one. Flat ratings alongside improving behavioural metrics and complaints that have shifted to smaller problems is the characteristic signature.
What is a then-test and how do I run one?
A then-test, or retrospective pretest, asks participants at the end of a period to rate both how things are now and how things were at the start, both judged from their current perspective. Because both ratings use today's internal scale, the difference between them cancels out response shift. Run it alongside the conventional change score and treat the gap between the two as your estimate of the shift.
Is the then-test reliable?
It solves response shift and introduces recall bias, so it is reliable only when used alongside the conventional score rather than instead of it. In one longitudinal study of 101 hospitalised older adults, agreement between conventional and perceived change had intraclass correlations of 0.34 and 0.40, rising to 0.98 and 0.90 once perceived change was adjusted for recall bias — evidence both that the disagreement is large and that it is structured enough to model.
How is response shift different from panel conditioning?
Panel conditioning is caused by repeated participation in research: people learn how to answer and their responses drift as a result of being measured. Response shift is caused by experience with the thing being rated, and it happens to first-time respondents who have never been surveyed. Swapping in a fresh cohort removes conditioning and leaves response shift entirely intact.
Can I design a tracker that is immune to this?
No instrument that asks for a subjective rating is immune, because the reference is supplied by the respondent. What you can do is include behavioural anchor items whose referents cannot drift, ask the retrospective version of your key ratings, compare tenured users against fresh cohorts each wave, and report the estimated shift explicitly rather than folding it silently into the change score.
Related Resources
- Structured Questions in AI Interviews — the six question types used to build a then-test wave.
- Just-Noticeable Difference: The Smallest Change Users Can Perceive — the threshold assumption this article relaxes.
- Panel Conditioning: Repeat Participants and Unreliable Data — the neighbouring effect caused by measurement rather than experience.
- Longitudinal Research: Tracking Behaviour and Attitudes Over Time — the study design this problem lives inside.
- Is 4.1 Good? Internal Benchmarks and Percentile Norms — why your own history is not automatically a stable comparison.
- Split-Ballot Experiments: How Much of Your Number Is the Question? — the distinct problem of a quantity that has no question-independent value.
Related Articles
Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms
A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.
Longitudinal Research: How to Track User Behavior and Attitudes Over Time
Longitudinal research captures how users change over time — not just a snapshot. This guide explains panel studies, cohort studies, and how AI-moderated interviews make multi-wave research feasible for any team.
Panel Conditioning: Why Your Most Reliable Participants Give You the Least Reliable Data (2026)
Panel conditioning is the measurement error you create by asking the same people again. Government statistical agencies have measured it for seventy years and it moves headline numbers by a full percentage point. Here is how to detect it in a product research panel and design around it.
Split-Ballot Experiments: How Much of Your Number Is the Question?
Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)
A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.