Back to docs
Analysis & Synthesis

Your Finding Is Only Valid Inside the Range You Measured (2026)

Every research finding has a calibrated range, and the decision you need it for is usually outside that range. Hydrologists have a name and a discipline for this problem.

A research finding is a relationship fitted to a set of observations. It is trustworthy over the span of conditions you actually observed, and it is a guess outside that span. This is obvious when stated and almost never stated. The discipline that handles it best is stream gauging, where the stakes make vagueness untenable, and where the structure of the problem is exactly the structure of yours: the measurements are easy to take in ordinary conditions, and the number everyone needs is the one from the extraordinary condition nobody measured.

The short answer

Write down the range of conditions your study actually covered - price points shown, team sizes interviewed, tenure, usage intensity, market - and publish it next to the finding. Any prediction outside that range is an extrapolation, and should be labelled as one rather than reported as an estimate. When a decision depends on a region you never measured, the correct response is to go and measure there, not to extend the line.

The hydrology case: the rating curve

River discharge is almost never measured directly in real time. Instead, hydrologists measure water height, which is cheap and continuous, and convert it to flow using a fitted relationship called a rating curve, built from paired measurements of height and flow.

The USGS states the limit of that curve without hedging: "The rating curve is considered accurate only over the range for which discharge measurements have been made." And it states what happens when you need a number outside that range: "Estimates of peak flows, which are outside the range of the established rating curve, may be made by an extrapolation of the rating curve to the peak stage."

Now notice the structural trap, because it is the whole point. You build the curve by going out and measuring flow directly, which is practical during normal and moderate conditions. You cannot easily do it at the flood peak - that is dangerous, brief, and happens at three in the morning. So your measurements cluster in the ordinary range. And the single number the public most needs, the peak discharge of the flood, falls outside it by definition. The calibration is densest exactly where the decision does not matter, and absent exactly where it does.

That is not a hydrology quirk. It is the general shape of empirical knowledge, and it describes your research programme precisely.

Your study has a gauged range too

Every study has boundaries that nobody wrote down. The useful exercise is to name them explicitly:

DimensionWhat you measuredWhat you are asked to predict
PriceReactions to 10 to 40 per seatWhether 79 per seat is acceptable
Team sizeInterviews with 3 to 25 seat accountsBehaviour of a 500 seat rollout
TenureUsers in their first 90 daysWhat year-three renewal depends on
Usage intensityWeekly active users who volunteeredUsers who log in twice a quarter
MarketFour English-speaking countriesLaunch in Japan and Brazil
LoadWorkflows with under 50 itemsCustomers importing 40,000 items

Each row is a rating curve. In each one, the decision the business actually faces sits outside the measured span.

A worked example: where the line stops being a line

You study onboarding completion against the number of required setup steps. Across the accounts you observed, step counts ran from 2 to 6, and completion fell from 92 percent to 76 percent. That is a clean, strong relationship: roughly 4 percentage points of completion lost per added step.

Someone now proposes a 14-step enterprise setup and asks what completion to expect. Extending the fitted line: 76 minus 8 more steps at 4 points each gives 44 percent. The team declares the plan unviable.

Six months later a 14-step enterprise flow ships anyway, and completion comes in at 61 percent. The linear model was wrong by 17 points, and crucially it was wrong in a structured, predictable direction - not noisy, but biased. The reason is that the mechanism changed outside the measured range. In 2-to-6-step accounts, setup is done by the end user in one sitting, and each step is a chance to abandon. At 14 steps, setup is done by a designated admin across several sessions as part of a paid rollout, and abandonment is governed by a completely different force. The curve bends because the thing generating the curve is no longer the same thing.

This is the failure mode to internalise: extrapolation error is usually not random, and it is usually not a widening of your confidence interval. It is a systematic error introduced by a mechanism change you could not observe, which is why your error bars do not warn you about it. A confidence interval computed inside the measured range describes sampling noise, not the risk of being outside the range at all.

The second failure: the relationship expires

The other way a fitted relationship stops being valid is that time passes. Hydrology names this too. Milly and colleagues published a paper in Science (2008; volume 319, pages 573-4) titled "Stationarity Is Dead: Whither Water Management?", arguing that the long-standing practice of designing infrastructure on the assumption that the statistical properties of the past will hold in future had become untenable under a changing climate. A river's rating curve also physically changes when the channel shifts after a flood, which is why gauges are re-measured rather than trusted indefinitely.

Your relationships expire for the same kind of reason. An onboarding model fitted before you shipped an AI assistant, a conversion curve measured before a pricing change, a support-volume baseline from before a competitor exited the market - each of these was a valid measurement of a system that no longer exists. The dangerous property is that an expired model keeps producing confident, well-formatted numbers. Nothing in the output announces that the channel has moved.

There is a related and underappreciated finding on uncertainty itself. Coxon, Freer, Westerberg, Wagener, Woods and Smith applied a discharge uncertainty framework across 500 UK gauging stations (Water Resources Research, 2015; volume 51, pages 5531-5546) and found that local conditions dominate in determining the magnitude of discharge uncertainty. The practical lesson for research is that you cannot borrow somebody else's error bar. An industry benchmark for how far a finding generalises does not tell you how far yours does, because the dominant term is specific to your situation.

What to report instead

The honest version costs very little:

  1. Publish the range. One line under the finding: the span of each dimension your sample actually covered.
  2. Label extrapolations as extrapolations. Not a lower-confidence estimate. A different kind of claim.
  3. State the suspected mechanism change. If you must extrapolate, say which force you expect to take over outside the range, and in which direction it moves the answer. This converts an invisible risk into a testable prediction.
  4. Date the finding. Record what the system looked like when you measured it, so a future reader can tell whether the channel has shifted.
  5. Re-measure instead of re-deriving. When a decision hangs on an unmeasured region, the correct next step is a small targeted study in that region.

How Koji handles this

Point 5 is where the economics have genuinely changed, and it is the reason this article is actionable rather than merely cautionary. For most of the history of user research, "go and measure the unmeasured segment" meant weeks of recruiting and scheduling, so extending the line was the only realistic option. When a study takes hours instead, extrapolation stops being a necessary evil and becomes a choice. If your pricing evidence covers 10 to 40 and the decision is about 79, Koji lets you run the study at 79 rather than argue about the slope.

Koji also makes the gauged range explicit rather than implicit. Because a Koji study defines its questions as first-class objects, the measured span is recorded in the study itself: a scale question carries its minimum, maximum and point labels, and single_choice, multiple_choice and ranking questions carry the exact options presented. You can therefore state precisely which price points, which ranges and which alternatives any finding was actually built on, instead of reconstructing it later from a screenshot. Those are four of Koji's six structured question types, alongside open_ended and yes_no.

For the expiry problem, Koji's questions carry stable identifiers that persist across studies and templates, which is what makes honest cross-study comparison possible. Re-running the same instrument next quarter and comparing like with like is the direct test of whether a relationship still holds - the research equivalent of re-measuring the gauge after the flood. And because Koji's AI interviewer probes open-ended answers in the moment, a participant who sits at the edge of your range can be asked why their situation differs, which is often how you discover the mechanism change before it ambushes a forecast.

Common mistakes

  • Reporting an extrapolation with a confidence interval. The interval describes noise inside the measured range and says nothing about the risk of leaving it. Presenting the two together implies a precision that does not exist.
  • Assuming the error is symmetric. Mechanism changes push in a direction. The onboarding example was wrong by 17 points, all on one side.
  • Treating a strong fit as licence to extend. The 2-to-6-step relationship was strong, clean and real. Strength inside the range tells you nothing about validity outside it.
  • Borrowing someone else's generalisation limits. Uncertainty is dominated by local conditions, so an industry rule of thumb is not a substitute for knowing your own span.
  • Letting a baseline outlive the system it described. Date your findings and re-measure after anything that changes the mechanism.
  • Never recording the range at all. This is the most common version, and it makes every later reader an unwitting extrapolator.

Frequently asked questions

What does it mean for a finding to have a calibrated range?

It means the relationship was fitted using observations that spanned a particular set of conditions - certain price points, team sizes, tenures or markets - and it is only supported across that span. Outside it, the finding is an extrapolation rather than a measurement.

Why is extrapolation error not just a wider confidence interval?

Because a confidence interval describes sampling variability inside the range you measured. Extrapolation error typically comes from a mechanism change outside that range, which is a systematic bias in a particular direction. Your error bars are computed from data that cannot contain any information about it.

How do I know whether a mechanism has changed outside my range?

You usually cannot know from the data alone, which is why you should state the suspected change as an explicit prediction. Ask who performs the behaviour at the extreme, and whether it is the same actor under the same constraints. In the onboarding example, setup shifted from an end user in one sitting to a designated admin across sessions.

What is non-stationarity and why does it matter for research?

Non-stationarity means the statistical properties of a system change over time, so past measurements stop describing the future. Milly and colleagues argued in Science in 2008 that water management could no longer assume stationarity. Research findings expire the same way after a pricing change, a major feature launch or a market shift.

Can I use an industry benchmark to judge how far my findings generalise?

Not reliably. Work applying an uncertainty framework across 500 UK gauging stations found that local conditions dominate the magnitude of uncertainty, which means the limits on generalisation are specific to your situation rather than borrowable from a published figure.

What should I do when a decision depends on a range I never measured?

Measure it. A small targeted study in the unmeasured region beats any amount of argument about the slope, and modern AI-moderated research makes that cheap enough to be the default rather than the luxury option.

Related Resources

Related Articles

Model Version Drift: What Happens to Your Research When the AI Changes Mid-Study (2026)

When the model behind your AI moderator or analyst is upgraded, your measuring instrument changed. The evidence, the three layers of drift, the bridge sample method, and how to make model version part of your method section.

How Many Interviews Are Enough? A Guide to Sample Size

Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.

Why Your Quarterly Metric Shows a Trend That Is Not There (2026)

Undersampling does not blur a cycle, it counterfeits a different one. How the gap between your measurement waves manufactures smooth trends, flat lines, and reversed directions - and the three-question test that catches it.

Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)

Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Activating Research Insights: Turn Findings Into Product Decisions

A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.