{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-01T10:09:01.836Z"},"content":[{"type":"documentation","id":"281f9058-b40d-45f6-a854-d9c9be501cf0","slug":"extrapolating-beyond-observed-range","title":"Your Finding Is Only Valid Inside the Range You Measured (2026)","url":"https://www.koji.so/docs/extrapolating-beyond-observed-range","summary":"Every fitted relationship is valid only across the range of conditions observed. The USGS states that a stream rating curve is accurate only over the range for which discharge measurements have been made, while peak flows outside that range must be estimated by extrapolation - and calibration is densest where decisions do not matter and absent where they do. Applied to research: name the measured span of price, team size, tenure, usage and market, label predictions outside it as extrapolations, and expect systematic rather than random error because mechanisms change outside the range (worked example: a 4-point-per-step onboarding relationship fitted on 2-6 steps mispredicts a 14-step flow by 17 points because setup shifts from end user to designated admin). Findings also expire: Milly et al. (Science, 2008) argued stationarity is dead, and Coxon et al. (2015, 500 UK gauging stations) found local conditions dominate uncertainty, so generalisation limits cannot be borrowed.","content":"A research finding is a relationship fitted to a set of observations. It is trustworthy over the span of conditions you actually observed, and it is a guess outside that span. This is obvious when stated and almost never stated. The discipline that handles it best is stream gauging, where the stakes make vagueness untenable, and where the structure of the problem is exactly the structure of yours: the measurements are easy to take in ordinary conditions, and the number everyone needs is the one from the extraordinary condition nobody measured.\n\n## The short answer\n\nWrite down the range of conditions your study actually covered - price points shown, team sizes interviewed, tenure, usage intensity, market - and publish it next to the finding. Any prediction outside that range is an extrapolation, and should be labelled as one rather than reported as an estimate. When a decision depends on a region you never measured, the correct response is to go and measure there, not to extend the line.\n\n## The hydrology case: the rating curve\n\nRiver discharge is almost never measured directly in real time. Instead, hydrologists measure water height, which is cheap and continuous, and convert it to flow using a fitted relationship called a rating curve, built from paired measurements of height and flow.\n\nThe USGS states the limit of that curve without hedging: \"The rating curve is considered accurate only over the range for which discharge measurements have been made.\" And it states what happens when you need a number outside that range: \"Estimates of peak flows, which are outside the range of the established rating curve, may be made by an extrapolation of the rating curve to the peak stage.\"\n\nNow notice the structural trap, because it is the whole point. You build the curve by going out and measuring flow directly, which is practical during normal and moderate conditions. You cannot easily do it at the flood peak - that is dangerous, brief, and happens at three in the morning. So your measurements cluster in the ordinary range. And the single number the public most needs, the peak discharge of the flood, falls outside it by definition. **The calibration is densest exactly where the decision does not matter, and absent exactly where it does.**\n\nThat is not a hydrology quirk. It is the general shape of empirical knowledge, and it describes your research programme precisely.\n\n## Your study has a gauged range too\n\nEvery study has boundaries that nobody wrote down. The useful exercise is to name them explicitly:\n\n| Dimension | What you measured | What you are asked to predict |\n| --- | --- | --- |\n| Price | Reactions to 10 to 40 per seat | Whether 79 per seat is acceptable |\n| Team size | Interviews with 3 to 25 seat accounts | Behaviour of a 500 seat rollout |\n| Tenure | Users in their first 90 days | What year-three renewal depends on |\n| Usage intensity | Weekly active users who volunteered | Users who log in twice a quarter |\n| Market | Four English-speaking countries | Launch in Japan and Brazil |\n| Load | Workflows with under 50 items | Customers importing 40,000 items |\n\nEach row is a rating curve. In each one, the decision the business actually faces sits outside the measured span.\n\n## A worked example: where the line stops being a line\n\nYou study onboarding completion against the number of required setup steps. Across the accounts you observed, step counts ran from 2 to 6, and completion fell from 92 percent to 76 percent. That is a clean, strong relationship: roughly 4 percentage points of completion lost per added step.\n\nSomeone now proposes a 14-step enterprise setup and asks what completion to expect. Extending the fitted line: 76 minus 8 more steps at 4 points each gives 44 percent. The team declares the plan unviable.\n\nSix months later a 14-step enterprise flow ships anyway, and completion comes in at 61 percent. The linear model was wrong by 17 points, and crucially it was wrong in a structured, predictable direction - not noisy, but biased. The reason is that the mechanism changed outside the measured range. In 2-to-6-step accounts, setup is done by the end user in one sitting, and each step is a chance to abandon. At 14 steps, setup is done by a designated admin across several sessions as part of a paid rollout, and abandonment is governed by a completely different force. The curve bends because the thing generating the curve is no longer the same thing.\n\nThis is the failure mode to internalise: **extrapolation error is usually not random, and it is usually not a widening of your confidence interval.** It is a systematic error introduced by a mechanism change you could not observe, which is why your error bars do not warn you about it. A confidence interval computed inside the measured range describes sampling noise, not the risk of being outside the range at all.\n\n## The second failure: the relationship expires\n\nThe other way a fitted relationship stops being valid is that time passes. Hydrology names this too. Milly and colleagues published a paper in Science (2008; volume 319, pages 573-4) titled \"Stationarity Is Dead: Whither Water Management?\", arguing that the long-standing practice of designing infrastructure on the assumption that the statistical properties of the past will hold in future had become untenable under a changing climate. A river's rating curve also physically changes when the channel shifts after a flood, which is why gauges are re-measured rather than trusted indefinitely.\n\nYour relationships expire for the same kind of reason. An onboarding model fitted before you shipped an AI assistant, a conversion curve measured before a pricing change, a support-volume baseline from before a competitor exited the market - each of these was a valid measurement of a system that no longer exists. The dangerous property is that an expired model keeps producing confident, well-formatted numbers. Nothing in the output announces that the channel has moved.\n\nThere is a related and underappreciated finding on uncertainty itself. Coxon, Freer, Westerberg, Wagener, Woods and Smith applied a discharge uncertainty framework across 500 UK gauging stations (Water Resources Research, 2015; volume 51, pages 5531-5546) and found that local conditions dominate in determining the magnitude of discharge uncertainty. The practical lesson for research is that you cannot borrow somebody else's error bar. An industry benchmark for how far a finding generalises does not tell you how far yours does, because the dominant term is specific to your situation.\n\n## What to report instead\n\nThe honest version costs very little:\n\n1. **Publish the range.** One line under the finding: the span of each dimension your sample actually covered.\n2. **Label extrapolations as extrapolations.** Not a lower-confidence estimate. A different kind of claim.\n3. **State the suspected mechanism change.** If you must extrapolate, say which force you expect to take over outside the range, and in which direction it moves the answer. This converts an invisible risk into a testable prediction.\n4. **Date the finding.** Record what the system looked like when you measured it, so a future reader can tell whether the channel has shifted.\n5. **Re-measure instead of re-deriving.** When a decision hangs on an unmeasured region, the correct next step is a small targeted study in that region.\n\n## How Koji handles this\n\nPoint 5 is where the economics have genuinely changed, and it is the reason this article is actionable rather than merely cautionary. For most of the history of user research, \"go and measure the unmeasured segment\" meant weeks of recruiting and scheduling, so extending the line was the only realistic option. When a study takes hours instead, extrapolation stops being a necessary evil and becomes a choice. If your pricing evidence covers 10 to 40 and the decision is about 79, Koji lets you run the study at 79 rather than argue about the slope.\n\nKoji also makes the gauged range explicit rather than implicit. Because a Koji study defines its questions as first-class objects, the measured span is recorded in the study itself: a `scale` question carries its minimum, maximum and point labels, and `single_choice`, `multiple_choice` and `ranking` questions carry the exact options presented. You can therefore state precisely which price points, which ranges and which alternatives any finding was actually built on, instead of reconstructing it later from a screenshot. Those are four of Koji's six structured question types, alongside open_ended and yes_no.\n\nFor the expiry problem, Koji's questions carry stable identifiers that persist across studies and templates, which is what makes honest cross-study comparison possible. Re-running the same instrument next quarter and comparing like with like is the direct test of whether a relationship still holds - the research equivalent of re-measuring the gauge after the flood. And because Koji's AI interviewer probes open-ended answers in the moment, a participant who sits at the edge of your range can be asked why their situation differs, which is often how you discover the mechanism change before it ambushes a forecast.\n\n## Common mistakes\n\n- **Reporting an extrapolation with a confidence interval.** The interval describes noise inside the measured range and says nothing about the risk of leaving it. Presenting the two together implies a precision that does not exist.\n- **Assuming the error is symmetric.** Mechanism changes push in a direction. The onboarding example was wrong by 17 points, all on one side.\n- **Treating a strong fit as licence to extend.** The 2-to-6-step relationship was strong, clean and real. Strength inside the range tells you nothing about validity outside it.\n- **Borrowing someone else's generalisation limits.** Uncertainty is dominated by local conditions, so an industry rule of thumb is not a substitute for knowing your own span.\n- **Letting a baseline outlive the system it described.** Date your findings and re-measure after anything that changes the mechanism.\n- **Never recording the range at all.** This is the most common version, and it makes every later reader an unwitting extrapolator.\n\n## Frequently asked questions\n\n### What does it mean for a finding to have a calibrated range?\n\nIt means the relationship was fitted using observations that spanned a particular set of conditions - certain price points, team sizes, tenures or markets - and it is only supported across that span. Outside it, the finding is an extrapolation rather than a measurement.\n\n### Why is extrapolation error not just a wider confidence interval?\n\nBecause a confidence interval describes sampling variability inside the range you measured. Extrapolation error typically comes from a mechanism change outside that range, which is a systematic bias in a particular direction. Your error bars are computed from data that cannot contain any information about it.\n\n### How do I know whether a mechanism has changed outside my range?\n\nYou usually cannot know from the data alone, which is why you should state the suspected change as an explicit prediction. Ask who performs the behaviour at the extreme, and whether it is the same actor under the same constraints. In the onboarding example, setup shifted from an end user in one sitting to a designated admin across sessions.\n\n### What is non-stationarity and why does it matter for research?\n\nNon-stationarity means the statistical properties of a system change over time, so past measurements stop describing the future. Milly and colleagues argued in Science in 2008 that water management could no longer assume stationarity. Research findings expire the same way after a pricing change, a major feature launch or a market shift.\n\n### Can I use an industry benchmark to judge how far my findings generalise?\n\nNot reliably. Work applying an uncertainty framework across 500 UK gauging stations found that local conditions dominate the magnitude of uncertainty, which means the limits on generalisation are specific to your situation rather than borrowable from a published figure.\n\n### What should I do when a decision depends on a range I never measured?\n\nMeasure it. A small targeted study in the unmeasured region beats any amount of argument about the slope, and modern AI-moderated research makes that cheap enough to be the default rather than the luxury option.\n\n## Related Resources\n\n- [Expected Value of Information](/docs/value-of-information-research-decisions) - deciding whether the targeted study in the unmeasured region is worth running\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - why a scale that worked in one segment may not transfer to another\n- [Why Your Quarterly Metric Shows a Trend That Is Not There](/docs/measurement-interval-false-trend) - a related artifact that survives inside the measured range\n- [Model Version Drift](/docs/ai-model-version-drift-research) - what happens when the instrument itself stops being stationary mid-study\n- [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) - sample size, which governs precision rather than coverage of conditions\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and how scale and choice options record your measured range\n","category":"Analysis & Synthesis","lastModified":"2026-10-01T03:39:39.487427+00:00","metaTitle":"The Range Your Research Is Actually Valid For","metaDescription":"Findings hold only across the conditions you measured, and the decision is usually outside them. How to report your range honestly.","keywords":["extrapolating research findings","range of validity research","non-stationarity research","rating curve extrapolation","generalizing user research","research findings expire"],"aiSummary":"Every fitted relationship is valid only across the range of conditions observed. The USGS states that a stream rating curve is accurate only over the range for which discharge measurements have been made, while peak flows outside that range must be estimated by extrapolation - and calibration is densest where decisions do not matter and absent where they do. Applied to research: name the measured span of price, team size, tenure, usage and market, label predictions outside it as extrapolations, and expect systematic rather than random error because mechanisms change outside the range (worked example: a 4-point-per-step onboarding relationship fitted on 2-6 steps mispredicts a 14-step flow by 17 points because setup shifts from end user to designated admin). Findings also expire: Milly et al. (Science, 2008) argued stationarity is dead, and Coxon et al. (2015, 500 UK gauging stations) found local conditions dominate uncertainty, so generalisation limits cannot be borrowed.","aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}