Why Anything You Measure Mid-Flight Looks Longer Than It Is (2026)
Sampling items that are still in progress oversamples long ones. Learn why snapshot duration estimates are inflated by variance over mean, and how to fix the frame.
Anything you measure while it is still running looks longer than it really is.
The short answer
If you estimate how long onboarding takes by looking at the accounts that are in onboarding right now, or how long support takes by sampling the tickets sitting in the queue today, your estimate is too high by a precise and predictable amount: the variance of the true durations divided by their mean. In the worked example below, a true average of 10.4 days gets reported as 26.2 days. That is an error of two and a half times, and not one measurement was taken incorrectly.
This is called length-biased sampling, and the version of it that bites product teams is known as the inspection paradox. It is arithmetic, not bad luck. No amount of extra data fixes it, because every additional observation is drawn from the same distorted frame. What fixes it is changing what you sample.
Why a snapshot oversamples long things
The mechanism is almost too simple to believe. A long process is in progress for more days than a short one, so on any given day it is more likely to be caught in a snapshot. A process that takes 30 days is available to be sampled on 30 different days. A process that takes 2 days is available on 2. Sample "what is happening now" and you have quietly weighted every case by its own duration.
Renewal theory, the branch of probability that studies recurring events, states the result plainly. As the standard formulation puts it, "A curious feature of renewal processes is that if we wait some predetermined time t and then observe how large the renewal interval containing t is, we should expect it to be typically larger than a renewal interval of average size." The explanation is the part worth memorising: "The resolution of the paradox is that our sampled distribution at time t is size-biased (see sampling bias), in that the likelihood an interval is chosen is proportional to its size."
The classic illustration is transport. In the same treatment: "A vivid example is the bus waiting time paradox: For a given random distribution of bus arrivals, the average rider at a bus stop observes more delays than the average operator of the buses." Neither party is wrong. The operator averages over buses; the rider averages over riders, and riders pile up during the long gaps. Two honest averages of the same timetable disagree because they use different sampling frames.
Medicine formalised the same effect under a different name. Cancer screening research defines length time bias as "an overestimation of survival duration due to the relative excess of cases detected that are asymptomatically slowly progressing, while fast progressing cases are detected after giving symptoms." The reason is structural: "As a result, if the same number of slow-growing and fast-growing tumors appear in a year, the screening test detects more slow-growers than fast-growers." A screening programme can look like it extends life when all it has done is preferentially find the slow cases.
Statisticians who design studies this way name the problem explicitly. Writing on cross-sectional prevalent cohort designs in Statistical Methods in Medical Research, Liu, Shen, Ning and Qin note that "The sampling scheme in such design gives rise to length-biased data that require specialized analysis strategy but can improve study efficiency." That last clause matters, and we come back to it: length-biased data is not garbage. It is informative data on the wrong scale, and it can be corrected.
The arithmetic, with a number you can check
Take a deliberately simple onboarding process. Seventy percent of accounts finish in 2 days. Thirty percent get stuck and take 30 days.
The true average is straightforward: 0.7 times 2, plus 0.3 times 30, which is 10.4 days. That is the number you want.
Now sample the accounts that are mid-onboarding on an arbitrary Tuesday. The chance of catching a given account is proportional to how long it stays in that state, so the weights are 0.7 times 2 for the fast group and 0.3 times 30 for the slow group. Those are 1.4 and 9, out of a total of 10.4. So the slow group, which is 30 percent of accounts, makes up 9 divided by 10.4, or about 87 percent of your sample. The fast group, 70 percent of all accounts, is only 13.46 percent of what you see.
The mean you would report is 26.2 days against a truth of 10.4. And there is a closed form that predicts it exactly. The observed mean equals the true mean plus the variance divided by the true mean. The variance here is 164.64, so 10.4 plus 164.64 divided by 10.4 gives 26.2308 - identical to the direct calculation.
That formula is the practical takeaway, because it tells you when to worry. The inflation term is variance over mean, so the penalty scales with how spread out your durations are. If every account onboards in exactly 10 days, the variance is zero and a snapshot is perfectly accurate. The more your durations vary, the more a snapshot lies, and the direction is always the same: upward, never downward. Processes with a long tail - onboarding, enterprise procurement, escalated tickets, research recruitment - are precisely the processes where snapshots are least trustworthy.
Where this shows up in product and customer research
The pattern is easy to recognise once you have the shape of it.
Estimating cycle time from open work. "Our average time to resolve is 9 days" computed from currently open tickets is length-biased. Tickets that resolve in an hour are almost never in the open set. Compute it from tickets that closed in a window instead.
Interviewing people who are currently in a flow. If you recruit participants who are mid-onboarding, mid-migration or mid-evaluation, you have oversampled the people for whom that stage is dragging. Their frustration is real and worth hearing, but it is not the average experience, and treating it as such will send you to fix the wrong step.
Asking how long something took. A question like "how long did setup take you?" fielded to active users inherits the same bias, because users still in setup are disproportionately the slow ones. The wording is fine; the frame is not.
Reading session or visit duration. Sampling sessions that are open at a moment in time oversamples long sessions. Average duration computed that way describes the tail, not the middle.
Judging a research panel by who is available. Participants who are always available to take a study are, by construction, people with long gaps between commitments. Convenience sampling from an always-on panel is a size-biased draw on free time.
Why this is not survivorship bias
These two get confused constantly, and the fix for one does not fix the other.
Survivorship bias is about who is missing from your frame. You interview current customers, the churned ones are gone, and you never hear the disconfirming story. It is a question of membership. The remedy is to go and find the people who left.
Length bias is about weighting inside a frame where nobody is missing. Every account is eligible for your onboarding snapshot. The problem is that they are not equally likely to be caught, and the probability is proportional to duration. You could have a perfect, complete, zero-attrition population and still get 26.2 instead of 10.4.
A short test distinguishes them. If the cases you are missing would change the answer, that is survivorship. If you are seeing every kind of case but the long ones show up more often than they should, that is length bias. They can of course occur together, which is why it helps to name them separately. For the membership problem, see our guide to survivorship bias in customer research; for what tenure does to a churn number, see the churn hazard curve.
The three fixes
Switch from a snapshot to a cohort. Instead of "cases in progress now", take "all cases that started in March" and follow them to completion. This removes the bias entirely rather than adjusting for it, and it is the only fix that requires no assumptions. The cost is waiting for the slowest case to finish.
Reweight by the inverse of duration. If you must use a snapshot, weight each observation by one over its duration to undo the size-biasing. This recovers the underlying distribution and is the standard correction in prevalent cohort analysis. It needs each case duration recorded, and it inflates the influence of the few short cases you did catch, so confidence intervals widen.
Report the right quantity to the right audience. The rider and the operator both needed a number, and their numbers legitimately differ. If you want to know "how long does onboarding take", use the cohort mean. If you want to know "what is the typical experience of someone currently stuck", the length-biased figure is the honest answer to that question. State which one you computed. Most reporting disasters here come from computing the second and labelling it the first.
One more note on summarising skewed durations at all. Nielsen Norman Group's quantitative glossary warns that with "skewed distributions like task times, the median or the geometric mean may be more appropriate" than the arithmetic mean, and observes that "Time-on-task data is often skewed." Fixing the sampling frame and then reporting a mean that the distribution cannot support solves only half the problem.
How Koji helps
The reason teams reach for snapshots is that cohort research is expensive to run. Waiting for a March cohort to finish, then recruiting and interviewing a properly framed sample, has traditionally meant weeks of scheduling. That cost is what makes "just look at who is in the flow now" so tempting, and Koji exists to remove it.
Koji runs AI-moderated interviews that launch in minutes and run in parallel, so a cohort-based study stops being a quarterly project. You can define the frame you actually want - everyone who started onboarding in a given window, including the people who finished in two days and would never appear in a snapshot - and field it to all of them at once. Voice interviews reach the fast finishers who would not book a 45-minute call for a process they barely noticed, which is exactly the group length bias hides.
Koji's structured questions are what make the correction computable. Alongside open-ended conversation, Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so you can capture the duration itself as a scale or single_choice item while the AI interviewer probes the story behind it. Recording duration as structured data for every participant is the precondition for inverse-duration weighting; you cannot reweight what you only have as prose. Koji's automatic thematic analysis then reports the qualitative findings against those durations, so a theme that only appears among slow cases is visible as such rather than blended into one average.
Because Koji's AI consultants are customisable, you can also brief the interviewer to establish start and end dates explicitly rather than accepting "a while ago". And because real-time reporting updates as interviews land, you can watch whether your fast-finisher segment is actually filling up, instead of discovering at analysis time that your sample is 87 percent stuck accounts. Traditional survey tools like SurveyMonkey will happily average whatever arrives; an AI-native platform like Koji lets you specify the frame, verify it as it fills, and keep the duration data structured enough to correct.
Common mistakes
Adding more data to a biased snapshot. A larger snapshot converges more precisely on the wrong number. Precision is not accuracy, and here they move in opposite directions.
Assuming the bias is small because the process is short. The inflation is variance over mean, not mean. A process averaging 3 days with a handful of 60-day stragglers can be badly length-biased.
Treating the median as immune. The median of a length-biased sample is also shifted upward. It is more robust to the tail, but robustness is not correction.
Mixing frames in one chart. Plotting snapshot-derived durations next to cohort-derived durations produces a trend that is an artefact of how each point was collected.
Forgetting that the fix is free at design time. Choosing a cohort frame before you collect costs nothing. Correcting a snapshot afterwards costs assumptions and precision.
Frequently asked questions
Is length-biased sampling the same as the inspection paradox?
They are the same phenomenon under two names, with a small difference in emphasis. Length-biased sampling describes the distorted distribution you end up with when selection probability is proportional to duration. The inspection paradox describes the surprising consequence, that an interval observed at an arbitrary moment is on average longer than a randomly chosen interval. If you understand one, you understand the other.
How do I know whether my duration estimate is affected?
Ask one question: did each case have an equal chance of entering my sample, or did cases that lasted longer have more opportunities to be included? If you selected on being in progress, in the queue, currently active, or still open, the answer is the second one and your estimate is inflated. Cases selected by their start date are unaffected.
Can I just use the median instead of the mean?
No, though it helps. The whole sampled distribution is shifted, not just its mean, so the median of a length-biased sample is also too high. The median is less sensitive to the extreme tail, which makes the damage smaller, but it does not remove the bias. Only reweighting or a cohort frame does that.
How large can the error actually get?
Unbounded in principle, because the inflation term is the variance divided by the mean and variance has no upper limit. The worked example in this guide produces a factor of about 2.5 from a modest two-group process. Durations with a heavy tail, such as enterprise deals or escalated tickets, can be inflated much further.
Does this affect qualitative research or only metrics?
Both, and the qualitative case is more insidious because there is no number to sanity-check. If you recruit participants who are currently mid-process, you have systematically oversampled people for whom that process is slow, and their themes will dominate your analysis. The transcripts are accurate; the sample is weighted.
What is the single fastest fix?
Change the selection rule from "in progress now" to "started during a fixed window", then follow those cases to completion. It requires no statistical correction and no assumptions about the duration distribution. Koji makes this practical by letting you field the study to an entire start-date cohort at once rather than interviewing whoever happens to be available.
Related Resources
- Survivorship Bias in Customer Research - the membership problem this guide is often confused with
- The Churn Hazard Curve - what tenure does to a single churn rate
- Nonresponse Bias - when who answers is not who you asked
- Capture-Recapture for Theme Coverage - estimating what your study never found
- Structured Questions in AI Interviews - the six question types that make corrections computable
- How Many Interviews Are Enough? - sample size, once your frame is right
Related Articles
Capture-Recapture for Research: How to Estimate the Themes Your Study Never Found (2026)
Two independent coding passes turn "no new themes" into a number: the overlap between them estimates how many themes neither pass ever reached.
The Churn Hazard Curve: Why One Churn Rate Hides Three Different Problems (2026)
Your monthly churn rate averages three unrelated problems into one number. Learn to plot the churn hazard by tenure, read the three regimes, and avoid the sorting trap that makes a flattening curve look like product-market fit.
How Many Interviews Are Enough? A Guide to Sample Size
Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.
Nonresponse Bias: How Missing Respondents Skew Your Data
Nonresponse bias occurs when the people who do not answer your survey differ systematically from those who do. Learn why a low response rate is not the same as bias, how to detect it, and how to reduce it.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.