Back to docs
Research Methods

Why Your Newest Cohort Always Looks Best: Reporting Lag in Customer Feedback (2026)

A recent cohort shows fewer complaints because it has had less time to report them, not because it is healthier. How to build a development triangle for feedback data and compare cohorts at equal maturity.

Your newest cohort is not your healthiest cohort. It is your least developed one, and on most dashboards those two things are impossible to tell apart.

Insurance actuaries have dealt with this for about a century and gave it a name: incurred but not reported. A loss that has already happened but has not yet been filed is still a loss, and an insurer that read this month is thin claim count as good news would be insolvent inside a few years. Customer feedback behaves exactly the same way, and almost nobody corrects for it.

The short answer

  • Complaints, cancellations and support tickets arrive with a delay. A recent cohort has had less time to report, so it shows fewer problems per account no matter how good it actually is.
  • This does not merely add noise to a trend. It can reverse the sign of the trend, so a product that is getting worse reads as a steady improvement.
  • The fix is a development triangle: measure every cohort at the same age, estimate what fraction of its eventual feedback has arrived by that age, and scale up to ultimate before comparing.
  • Never compare a two-month-old cohort with a fourteen-month-old one using a to-date number. Compare at equal maturity, or compare developed estimates.

Incurred but not reported

The insurance term is precise. According to Wikipedia, "incurred but not reported (IBNR) claims are the amount owed by an insurer to all valid claimants who have had a covered loss but have not yet reported it." The critical consequence follows immediately: "Since the insurer knows neither how many of these losses have occurred, nor the severity of each loss, IBNR is necessarily an estimate."

Actuaries then define the number they actually care about. "The sum of IBNR losses plus reported losses yields an estimate of the total eventual liabilities the insurer will cover, known as ultimate losses." Reported losses are an observation. Ultimate losses are the answer. No serious insurer confuses the two.

There is a second distinction worth importing. Pure IBNR covers claims nobody has told you about yet. Incurred but not enough reported, or IBNER, "refers to development on reported claims" - the tickets you already have that will turn out to be bigger than they currently look. Research data has both. Some customers have not complained yet, and some have complained mildly about something that will escalate.

Cancer registries publish their lag, and it is 22 months

If you think a reporting delay is a small correction, look at a field that measures it formally. The US National Cancer Institute runs delay-adjusted incidence models precisely because raw counts mislead. In their own words: "Timely and accurate calculation of cancer incidence rates is hampered by reporting delay, the time elapsed before a diagnosed cancer case is reported to the cancer registries."

The scale of the allowance is striking. NCI notes that the SEER program "allows a standard delay of 22 months between the end of the diagnosis year and the time the cancers are first reported to the NCI in November, almost two years later." And the size of the shortfall is published: "The submissions for the most recent diagnosis year are, in general, about four percent below the number of cancers that will be submitted for that year eventually, although this varies by cancer site and other factors."

Four percent sounds tolerable until you notice it applies to the most recent year only, which is exactly the year every trend line ends on. The purpose of the adjustment is stated plainly: "The idea behind modeling reporting delay is to adjust the current case count to account for anticipated future corrections (both additions and deletions) to the data."

The failure mode is a reversed sign, not a rounding error

Here is the finding that should worry anyone running a feedback dashboard. In a 2025 study of registry reporting delay, Chen, Midthune, Zou, Miller and Feuer report that "delay-adjusted rates exhibit a stabilized trend, in contrast to the rapid decline seen in observed (unadjusted) rates" (Cancer Epidemiol Biomarkers Prev, 2025;34(10):1710-1721).

Read that again in product terms. The uncorrected series showed a rapid decline. The corrected series showed no decline at all. The decline was an artefact of the lag. If you had shipped a strategy on the strength of that falling line, you would have been responding to a measurement property rather than to anything happening in the world.

A development triangle for feedback data

The actuarial tool for this is the chain-ladder method, "a prominent actuarial loss reserving technique." You arrange your data so that, as Wikipedia puts it, "Losses (either reported or paid) are compiled into a triangle, where the rows represent accident years and the columns represent valuation dates." Swap accident years for signup cohorts and valuation dates for months since signup, and the machinery transfers without modification.

Suppose you track complaints per 100 accounts, cumulatively, by months since signup. Six monthly cohorts, A the oldest through F the newest:

CohortM1M2M3M4M5M6
A8.013.016.018.019.020.0
B8.013.016.018.019.0
C8.013.016.018.0
D8.814.317.6
E9.615.6
F10.0

The diagonal edge of that triangle is what your dashboard shows you today: the latest figure for each cohort. Read it off - 20.0, 19.0, 18.0, 17.6, 15.6, 10.0 - and you have a clean, monotonic, halving improvement. Complaints per 100 accounts have apparently fallen from 20 to 10 across six cohorts. Somebody is getting promoted.

Now develop it. From the mature cohorts you can read the pattern of how feedback accumulates. Cohort A reached 8.0 of its eventual 20.0 by month one, which is 40 percent. By month two it was at 13.0, or 65 percent. The full pattern is 40, 65, 80, 90, 95, 100 percent. These proportions are what actuaries capture as "age-to-age factors, also called loss development factors (LDFs) or link ratios," which "represent the ratio of loss amounts from one valuation date to another."

Divide each cohort latest figure by the fraction of ultimate that its age implies:

CohortAgeTo-datePercent of ultimate reportedDeveloped ultimate
A6 months20.0100%20.0
B5 months19.095%20.0
C4 months18.090%20.0
D3 months17.680%22.0
E2 months15.665%24.0
F1 month10.040%25.0

The raw series fell from 20.0 to 10.0. The developed series rises from 20.0 to 25.0. Your product is getting worse by 25 percent and the dashboard reported a 50 percent improvement. The lag did not blur the signal, it inverted it.

Notice that the table validates its own construction. Cohorts A, B and C all develop to exactly 20.0 from three different ages and three different to-date numbers. That agreement is the evidence that the development pattern is real. If your oldest cohorts do not converge when developed, your pattern is wrong and you should not trust the young cohorts either.

The month-one multiplier is the number to remember

In that example a one-month-old cohort has produced 40 percent of the feedback it will eventually produce, so its complaint count needs multiplying by 2.5. That single number reframes every early read. A new release whose first month looks 30 percent better than the last release is not 30 percent better. It has reported 40 percent of its verdict.

Your own multiplier will differ, and finding it is a small piece of work with a large payoff: take cohorts that are old enough to be effectively complete, compute the cumulative fraction reported at each month of age, and keep that curve where the dashboard lives.

Compare at equal maturity

The cleanest discipline is to refuse the comparison the lag corrupts. Raluca Budiu of Nielsen Norman Group makes the general point about benchmark studies: "To soundly compare the results from the two studies, the company needs to compare apples to apples." She continues, "In other words, the studies need to have the same methodology and collect the same metrics," warning that otherwise a difference "could be due not to a better design, but to the way in which the study was conducted."

Cohort age is part of the methodology. A cohort observed for two months and a cohort observed for fourteen are not the same measurement, and setting them side by side is precisely the apples-to-oranges error. The practical rule: either truncate every cohort to the age of your youngest one, or develop them all to ultimate. Never mix.

Truncation is the safer of the two because it needs no model. If your newest cohort is two months old, compare everyone at month two: A through F read 13.0, 13.0, 13.0, 14.3, 15.6 and, for F, not yet available. The deterioration is visible immediately and you did not have to estimate anything.

The chain-ladder assumption, stated honestly

The method has one load-bearing premise, and Wikipedia states it directly: "The primary underlying assumption of the chain-ladder method is that historical loss development patterns are indicative of future loss development patterns."

That assumption breaks in ways product teams should recognise. If you make complaining easier - a new in-app feedback widget, a support chat, a shorter form - the reporting curve steepens and your old development factors will now overstate ultimate. If you add a friction step or retire a channel, it flattens and they will understate. Any change to how feedback reaches you changes the lag, which is why the honest version of this analysis ships with a note about what changed in the collection pipeline and when.

The modern approach: shorten the lag instead of only modelling it

Actuaries model reporting delay because they cannot phone every policyholder and ask whether anything has gone wrong. Product teams can, and this is where the calculus differs from insurance.

A development triangle is a correction applied to slow data. The stronger move is to stop generating slow data. Where legacy survey tools like SurveyMonkey collect a response only when a customer chooses to fill in a form - which is exactly the delay-generating step - Koji runs AI-moderated interviews you can put in front of a cohort on demand, so a two-week-old cohort can be asked directly rather than waited on. Instead of inferring the unreported from the reported, you go and collect it.

Three things make this practical rather than aspirational:

  • Speed removes most of the lag. Koji interviews run in parallel and are analysed as they complete, so a fresh cohort read takes hours rather than the months a reporting curve needs to mature. Teams using AI-assisted research tools commonly report far faster time-to-insight, and in this context speed is not a convenience, it is bias reduction.
  • Structured questions give you a stable denominator. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so every participant in a cohort answers the same scale item. A scale distribution collected from everyone at week two is directly comparable with the same item at week two of the next cohort. Unsolicited complaint counts never give you that, because the denominator is whoever happened to speak up.
  • Proactive beats passive. Koji AI interviews, including voice interviews, ask the cohort rather than waiting for it. A customer who would have complained in month five, or never, answers in week two.

The discipline still matters. Run the same Koji study at the same cohort age each time, keep the structured items fixed, and you have equal-maturity comparison by construction. Koji automatic thematic analysis then gives you themes for each cohort read without a week of manual coding, and real-time reporting means the month-two read exists while month two is still actionable.

Frequently asked questions

What is reporting lag in customer feedback?

Reporting lag is the delay between a customer experiencing a problem and that problem reaching you as a complaint, ticket, review or cancellation. Because of it, a recently acquired cohort has systematically fewer recorded problems per account than an older one, independent of product quality. Insurance calls the unrecorded remainder incurred but not reported.

Why does my newest cohort always look better?

Because it has had less time to report. If a cohort produces 40 percent of its eventual complaints in month one, then a one-month-old cohort shows 40 percent of its true problem rate while a mature cohort shows 100 percent of its own. Plotting both on the same axis produces an apparent improvement that is entirely an artefact of age.

What is a development triangle?

A table with cohorts as rows and age in periods as columns, holding the cumulative feedback each cohort had produced by each age. Mature rows reveal what fraction of eventual feedback arrives by each age, and those fractions let you project incomplete rows to their ultimate value. It is the standard actuarial chain-ladder layout applied to research data.

How many months of data do I need before comparing cohorts?

Enough for at least two or three cohorts to be effectively complete, so you can derive a development pattern and check that those cohorts converge to the same developed value. Until then, do not project: truncate every cohort to the age of the youngest and compare at equal maturity, which requires no model at all.

Does this apply to NPS and satisfaction scores too?

Yes, whenever response timing is tied to tenure. If detractors take longer to respond, or churn before responding, a young cohort score is measured on a different and more favourable mix of people than a mature one. The safeguard is the same: fix the age at which you survey each cohort and compare only like ages.

How does Koji help with reporting lag?

Koji lets you ask a cohort directly instead of waiting for it to report. AI-moderated and voice interviews can be run against a two-week-old cohort on demand, structured questions keep the denominator identical across cohort reads, and automatic thematic analysis plus real-time reporting deliver the result while it is still early enough to matter. Running the same study at the same cohort age gives equal-maturity comparison by design.

Related Resources

Related Articles

Tenure, Calendar, or Vintage: The Three Effects Hiding in Every Cohort Chart (2026)

Every cohort chart contains three clocks at once: time since signup, calendar date, and signup vintage. Learn the grid-reading diagnostic that tells them apart on sight.

Cohort Analysis: How to Read Retention and Find the "Why" (2026)

Cohort analysis groups users by a shared starting point and tracks their behavior over time, revealing retention patterns that aggregate metrics hide. This guide explains how to build and read cohort tables, interpret the retention curve, and pair the numbers with qualitative research to explain them.

The Moving Average on Your Dashboard Is Hiding the Week That Mattered (2026)

Smoothing conserves the area under an event and destroys its height - and every alert threshold you own is a height. The arithmetic of what a rolling average deletes, how late it reports, and why it cannot give you a number for now.

How Long Is User Research Valid? Insight Decay and When to Re-Run a Study

Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.