{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-26T10:27:05.163Z"},"content":[{"type":"documentation","id":"a5a0a702-0bff-4940-9648-66ef8e67805e","slug":"reporting-lag-immature-cohort-feedback","title":"Why Your Newest Cohort Always Looks Best: Reporting Lag in Customer Feedback (2026)","url":"https://www.koji.so/docs/reporting-lag-immature-cohort-feedback","summary":"A recent cohort records fewer complaints per account because feedback arrives with a delay, not because it is healthier. Borrowing the actuarial concepts of IBNR and the chain-ladder development triangle, this guide shows how reporting lag can reverse the sign of a trend, how to derive a development pattern that validates itself, and why comparing cohorts at equal maturity is the safest discipline.","content":"Your newest cohort is not your healthiest cohort. It is your least developed one, and on most dashboards those two things are impossible to tell apart.\n\nInsurance actuaries have dealt with this for about a century and gave it a name: incurred but not reported. A loss that has already happened but has not yet been filed is still a loss, and an insurer that read this month is thin claim count as good news would be insolvent inside a few years. Customer feedback behaves exactly the same way, and almost nobody corrects for it.\n\n## The short answer\n\n- Complaints, cancellations and support tickets arrive with a delay. A recent cohort has had less time to report, so it shows fewer problems per account no matter how good it actually is.\n- This does not merely add noise to a trend. It can reverse the sign of the trend, so a product that is getting worse reads as a steady improvement.\n- The fix is a development triangle: measure every cohort at the same age, estimate what fraction of its eventual feedback has arrived by that age, and scale up to ultimate before comparing.\n- Never compare a two-month-old cohort with a fourteen-month-old one using a to-date number. Compare at equal maturity, or compare developed estimates.\n\n## Incurred but not reported\n\nThe insurance term is precise. According to Wikipedia, \"incurred but not reported (IBNR) claims are the amount owed by an insurer to all valid claimants who have had a covered loss but have not yet reported it.\" The critical consequence follows immediately: \"Since the insurer knows neither how many of these losses have occurred, nor the severity of each loss, IBNR is necessarily an estimate.\"\n\nActuaries then define the number they actually care about. \"The sum of IBNR losses plus reported losses yields an estimate of the total eventual liabilities the insurer will cover, known as ultimate losses.\" Reported losses are an observation. Ultimate losses are the answer. No serious insurer confuses the two.\n\nThere is a second distinction worth importing. Pure IBNR covers claims nobody has told you about yet. Incurred but not enough reported, or IBNER, \"refers to development on reported claims\" - the tickets you already have that will turn out to be bigger than they currently look. Research data has both. Some customers have not complained yet, and some have complained mildly about something that will escalate.\n\n## Cancer registries publish their lag, and it is 22 months\n\nIf you think a reporting delay is a small correction, look at a field that measures it formally. The US National Cancer Institute runs delay-adjusted incidence models precisely because raw counts mislead. In their own words: \"Timely and accurate calculation of cancer incidence rates is hampered by reporting delay, the time elapsed before a diagnosed cancer case is reported to the cancer registries.\"\n\nThe scale of the allowance is striking. NCI notes that the SEER program \"allows a standard delay of 22 months between the end of the diagnosis year and the time the cancers are first reported to the NCI in November, almost two years later.\" And the size of the shortfall is published: \"The submissions for the most recent diagnosis year are, in general, about four percent below the number of cancers that will be submitted for that year eventually, although this varies by cancer site and other factors.\"\n\nFour percent sounds tolerable until you notice it applies to the most recent year only, which is exactly the year every trend line ends on. The purpose of the adjustment is stated plainly: \"The idea behind modeling reporting delay is to adjust the current case count to account for anticipated future corrections (both additions and deletions) to the data.\"\n\n## The failure mode is a reversed sign, not a rounding error\n\nHere is the finding that should worry anyone running a feedback dashboard. In a 2025 study of registry reporting delay, Chen, Midthune, Zou, Miller and Feuer report that \"delay-adjusted rates exhibit a stabilized trend, in contrast to the rapid decline seen in observed (unadjusted) rates\" (Cancer Epidemiol Biomarkers Prev, 2025;34(10):1710-1721).\n\nRead that again in product terms. The uncorrected series showed a rapid decline. The corrected series showed no decline at all. The decline was an artefact of the lag. If you had shipped a strategy on the strength of that falling line, you would have been responding to a measurement property rather than to anything happening in the world.\n\n## A development triangle for feedback data\n\nThe actuarial tool for this is the chain-ladder method, \"a prominent actuarial loss reserving technique.\" You arrange your data so that, as Wikipedia puts it, \"Losses (either reported or paid) are compiled into a triangle, where the rows represent accident years and the columns represent valuation dates.\" Swap accident years for signup cohorts and valuation dates for months since signup, and the machinery transfers without modification.\n\nSuppose you track complaints per 100 accounts, cumulatively, by months since signup. Six monthly cohorts, A the oldest through F the newest:\n\n| Cohort | M1 | M2 | M3 | M4 | M5 | M6 |\n|---|---|---|---|---|---|---|\n| A | 8.0 | 13.0 | 16.0 | 18.0 | 19.0 | 20.0 |\n| B | 8.0 | 13.0 | 16.0 | 18.0 | 19.0 | |\n| C | 8.0 | 13.0 | 16.0 | 18.0 | | |\n| D | 8.8 | 14.3 | 17.6 | | | |\n| E | 9.6 | 15.6 | | | | |\n| F | 10.0 | | | | | |\n\nThe diagonal edge of that triangle is what your dashboard shows you today: the latest figure for each cohort. Read it off - 20.0, 19.0, 18.0, 17.6, 15.6, 10.0 - and you have a clean, monotonic, halving improvement. Complaints per 100 accounts have apparently fallen from 20 to 10 across six cohorts. Somebody is getting promoted.\n\nNow develop it. From the mature cohorts you can read the pattern of how feedback accumulates. Cohort A reached 8.0 of its eventual 20.0 by month one, which is 40 percent. By month two it was at 13.0, or 65 percent. The full pattern is 40, 65, 80, 90, 95, 100 percent. These proportions are what actuaries capture as \"age-to-age factors, also called loss development factors (LDFs) or link ratios,\" which \"represent the ratio of loss amounts from one valuation date to another.\"\n\nDivide each cohort latest figure by the fraction of ultimate that its age implies:\n\n| Cohort | Age | To-date | Percent of ultimate reported | Developed ultimate |\n|---|---|---|---|---|\n| A | 6 months | 20.0 | 100% | 20.0 |\n| B | 5 months | 19.0 | 95% | 20.0 |\n| C | 4 months | 18.0 | 90% | 20.0 |\n| D | 3 months | 17.6 | 80% | 22.0 |\n| E | 2 months | 15.6 | 65% | 24.0 |\n| F | 1 month | 10.0 | 40% | 25.0 |\n\nThe raw series fell from 20.0 to 10.0. The developed series rises from 20.0 to 25.0. Your product is getting worse by 25 percent and the dashboard reported a 50 percent improvement. The lag did not blur the signal, it inverted it.\n\nNotice that the table validates its own construction. Cohorts A, B and C all develop to exactly 20.0 from three different ages and three different to-date numbers. That agreement is the evidence that the development pattern is real. If your oldest cohorts do not converge when developed, your pattern is wrong and you should not trust the young cohorts either.\n\n## The month-one multiplier is the number to remember\n\nIn that example a one-month-old cohort has produced 40 percent of the feedback it will eventually produce, so its complaint count needs multiplying by 2.5. That single number reframes every early read. A new release whose first month looks 30 percent better than the last release is not 30 percent better. It has reported 40 percent of its verdict.\n\nYour own multiplier will differ, and finding it is a small piece of work with a large payoff: take cohorts that are old enough to be effectively complete, compute the cumulative fraction reported at each month of age, and keep that curve where the dashboard lives.\n\n## Compare at equal maturity\n\nThe cleanest discipline is to refuse the comparison the lag corrupts. Raluca Budiu of Nielsen Norman Group makes the general point about benchmark studies: \"To soundly compare the results from the two studies, the company needs to compare apples to apples.\" She continues, \"In other words, the studies need to have the same methodology and collect the same metrics,\" warning that otherwise a difference \"could be due not to a better design, but to the way in which the study was conducted.\"\n\nCohort age is part of the methodology. A cohort observed for two months and a cohort observed for fourteen are not the same measurement, and setting them side by side is precisely the apples-to-oranges error. The practical rule: either truncate every cohort to the age of your youngest one, or develop them all to ultimate. Never mix.\n\nTruncation is the safer of the two because it needs no model. If your newest cohort is two months old, compare everyone at month two: A through F read 13.0, 13.0, 13.0, 14.3, 15.6 and, for F, not yet available. The deterioration is visible immediately and you did not have to estimate anything.\n\n## The chain-ladder assumption, stated honestly\n\nThe method has one load-bearing premise, and Wikipedia states it directly: \"The primary underlying assumption of the chain-ladder method is that historical loss development patterns are indicative of future loss development patterns.\"\n\nThat assumption breaks in ways product teams should recognise. If you make complaining easier - a new in-app feedback widget, a support chat, a shorter form - the reporting curve steepens and your old development factors will now overstate ultimate. If you add a friction step or retire a channel, it flattens and they will understate. Any change to how feedback reaches you changes the lag, which is why the honest version of this analysis ships with a note about what changed in the collection pipeline and when.\n\n## The modern approach: shorten the lag instead of only modelling it\n\nActuaries model reporting delay because they cannot phone every policyholder and ask whether anything has gone wrong. Product teams can, and this is where the calculus differs from insurance.\n\nA development triangle is a correction applied to slow data. The stronger move is to stop generating slow data. Where legacy survey tools like SurveyMonkey collect a response only when a customer chooses to fill in a form - which is exactly the delay-generating step - Koji runs AI-moderated interviews you can put in front of a cohort on demand, so a two-week-old cohort can be asked directly rather than waited on. Instead of inferring the unreported from the reported, you go and collect it.\n\nThree things make this practical rather than aspirational:\n\n- **Speed removes most of the lag.** Koji interviews run in parallel and are analysed as they complete, so a fresh cohort read takes hours rather than the months a reporting curve needs to mature. Teams using AI-assisted research tools commonly report far faster time-to-insight, and in this context speed is not a convenience, it is bias reduction.\n- **Structured questions give you a stable denominator.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - so every participant in a cohort answers the same scale item. A scale distribution collected from everyone at week two is directly comparable with the same item at week two of the next cohort. Unsolicited complaint counts never give you that, because the denominator is whoever happened to speak up.\n- **Proactive beats passive.** Koji AI interviews, including voice interviews, ask the cohort rather than waiting for it. A customer who would have complained in month five, or never, answers in week two.\n\nThe discipline still matters. Run the same Koji study at the same cohort age each time, keep the structured items fixed, and you have equal-maturity comparison by construction. Koji automatic thematic analysis then gives you themes for each cohort read without a week of manual coding, and real-time reporting means the month-two read exists while month two is still actionable.\n\n## Frequently asked questions\n\n### What is reporting lag in customer feedback?\n\nReporting lag is the delay between a customer experiencing a problem and that problem reaching you as a complaint, ticket, review or cancellation. Because of it, a recently acquired cohort has systematically fewer recorded problems per account than an older one, independent of product quality. Insurance calls the unrecorded remainder incurred but not reported.\n\n### Why does my newest cohort always look better?\n\nBecause it has had less time to report. If a cohort produces 40 percent of its eventual complaints in month one, then a one-month-old cohort shows 40 percent of its true problem rate while a mature cohort shows 100 percent of its own. Plotting both on the same axis produces an apparent improvement that is entirely an artefact of age.\n\n### What is a development triangle?\n\nA table with cohorts as rows and age in periods as columns, holding the cumulative feedback each cohort had produced by each age. Mature rows reveal what fraction of eventual feedback arrives by each age, and those fractions let you project incomplete rows to their ultimate value. It is the standard actuarial chain-ladder layout applied to research data.\n\n### How many months of data do I need before comparing cohorts?\n\nEnough for at least two or three cohorts to be effectively complete, so you can derive a development pattern and check that those cohorts converge to the same developed value. Until then, do not project: truncate every cohort to the age of the youngest and compare at equal maturity, which requires no model at all.\n\n### Does this apply to NPS and satisfaction scores too?\n\nYes, whenever response timing is tied to tenure. If detractors take longer to respond, or churn before responding, a young cohort score is measured on a different and more favourable mix of people than a mature one. The safeguard is the same: fix the age at which you survey each cohort and compare only like ages.\n\n### How does Koji help with reporting lag?\n\nKoji lets you ask a cohort directly instead of waiting for it to report. AI-moderated and voice interviews can be run against a two-week-old cohort on demand, structured questions keep the denominator identical across cohort reads, and automatic thematic analysis plus real-time reporting deliver the result while it is still early enough to matter. Running the same study at the same cohort age gives equal-maturity comparison by design.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and why a fixed denominator survives cohort comparison\n- [Cohort Analysis: How to Read Retention and Find the Why](/docs/cohort-analysis-guide) - the dashboard this article is warning you about\n- [Tenure, Calendar, or Vintage](/docs/age-period-cohort-effects-product-research) - the other reason two cohorts are not comparable\n- [The Moving Average Is Hiding the Week That Mattered](/docs/moving-average-hides-events-research) - a second way a smooth line conceals an event\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - who never reports at all, as distinct from who reports late\n- [How Long Is User Research Valid?](/docs/research-refresh-cadence) - when a finding has aged out rather than under-reported","category":"Research Methods","lastModified":"2026-09-26T03:35:12.226068+00:00","metaTitle":"Reporting Lag in Customer Feedback: Why New Cohorts Look Best (2026)","metaDescription":"Your newest cohort looks best because its complaints have not arrived yet. Build a development triangle and compare cohorts at equal maturity.","keywords":["reporting lag customer feedback","incurred but not reported","development triangle","chain ladder method","immature cohort comparison","cohort maturity bias","delay adjusted rates"],"aiSummary":"A recent cohort records fewer complaints per account because feedback arrives with a delay, not because it is healthier. Borrowing the actuarial concepts of IBNR and the chain-ladder development triangle, this guide shows how reporting lag can reverse the sign of a trend, how to derive a development pattern that validates itself, and why comparing cohorts at equal maturity is the safest discipline.","aiPrerequisites":["Basic familiarity with cohort or retention charts","Ability to group customers by signup period"],"aiLearningOutcomes":["Explain reporting lag and why it biases recent cohorts","Build a development triangle from cohort feedback data","Derive loss development factors and project to ultimate","Compare cohorts at equal maturity instead of to-date","Recognise when a development pattern has been invalidated"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}