Back to docs
Research Methods

Three Stories, One Grid: The Question Your Cohort Data Cannot Answer (2026)

Tenure, calendar and vintage are exactly collinear, so infinitely many contradictory explanations fit a cohort grid perfectly. No amount of data resolves it. Interviews do.

Two teams can look at the same cohort grid, reach opposite conclusions, and both fit the data perfectly — to the last decimal place, with zero residual. This is not a sampling problem, a power problem or a modelling-skill problem. It is a structural property of tenure, calendar time and signup vintage, and no quantity of additional data will fix it, because the three quantities are locked together by an exact equation.

The practical consequence is uncomfortable and worth stating plainly: if your analysis reports how much of a retention decline is due to tenure versus product changes versus cohort quality, that split came from an assumption somebody made, not from the data. Usually nobody in the room knows which assumption, because a piece of software chose it.

The demonstration

Here is a cohort grid. Rows are tenure in quarters, columns are the calendar quarter of measurement, cells are an engagement score.

Q1Q2Q3
0 quarters old60.060.060.0
1 quarter old54.054.054.0
2 quarters old48.048.048.0

Story A — it is tenure. Engagement falls 6 points for every quarter a customer has been with you. The calendar has no effect. Cohort quality is identical throughout.

Conclusion: we have a lifecycle problem. Fund habit formation and month-four re-engagement.

Story B — it is everything except tenure. Tenure has no effect at all. Instead the product got 6 points worse every calendar quarter, while each successive signup cohort arrived 6 points better than the one before it.

Conclusion: we have a product regression and a marketing team that is doing unusually well. Stop the regression; do not touch onboarding.

EffectStory AStory B
Tenure (per quarter)-6.00.0
Calendar (per quarter)0.0-6.0
Vintage (per later cohort)0.0+6.0

Now check them against the grid. Story B's prediction for a two-quarter-old account in Q3: base 60, tenure 0, calendar -12, and that account signed up in Q1 so it is the middle vintage with effect 0. Total 48. The grid says 48. Do this for all nine cells and the maximum discrepancy between Story A and Story B is exactly 0.0000000000.

Two irreconcilable business conclusions. Identical fit. And these are only two members of an infinite family: pick any number d, add d per quarter to the tenure effects, subtract d per quarter from the calendar effects, add d per cohort to the vintage effects, and every predicted cell is unchanged.

Why this happens, in one line

A customer's signup vintage is the date you measured them minus how long they have been with you. Ryder stated it in the founding paper: "If t is the time of occurrence and a is the age at that time, then the observations for age a, time t, apply (approximately) to the cohort born in year t-a."

So vintage = calendar - tenure, exactly, by definition, always. The three variables are not merely correlated — correlation could be broken with more data or a better sample. They are perfectly collinear by construction, which means the model has one more parameter than the data can ever pin down. Add a linear trend to one effect and you can always compensate in the other two.

The methodological literature is unusually blunt about this. Reviewing the field in the Annals of Human Biology (2020, volume 47, pages 208-217), Andrew Bell writes that "exact collinearity between these three (Age = Year - Birth Year) leads to difficulty estimating these effects," that "it is thus impossible to estimate linear components of these effects without strong assumptions about at least one of these," and — the sentence to put in front of anyone selling you an APC model — that "attempts to 'solve' this identification problem without strong assumptions are, in fact, making hidden unintended assumptions." His conclusion acknowledges "there is a 'line of solutions' of possible combinations of APC effects, and not a single answer that can be estimated empirically," and states flatly that "mechanical solutions to the identification problem do not work."

Tu and colleagues put the same point in regression terms in Epidemiology (2012, volume 23, pages 583-593): "as these 3 variables are perfectly collinear by definition, regression coefficients in a general linear model are not unique."

The methods that claim to solve it

Because the problem is old and the demand for an answer is high, a series of methods have been proposed that appear to return all three effects. The most widely used in recent years is the intrinsic estimator. It does return a unique answer — but the uniqueness comes from a constraint it imposes, not from the data.

Luo's assessment in Demography (volume 50, issue 6, pages 1945-1967) found that the intrinsic estimator "implicitly assumes a constraint on the linear age, period, and cohort effects." That constraint "not only depends on the number of age, period, and cohort categories but also has nontrivial implications for estimation" - and the verdict is unambiguous: "because this assumption is extremely difficult, if not impossible, to verify in empirical research, IE cannot and should not be used to estimate age, period, and cohort effects." The exchange that followed in the same journal is worth knowing about, and one line from it generalises to every method in the family: all APC models "provide just one possible solution from the infinite number of solutions" (Masters and colleagues, Demography, 2014, page 2066).

The pattern is consistent. Every technique that returns three clean numbers has smuggled in a fourth input. The techniques differ in how visible the smuggling is:

ApproachWhat it actually assumesIs the assumption visible?
Drop one effect from the modelThat effect is exactly zeroYes — the most honest option
Constrain two adjacent categories to be equalThose two groups differ only by chanceYes, if you state which two
Intrinsic estimatorA constraint determined by the number of categories in your tableNo — and it changes if you re-bin the data
Proxy one clock with a covariateThe proxy captures that clock and nothing elsePartly

The one that should worry you is the third row. A constraint that depends on how many tenure buckets you happened to create is a constraint nobody chose and nobody can defend, and re-binning quarterly data into halves changes the answer while leaving the underlying reality untouched.

What this is not

This is a different failure from two neighbours that sound similar, and the distinction determines the remedy.

It is not that there is no true answer. For some quantities — a stated willingness to pay, an attitude rating — the number genuinely does not exist independently of the question that produced it, so two different wordings yield two correct and incompatible answers. That is the subject of split-ballot experiments. The APC case is the opposite: there is a fact of the matter. Your product either did degrade in Q3 or it did not. The true values exist and are perfectly well defined — they are simply not recoverable from this data by any method.

It is not measurement error. More respondents, cleaner instrumentation and better sampling all improve your estimates of the grid cells. They do nothing whatsoever to the identification problem, because the problem is in the design matrix, not in the cells. A grid measured with infinite precision has exactly the same infinite family of solutions.

That is what makes this a genuinely different class of problem from most analysis failures: it survives every fix that normally works.

What to do instead

The literature's own recommendation is not "give up." It is to stop pretending the constraint came from the data, and to source it from somewhere defensible instead. Bell's recommendations are to consider the non-linearities around the linear effects and to state strong and explicit theory-based assumptions.

In practice, four things work:

1. Report the non-linear features, which are identified. The linear trends are unrecoverable, but departures from them are not. A sharp one-quarter dip that hits every tenure band simultaneously is a real, identifiable period effect. A single cohort that sits below its neighbours is a real, identifiable vintage effect. Say what the data can support: "there is a distinct Q3 shock affecting all tenures" is defensible; "42% of the decline is tenure-driven" is not.

2. Set one effect from outside the data. If you can establish the size of one clock independently, the other two become identified. This is the only genuine solution, and it requires evidence from outside the grid.

3. Publish the constraint as a sentence, not a setting. "We assume the product did not change over this window" is an assumption a stakeholder can dispute — which is the point. "We used the intrinsic estimator" is not.

4. Show the range. Because the solutions form a line, you can compute what the answer would be under several plausible constraints and report the interval. A finding that survives every reasonable constraint is trustworthy. A finding that flips between them was never a finding.

The move that actually resolves it

Step 2 above is where this stops being a statistics problem and becomes a research problem — and where the deadlock breaks completely.

The identification problem is a property of aggregate data. Every row of your grid is a count of anonymous accounts, and an account cannot tell you which clock moved it. A person can. A customer occupies exactly one cell, but they carry a memory that spans all three axes, and the three explanations produce completely different sentences:

  • "It changed in the spring — it used to sync automatically and then it stopped." → a period effect, with a date you can check against a release log.
  • "Honestly I used it constantly for a month and then I'd got what I needed from it." → a tenure effect.
  • "I signed up for the free migration offer and I never really got it working the way I expected." → a cohort effect, naming its own campaign.

No respondent needs to understand collinearity. They just need to be asked when, and what changed. The untestable constraint becomes a testable question the moment you talk to somebody — and this is the rare case where qualitative work is not a supplement to the quantitative analysis but the only thing that can complete it.

How Koji makes the constraint an empirical question

The reason teams reach for a model instead is cost. Doing this properly means interviewing several cohorts at matched tenure, repeatedly, and asking a question — "did the product change for you, and when?" — that only pays off when asked across the whole grid. At traditional interview economics that is a quarter of work for one parameter, so the model wins by default.

Platforms like Koji invert that trade:

  • Interview across the grid, not at one point. Import cohort lists and run one guide against all of them. AI-moderated interviews in voice or text need no moderator and no scheduling, so covering five cohorts at matched tenure is a days-long exercise.
  • Recover the date, not just the complaint. Koji's AI asks follow-up questions automatically, so "it got worse" becomes "when did you first notice?" — turning an unusable sentiment into a dated, checkable period effect.
  • Separate the clocks in the instrument itself. Ask yes_no whether the product has changed for them; single_choice for signup vintage and acquisition offer; scale for current engagement; ranking for what drives their usage now versus at signup; multiple_choice for which changes they noticed; open_ended for the account in their own words. Six question types, three clocks, one study — see the structured questions guide.
  • Build the evidence prospectively. A standing study that captures each cohort at matched tenure as it arrives means the external constraint is being accumulated continuously, rather than reconstructed from memory after the argument starts.
  • Report both. The grid gives you the identified non-linear features; the interviews give you the linear constraint. Together they are identified. Neither is on its own.

This is the strongest argument in this entire cluster for research over analytics, and it is not a matter of taste. The analytics cannot answer the question — provably, mathematically, regardless of scale. Ten million more rows leave the answer exactly as undetermined as it is today. Forty conversations resolve it.

A working checklist

  • Write out the grid and confirm that vintage equals calendar minus tenure in your own units.
  • Before accepting any three-way attribution, ask which constraint produced it.
  • Refuse attributions from methods whose constraint depends on your bin count.
  • Report identified non-linear features; label linear splits as assumption-dependent.
  • Compute the answer under at least two plausible constraints and publish the range.
  • Interview matched-tenure cohorts to source the constraint from outside the data.
  • Record acquisition vintage at signup so the third axis exists at all.

Frequently asked questions

What is the age-period-cohort identification problem?

It is the fact that age, period and cohort are exactly linearly dependent — cohort equals period minus age — so a linear model containing all three has infinitely many solutions that fit the data identically. Any estimate of how a trend divides among the three comes from a constraint imposed by the analyst or the software, not from the data. In product terms: tenure, calendar date and signup vintage cannot all three be estimated from a cohort grid.

Can more data solve the identification problem?

No. This is the property that distinguishes it from almost every other analysis problem. More respondents, longer time series and finer measurement all improve your estimates of the individual cells, but the collinearity is in the structure of the design rather than in the data, so the family of equally good solutions remains infinite no matter how much you collect.

Are age-period-cohort models useless then?

Not useless, but limited in a specific way. They can identify non-linear features — a one-off shock affecting all tenures, or a single cohort that sits out of line with its neighbours — and those findings are trustworthy. What they cannot identify is the linear trend in each effect. Use them for the bumps, not for the slopes, and treat any percentage attribution across the three as assumption-dependent.

What is wrong with the intrinsic estimator?

It returns a unique answer by imposing a constraint that depends on the number of age, period and cohort categories in the table. Luo's 2013 assessment in Demography found that this assumption is extremely difficult, if not impossible, to verify empirically. The practical tell is that re-binning your data — quarters into halves, say — changes the estimated effects even though nothing about the underlying reality has changed.

How do interviews solve what statistics cannot?

Because the deadlock is a property of aggregate data, not of reality. An anonymous row in a grid cannot report which clock moved it, but a customer can: they know whether the product changed, roughly when, whether they simply exhausted their use for it, and what they were promised at signup. That testimony provides the external information that identifies the model. It is one of the few situations where qualitative work is not complementary to the analysis but strictly necessary to complete it.

What should I do if a stakeholder demands a three-way split anyway?

Give them the split together with the sentence that produced it — for example, "assuming the product did not materially change during this window, 6 points per quarter is tenure." Stakeholders can argue with a sentence, and often will, which surfaces the real disagreement immediately. Then show how the answer changes under one or two alternative assumptions. If the conclusion is stable across all of them, you have something; if it flips, you have learned that the question needs fieldwork rather than more analysis.

Related Resources

Related Articles

Cohort Analysis: How to Read Retention and Find the "Why" (2026)

Cohort analysis groups users by a shared starting point and tracks their behavior over time, revealing retention patterns that aggregate metrics hide. This guide explains how to build and read cohort tables, interpret the retention curve, and pair the numbers with qualitative research to explain them.

Competing Risks: Why Your Retention Curve Overstates the Churn You Care About (2026)

Your retention curve treats acquisitions, downgrades and payment failures as if those accounts were still at risk of cancelling. That inflates the number. Here is the correction, the size of the error, and the interview that produces the missing field.

Same Data, Different Answers: The Many-Analysts Problem in Product Research

When 73 teams analyzed identical data to test one hypothesis, over 95 percent of the variance in their results was unexplained. Your analysis is one draw from a distribution you never see.

Split-Ballot Experiments: How Much of Your Number Is the Question?

Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.