{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T17:52:31.832Z"},"content":[{"type":"documentation","id":"dfa9dc15-2690-4f4c-9111-62d2d3e707d3","slug":"ergodicity-average-user-trajectory","title":"Why 100 Customers Once Is Not One Customer 100 Times (2026)","url":"https://www.koji.so/docs/ergodicity-average-user-trajectory","summary":"Averaging across customers at one moment (the ensemble average) and following one customer over time (the time average) converge only if the process is ergodic, which requires that all individuals follow the same process and that it does not change over time. Human behaviour rarely satisfies either. Fisher, Medaglia and Jeronimus (PNAS, 2018) found within-individual variance two to four times larger than between-individual variance across six studies with a repeated-measure design. A worked steady-state panel shows the failure exactly: four customers with tenures of 1 to 4 quarters score 9, 8, 7, 6, so the cross-sectional average is 7.5 with slope zero every quarter, while each individual declines 1 point per quarter. Levels agree, slopes do not, and larger samples do not help. Robinson (1950) found illiteracy and nativity correlated minus 0.53 at state level and plus 0.12 individually, an outright sign reversal.","content":"**Short answer:** averaging across many customers at one moment and following one customer across many moments answer different questions, and they agree only if your process is ergodic. Human behaviour almost never is. A 2018 study in the Proceedings of the National Academy of Sciences analysed six studies with a repeated-measure design and found that \"the variance around the expected value was two to four times larger within individuals than within groups\", concluding that \"Only for ergodic processes will inferences based on group-level data generalize to individual experience or behavior.\"\n\nThe practical consequence is sharper than it sounds. Your satisfaction dashboard can report the correct average, every quarter, while every single customer in it is getting less satisfied.\n\n## The assumption nobody states\n\nTwo averages get used interchangeably in product research and they are not the same object.\n\n- The **ensemble average** is what you get by measuring many customers at one point in time. This is a survey, an NPS wave, a cross-sectional study.\n- The **time average** is what you get by measuring one customer repeatedly over a long stretch. This is a diary study, a cohort trace, one account's history.\n\nErgodicity is the technical name for the condition under which those two converge. It requires two things that are easy to state and rarely true of people: every individual must follow the same process, and that process must not change over time.\n\nNeither holds for customers. Different accounts genuinely behave differently, not just noisily, and their behaviour changes as they learn your product, change jobs, grow, or start evaluating a competitor. The PNAS authors put the conclusion plainly: \"Because human social and psychological processes typically have an individually variable and time-varying nature, they are unlikely to be ergodic.\"\n\nOnce that is on the table, a whole category of everyday research claim becomes suspect. *Our average user logs in 3 times a week* does not mean a typical user logs in 3 times a week. It means the population, sliced at one instant, averages 3. Those are different sentences and only one of them is supported.\n\n## A panel whose average is right and whose trend is backwards\n\nHere is the cleanest demonstration, small enough to check by hand.\n\nSuppose satisfaction declines by exactly 1 point per quarter for every customer you have, starting at 9 when they onboard. Suppose also that you are in a steady state: you run a panel of four active customers, one churns each quarter after four quarters, and one new customer joins each quarter.\n\nAt any moment your four active customers have tenures of 1, 2, 3 and 4 quarters, which means their satisfaction scores are 9, 8, 7 and 6.\n\n| View | Data | Average | Slope |\n| --- | --- | --- | --- |\n| Cross-sectional, any quarter | 9, 8, 7, 6 | 7.5 | 0 per quarter |\n| One customer's own history | 9, 8, 7, 6 | 7.5 | minus 1 per quarter |\n\nLook carefully at what agrees and what does not. The ensemble average is 7.5. The time average of any individual customer is also 7.5. The levels match perfectly, so a mean-based dashboard looks completely healthy and stable.\n\nThe slopes do not match at all. Measured across customers, satisfaction is flat forever. Measured within any customer, satisfaction falls 3 points over their lifetime. A quarterly tracking study on this population would report *no change in satisfaction* for as long as you cared to run it, while describing a product that disappoints everyone who uses it for a year.\n\nThis is not a sampling problem and a bigger sample does not touch it. Interview 4,000 customers instead of 4 and the cross-sectional slope is still zero, measured with beautiful precision. The information about direction is not in the data, because the design threw it away.\n\n## The ecological fallacy is the same error, one level up\n\nSociology hit this before product teams did, and the canonical example is worth knowing because the sign actually flips rather than merely flattening.\n\nW. S. Robinson's 1950 analysis of United States census data compared illiteracy and immigrant status. At the state level the two were associated with a correlation of minus 0.53, which reads as immigrants being more literate. At the individual level the correlation between illiteracy and being foreign-born was plus 0.12, the opposite direction. The aggregate relationship was real, and it was produced by where immigrants settled rather than by who they were: they moved to states whose native populations were more literate. A 2011 re-examination corrected the state-level figure to minus 0.46, which changes the number and not the lesson.\n\nTwo correlations, opposite signs, same underlying people. Cross-level inference is not a rounding error.\n\nThe product version happens constantly. Accounts on your enterprise plan report higher satisfaction than accounts on your starter plan, so you conclude that upgrading raises satisfaction. What you may actually have measured is which kinds of company buy which plan. The only design that separates those is following the same accounts through an upgrade.\n\n## What the evidence says about the size of the problem\n\nThe strongest empirical statement available is the PNAS finding quoted at the top: within-individual variance ran two to four times larger than between-individual variance across the studies examined. Sit with the direction of that ratio. The variation inside one person over time is larger than the variation between people, which means the group average is describing the smaller and less consequential source of difference.\n\nThat is a measured claim about psychological and health processes, and the authors frame the generalisability failure as a threat to human subjects research rather than as a curiosity. Whether the same ratio holds for your product metrics is an empirical question about your product, not something to assume in either direction. But the burden of proof runs the way most teams assume it does not: an average across customers is the thing that requires justification before it can be read as a statement about a customer.\n\n## When a cross-sectional average is fine\n\nThis is not an argument that surveys are worthless. Ensemble averages answer ensemble questions correctly, and plenty of real decisions are ensemble decisions.\n\nUse a cross-sectional average when:\n\n- The question is genuinely about the population right now. *What share of our base is affected by this bug* is an ensemble question and a survey answers it properly.\n- You are sizing a market, a segment or a support load.\n- You are comparing two groups measured the same way at the same time, where the shared measurement conditions cancel.\n\nInsist on within-person data when:\n\n- The claim contains a direction: improving, declining, learning, adopting, churning.\n- The claim is about what happens to a single customer, which includes almost every onboarding and lifecycle claim.\n- You are attributing a change to something you did.\n- The population composition is itself changing, which is when cross-sectional trends are least trustworthy.\n\nThe honest version of the flat-satisfaction dashboard above is two numbers reported side by side: the population level, and the within-customer slope. They are not in conflict. They are answers to different questions, and a report that gives only the first is not wrong so much as silent on the thing the reader wants to know. The [cross-sectional versus longitudinal](/docs/cross-sectional-vs-longitudinal-study) trade-off covers how to choose a design; this is about what your chosen design can and cannot say afterwards.\n\n## How Koji handles this\n\nWithin-person measurement has always been the expensive design, which is why teams default to the cross-sectional one and then read it as though it were longitudinal. Koji changes the cost structure rather than the statistics.\n\n- **Repeat interviews with the same participants are cheap to run.** Because Koji's AI consultant conducts interviews concurrently and on the participant's schedule, a wave 2 and wave 3 with the same accounts is a configuration change, not a new research project. That is what makes a time average obtainable at all.\n- **Structured questions give you comparable values across waves.** Koji supports open_ended, scale, single_choice, multiple_choice, ranking and yes_no. The scale and ranking types produce the numeric series you need to compute a within-person slope, while open_ended captures why it moved.\n- **Stable question identifiers keep answers traceable across studies**, so wave 3 answers line up against wave 1 answers per participant instead of being compared as aggregate summaries.\n- **Automatic thematic analysis runs per participant as well as per study**, so a theme that is intensifying for individual customers does not get flattened into a stable overall percentage.\n- **Real-time reporting shows distributions rather than only means**, which is the first line of defence: a flat mean with a widening spread is visible immediately instead of after somebody thinks to check.\n- **Customizable AI consultants can carry a participant's earlier answers into the next conversation**, which is how you get a genuine trajectory rather than four unrelated snapshots of the same person.\n\nThe point is not that Koji makes averages honest. It is that the within-person design, the one that can actually support a directional claim, stops being the expensive option.\n\n## Common mistakes\n\n- **Reading an average as a typical individual.** The average customer is a construct; no customer is obliged to resemble it, and in a heterogeneous population most do not.\n- **Concluding a trend from repeated cross-sections.** Repeating a survey quarterly gives you a series of ensemble averages. If your population composition is drifting, that series can move in the opposite direction from every individual in it.\n- **Treating a bigger sample as a fix.** Sample size reduces sampling error. It does nothing about the cross-level inference problem, and a very large study will state the wrong conclusion with narrow confidence intervals.\n- **Comparing cohorts and calling it longitudinal.** Different cohorts are different people. That is a between-person comparison wearing a time axis, and [survivorship](/docs/survivorship-bias-customer-research) makes it worse, because the customers who stayed are not the customers who left.\n- **Assuming stationarity because the average is stable.** A stable mean is consistent with every individual changing, as the worked table shows. Stability of the aggregate is not evidence of stability underneath it.\n\n## Frequently asked questions\n\n### Does this mean cross-sectional research is useless?\n\nNot at all. It means cross-sectional research answers population-level questions, and you should only ask it those. What proportion of our customers hit this problem is answered correctly by a survey. Whether that problem gets worse the longer someone uses our product is not, at any sample size. The failure is not the method, it is reading one kind of answer as though it were the other.\n\n### How do I tell whether my process is ergodic?\n\nYou test it rather than assume it, and the test needs within-person data, which is the catch. Ergodicity requires that all individuals follow the same process and that the process does not change over time. In practice, measure a subset of customers repeatedly, compute each person's own average and slope, then compare the spread of those individual slopes against your cross-sectional estimate. If individual slopes vary widely, or differ in sign from the aggregate, you have your answer and you should stop generalising from the average.\n\n### Is this just survivorship bias?\n\nThey are related but distinct, and both can be present. Survivorship bias is about who is missing from your data because they left. Non-ergodicity is a problem even with zero attrition: the worked example in this article has a complete, balanced panel and the cross-sectional slope is still zero while every individual declines. Selective churn then makes it worse, because the customers who remain are the ones the product suited.\n\n### How many time points do I need per person?\n\nThree is the practical minimum, because two points give you a difference with no way to distinguish a trend from noise, and any single pair of measurements can be an artefact of when you happened to ask. Three or more lets you see whether a direction persists. What matters more than the count is the spacing: intervals should be short enough that change is visible and long enough that real change has had time to happen.\n\n### Does a larger sample fix the problem?\n\nNo, and this is the most expensive misunderstanding in this article. A larger sample shrinks the uncertainty around your ensemble average, so a cross-sectional design with 4,000 participants gives you a very precise estimate of a quantity that still does not describe any individual. Precision and relevance are independent properties. The fix is a different design, not more of the same one.\n\n### How does Koji support within-person measurement?\n\nRun the same study again with the same participants. Koji's AI consultant handles the interviews concurrently, so repeat waves cost roughly what the first one did, and structured scale and ranking questions produce comparable numeric answers across waves that you can difference per participant. The consultant can also be configured to reference a participant's earlier answers, which turns a set of snapshots into an actual trajectory and lets Koji's reporting show individual slopes alongside the population average.\n\n## Related Resources\n\n- [Structured Questions: The Complete Guide](/docs/structured-questions-guide) - the scale and ranking types that make a within-person slope computable\n- [Cross-Sectional vs Longitudinal Study](/docs/cross-sectional-vs-longitudinal-study) - choosing between the two designs in the first place\n- [Longitudinal Research](/docs/longitudinal-research-guide) - how to run repeated measurement without losing your panel\n- [Diary Studies](/docs/diary-study-guide) - the highest-resolution way to collect a time average\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - why the customers who remain are not a fair sample\n- [Average Customer Lifetime Is Not 1 Divided by Your Churn Rate](/docs/customer-lifetime-mtbf-ltv-formula) - the same cross-level error inside a single formula\n","category":"Research Methods","lastModified":"2026-09-28T04:00:44.319912+00:00","metaTitle":"Ergodicity in Customer Research: Why the Average User Is Not a User (2026)","metaDescription":"Cross-sectional averages cannot support directional claims. Within-person variance runs two to four times larger than between-person.","keywords":["ergodicity research","average user fallacy","time average vs ensemble average","within-person vs between-person","ecological fallacy","cross-level inference"],"aiSummary":"Averaging across customers at one moment (the ensemble average) and following one customer over time (the time average) converge only if the process is ergodic, which requires that all individuals follow the same process and that it does not change over time. Human behaviour rarely satisfies either. Fisher, Medaglia and Jeronimus (PNAS, 2018) found within-individual variance two to four times larger than between-individual variance across six studies with a repeated-measure design. A worked steady-state panel shows the failure exactly: four customers with tenures of 1 to 4 quarters score 9, 8, 7, 6, so the cross-sectional average is 7.5 with slope zero every quarter, while each individual declines 1 point per quarter. Levels agree, slopes do not, and larger samples do not help. Robinson (1950) found illiteracy and nativity correlated minus 0.53 at state level and plus 0.12 individually, an outright sign reversal.","aiPrerequisites":["Basic understanding of averages and correlation","Familiarity with survey and cohort research designs"],"aiLearningOutcomes":["Distinguish an ensemble average from a time average","State the two conditions ergodicity requires","Explain why a larger sample cannot fix a cross-level inference error","Decide when a directional claim requires within-person data"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}