{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-26T10:21:19.533Z"},"content":[{"type":"documentation","id":"177e2889-e27e-4810-b57f-5634c6b45108","slug":"metric-hyperstability-effort-adjusted-feedback","title":"Why Your Satisfaction Score Stays Flat While Satisfaction Falls (2026)","url":"https://www.koji.so/docs/metric-hyperstability-effort-adjusted-feedback","summary":"Metrics of the form response-per-unit-of-outreach measure true sentiment multiplied by collection efficiency. Because efficiency rises as sentiment falls - the dissatisfied churn and stop answering - such indices are hyperstable and decline far more slowly than reality. Fisheries science documented this when catch-per-unit-effort rose as the northern cod collapsed to 1 percent of historic levels. This guide gives the power-law arithmetic, the diagnostic of a flat score over a falling response rate, and the fishery-independent survey as the fix.","content":"In 1992 Canada closed the northern cod fishery. The stock had fallen to about one percent of its historic level. For years beforehand, the catch rates had looked fine.\n\nThat is not a story about bad data collection. The catch rate data were accurate. The problem was structural: catch per unit effort is the product of how many fish there are and how efficiently you can find them, and as the fish became scarcer they packed into denser aggregations that were easier to fish. One term fell while the other rose, and the index they multiply into barely moved.\n\nYour feedback metrics have the same shape, for the same reason.\n\n## The short answer\n\n- Any metric of the form response-per-unit-of-outreach is an effort-based index. It measures true sentiment multiplied by your current collection efficiency, not sentiment alone.\n- These indices are hyperstable: they decline more slowly than the thing they track. In fisheries this repeatedly hid collapses until the stock was gone.\n- Collection efficiency tends to rise exactly when sentiment falls, because unhappy users churn, disengage and stop answering, leaving a sample that is progressively easier and more flattering to survey. The two effects cancel.\n- The fix that worked in fisheries is a survey with a fixed design, run independently of where the signal is easy to get. In research terms: a proactive study with a fixed sample frame, not inbound volume.\n\n## What an effort-based index is\n\nFisheries science defines the quantity precisely. Catch per unit effort, as Wikipedia puts it, \"is an indirect measure of the abundance of a target species.\" The inferential step managers rely on is that \"Changes in the catch per unit effort are inferred to signify changes to the target species true abundance,\" and the textbook reading is that \"A decreasing CPUE indicates overexploitation, while an unchanging CPUE indicates sustainable harvesting.\"\n\nThat last sentence is the trap, and the same source flags the reason: \"CPUE is also often nonlinearly related to abundance, making interpretation more difficult.\"\n\nNow list the metrics a product team watches that have this exact form. Support tickets per thousand active users. Complaints per release. Negative reviews per month. Survey score per campaign sent. Every one of them is a catch divided by an effort, and every one inherits the nonlinearity.\n\n## Hyperstability, and the cod that got easier to catch as they vanished\n\nThe failure mode has a name. A relationship is hyperstable when the index stays high while true abundance declines, so the index systematically overstates what remains. The canonical demonstration is Rose and Kulka 1999, whose title states the finding outright: \"Hyperaggregation of fish and fisheries: how catch-per-unit-effort increased as the northern cod (Gadus morhua) declined\" (Canadian Journal of Fisheries and Aquatic Sciences, 1999).\n\nCatch per unit effort increased. The stock was collapsing. Both are true, and the mechanism is that the fish concentrated and the fleet concentrated with them.\n\nThe outcome is documented plainly. Wikipedia records that \"In 1992, Northern cod populations fell to 1% of historic levels, in large part from decades of overfishing,\" and that across stocks spawning biomass had decreased by at least 75% in all stocks, by 90% in three of the six stocks, and by 99% for northern cod, once the largest cod fishery in the world. The human cost: \"Approximately 37,000 fishermen and fish plant workers lost their jobs by the collapse of the cod fisheries.\"\n\nThere is also a sentence in that record that belongs on the wall of every analytics team: \"The previous increases in catches had been wrongly thought to be caused by 'the stock growing' but were really caused by new technology such as trawlers.\" A rising number was read as a healthier population when it was actually a more efficient instrument.\n\nThe problem is not historical. Charbonneau and colleagues asked in 2025 whether catch-per-unit-effort data are masking the magnitude of steelhead declines in an inland recreational fishery, and found a hyperstable relationship (Transactions of the American Fisheries Society, 2025;154(4):339-351). They report that \"when the population declined by 50%, CPUE decreased by only 40%\", which on its own already understates the loss by a fifth.\n\n## The arithmetic, and how much it hides\n\nHyperstability is usually modelled as a power relationship: the index is proportional to abundance raised to some exponent beta, where beta below 1 means hyperstability and beta equal to 1 is the proportionality everyone assumes. Take beta equal to 0.7 as an illustration and work out what an index would report at each true level.\n\n| True share remaining | True decline | Index reads | Apparent decline | Overestimate |\n|---|---|---|---|---|\n| 1.00 | 0% | 1.000 | 0% | 1.00x |\n| 0.90 | 10% | 0.929 | 7% | 1.03x |\n| 0.70 | 30% | 0.779 | 22% | 1.11x |\n| 0.50 | 50% | 0.616 | 38% | 1.23x |\n| 0.25 | 75% | 0.379 | 62% | 1.52x |\n| 0.10 | 90% | 0.200 | 80% | 2.00x |\n| 0.01 | 99% | 0.040 | 96% | 3.98x |\n\nTwo things in that table are worth dwelling on.\n\nFirst, it agrees with the field data. At a true 50 percent decline the model predicts the index falls 38 percent; Charbonneau and colleagues measured a 40 percent fall for a 50 percent decline. An exponent chosen as a round illustration reproduces an independently published observation to within two points, which is reason to take the shape seriously.\n\nSecond, the error is smallest when you could still act and largest when it is too late. At a 10 percent decline the index is off by 3 percent, which no dashboard would ever flag. At a 90 percent decline it still reports twice as much as remains. This is the opposite of a useful early-warning system: the metric is most trustworthy exactly when there is nothing to warn about, and most flattering when the situation is dire. At the cod stock actual 99 percent decline, a beta of 0.7 index would report four times the truth.\n\n## Why your collection efficiency rises as sentiment falls\n\nThe fisheries mechanism is aggregation. The research mechanism is attrition, and it is if anything stronger.\n\nAs satisfaction falls, the people most dissatisfied do several things that raise your apparent numbers. They churn, removing themselves from the denominator of active users. They disengage, so they stop opening the emails that carry your surveys. They stop answering, because answering a survey is a cooperative act and they are no longer cooperative. What remains in your sample is progressively enriched for the people least likely to tell you anything is wrong.\n\nSo your response rate falls while your average score holds. That is the signature. A flat score with a falling response rate is not stability, it is an index whose two terms are moving in opposite directions - and the flatness is the arithmetic of the cancellation, not evidence about your product.\n\nThis is also why the standard reassurance is wrong. \"Our score has not moved\" is only informative if collection efficiency has not moved. Nobody checks.\n\nTherese Fessenden of Nielsen Norman Group makes the general limitation explicit for the most common such metric: \"NPS, like all quantitative metrics, tells you how the experience is perceived but not why,\" and \"When used by itself, NPS, like any subjective metric, is fairly limited, variable in different geographic and industrial contexts, and far from being a good summary of the overall user experience.\" The hyperstability argument sharpens that: used by itself, and without its effort term, it can be not merely limited but directionally wrong.\n\n## What fisheries science actually did about it\n\nManagers did not solve this by analysing catch rates harder. They built a second, independent instrument: the fishery-independent survey. A research vessel tows the same stations, with the same gear, for the same duration, every year, regardless of where the fish are easy to catch that season. It catches far fewer fish than the fleet and it is far more expensive per fish. It is also the series you can actually trust, because its effort is held constant by design rather than optimised by participants.\n\nThe translation is direct.\n\n- **Inbound feedback is your fishery-dependent index.** Tickets, reviews, unsolicited emails and always-on surveys are collected wherever collection is easy. Effort is not constant and nobody is measuring it.\n- **A study with a fixed sampling design is your research survey.** Same sample frame, same questions, same recruitment rule, same cadence, run whether or not people feel like talking to you.\n\nFour practical rules follow. Always publish the effort term next to the index, so a score never appears without its response rate and sample size. Watch the ratio rather than the level, because a stable score over a halving response rate is a red flag and not a green one. Sample the frame, not the respondents - draw from all eligible customers and chase non-responders, instead of reporting on whoever arrived. And keep one series methodologically frozen, since a frozen imperfect instrument beats a drifting refined one for trend detection.\n\n## The modern approach: make the fixed-design survey cheap\n\nThe reason teams lean on inbound feedback is not ignorance. It is that the fishery-independent equivalent has always been expensive. Running a properly framed study every quarter, with real recruitment of reluctant participants, used to mean weeks of scheduling and manual analysis, so teams substituted the cheap hyperstable index and hoped.\n\nThat trade is what AI-native research changes. Koji makes the fixed-design study the affordable option rather than the aspirational one:\n\n- **Constant effort by construction.** Because Koji AI-moderated interviews run on demand and in parallel rather than through a scheduling calendar, you can execute the same study, on the same frame, at the same cadence, without the effort term drifting to match whoever happened to be available.\n- **Structured questions hold the instrument still.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and keeping the closed items identical across waves is what makes two waves comparable at all. A scale item asked the same way each quarter is the research trawl towing the same station.\n- **Reaching the disengaged.** The people who stop answering surveys are precisely the ones holding your signal. Koji voice interviews and conversational AI interviews are a lower-friction ask than a long form, which recovers part of the sample that attrition would otherwise remove - and recovering non-responders is the single highest-value thing you can do to an effort-based index.\n- **The why, not just the level.** Where legacy survey tools like SurveyMonkey return a score and leave you to guess, Koji automatic thematic analysis explains what moved, and real-time reporting surfaces it during the wave. A score that drops 3 points is ambiguous; a score that drops 3 points with a named, quoted reason is actionable.\n- **Customisable AI consultants keep the frame honest.** You can hold the interview objective fixed across waves while still probing adaptively within it, so comparability and depth stop being a trade-off.\n\nYou do not need a PhD in survey methodology to apply the lesson. You need to stop reading an effort-based index as if effort were constant, and to own one series where it genuinely is.\n\n## Frequently asked questions\n\n### What is hyperstability in a metric?\n\nHyperstability is when an indicator declines more slowly than the quantity it is supposed to track, so it systematically overstates what remains. It arises whenever the indicator is a ratio of outcome to effort and the efficiency of that effort rises as the underlying quantity falls. Fisheries science identified it in catch-per-unit-effort data, where it repeatedly concealed stock collapses.\n\n### Which of my metrics are effort-based indices?\n\nAnything expressed as a count per unit of outreach or activity: support tickets per thousand active users, complaints per release, negative reviews per month, or a survey score computed from whoever responded. If the denominator reflects how hard you tried or how many people chose to engage, rather than a fixed population, the metric carries an effort term.\n\n### How can I tell if my satisfaction score is hyperstable?\n\nPlot the response rate on the same chart as the score. If the score is flat or improving while the response rate declines, treat the flatness as suspect, because the two terms are moving in opposite directions. Then compare against a fixed-design measurement on the full eligible population; a gap between the two is the size of your hyperstability.\n\n### Is this the same as survivorship bias?\n\nThey are related but distinct. Survivorship bias is about which people reach you at all. Hyperstability is about the arithmetic of an index built from two moving terms, where collection efficiency rises as sentiment falls and cancels the decline. Survivorship explains part of why efficiency rises; hyperstability describes what that does to the trend line.\n\n### Should I stop tracking NPS or CSAT?\n\nNo, but stop reading them as unbiased levels. Publish the response rate and sample size alongside every score, hold the question wording and sampling rule constant across waves, and add one fixed-design study drawn from the whole eligible population. Use the always-on score for speed and the fixed-design study for direction.\n\n### How does Koji help detect a collapsing signal?\n\nKoji turns the fixed-design study into something you can run every cycle: AI-moderated and voice interviews execute in parallel without scheduling, structured questions keep the instrument identical across waves, and lower-friction formats recover non-responders who would otherwise drop out of the sample. Automatic thematic analysis and real-time reporting then tell you why a number moved, not just that it did.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and holding an instrument constant across waves\n- [The Moving Average Is Hiding the Week That Mattered](/docs/moving-average-hides-events-research) - a second mechanism that flattens a real signal\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - why the disengaged leave your sample in the first place\n- [Sampling Bias: Types, Examples, and How to Avoid It](/docs/sampling-bias-research) - the broader family this belongs to\n- [Is 4.1 Good? Internal Benchmarks and Percentile Norms](/docs/internal-benchmarks-percentile-norms) - comparing a score against something other than itself\n- [Why Your Newest Cohort Always Looks Best](/docs/reporting-lag-immature-cohort-feedback) - the companion timing artefact in cohort data\n- [Convenience Sampling: When Fast and Cheap Is the Right Call](/docs/convenience-sampling-guide) - when an easy sample is and is not acceptable","category":"Research Methods","lastModified":"2026-09-26T03:35:12.226068+00:00","metaTitle":"Metric Hyperstability: Flat Scores Hiding Real Declines (2026)","metaDescription":"Inbound feedback is a catch-per-unit-effort index: it stays flat while satisfaction collapses, because your sampling re-targets the happy.","keywords":["metric hyperstability","catch per unit effort","satisfaction score flat","effort adjusted metric","response rate bias","NPS misleading","fixed design survey"],"aiSummary":"Metrics of the form response-per-unit-of-outreach measure true sentiment multiplied by collection efficiency. Because efficiency rises as sentiment falls - the dissatisfied churn and stop answering - such indices are hyperstable and decline far more slowly than reality. Fisheries science documented this when catch-per-unit-effort rose as the northern cod collapsed to 1 percent of historic levels. This guide gives the power-law arithmetic, the diagnostic of a flat score over a falling response rate, and the fishery-independent survey as the fix.","aiPrerequisites":["Familiarity with a satisfaction or feedback metric such as NPS or CSAT","Basic understanding of response rates"],"aiLearningOutcomes":["Identify which of your metrics are effort-based indices","Explain hyperstability and why it hides declines","Diagnose it from a flat score against a falling response rate","Distinguish it from survivorship bias","Design a fixed-design study that holds effort constant"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}