Back to docs
Research Methods

Why Your Satisfaction Score Stays Flat While Satisfaction Falls (2026)

Inbound feedback is a catch-per-unit-effort index, and those are famously hyperstable: they hold steady while the thing they measure collapses. The fisheries fix is a survey with a fixed design.

In 1992 Canada closed the northern cod fishery. The stock had fallen to about one percent of its historic level. For years beforehand, the catch rates had looked fine.

That is not a story about bad data collection. The catch rate data were accurate. The problem was structural: catch per unit effort is the product of how many fish there are and how efficiently you can find them, and as the fish became scarcer they packed into denser aggregations that were easier to fish. One term fell while the other rose, and the index they multiply into barely moved.

Your feedback metrics have the same shape, for the same reason.

The short answer

  • Any metric of the form response-per-unit-of-outreach is an effort-based index. It measures true sentiment multiplied by your current collection efficiency, not sentiment alone.
  • These indices are hyperstable: they decline more slowly than the thing they track. In fisheries this repeatedly hid collapses until the stock was gone.
  • Collection efficiency tends to rise exactly when sentiment falls, because unhappy users churn, disengage and stop answering, leaving a sample that is progressively easier and more flattering to survey. The two effects cancel.
  • The fix that worked in fisheries is a survey with a fixed design, run independently of where the signal is easy to get. In research terms: a proactive study with a fixed sample frame, not inbound volume.

What an effort-based index is

Fisheries science defines the quantity precisely. Catch per unit effort, as Wikipedia puts it, "is an indirect measure of the abundance of a target species." The inferential step managers rely on is that "Changes in the catch per unit effort are inferred to signify changes to the target species true abundance," and the textbook reading is that "A decreasing CPUE indicates overexploitation, while an unchanging CPUE indicates sustainable harvesting."

That last sentence is the trap, and the same source flags the reason: "CPUE is also often nonlinearly related to abundance, making interpretation more difficult."

Now list the metrics a product team watches that have this exact form. Support tickets per thousand active users. Complaints per release. Negative reviews per month. Survey score per campaign sent. Every one of them is a catch divided by an effort, and every one inherits the nonlinearity.

Hyperstability, and the cod that got easier to catch as they vanished

The failure mode has a name. A relationship is hyperstable when the index stays high while true abundance declines, so the index systematically overstates what remains. The canonical demonstration is Rose and Kulka 1999, whose title states the finding outright: "Hyperaggregation of fish and fisheries: how catch-per-unit-effort increased as the northern cod (Gadus morhua) declined" (Canadian Journal of Fisheries and Aquatic Sciences, 1999).

Catch per unit effort increased. The stock was collapsing. Both are true, and the mechanism is that the fish concentrated and the fleet concentrated with them.

The outcome is documented plainly. Wikipedia records that "In 1992, Northern cod populations fell to 1% of historic levels, in large part from decades of overfishing," and that across stocks spawning biomass had decreased by at least 75% in all stocks, by 90% in three of the six stocks, and by 99% for northern cod, once the largest cod fishery in the world. The human cost: "Approximately 37,000 fishermen and fish plant workers lost their jobs by the collapse of the cod fisheries."

There is also a sentence in that record that belongs on the wall of every analytics team: "The previous increases in catches had been wrongly thought to be caused by 'the stock growing' but were really caused by new technology such as trawlers." A rising number was read as a healthier population when it was actually a more efficient instrument.

The problem is not historical. Charbonneau and colleagues asked in 2025 whether catch-per-unit-effort data are masking the magnitude of steelhead declines in an inland recreational fishery, and found a hyperstable relationship (Transactions of the American Fisheries Society, 2025;154(4):339-351). They report that "when the population declined by 50%, CPUE decreased by only 40%", which on its own already understates the loss by a fifth.

The arithmetic, and how much it hides

Hyperstability is usually modelled as a power relationship: the index is proportional to abundance raised to some exponent beta, where beta below 1 means hyperstability and beta equal to 1 is the proportionality everyone assumes. Take beta equal to 0.7 as an illustration and work out what an index would report at each true level.

True share remainingTrue declineIndex readsApparent declineOverestimate
1.000%1.0000%1.00x
0.9010%0.9297%1.03x
0.7030%0.77922%1.11x
0.5050%0.61638%1.23x
0.2575%0.37962%1.52x
0.1090%0.20080%2.00x
0.0199%0.04096%3.98x

Two things in that table are worth dwelling on.

First, it agrees with the field data. At a true 50 percent decline the model predicts the index falls 38 percent; Charbonneau and colleagues measured a 40 percent fall for a 50 percent decline. An exponent chosen as a round illustration reproduces an independently published observation to within two points, which is reason to take the shape seriously.

Second, the error is smallest when you could still act and largest when it is too late. At a 10 percent decline the index is off by 3 percent, which no dashboard would ever flag. At a 90 percent decline it still reports twice as much as remains. This is the opposite of a useful early-warning system: the metric is most trustworthy exactly when there is nothing to warn about, and most flattering when the situation is dire. At the cod stock actual 99 percent decline, a beta of 0.7 index would report four times the truth.

Why your collection efficiency rises as sentiment falls

The fisheries mechanism is aggregation. The research mechanism is attrition, and it is if anything stronger.

As satisfaction falls, the people most dissatisfied do several things that raise your apparent numbers. They churn, removing themselves from the denominator of active users. They disengage, so they stop opening the emails that carry your surveys. They stop answering, because answering a survey is a cooperative act and they are no longer cooperative. What remains in your sample is progressively enriched for the people least likely to tell you anything is wrong.

So your response rate falls while your average score holds. That is the signature. A flat score with a falling response rate is not stability, it is an index whose two terms are moving in opposite directions - and the flatness is the arithmetic of the cancellation, not evidence about your product.

This is also why the standard reassurance is wrong. "Our score has not moved" is only informative if collection efficiency has not moved. Nobody checks.

Therese Fessenden of Nielsen Norman Group makes the general limitation explicit for the most common such metric: "NPS, like all quantitative metrics, tells you how the experience is perceived but not why," and "When used by itself, NPS, like any subjective metric, is fairly limited, variable in different geographic and industrial contexts, and far from being a good summary of the overall user experience." The hyperstability argument sharpens that: used by itself, and without its effort term, it can be not merely limited but directionally wrong.

What fisheries science actually did about it

Managers did not solve this by analysing catch rates harder. They built a second, independent instrument: the fishery-independent survey. A research vessel tows the same stations, with the same gear, for the same duration, every year, regardless of where the fish are easy to catch that season. It catches far fewer fish than the fleet and it is far more expensive per fish. It is also the series you can actually trust, because its effort is held constant by design rather than optimised by participants.

The translation is direct.

  • Inbound feedback is your fishery-dependent index. Tickets, reviews, unsolicited emails and always-on surveys are collected wherever collection is easy. Effort is not constant and nobody is measuring it.
  • A study with a fixed sampling design is your research survey. Same sample frame, same questions, same recruitment rule, same cadence, run whether or not people feel like talking to you.

Four practical rules follow. Always publish the effort term next to the index, so a score never appears without its response rate and sample size. Watch the ratio rather than the level, because a stable score over a halving response rate is a red flag and not a green one. Sample the frame, not the respondents - draw from all eligible customers and chase non-responders, instead of reporting on whoever arrived. And keep one series methodologically frozen, since a frozen imperfect instrument beats a drifting refined one for trend detection.

The modern approach: make the fixed-design survey cheap

The reason teams lean on inbound feedback is not ignorance. It is that the fishery-independent equivalent has always been expensive. Running a properly framed study every quarter, with real recruitment of reluctant participants, used to mean weeks of scheduling and manual analysis, so teams substituted the cheap hyperstable index and hoped.

That trade is what AI-native research changes. Koji makes the fixed-design study the affordable option rather than the aspirational one:

  • Constant effort by construction. Because Koji AI-moderated interviews run on demand and in parallel rather than through a scheduling calendar, you can execute the same study, on the same frame, at the same cadence, without the effort term drifting to match whoever happened to be available.
  • Structured questions hold the instrument still. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and keeping the closed items identical across waves is what makes two waves comparable at all. A scale item asked the same way each quarter is the research trawl towing the same station.
  • Reaching the disengaged. The people who stop answering surveys are precisely the ones holding your signal. Koji voice interviews and conversational AI interviews are a lower-friction ask than a long form, which recovers part of the sample that attrition would otherwise remove - and recovering non-responders is the single highest-value thing you can do to an effort-based index.
  • The why, not just the level. Where legacy survey tools like SurveyMonkey return a score and leave you to guess, Koji automatic thematic analysis explains what moved, and real-time reporting surfaces it during the wave. A score that drops 3 points is ambiguous; a score that drops 3 points with a named, quoted reason is actionable.
  • Customisable AI consultants keep the frame honest. You can hold the interview objective fixed across waves while still probing adaptively within it, so comparability and depth stop being a trade-off.

You do not need a PhD in survey methodology to apply the lesson. You need to stop reading an effort-based index as if effort were constant, and to own one series where it genuinely is.

Frequently asked questions

What is hyperstability in a metric?

Hyperstability is when an indicator declines more slowly than the quantity it is supposed to track, so it systematically overstates what remains. It arises whenever the indicator is a ratio of outcome to effort and the efficiency of that effort rises as the underlying quantity falls. Fisheries science identified it in catch-per-unit-effort data, where it repeatedly concealed stock collapses.

Which of my metrics are effort-based indices?

Anything expressed as a count per unit of outreach or activity: support tickets per thousand active users, complaints per release, negative reviews per month, or a survey score computed from whoever responded. If the denominator reflects how hard you tried or how many people chose to engage, rather than a fixed population, the metric carries an effort term.

How can I tell if my satisfaction score is hyperstable?

Plot the response rate on the same chart as the score. If the score is flat or improving while the response rate declines, treat the flatness as suspect, because the two terms are moving in opposite directions. Then compare against a fixed-design measurement on the full eligible population; a gap between the two is the size of your hyperstability.

Is this the same as survivorship bias?

They are related but distinct. Survivorship bias is about which people reach you at all. Hyperstability is about the arithmetic of an index built from two moving terms, where collection efficiency rises as sentiment falls and cancels the decline. Survivorship explains part of why efficiency rises; hyperstability describes what that does to the trend line.

Should I stop tracking NPS or CSAT?

No, but stop reading them as unbiased levels. Publish the response rate and sample size alongside every score, hold the question wording and sampling rule constant across waves, and add one fixed-design study drawn from the whole eligible population. Use the always-on score for speed and the fixed-design study for direction.

How does Koji help detect a collapsing signal?

Koji turns the fixed-design study into something you can run every cycle: AI-moderated and voice interviews execute in parallel without scheduling, structured questions keep the instrument identical across waves, and lower-friction formats recover non-responders who would otherwise drop out of the sample. Automatic thematic analysis and real-time reporting then tell you why a number moved, not just that it did.

Related Resources

Related Articles

Convenience Sampling: When Fast and Cheap Is the Right Call (2026)

A practical guide to convenience sampling — what it is, its advantages and hidden biases, when it is acceptable, how to reduce bias, and how AI-native research makes rigor almost as fast as convenience.

Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms

A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.

The Moving Average on Your Dashboard Is Hiding the Week That Mattered (2026)

Smoothing conserves the area under an event and destroys its height - and every alert threshold you own is a height. The arithmetic of what a rolling average deletes, how late it reports, and why it cannot give you a number for now.

Sampling Bias: Types, Examples, and How to Avoid It

Sampling bias is when some people in your population are systematically more likely to end up in your sample than others — quietly invalidating your findings. Learn the six main types, classic examples, and how to build a representative sample at scale.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.