Credibility Weighting: How Much of a Small Segment Score Should You Believe? (2026)
An eight-person segment scoring 4.6 against a 4.1 average is neither 4.6 nor unusable. Credibility weighting gives you the exact weight to apply, using a formula actuaries have relied on since 1918.
Your enterprise segment scored 4.6 out of 5. Eight people answered. The company-wide average is 4.1. The honest number to put in the deck is neither 4.6 nor we do not have enough data - it is a weighted blend of the two, and there is a century-old formula that tells you the weight.
That formula is credibility weighting, and it is the single most useful piece of arithmetic for anyone who reports numbers broken out by segment. It answers a question that sample-size rules of thumb cannot: not is this segment big enough, but given that it is this big, how much of its own number should I use?
The answer, stated first
Report a blend of the segment's own score and the overall average:
Estimate = Overall average + Z x (Segment score - Overall average)
where Z is between 0 and 1 and rises with the segment's sample size. The standard form is:
Z = n / (n + k)
n is the number of responses in the segment. k is a constant you estimate from your own data - it is the point at which you trust the segment and the overall average equally. When n equals k, Z is 0.5 and you split the difference.
For the 4.6 example, if k works out to 23 for your satisfaction metric, then Z = 8 / (8 + 23) = 0.26, and the reported estimate is 4.1 + 0.26 x (4.6 - 4.1) = 4.23. The segment is probably above average. It is almost certainly not at 4.6.
Where this comes from, and why it is not a hack
This is not a heuristic somebody invented for dashboards. It is the foundation of insurance ratemaking, where the identical problem appears: a single business has three years of claims, the class it belongs to has thirty thousand, and the premium has to reflect both.
L. H. Longley-Cook set out the elementary version for the Casualty Actuarial Society in 1962, in a report written because the existing literature was, in his words, "difficult to follow without a knowledge of the subject". He opens the section on the meaning of credibility with a line from Arthur L. Bailey to the effect that the basis for these credibility formulas had been a profound mystery to most people who encountered them - which remains a fair description of how segment-level numbers are produced in most companies today.
The logic he lays out is worth following, because it is entirely common sense. If the new data are extensive enough to stand alone, use them. If they are so thin as to be meaningless, use the overall figure. And in between, as Longley-Cook writes, the answer "must lie between" the two, which expressed mathematically means:
P = P0 (1 - Z) + P1 Z, or equivalently P = P0 + Z (P1 - P0)
where P0 is the overall figure, P1 is the segment's own, and Z "is called the credibility assigned to the new data."
The particular form Z = n / (n + k) he attributes to Albert W. Whitney, in a 1918 paper for the same society. The idea that you should partially believe a small sample is older than almost every research method your team uses.
The table that shows the shape
Longley-Cook derives an illustrative credibility curve by assuming the volume of data for full credibility is twice a reference volume. His published column and the resulting scaled credibility look like this:
| Volume of new data (% of full) | Indicated credibility | Scaled to 1.00 |
|---|---|---|
| 100% | 0.67 | 1.00 |
| 80% | 0.62 | 0.92 |
| 50% | 0.50 | 0.75 |
| 20% | 0.29 | 0.43 |
| 10% | 0.17 | 0.25 |
| 0% | 0 | 0 |
Every value in the middle column reproduces exactly from n/(n+k): at 50 per cent the new data equal the reference volume, so Z = 1/(1+1) = 0.50; at 20 per cent, 0.4/1.4 = 0.29; at 10 per cent, 0.2/1.2 = 0.17. The curve is steep at the bottom and flat at the top, which is the behaviour you want. Going from 8 respondents to 16 buys you a lot. Going from 800 to 816 buys you nothing.
k is not a universal number, and that is the whole point
The most common mistake is to look for a single threshold - you need 30 per segment - and apply it everywhere. The credibility constant is not a property of good practice. It is a property of your metric, and it has a definition:
k = (variation within segments) / (variation between segments)
In the standard normal model, k is the within-segment variance divided by the between-segment variance. That ratio is the entire question, and it explains why no universal threshold can exist:
- When your segments genuinely differ a lot from each other, between-segment variation is large, k is small, and even a tiny segment's own number is worth believing.
- When your segments barely differ, between-segment variation is small, k is large, and even a fairly big segment gets pulled hard toward the average - correctly, because there is not much real difference out there to find.
Concretely, for a 1-5 satisfaction scale with a within-segment standard deviation of 1.2:
| Between-segment SD | k | Z at n=8 | Z at n=50 |
|---|---|---|---|
| 0.40 (segments really differ) | 9.0 | 0.47 | 0.85 |
| 0.25 (segments barely differ) | 23.0 | 0.26 | 0.68 |
Same metric, same sample sizes, weights that differ by nearly a factor of two. A published benchmark for minimum segment size cannot know which of these you are in. Your own data can.
For a real worked example, Gelman and Price fitted a hierarchical model to county radon measurements and reported within- and between-county standard deviations of 1.0 and 0.7. That gives k = 1.0 / 0.49 = 2.04 - remarkably small, because counties genuinely differ in radon. In that setting a county with only 8 measurements still earns Z = 0.80. Radon is real and varies by place. Customer satisfaction across your B2B segments usually varies far less, which is exactly why your segment differences deserve more scepticism than a geologist's.
How to estimate k without a statistics team
You need two numbers from data you already have:
- Within-segment variation. Take the standard deviation of individual responses inside each segment, and average across segments. Square it.
- Between-segment variation. Take the segment means, compute their standard deviation, and square it. Then subtract the sampling noise you expect from finite segments - the average of (within-variance / n) across segments. What is left is the real between-segment variance.
Divide the first by the second. That is k. If step 2 comes out at or below zero, you have just learned something important: there is no detectable real variation between your segments at all, and every difference in your dashboard is noise. That is a finding, not a failure.
What this replaces
Most teams currently handle small segments with a cutoff: below some n, drop the row. That rule appears in good guidance for good reasons - a norm bank built from tiny waves really does become noise, and there are separate and legitimate reasons to suppress small cells when you publish, which is a privacy question rather than a precision one.
But a cutoff throws away information on one side of the line and pretends to certainty on the other. A segment of 29 is not worthless and a segment of 31 is not trustworthy. Credibility weighting replaces the cliff with a ramp. You keep every segment, and each one is reported with the weight it has earned.
It also changes what you argue about in the readout. Instead of can we even show this segment, the conversation becomes this segment is at 4.23 with a credibility of 0.26, so most of what you are looking at is the company average - if you want to act on it, we need about thirty more interviews. That is an answerable request.
The thing credibility weighting does not do
It does not tell you the segment is different. It gives you a better point estimate assuming your segments are all drawn from the same underlying population - the same kind of thing, differing only by chance and by a real but modest amount. When a segment is genuinely different in kind, the blend pulls it toward a mean it does not belong to, and it does so silently. That failure is serious enough to deserve its own treatment, and it is covered in when shrinkage hides the one segment that actually changed.
It also does not license ranking. Blending fixes each estimate individually and still leaves the ordering of your segments confounded with their sizes, in a way that surprises almost everyone - see why the top and bottom segments in your report are both the smallest ones.
How Koji makes this practical
Credibility weighting needs two things that traditional research tooling makes expensive: enough responses per segment that k is estimable at all, and per-respondent data rather than pre-aggregated toplines.
Koji's AI-moderated interviews change the first. Because the interview is conducted by AI rather than a scheduled human moderator, running 200 conversations across twelve segments costs roughly what running 20 costs in calendar time - there is no recruiter bottleneck, no scheduling, and no transcription queue. Segments that were previously represented by we spoke to three of them become segments with real n. That is the difference between a k you can estimate and a k you have to guess.
The second is where Koji's structured questions matter. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and the closed types produce clean per-respondent numeric values attached to each interview, not just a chart. A scale question gives you the individual ratings you need for the within-segment variance; the segment identity comes from your own metadata or from single_choice and multiple_choice screening questions asked in the same conversation. You can export the respondent-level table and compute k directly.
What Koji adds that a survey tool cannot is the why attached to each number. When a segment's credibility-weighted score sits meaningfully above the average, the open_ended responses and the AI's follow-up probes in that segment are right there, and you can read what those eight people actually said rather than inferring a story from a mean. A ranking question tells you what that segment prioritises differently; a yes_no question gives you a clean incidence to weight. The arithmetic tells you how much to believe the number. The transcripts tell you what the number is about.
Koji's automatic analysis also surfaces segment-level breakdowns without a separate analyst pass, which means the credibility question comes up while the study is still open and you can still fix it by collecting more interviews in the thin segments.
A working procedure
- Pull respondent-level data, not toplines.
- Compute within-segment variance and real between-segment variance; divide to get k.
- If between-segment variance is zero or negative, stop and report that segments do not differ.
- Compute Z = n/(n+k) for every segment.
- Report the blended estimate, and report Z next to it.
- For any segment where Z is below roughly 0.5, treat the number as provisional and say what n would fix it - the arithmetic in how much data a segment needs before its own number is enough gives you the target.
Frequently asked questions
Is credibility weighting the same as Bayesian shrinkage?
They coincide. The credibility estimate Z x segment + (1 - Z) x overall is exactly the posterior mean of a hierarchical normal model, with k equal to the ratio of within- to between-segment variance. Actuaries arrived at the formula from ratemaking practice decades before it was framed as empirical Bayes, which is why the same quantity has two names. If your organisation is more comfortable with one vocabulary than the other, use it - the number is identical.
Does this mean I should never report a raw segment average?
Report it when the segment has enough data that Z is close to 1, which is what "enough" actually means. Below that, the raw average is not wrong so much as overconfident: it is an unbiased estimate with a standard error nobody looks at. The blend trades a little bias for a large reduction in error, which is the trade you want when the number is going into a decision.
What if my segments are not comparable at all?
Then the method does not apply, and you should not be pooling. Credibility weighting assumes the segments are variations on a common population. If one "segment" is a different product in a different country sold on a different contract, it does not belong in the same blend. Splitting into separate pools, each internally comparable, is the correct response.
How many segments do I need before I can estimate k?
Practically, around eight to ten segments before the between-segment variance is itself estimated well enough to be useful. With fewer, k is noisy and you are better off choosing a conservative k by judgement and stating that you did. With Koji, adding segments is usually a matter of asking one more single_choice screening question, so this constraint binds less than it used to.
Should I weight by number of interviews or by something else?
Use whatever unit your variance was computed in. If the metric is a per-respondent rating, n is respondents. If it is a per-event rate, n is events. The classical actuarial version counts claims for exactly this reason - the unit of the count must match the unit of the random variation.
Can I apply this to qualitative findings?
Not to the arithmetic, but the underlying discipline transfers. A theme mentioned by two of eight people in a small segment is weak evidence about that segment and moderate evidence about your customers generally. Koji's analysis makes this tractable because the open_ended responses are indexed alongside the structured ones, so you can check whether a theme that looks segment-specific is actually present everywhere.
Related Resources
- Structured questions guide - the six question types and when to use each
- Why the top and bottom segments in your report are both the smallest ones - the ranking problem this does not fix
- How much data a segment needs before its own number is enough - the full-credibility arithmetic
- When shrinkage hides the one segment that actually changed - the failure mode of the blend
- Is 4.1 good? Internal benchmarks and percentile norms - what to compare a segment against
- Regression to the mean - the temporal cousin of this problem
Related Articles
When Shrinkage Hides the One Segment That Actually Changed (2026)
Every small-segment correction assumes your segments are draws from one population. When a segment genuinely breaks away, the correction pulls it back toward a mean it no longer belongs to, silently.
How Much Data a Segment Needs Before Its Own Number Is Enough (2026)
How many people do I need per segment has an exact answer, it predates modern market research, and it is not 30. Here is the formula, the classical table, and the translation to research metrics.
Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms
A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.
Regression to the Mean: Why Your Fix Looks Like It Worked (2026)
Regression to the mean makes ordinary noise look like a successful intervention. Learn the formula that predicts how much of your improvement is arithmetic, the five product-research traps it hides in, and the designs that separate a real win from a bounce-back.
Why the Top and Bottom Segments in Your Report Are Both the Smallest Ones (2026)
Rank your segments by score and the top and bottom of the list fill up with your smallest segments. This is arithmetic, not bad luck, and no multiplicity correction touches it.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.