The Average of a Ratio Is Not the Ratio of the Averages (2026)
Averaging percentages across segments, or averaging through a curve, produces a number that describes nobody. Two distinct failure modes, two different fixes.
There are two ways to average a research metric and get a number that describes none of your users. They look identical in a report and they have completely different fixes.
The short answer
If you average the conversion rates of your segments, you get the average of a ratio. If you pool everyone and divide once, you get the ratio of the averages. These are different numbers, and the gap can be enormous: in the worked example below, an unweighted average of two segment rates reports 47.50 percent where the true pooled rate is 22.12 percent. That is an overstatement of 2.1 times, from data containing no errors at all.
A second, separate problem appears whenever you average something and then push it through a curve, or push values through a curve and then average. Jensen's inequality says these two operations disagree for any function that is not a straight line. For task times, ratings, and anything with a long tail, the arithmetic mean sits well above the typical value: 82.50 seconds against a geometric mean of 23.40 seconds in the example below.
The first problem is about weights. The second is about shape. Diagnosing which one you have determines the fix, so this guide treats them separately.
Failure one: averaging rates without weights
The rule is simple. A percentage is a fraction with a denominator, and you cannot average fractions by averaging their numerators and denominators separately. A segment of 20 users and a segment of 500 users do not get equal votes in a company-wide rate, but an unweighted mean gives them exactly that.
Take two segments. Enterprise has 20 users and 15 of them activate, a 75 percent rate. Self-serve has 500 users and 100 of them activate, a 20 percent rate.
Average the two rates and you get 47.50 percent. Pool the users and you get 115 activations out of 520 people, which is 22.12 percent. The unweighted figure overstates the truth by a factor of 2.148, and it does so because the 20-person segment was allowed to contribute as much as the 500-person one.
The most famous instance of this in the research literature runs the other direction, and is worth knowing because it shows the reversal can be complete rather than merely large. The 1973 graduate admissions figures at the University of California, Berkeley are the standard teaching case for Simpson's paradox. In the aggregate, 8,442 men applied and 44 percent were admitted, against 4,321 women of whom 35 percent were admitted, a total of 12,763 applicants at an overall rate of 41 percent. As the standard account records it, "The admission figures for the fall of 1973 showed that men applying were more likely than women to be admitted, and the difference was so large that it was unlikely to be due to chance."
Then look department by department. Across the six largest departments - 2,691 men admitted at 45 percent and 1,835 women admitted at 30 percent - women were admitted at an equal or higher rate in four of the six. In Department A, 108 women applied and 82 percent were admitted, against 62 percent of 825 men. The reversal happens because women applied disproportionately to the departments that rejected nearly everyone: Department F admitted 6 percent of men and 7 percent of women. The same analysis, the same source notes, found a "small but statistically significant bias in favor of women" once the data were pooled and corrected.
Nothing in that dataset is mismeasured. The aggregate and the per-department views disagree because applicants are distributed unevenly across departments with very different base rates. That is the shape of every weighting failure you will meet in product research: a segment mix that differs from the mix you implicitly assumed when you took a plain average.
Failure two: averaging through a curve
The second failure has nothing to do with weights and survives even when every group is the same size. It comes from nonlinearity.
Jensen's inequality, in its simplest form, states that "the convex transformation of a mean is less than or equal to the mean applied after convex transformation", with the opposite inequality holding for concave transformations. In plain terms, transforming the average is not the same as averaging the transformed values, and the direction of the gap is fixed by whether the curve bends up or down. Equality holds only in the degenerate cases - when every value is identical, or when the transformation is linear on the range of your data. Averaging is safe through addition and multiplication by a constant. It is unsafe through anything that bends.
Most research quantities bend. Task times are bounded below by zero and unbounded above, so their distribution is right-skewed, and the arithmetic mean is pulled toward the slow tail. Take four users completing a task in 10, 10, 10 and 300 seconds. The arithmetic mean is 82.50 seconds. The geometric mean is 23.40 seconds, about 28 percent of the arithmetic figure. The median is 10 seconds. Three of the four users finished in ten seconds, and the headline "average time on task: 82.5 seconds" describes not one of them.
Nielsen Norman Group's quantitative glossary is explicit about this: with "skewed distributions like task times, the median or the geometric mean may be more appropriate" than the arithmetic mean, because "Time-on-task data is often skewed." The same caution applies to any rate expressed as a reciprocal. Averaging speeds, throughputs, or per-unit costs with a plain arithmetic mean silently assumes the denominators are equal, and they rarely are.
Telling the two apart
The diagnostic takes one question each.
For weighting: do my groups have different sizes, and did I give each group one vote? If yes, you have failure one. Recompute by pooling the raw numerators and denominators, then compare. If the two numbers differ materially, report the pooled figure and, if the segment differences matter, report them separately rather than averaging them away.
For shape: is the quantity I am averaging a duration, a rate, a ratio, a count with a long tail, or anything I would plot on a log scale? If yes, you have failure two. Check the mean against the median. A large gap between them is the signature of a distribution that the arithmetic mean cannot summarise.
It is entirely possible to have both at once, which is why the order matters: fix the weights first, because pooling correctly changes which distribution you are then trying to summarise.
Which summary to report
Pooled rate for anything with a denominator. One numerator, one denominator, one division, done last.
Segment rates shown side by side when the segments genuinely differ. The mistake is not looking at segments; it is collapsing them with a plain average. If Enterprise activates at 75 percent and self-serve at 20 percent, that difference is the finding. Averaging it produces 47.50 percent, which erases the only useful thing in the data.
Median or geometric mean for skewed durations, with the distribution shown. For task times a median plus an interquartile range communicates more than any single mean.
Explicit weights when you do need a single composite. If a weighted average is genuinely what the decision needs, state the weights. A weighted average whose weights are written down is a defensible estimate; an unweighted one is an accidental claim that all your segments are the same size.
Note that this is a different question from whether the underlying scale supports averaging at all. For the ordinal case, see our guide on whether you can average Likert scale data. For why your smallest segments keep appearing at the top and bottom of ranked reports, see the segment ranking sample size artefact.
How Koji helps
Both failures are made worse by thin data. When a segment has 20 people in it, you are forced to choose between an unweighted average that overweights it and a pooled figure that buries it - and neither option recovers what that segment actually thinks. The real fix is enough coverage in every segment to report them separately, which is a cost problem, and cost is what Koji changes.
Koji runs AI-moderated interviews in parallel, so filling a thin segment stops meaning weeks of scheduling. A study that would have yielded 20 enterprise responses can reach a sample large enough to stand on its own, which means you can show segments side by side instead of collapsing them into one misleading mean. Voice interviews widen that reach further, because they capture the participants who will talk for six minutes but never complete a form.
Koji's structured questions are what keep the arithmetic honest. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and the distinction matters directly here. A yes_no or single_choice item gives you a clean numerator and denominator per respondent, so a pooled rate is computable rather than reconstructed from prose. A scale item preserves the full distribution instead of only its mean, so you can see the skew before you choose a summary. A ranking item records order rather than an averaged score, which sidesteps a whole family of averaging artefacts.
Koji's real-time reporting shows segment counts as responses arrive, so an unbalanced mix is visible while you can still fix it, not after you have written the summary. Automatic thematic analysis reports themes per segment rather than pooling them into one list, which is the qualitative version of the same discipline. And because Koji's AI consultants are customisable, you can brief the interviewer to capture the segment-defining attribute explicitly, so nobody has to infer group membership at analysis time. Traditional survey tools like SurveyMonkey will compute whatever average you click; an AI-native platform like Koji gets you the per-segment depth that makes the correct average possible in the first place.
Common mistakes
Averaging percentages that came from different denominators. The single most common version of failure one. If the denominators differ, the mean of the percentages is not a percentage of anything.
Averaging an average. A mean of segment means inherits every weighting error in the segments beneath it, and compounds them one level up.
Reporting a mean for a long-tailed duration. The arithmetic mean of task times is dominated by whoever struggled most. It is a real number about a real sample, and it describes nobody.
Assuming equal group sizes because the groups look comparable. Two segments that both "matter" to the business can differ by an order of magnitude in count.
Fixing shape before weights. Switching to a median does not repair an unweighted mix. Pool correctly first, then choose the summary.
Treating a reversal as a data error. When aggregate and segment views disagree, both are usually computed correctly. The disagreement is information about your segment mix.
Frequently asked questions
When is it safe to average percentages directly?
Only when every group has the same denominator, or when you attach explicit weights proportional to those denominators. With equal group sizes an unweighted mean of rates equals the pooled rate exactly. With unequal sizes it does not, and the gap grows with the imbalance. If you cannot state the denominators, you cannot defend the average.
Is this the same as Simpson's paradox?
Simpson's paradox is the most dramatic form of the weighting failure, where the aggregate comparison actually reverses direction once you look within groups. The weighting problem is broader: most of the time the numbers are merely wrong in magnitude rather than reversed in sign. The Berkeley admissions case is the reversal version, and it is the one worth memorising because it proves the effect is not a rounding concern.
Why does the geometric mean fall below the arithmetic mean?
Because the logarithm bends downward, and Jensen's inequality fixes the direction of the gap for any curve that bends. Taking logs, averaging, and converting back gives a value that is always less than or equal to the arithmetic mean, with equality only when every value is identical. For right-skewed data like task times, that lower number is usually the better description of a typical user.
Should I use the median or the geometric mean for task times?
Either is defensible and both beat the arithmetic mean. The median is easier to explain and more robust to a single extreme value. The geometric mean uses all the data and behaves better when you need to compare ratios across studies. Reporting the median alongside the distribution is the safest default for a mixed audience.
Does this affect qualitative analysis too?
Yes, in the form of theme counts. If one segment contributed 40 interviews and another contributed 5, a combined theme frequency list is an unweighted average in disguise, and the larger segment sets the agenda. Koji reports themes per segment for exactly this reason, so prevalence in a small segment stays visible instead of being diluted.
What should I do when the pooled and segment views disagree?
Report both, and treat the disagreement as the finding rather than a problem to resolve. A pooled rate answers what happens across the business; segment rates answer who it happens to. When they point in different directions, the cause is almost always an uneven segment mix, and that mix is usually something the team can act on.
Related Resources
- Can You Average Likert Scale Data? - whether the scale supports an average at all
- Why Complaint Counts Cannot Become Rates - the missing denominator problem
- Why the Top and Bottom Segments Are Both the Smallest - small samples at the extremes of a ranked report
- Why Adding One Option Can Reverse Your Ranking Results - averaging ranks, and what it does
- Structured Questions in AI Interviews - the six question types and the arithmetic each one supports
- The Complete Guide to Thematic Analysis - reporting themes without pooling them away
Related Articles
Why Adding One Option Can Reverse Your Ranking Results (2026)
Average rank, the default summary for ranking questions, depends on which other options are in the list. A worked example of a reversal no respondent caused, and the first-place and pairwise summaries that stay stable.
Can You Average Likert Scale Data? What the Evidence Actually Says (2026)
Yes, in most situations, and the tests will behave. But robustness is about p-values, not meaning: the median can freeze while real change happens, and the mean can rank two groups in the opposite order from every other summary.
Why Complaint Counts Cannot Become Rates (And What to Compute Instead)
A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.
Why the Top and Bottom Segments in Your Report Are Both the Smallest Ones (2026)
Rank your segments by score and the top and bottom of the list fill up with your smallest segments. This is arithmetic, not bad luck, and no multiplicity correction touches it.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.