Back to docs
Analysis & Synthesis

Digit Heaping: Why Self-Reported Numbers Pile Up on Fives and Tens (2026)

Self-reported counts cluster on multiples of 5 and 10. Heaping is measurable with Whipple's index from data you already hold, and it breaks any threshold placed on a round number.

When you ask people for a number, most of them give you a rounded one. Ask how many times someone used a feature last month and the answers cluster on 5, 10, 20 and 30. This is called heaping, it is not random noise that averages out, and it is measurable from data you already hold. It is a resolution limit with a direction.

Heaping matters because the two things product teams do most with self-reported counts are exactly the two things heaping ruins: setting a cutoff (users who do this 10 or more times a month) and comparing a small difference between segments. A threshold placed on a heap point is maximally ambiguous, and a difference smaller than the rounding grain is not a finding at all.

What heaping is, and why it is not just noise

Heaping happens because people do not retrieve counts, they estimate them. Asked for a frequency over any period longer than a few days, a respondent reconstructs a rough rate and reports a round multiple. Some respondents genuinely do count, which is what makes the resulting distribution a mixture of two reporting behaviours rather than one noisy signal.

Wang and Heitjan put the mixture plainly in Statistics in Medicine in 2008, modelling self-reported cigarette counts: "Some subjects report exact cigarette counts, whereas others report rounded-off counts, particularly multiples of 20, 10 or 5". The multiples they name correspond to a pack, a half pack and a quarter pack. The grid people round to is set by the units in which they already think about the behaviour, not by the units you asked for.

Crawford, Weiss and Suchard, writing in The Annals of Applied Statistics in 2015 on self-reported counts of sexual partners, state both halves of the problem in a single sentence: "Respondents may misremember or round to a nearby multiple of 5 or 10. This phenomenon is called heaping, and the error inherent in heaped self-reported numbers can bias estimation."

Two features make this different from ordinary measurement error. Heaping is structured: the errors are not independent draws around the truth, they are displacements onto a fixed grid. And its damage is positional: because the grid is dense at round numbers and absent between them, how badly you are hurt depends entirely on where your analysis cuts the number line.

This is not recall bias, and the distinction is practical. Recall bias is a failure of memory, where an event is forgotten or misdated. Heaping survives perfect memory. A respondent who knows with certainty that they logged in eleven times may still answer about ten, because the question invited an estimate and ten is the conventional way to express that magnitude. The fix for forgetting is a shorter reference period. The fix for heaping is a different question format and a different analysis. Teams that confuse the two shorten the recall window and then wonder why the answers still pile up on multiples of five.

Where this shows up in product research

You are collecting heaped numbers more often than you think. The pattern appears wherever a question asks for a quantity the respondent has never had a reason to count.

QuestionTypical heap pointsWhat it breaks
How many hours a week do you spend in the tool?1, 2, 5, 10, 20, 40Any hours-based ROI claim
How many times did you do X last month?5, 10, 20, 30Engagement thresholds
How many people on your team use this?5, 10, 20, 50, 100Seat and segment sizing
How long have you used the product?6 months, 1 year, 2 years, 5 yearsTenure cohorts
What percentage of your work involves X?10, 25, 50, 75, 90Any share-of-work estimate

The last row heaps on a different grid. People answer in quarters and tenths, so 25, 50 and 75 absorb enormous mass. A reported about half is compatible with anything from a third to two thirds, which means a segment gap between 45 percent and 55 percent is very likely manufactured by the grid rather than observed in the world.

Tenure questions heap for an additional reason: the available answers are calendar-shaped. Nobody says fourteen months, they say about a year. If your cohort boundary sits at twelve months, you are splitting a heap, which is the worst available place to cut.

Measure it before you argue about it

Demographers have measured this for a century, because census ages heap badly. The standard instrument is Whipple's index. In its classic form you take everyone reporting an age between 23 and 62 inclusive, count those whose reported age ends in 0 or 5, divide by the total in that range, and multiply by 5. Under no heaping, one age in five ends in 0 or 5, so a clean dataset reads 100. The United Nations publishes interpretation bands: below 105 is highly accurate, 105 to 109.9 fairly accurate, 110 to 124.9 approximate, 125 to 174.9 rough, and 175 or above very rough.

The same calculation works on any self-reported count. Take the share of answers ending in 0 or 5, divide by 0.2, multiply by 100.

A worked example. Four hundred people answer how many times did you use the export feature last month. Under no heaping you would expect about 20 percent of answers, roughly 80 of them, to end in 0 or 5. Suppose 210 do.

  • Observed share ending in 0 or 5: 210 of 400, which is 52.5 percent.
  • Divide by the 20 percent expected: 2.625.
  • Index: 262.5.

That is far above 175, so on the UN scale this data is very rough. The number to carry forward is not the mean. It is the grain: answers are being reported to the nearest 5, so any difference smaller than 5 units between two groups sits inside one rounding step and cannot be separated from a difference in reporting habit.

The threshold trap, with the arithmetic

Suppose you define an engaged user as one who performs an action 10 or more times a month, and you measure it by asking. Take 500 users whose true counts are spread evenly across 8, 9, 10, 11 and 12, so 100 users at each value. The true number of engaged users is 300, the ones at 10, 11 and 12.

Now let 60 percent of users round to the nearest multiple of 5 and 40 percent report exactly. Everyone between 8 and 12 who rounds lands on 10.

  • Rounders: 60 at each of the five values, 300 in total, all reporting 10. Every one of them passes a 10-or-more test.
  • Exact reporters: 40 at each value. The 80 at 8 and 9 fail. The 120 at 10, 11 and 12 pass.
  • Measured engaged users: 300 plus 120, which is 420, against a true 300. You overstate engagement by 40 percent.

Now move the threshold one step up, to 11 or more, and change nothing else. The true count is 200, the users at 11 and 12.

  • All 300 rounders report 10, and now all 300 fail.
  • Only the exact reporters at 11 and 12 pass, which is 80 people.
  • Measured: 80 against a true 200. You understate by 60 percent.

Same users, same answers, same honesty, a one-unit change in the threshold, and the error swings from plus 40 percent to minus 60 percent. The direction is set by which side of the heap your cutoff falls on. This is the most expensive consequence of heaping and it is invisible in every summary statistic: the mean barely moves, the sample size is unchanged, and nothing in the analysis flags it.

The same mechanism quietly suppresses change over time. If a true shift from 11 to 13 sessions leaves both values rounding to 10, your tracking study reports no movement at all.

What to do about it

  • Compute the index before the analysis. One query over terminal digits tells you the grain you are entitled to claim. Report it next to the estimate.
  • Never put a threshold on or beside a heap point. If 10 is a heap, define the cutoff at 12 or 13, where almost nobody is displaced, or build the cutoff on observed behaviour rather than a reported count.
  • Do not publish a difference smaller than the grain. If the data rounds to 5, a reported gap of 2 between segments is not a result.
  • Prefer a bounded, countable question. How many times in the last 7 days invites counting. How many times in a typical month forces estimation and guarantees heaping.
  • Ask for the number and then probe it. The follow-up you said about ten, walk me through last week converts an estimate into an episode count, and it tells you whether the ten was counted or constructed.
  • Treat the digit pattern as a data-quality signal. A heaping index that jumps between two waves of the same survey means the reporting behaviour changed, which you want to know before you read the trend. Heaping belongs in the same error budget as sampling error, alongside the other components of total survey error.

How Koji handles this

Heaping is a question-design and analysis problem, so the leverage sits in how the question is asked and what gets stored next to the answer.

  • Six structured question types, so a count is stored as a number. Koji ships six structured question types: open_ended, scale, single_choice, multiple_choice, ranking and yes_no. A scale question stores a real numeric value, which is what makes a terminal-digit audit possible at all. A count buried in free text cannot be audited for heaping until somebody extracts it.
  • The anchored follow-up turns an estimate into an episode. Every Koji question carries a probing configuration, and for scale questions an anchor option makes the AI interviewer follow a number with a request to account for it. That is the one move that separates a counted ten from a constructed ten, and it runs on every interview rather than only when a human moderator remembers to ask.
  • Both the number and the reasoning are kept for the same question. Koji's analysis stores a structured value and a qualitative answer per question, so the round number and the sentence that produced it stay attached. When an answer looks heaped, you can read how it was reached instead of guessing.
  • Scale answers render as a distribution, not just a mean. Koji visualises a scale question as a distribution, which is where heaping becomes visible to the naked eye. A mean hides it completely. A histogram with spikes on 5, 10 and 20 shows you the grid.
  • Traceability back to the transcript. Each extracted answer records which messages it came from, so a suspicious count can be checked against what the participant actually said.

Running a hundred interviews that each ask for a count and then probe it is not realistic by hand, which is the practical reason heaping goes unexamined on most teams. A legacy survey tool like SurveyMonkey can collect the number but cannot ask the follow-up that reveals how it was produced. An AI-native platform asks it every time.

Frequently asked questions

What is digit heaping in survey data?

Digit heaping is the tendency for self-reported numbers to cluster on round values, usually multiples of 5 and 10, because respondents estimate rather than count. It produces a distribution that mixes exact reporters with rounders, so the error is structured displacement onto a grid rather than independent noise around the truth.

How do I know if my data is heaped?

Compute the share of answers whose last digit is 0 or 5, divide it by 0.2, and multiply by 100. That is Whipple's index applied to a count. A reading near 100 means no heaping, while the United Nations treats 125 to 174.9 as rough and 175 or above as very rough. The calculation needs no new fieldwork, only the answers you already have.

Is heaping the same as recall bias?

No. Recall bias is a memory failure, where an event is forgotten or misdated. Heaping survives perfect memory: someone who knows the true count was eleven may still answer about ten, because the question invited an estimate. The fixes differ, so shortening the recall window will not remove heaping.

Why does heaping matter if the average is about right?

Because the average is rarely what you act on. Heaping does most of its damage at thresholds and in small comparisons. A cutoff placed just below a heap point can overstate the qualifying population by around 40 percent, while moving that same cutoff one unit above the heap can understate it by around 60 percent, on identical data.

How do I stop respondents rounding?

You cannot eliminate rounding, but you can reduce it by making counting feasible and by probing the answer. Ask about a short bounded period rather than a typical month, and follow the number with a request to walk through that period. Koji's anchored probing does this automatically on scale questions.

Should I throw out heaped answers?

No. Heaping is a property of the measurement rather than a sign of a bad respondent, and discarding rounders would remove a large and non-random share of your sample. Report the grain instead, avoid thresholds on heap points, and keep differences smaller than the rounding step out of your conclusions.

Related Resources

Related Articles

Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)

Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Question Types: The Complete Guide to 14 Question Types with Examples (2026)

A complete reference of every survey question type — open-ended, closed-ended, Likert, matrix, ranking, semantic differential, and more. When to use each, real examples, common pitfalls, and the AI-native approach that combines them all in one conversation.

Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)

Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.

Activating Research Insights: Turn Findings Into Product Decisions

A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.

How to Analyze Open-Ended Survey Responses with AI (2026 Guide)

Stop manually coding free-text survey responses. Learn how AI analyzes open-ended answers at scale — surfacing themes, sentiment, and quotes in minutes, plus why an AI interview captures 10x more depth than any survey can.