Back to docs
Research Methods

Often, Sometimes, Rarely: Why Verbal Frequency Options Are Not Comparable (2026)

Vague frequency words carry different numeric meanings for different respondents, and the shift tracks the groups you are comparing. A measured study shows the bias can hide a real difference rather than invent one.

The word often is not a quantity. Its numeric meaning is set by the respondent, calibrated against their own experience, and it shifts systematically between the very groups you are trying to compare. When two segments differ on a verbal frequency scale, part of that difference is behaviour and part is vocabulary, and nothing in the output tells you the ratio.

This is not a quibble about wording. It is a validity problem with a measured size and a known direction, and in the best available study the bias ran in the conservative direction, hiding a real difference rather than inventing one. That makes it a producer of quiet false negatives, which are much harder to catch than surprising results.

The study that actually measured the mapping

Schneider and Stone, publishing in Quality of Life Research in 2016, ran the experiment the problem demands. Six hundred respondents rated how often they had positive and negative experiences using vague quantifiers - never, rarely, sometimes, often, always - and then gave open-ended numeric frequency counts for the same items. With both the word and the number in hand, the researchers could ask what each word actually meant to each person, and whether that meaning differed by who the person was.

It did, in two ways.

Respondents with a chronic medical condition assigned a higher numeric frequency to the same vague quantifier for negative experiences than respondents without one. The effect did not appear for positive experiences. That asymmetry is the tell: the quantifier is being calibrated against the respondent's own base rate for that kind of event. If unpleasant things happen to you regularly, it takes more of them before you will call them often.

Older respondents gave more extreme numbers at both ends of the scale than younger respondents, lower at the low end and somewhat higher at the high end.

The conclusion is stated plainly: "The results suggest that people with different medical backgrounds and age do not interpret vague frequency quantifiers on a QoL scale in the same way."

Now the part that matters most for anyone running segment comparisons. The authors re-estimated the group differences after correcting for the differing interpretations, and reported that "After adjusting for these effects, differences in QoL became somewhat more pronounced between medical status groups, but not between age groups."

Read that direction carefully. The uncorrected verbal scale understated the real gap between people with and without a condition. Because the affected group needs more events before reaching for the same word, their greater burden was partly absorbed by their own vocabulary, and the scale reported them as closer to everyone else than they were.

A bias that shrinks a difference is more dangerous than one that inflates it. An inflated difference produces a surprising result, and surprising results get challenged. A shrunken difference produces a null, and nulls are not challenged, they are filed. A team using verbal frequency options to compare a power-user segment against a new-user segment should expect the scale to work against them, and should not read a flat result as evidence of no difference.

Even a standardised, regulated vocabulary fails

If vague quantifiers could be made to work by defining them, drug labelling would have solved it. The European Union publishes a standard vocabulary for side-effect frequency with numeric bands attached: very rare is below 0.01 percent, rare is 0.01 to 0.1 percent, uncommon is 0.1 to 1 percent, common is 1 to 10 percent, and very common is above 10 percent.

Berry, Knapp and Raynor tested whether those descriptors transmit the risk they are defined to mean, and reported in The Lancet in 2002 that "qualitative descriptions led to gross overestimation of risk".

The follow-up trial put numbers on it. Knapp, Raynor and Berry, writing in Quality and Safety in Health Care in 2004, gave 120 adults information about a medicine side effect in either verbal or numerical form. For constipation, presented either as common or as 2.5 percent, the mean estimated likelihood was 34.2 percent in the verbal group against 8.1 percent in the numerical group. For pancreatitis, presented either as rare or as 0.04 percent, it was 18 percent against 2.1 percent.

Work through what that means for the verbal condition:

  • Told common for an event stated at 2.5 percent, people estimated 34.2 percent - about 13.7 times the stated figure, and more than three times the top of the 1-to-10-percent band the word is defined to occupy.
  • Told rare for an event stated at 0.04 percent, people estimated 18 percent - roughly 450 times the stated figure.

The authors conclude that the use of verbal descriptors "leads to overestimation of the level of harm and may lead patients to make inappropriate decisions about whether or not they take the medicine."

Be precise about what transfers. In this trial the words were presented to respondents as descriptions to interpret, not offered as response options to choose from, so these specific magnitudes are not a measurement of answer-option error. What transfers is the harder structural point: this is the most carefully standardised verbal frequency vocabulary in existence, published in regulation with numeric bands attached, and it still failed to convey the quantity it was defined to mean. An ad hoc product survey offering rarely, sometimes, often has no definitions, no bands and no regulator. There is no reason to expect it to do better. For the size of the error in response options specifically, Schneider and Stone is the study to cite.

The tension worth resolving: labels help and labels hurt

Maria Rosala, writing for Nielsen Norman Group, notes that "While research has found that people find it easier to comprehend word-labeled scales compared to unlabeled ones, it can be hard to come up with the right word to describe an intermediate point on a scale." On scales left unlabelled, she observes that "because the scale is not labeled, each option could be interpreted differently across multiple respondents".

Set beside Schneider and Stone, those two statements look contradictory. One locates differential interpretation in the absence of labels. The other finds it in their presence. Both are correct, because they describe different properties of a scale.

  • Comprehension is whether a respondent knows what to do with the scale. Words help here, substantially. A labelled scale produces fewer missing and fewer mis-keyed answers.
  • Calibration is whether two respondents who choose the same option meant the same magnitude. Words do not help here, and vague quantifiers actively hurt, because each word imports the respondent's private reference frame.

The prescription falls straight out: label for comprehension, measure for calibration. Use words so that people can answer, and collect something numeric so that you can find out what the words meant. That is precisely what Schneider and Stone recommend: "Open-ended numeric frequency reports may be useful to detect and potentially correct for differences in the meaning of vague quantifiers."

What to ask instead

Instead ofAsk
How often do you use search? Rarely / Sometimes / OftenAbout how many times did you use search in the last 7 days?
How frequently does this error occur? Occasionally / RegularlyHow many times did you hit this error in the last 7 days?
Do you review reports often?In the last month, how many reports did you open?
How satisfied are you? Somewhat / VeryA 1 to 5 scale with only the two endpoints labelled

Four rules that follow:

  • Use absolute frequencies with explicit units and a bounded window. Times in the last 7 days is answerable. Often is not.
  • Anchor the endpoints and leave the interior numeric. This keeps the comprehension benefit of words where it is unambiguous, at the extremes, without asking the middle of the scale to carry a shared definition it cannot carry.
  • Collect the word and the number together, at least on a subsample. This is the only way to detect a shifted mapping, and it converts an untestable assumption into a measurement.
  • Never compare a verbal frequency scale across segments without evidence that the mapping is stable. That assumption has a name and a test suite - see measurement invariance. Non-invariance in a verbal frequency item is often exactly this problem, and the vague quantifier gives you a concrete, fixable cause rather than an abstract warning.

One honest caveat. Asking for a number buys calibration and costs you something else: self-reported counts cluster hard on multiples of 5 and 10, which caps the resolution of the answer. You are choosing which error to carry, not escaping error. See digit heaping for how to measure what that choice costs, and rating-scale reversals for a separate way group comparisons on ordinal scales go wrong.

How Koji handles this

  • Six structured question types, so you are not forced to choose. Koji ships six structured question types: open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Pairing a scale item with an open_ended numeric item for the same construct is the Schneider and Stone design, available as a standard study configuration rather than a custom research project.
  • Endpoint-only scale labels are a first-class option. A Koji scale question takes labels keyed to specific scale points, so you can label 1 and 5 and leave the interior numeric. The format that avoids the vague-quantifier trap is the format the tool makes easiest.
  • The AI interviewer can ask what the word meant. Koji questions carry a probing configuration, and for scale questions an anchor option prompts a follow-up on the number the participant gave. Asking you said you hit this sometimes, how many times was that last week is how a mapping gets measured, and it runs on every interview rather than only when a moderator remembers.
  • The word and the number are stored against the same question. Koji's analysis keeps a structured value alongside a qualitative answer per question, so the chosen option and the reasoning behind it stay linked for analysis.
  • Automatic thematic analysis reads the reasoning, not just the option. Where a verbal scale flattens two groups toward each other, the transcripts usually do not, because the described episodes differ even when the chosen word does not.

While a legacy survey tool like SurveyMonkey can only record which word a respondent picked, an AI-native platform like Koji can ask what they meant by it. That difference is the whole of the calibration problem.

Frequently asked questions

What is a vague quantifier in a survey?

A vague quantifier is a frequency or quantity word used as a response option, such as rarely, sometimes, often or regularly. It is vague because it names no quantity, so each respondent supplies their own numeric interpretation and two people choosing the same option may mean very different magnitudes.

Is it really a problem if everyone uses the same scale?

Yes, because the interpretation is not constant across people. Schneider and Stone found that respondents with a chronic condition attached higher numeric frequencies to the same quantifier for negative experiences than respondents without one, and that older respondents gave more extreme numbers at both ends. A shared scale does not produce a shared meaning.

Which direction does the bias run?

It can run either way, but the measured case was conservative. After adjusting for differing interpretations, the gap between medical status groups became more pronounced, meaning the raw verbal scale had understated a real difference. That is the dangerous direction, because it produces a null result that nobody thinks to challenge.

Should I just use numbers instead of words?

Numbers fix calibration but introduce their own artifact: self-reported counts pile up on multiples of 5 and 10, which limits the resolution you can claim. The strongest design collects both, a labelled option and an open numeric count, so that the mapping between them can be measured rather than assumed.

How should I label a rating scale?

Label the endpoints and leave the interior points numeric. Words at the extremes are close to unambiguous, while words for intermediate points ask respondents to share a definition they demonstrably do not share. Koji scale questions support labels on specific points, so endpoint-only labelling is straightforward.

How is this different from measurement invariance?

Measurement invariance is the framework that tests whether an instrument behaves the same way across groups, and it tells you that non-invariance exists. Vague quantifiers are one specific and fixable cause of it. The practical difference is that this cause can be measured directly by collecting a numeric count beside the word, without a psychometrics team.

Related Resources

Related Articles

Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)

Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.

When Relabelling the Scale Reverses Which Group Scores Higher (2026)

Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.

Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)

Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Question Types: The Complete Guide to 14 Question Types with Examples (2026)

A complete reference of every survey question type — open-ended, closed-ended, Likert, matrix, ranking, semantic differential, and more. When to use each, real examples, common pitfalls, and the AI-native approach that combines them all in one conversation.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.