{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-02T09:41:19.593Z"},"content":[{"type":"documentation","id":"d3018154-5840-4992-9e86-62eee1d0046e","slug":"vague-quantifiers-frequency-questions","title":"Often, Sometimes, Rarely: Why Verbal Frequency Options Are Not Comparable (2026)","url":"https://www.koji.so/docs/vague-quantifiers-frequency-questions","summary":"Vague frequency words such as often and sometimes carry respondent-specific numeric meanings that shift with medical status and age, so group comparisons on verbal frequency scales confound behaviour with vocabulary. In the measured case the bias was conservative, understating a real group difference. The fix is to label endpoints for comprehension and collect an open numeric count for calibration.","content":"**The word often is not a quantity.** Its numeric meaning is set by the respondent, calibrated against their own experience, and it shifts systematically between the very groups you are trying to compare. When two segments differ on a verbal frequency scale, part of that difference is behaviour and part is vocabulary, and nothing in the output tells you the ratio.\n\nThis is not a quibble about wording. It is a validity problem with a measured size and a known direction, and in the best available study the bias ran in the conservative direction, hiding a real difference rather than inventing one. That makes it a producer of quiet false negatives, which are much harder to catch than surprising results.\n\n## The study that actually measured the mapping\n\nSchneider and Stone, publishing in *Quality of Life Research* in 2016, ran the experiment the problem demands. Six hundred respondents rated how often they had positive and negative experiences using vague quantifiers - never, rarely, sometimes, often, always - and then gave open-ended numeric frequency counts for the same items. With both the word and the number in hand, the researchers could ask what each word actually meant to each person, and whether that meaning differed by who the person was.\n\nIt did, in two ways.\n\n**Respondents with a chronic medical condition assigned a higher numeric frequency to the same vague quantifier for negative experiences than respondents without one.** The effect did not appear for positive experiences. That asymmetry is the tell: the quantifier is being calibrated against the respondent's own base rate for that kind of event. If unpleasant things happen to you regularly, it takes more of them before you will call them *often*.\n\n**Older respondents gave more extreme numbers at both ends of the scale** than younger respondents, lower at the low end and somewhat higher at the high end.\n\nThe conclusion is stated plainly: \"The results suggest that people with different medical backgrounds and age do not interpret vague frequency quantifiers on a QoL scale in the same way.\"\n\nNow the part that matters most for anyone running segment comparisons. The authors re-estimated the group differences after correcting for the differing interpretations, and reported that \"After adjusting for these effects, differences in QoL became somewhat more pronounced between medical status groups, but not between age groups.\"\n\nRead that direction carefully. The uncorrected verbal scale **understated** the real gap between people with and without a condition. Because the affected group needs more events before reaching for the same word, their greater burden was partly absorbed by their own vocabulary, and the scale reported them as closer to everyone else than they were.\n\n**A bias that shrinks a difference is more dangerous than one that inflates it.** An inflated difference produces a surprising result, and surprising results get challenged. A shrunken difference produces a null, and nulls are not challenged, they are filed. A team using verbal frequency options to compare a power-user segment against a new-user segment should expect the scale to work against them, and should not read a flat result as evidence of no difference.\n\n## Even a standardised, regulated vocabulary fails\n\nIf vague quantifiers could be made to work by defining them, drug labelling would have solved it. The European Union publishes a standard vocabulary for side-effect frequency with numeric bands attached: very rare is below 0.01 percent, rare is 0.01 to 0.1 percent, uncommon is 0.1 to 1 percent, common is 1 to 10 percent, and very common is above 10 percent.\n\nBerry, Knapp and Raynor tested whether those descriptors transmit the risk they are defined to mean, and reported in *The Lancet* in 2002 that \"qualitative descriptions led to gross overestimation of risk\".\n\nThe follow-up trial put numbers on it. Knapp, Raynor and Berry, writing in *Quality and Safety in Health Care* in 2004, gave 120 adults information about a medicine side effect in either verbal or numerical form. For constipation, presented either as common or as 2.5 percent, the mean estimated likelihood was 34.2 percent in the verbal group against 8.1 percent in the numerical group. For pancreatitis, presented either as rare or as 0.04 percent, it was 18 percent against 2.1 percent.\n\nWork through what that means for the verbal condition:\n\n- Told *common* for an event stated at 2.5 percent, people estimated 34.2 percent - about 13.7 times the stated figure, and more than three times the top of the 1-to-10-percent band the word is defined to occupy.\n- Told *rare* for an event stated at 0.04 percent, people estimated 18 percent - roughly 450 times the stated figure.\n\nThe authors conclude that the use of verbal descriptors \"leads to overestimation of the level of harm and may lead patients to make inappropriate decisions about whether or not they take the medicine.\"\n\n**Be precise about what transfers.** In this trial the words were presented to respondents as descriptions to interpret, not offered as response options to choose from, so these specific magnitudes are not a measurement of answer-option error. What transfers is the harder structural point: this is the most carefully standardised verbal frequency vocabulary in existence, published in regulation with numeric bands attached, and it still failed to convey the quantity it was defined to mean. An ad hoc product survey offering *rarely, sometimes, often* has no definitions, no bands and no regulator. There is no reason to expect it to do better. For the size of the error in response options specifically, Schneider and Stone is the study to cite.\n\n## The tension worth resolving: labels help and labels hurt\n\nMaria Rosala, writing for Nielsen Norman Group, notes that \"While research has found that people find it easier to comprehend word-labeled scales compared to unlabeled ones, it can be hard to come up with the right word to describe an intermediate point on a scale.\" On scales left unlabelled, she observes that \"because the scale is not labeled, each option could be interpreted differently across multiple respondents\".\n\nSet beside Schneider and Stone, those two statements look contradictory. One locates differential interpretation in the *absence* of labels. The other finds it in their *presence*. Both are correct, because they describe different properties of a scale.\n\n- **Comprehension** is whether a respondent knows what to do with the scale. Words help here, substantially. A labelled scale produces fewer missing and fewer mis-keyed answers.\n- **Calibration** is whether two respondents who choose the same option meant the same magnitude. Words do not help here, and vague quantifiers actively hurt, because each word imports the respondent's private reference frame.\n\nThe prescription falls straight out: **label for comprehension, measure for calibration.** Use words so that people can answer, and collect something numeric so that you can find out what the words meant. That is precisely what Schneider and Stone recommend: \"Open-ended numeric frequency reports may be useful to detect and potentially correct for differences in the meaning of vague quantifiers.\"\n\n## What to ask instead\n\n| Instead of | Ask |\n| --- | --- |\n| How often do you use search? Rarely / Sometimes / Often | About how many times did you use search in the last 7 days? |\n| How frequently does this error occur? Occasionally / Regularly | How many times did you hit this error in the last 7 days? |\n| Do you review reports often? | In the last month, how many reports did you open? |\n| How satisfied are you? Somewhat / Very | A 1 to 5 scale with only the two endpoints labelled |\n\nFour rules that follow:\n\n- **Use absolute frequencies with explicit units and a bounded window.** *Times in the last 7 days* is answerable. *Often* is not.\n- **Anchor the endpoints and leave the interior numeric.** This keeps the comprehension benefit of words where it is unambiguous, at the extremes, without asking the middle of the scale to carry a shared definition it cannot carry.\n- **Collect the word and the number together, at least on a subsample.** This is the only way to detect a shifted mapping, and it converts an untestable assumption into a measurement.\n- **Never compare a verbal frequency scale across segments without evidence that the mapping is stable.** That assumption has a name and a test suite - see [measurement invariance](/docs/measurement-invariance-comparing-groups). Non-invariance in a verbal frequency item is often exactly this problem, and the vague quantifier gives you a concrete, fixable cause rather than an abstract warning.\n\nOne honest caveat. Asking for a number buys calibration and costs you something else: self-reported counts cluster hard on multiples of 5 and 10, which caps the resolution of the answer. You are choosing which error to carry, not escaping error. See [digit heaping](/docs/self-reported-number-heaping-rounding) for how to measure what that choice costs, and [rating-scale reversals](/docs/ordinal-scale-group-comparison-reversal) for a separate way group comparisons on ordinal scales go wrong.\n\n## How Koji handles this\n\n- **Six structured question types, so you are not forced to choose.** Koji ships six [structured question types](/docs/structured-questions-guide): open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Pairing a scale item with an open_ended numeric item for the same construct is the Schneider and Stone design, available as a standard study configuration rather than a custom research project.\n- **Endpoint-only scale labels are a first-class option.** A Koji scale question takes labels keyed to specific scale points, so you can label 1 and 5 and leave the interior numeric. The format that avoids the vague-quantifier trap is the format the tool makes easiest.\n- **The AI interviewer can ask what the word meant.** Koji questions carry a probing configuration, and for scale questions an anchor option prompts a follow-up on the number the participant gave. Asking *you said you hit this sometimes, how many times was that last week* is how a mapping gets measured, and it runs on every interview rather than only when a moderator remembers.\n- **The word and the number are stored against the same question.** Koji's analysis keeps a structured value alongside a qualitative answer per question, so the chosen option and the reasoning behind it stay linked for analysis.\n- **Automatic thematic analysis reads the reasoning, not just the option.** Where a verbal scale flattens two groups toward each other, the transcripts usually do not, because the described episodes differ even when the chosen word does not.\n\nWhile a legacy survey tool like SurveyMonkey can only record which word a respondent picked, an AI-native platform like Koji can ask what they meant by it. That difference is the whole of the calibration problem.\n\n## Frequently asked questions\n\n### What is a vague quantifier in a survey?\n\nA vague quantifier is a frequency or quantity word used as a response option, such as rarely, sometimes, often or regularly. It is vague because it names no quantity, so each respondent supplies their own numeric interpretation and two people choosing the same option may mean very different magnitudes.\n\n### Is it really a problem if everyone uses the same scale?\n\nYes, because the interpretation is not constant across people. Schneider and Stone found that respondents with a chronic condition attached higher numeric frequencies to the same quantifier for negative experiences than respondents without one, and that older respondents gave more extreme numbers at both ends. A shared scale does not produce a shared meaning.\n\n### Which direction does the bias run?\n\nIt can run either way, but the measured case was conservative. After adjusting for differing interpretations, the gap between medical status groups became more pronounced, meaning the raw verbal scale had understated a real difference. That is the dangerous direction, because it produces a null result that nobody thinks to challenge.\n\n### Should I just use numbers instead of words?\n\nNumbers fix calibration but introduce their own artifact: self-reported counts pile up on multiples of 5 and 10, which limits the resolution you can claim. The strongest design collects both, a labelled option and an open numeric count, so that the mapping between them can be measured rather than assumed.\n\n### How should I label a rating scale?\n\nLabel the endpoints and leave the interior points numeric. Words at the extremes are close to unambiguous, while words for intermediate points ask respondents to share a definition they demonstrably do not share. Koji scale questions support labels on specific points, so endpoint-only labelling is straightforward.\n\n### How is this different from measurement invariance?\n\nMeasurement invariance is the framework that tests whether an instrument behaves the same way across groups, and it tells you that non-invariance exists. Vague quantifiers are one specific and fixable cause of it. The practical difference is that this cause can be measured directly by collecting a numeric count beside the word, without a psychometrics team.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and how to pair a scale item with an open numeric one\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - the framework that tests whether a scale means the same thing across groups\n- [Digit Heaping](/docs/self-reported-number-heaping-rounding) - the cost of asking for a number instead of a word\n- [Rating Scale Reversals](/docs/ordinal-scale-group-comparison-reversal) - another way ordinal group comparisons mislead\n- [Survey Question Types](/docs/survey-question-types) - the full taxonomy of question formats\n- [Recall Bias](/docs/recall-bias) - why frequency questions are hard before the wording is even chosen\n","category":"Research Methods","lastModified":"2026-10-02T03:32:51.373658+00:00","metaTitle":"Vague Quantifiers: Why Often and Sometimes Are Not Comparable","metaDescription":"Verbal frequency options mean different numbers to different respondents, and the bias can hide a real group difference. What to ask instead.","keywords":["vague quantifiers","frequency question wording","often sometimes rarely survey","verbal versus numeric scale","response option design","rating scale labels","survey comparability"],"aiSummary":"Vague frequency words such as often and sometimes carry respondent-specific numeric meanings that shift with medical status and age, so group comparisons on verbal frequency scales confound behaviour with vocabulary. In the measured case the bias was conservative, understating a real group difference. The fix is to label endpoints for comprehension and collect an open numeric count for calibration.","aiPrerequisites":["Basic familiarity with survey question design","Access to a questionnaire using frequency response options"],"aiLearningOutcomes":["Explain why verbal frequency options are not comparable across respondents","Identify when a scale bias hides rather than inflates a difference","Rewrite vague frequency items as bounded absolute counts","Label rating scales for comprehension without sacrificing calibration"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}