Just-About-Right Scales: The Rating Question Whose Average Means Nothing (2026)
JAR scales measure signed distance from an optimum, so their mean is not a summary. How to field, report and interpret them without averaging away the answer.
Most rating scales run from bad to good, so the average means something: higher is better, and a move from 3.4 to 3.8 is progress. A just-about-right scale does not work that way. Its best answer sits in the middle, its two ends are opposite complaints, and the moment you take an average of it you produce a number that can point at the wrong action - or at no action at all - with complete confidence.
The answer, up front
A just-about-right (JAR) scale asks whether an attribute is at the level the respondent wants: much too little, too little, just about right, too much, much too much. Use it for any attribute with an optimum rather than a maximum - notification frequency, onboarding length, default zoom, email cadence, how much the assistant explains itself. Never report its mean. Report the percentage who chose just about right, plus the split of the remainder by direction. A JAR mean of exactly 3.0 is produced both by a product everybody is happy with and by a product that has divided its users into two camps who want opposite fixes, and nothing in the average distinguishes them.
Why the arithmetic breaks
The problem is not that JAR data is noisy. It is that the scale is not measuring one thing.
On a satisfaction scale, the numbers 1 through 5 are ordered on a single dimension: less of a good thing to more of it. On a JAR scale, 1 and 5 are both failures, and they are failures in opposite directions. Moving from 1 to 3 is an improvement; moving from 3 to 5 is an equal-sized deterioration; and the arithmetic mean treats those two moves as if they cancelled, because numerically they do.
The sensory literature, which has used these scales for product optimization for decades, states the constraint bluntly. Iserliyska, Dzhivoderova and Nikovska, writing in Current Trends in Natural Sciences (volume 6, issue 11, 2017), note that bipolar JAR scales "cannot be evaluated using linear approaches" because the ratings are not normally distributed and, in their words, "the scale has two directions."
The standard handling is to stop treating the five points as a number line and collapse them into three groups instead. In the method that paper sets out, "the JAR values are amalgamated into three groups" - the two too-little points, the just-right point, and the two too-much points - and everything downstream is computed on those three categories.
The 3.0 that means four different things
Here is the failure in its cleanest form. Three panels rate the same attribute on a 5-point JAR scale. All three produce a mean of exactly 3.00.
| Distribution | Mean | Standard deviation | Chose just about right |
|---|---|---|---|
| Everyone answers 3 | 3.00 | 0.000 | 100% |
| 30% answer 2, 40% answer 3, 30% answer 4 | 3.00 | 0.775 | 40% |
| 40% answer 1, 20% answer 3, 40% answer 5 | 3.00 | 1.789 | 20% |
The first product needs no change. The third has 80% of its users actively unhappy, split down the middle between much too little and much too much, and it is the one most likely to be broken into two segments with genuinely incompatible needs. The mean is identical in all three cases. The percentage in the just-right box - 100%, 40%, 20% - separates them immediately.
This is worse than an ordinary averaging problem, because the direction of the error is not random. Polarization pulls a JAR mean toward the midpoint, which is the value that reads as no action needed. The more divided your users are, the more reassuring the average looks.
What a real JAR distribution looks like
Constructed examples can feel rigged, so here is a measured one. In the orange juice study above, 81 consumers rated six commercial products, using a 9-point hedonic scale for overall liking and JAR scales for color, sweet taste, sour taste, bitter taste and amount of pulp. For one product, the bitterness responses collapsed to:
- not enough: 11.11%
- just about right: 11.11%
- too much: 77.78%
Only nine of 81 people thought the bitterness was right. That distribution is strongly one-directional, so in this particular case an average would have pointed the right way - but it would have understated the problem, because the 11.11% at the not enough end drag the mean back toward the middle. The percentage-in-the-JAR-box statistic does not have that failure mode: 11% is 11% regardless of how the rest distribute.
The same product's sweet taste ran the other way, with 67.90% saying not enough and 14.81% saying too much. Two attributes, two directions, one product. A mean per attribute would have compressed both stories into a single ambiguous digit each.
When to reach for a JAR scale, and when not to
The test is one question: is there such a thing as too much of this?
Use JAR when the attribute has an optimum. Notification frequency, tutorial length, default page density, how chatty an assistant is, animation speed, how often you prompt for feedback, the number of options in a menu. Every one of these has users on both sides.
Do not use JAR when the attribute has a maximum. Reliability, clarity, accuracy, security, speed of a background job. Nobody wants a too reliable product, and offering the option produces either confused respondents or a scale where one half is never used. For these, an ordinary agreement or satisfaction scale is correct; see the Likert scale guide.
Watch for attributes that look bounded and are not. Price is the classic. Too cheap is a real perception that signals low quality, which is why price research uses its own instruments rather than a satisfaction scale.
A practical warning: JAR items are attribute diagnostics, not overall verdicts. Ask for overall liking on a normal scale, then ask JAR items on the components. That pairing is what makes the analysis in penalty analysis possible, and without an overall liking score alongside them your JAR data can tell you what is off but never what it costs you.
Reporting JAR data without lying
Three numbers per attribute, and no mean:
- Percent just about right. The headline. This is the only figure that behaves like a score.
- Percent too little / percent too much. Reported separately, never netted. Netting them recreates the exact cancellation the mean commits.
- The direction of the larger group, stated as the action it implies.
A useful convention is to treat any non-JAR group above 20% of respondents as actionable and anything below it as background, a threshold that comes straight from the mean-drop plots used in sensory work. And keep the two directions visually separate in the readout - a diverging bar with too little running left and too much running right is honest in a way that a single average never is.
One caution on sample size. Because you are reporting three proportions rather than one mean, the precision you need is proportion precision: at 100 respondents, a 20% reading carries a margin of error of roughly plus or minus 8 points, which is wide enough that a 20% group and a 30% group are not reliably different. Size the study for the smallest split you intend to act on - the survey sample size guide has the arithmetic.
Running JAR studies in Koji
JAR scales are trivial to field and historically painful to interpret, because the response tells you the direction of the problem and nothing about its cause. A respondent who says notifications are too frequent has not told you whether the volume is wrong, the timing is wrong, or one specific notification type is wrong. Traditional survey tools collect the direction and stop there.
- Field the item as a
single_choicequestion with the five JAR points as explicit options, so the categories are fixed and collapse cleanly. Usescalefor the paired overall-liking question,multiple_choicewhen you need to know which contexts the complaint applies to,rankingto force a priority order across several off-target attributes, andyes_nofor the screening question that establishes whether the respondent has met the attribute at all. All six types are documented in the structured questions guide. - Pair every JAR item with an
open_endedprobe. The closed item records the direction; the open one records the reason. Koji fields both in a single pass, so you are not choosing between a countable answer and an explainable one. - Let the AI ask the follow-up you would have asked. Koji's interviewer probes automatically on the non-JAR answers: you said too frequent - which ones, and when? That single probe converts a direction into a fix, and it happens on every respondent rather than on the eight you had time to schedule calls with.
- Get the segment split without pre-planning it. Polarised JAR distributions are the signal that two populations are hiding in one average. Because Koji analyzes every transcript rather than sampling them, the two camps show up as distinct themes in the report rather than as an unexplained bimodal bar.
- Voice or text, one study definition. JAR items work in both Koji modes; voice interviews cost 3 credits and text interviews 1, so a directional-diagnostic study can run cheaply at text scale and be deepened selectively.
- No moderator, no scheduling. The reason most teams never chase the why behind a JAR distribution is that it costs a round of calls. Koji removes that step, which is what makes running the diagnostic on every respondent realistic rather than aspirational.
The comparison against a form builder is not about the widget. SurveyMonkey, Typeform and Qualtrics will all render five radio buttons. None of them will notice that 40% of your respondents said too much, ask each of them why, and hand you the reason with the number.
The inversion worth remembering
An anchored, well-built rating scale - the kind described in the attribute lexicon guide - is designed so that more of the attribute means a higher number, and the average of those numbers means something. A JAR scale is the case where that entire apparatus inverts. More is not better, the midpoint is the target, and the average is at its most reassuring precisely when your users are at their most divided.
Frequently asked questions
Should a JAR scale have 3, 5 or 7 points?
Five is the working default: two degrees of too little, just right, two degrees of too much. Three points loses the intensity information that makes the collapse meaningful, and seven invites false precision on a scale that will be collapsed to three groups anyway. The study cited above used a nine-point JAR variant, which is defensible for expert panels but heavy for consumer work.
Can I ever average a JAR scale?
Only if you have already established that the distribution is unimodal, and even then the mean adds nothing that percent-just-right does not tell you more clearly. The one legitimate use of the numeric values is as an input to penalty analysis, where they are collapsed to three groups first and the arithmetic is done on liking scores rather than on the JAR values themselves.
How is this different from an importance-performance gap?
An importance-performance analysis compares how much an attribute matters against how well you deliver it, and both of its axes run from low to high. A JAR scale has no high-is-good axis at all; it measures signed distance from a target. The two answer different questions, and importance-performance analysis is the right tool when your attributes have maxima rather than optima.
What if most people pick just about right for everything?
That is usually a symptom, not a result. Check whether the attribute genuinely has an optimum, whether respondents have used the feature recently enough to have a view, and whether the middle option is doing the work of no opinion. A high just-right rate on an attribute users have never noticed tells you about their attention, not your product.
Does a JAR scale suffer from acquiescence bias?
Less than an agree/disagree item, because there is no agreeable end to drift toward - both extremes are complaints. That is one of the format's real advantages, and it is why JAR items are worth considering when acquiescence bias is a live concern. It does carry its own midpoint pull, which is the tendency to pick the safe middle box.
Can I ask JAR questions in a voice interview?
Yes. Koji's AI interviewer reads the five options conversationally and records the structured answer, so voice and text responses aggregate into the same distribution. The advantage of voice here is that the follow-up probe on a non-JAR answer tends to produce a longer, more specific explanation than a typed one.
Related Resources
- Structured Questions in AI Interviews - the six question types and how JAR items map onto them
- Penalty Analysis - turning JAR distributions into a ranked, costed fix list
- Attribute Lexicons and Reference Anchors - making sure every rater means the same thing before you scale anything
- Likert Scale Questions - the right instrument for attributes with a maximum rather than an optimum
- Importance-Performance Analysis - prioritising attributes that run from low to high
- Survey Sample Size - sizing a study around proportions rather than means
Related Articles
Importance-Performance Analysis (IPA): The Priority Matrix Guide (2026)
How to run an importance-performance analysis: plot attribute importance against performance to find your fix-first priorities, avoid over-investing, and turn survey data into a decision.
Likert Scale Questions: How to Use Rating Scales in User Research
A complete guide to Likert scale questions in user research — what they are, when to use them, how to write them correctly, and how Koji's AI interviews take rating scales further by pairing quantitative scores with qualitative follow-up.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)
A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.
5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)
A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.
The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)
A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.