Penalty Analysis: Turning Just-About-Right Data Into a Ranked Fix List (2026)
Mean drop times incidence gives you the liking each off-target attribute costs. A fully worked, reproducible example and the ceiling on any fix it implies.
You have just-about-right data on eight attributes. Six of them have a substantial group of users saying too much or too little. Your roadmap has room for one. Which do you fix?
The instinct is to fix the loudest complaint - the attribute with the biggest off-target group. That instinct is wrong often enough to be dangerous, because the number of people who noticed something is not the same as the amount of damage it did. Penalty analysis is the arithmetic that separates the two, and it produces a ranked fix list with a cost attached to each row.
The answer, up front
Penalty analysis combines two numbers per attribute: the mean drop (how much lower overall liking runs among people who said the attribute was off-target, compared with people who said it was just right) and the incidence (what share of respondents said it was off-target). Multiply the direction-weighted mean drop by the incidence and you have the expected liking you are leaving on the table. Rank by that, not by incidence. In practice the two rankings disagree, and the attribute with 74% of users complaining can be worth less than the attribute with 22%.
You need two things to run it: a JAR item per attribute, and one overall liking score on a normal scale. If you have only the JAR items you can see direction but never cost.
The method, in three steps
The procedure is standard in sensory product development, where it is used to decide what to change in a reformulation. Iserliyska, Dzhivoderova and Nikovska set it out in three steps in Current Trends in Natural Sciences (volume 6, issue 11, 2017):
- Collapse. "the JAR values are amalgamated into three groups" - the too-little points, just-right, and the too-much points.
- Compute the drops. "The mean overall liking (rating) is calculated for each group. The penalties (or mean drops) are calculated as the differences between the means of the two non-JAR categories and the mean of the JAR category."
- Plot against incidence. "These values are plotted versus the percentage giving each response in a so called mean drop plot."
Step 2 is where the cost appears. The mean drop is a difference between two group means on the liking scale, not on the JAR scale - which is exactly why you need the overall liking question. The overall penalty per attribute is then the incidence-weighted average of the two directional drops.
A worked example you can check
The same paper fields six commercial orange juices with 81 consumers, overall liking on a 9-point hedonic scale and JAR items on color, sweet taste, sour taste, bitter taste and amount of pulp. This is its penalty table for one product:
| Attribute | Level | % | Mean overall liking | Mean drop | Penalty | p-value |
|---|---|---|---|---|---|---|
| Color | not enough | 11.11% | 3.11 | 0.40 | 0.18 | 0.749 |
| JAR | 58.02% | 3.51 | ||||
| too much | 30.86% | 3.40 | 0.11 | |||
| Sweet taste | not enough | 67.90% | 2.96 | 2.60 | 2.58 | 0.000 |
| JAR | 17.28% | 5.57 | ||||
| too much | 14.81% | 3.08 | 2.48 | |||
| Sour taste | not enough | 25.93% | 2.85 | 2.02 | 1.83 | 0.008 |
| JAR | 20.99% | 4.88 | ||||
| too much | 53.09% | 3.14 | 1.74 | |||
| Bitter taste | not enough | 11.11% | 3.44 | 2.66 | 3.01 | 0.001 |
| JAR | 11.11% | 6.11 | ||||
| too much | 77.78% | 3.04 | 3.06 | |||
| Amount of pulp | not enough | 62.96% | 3.21 | 0.92 | 0.96 | 0.143 |
| JAR | 25.93% | 4.14 | ||||
| too much | 11.11% | 3.00 | 1.14 |
Two things are worth doing with a table like this before you trust it.
Check that it closes. Each attribute partitions the same 81 people, so the liking totals must agree across attributes. They do: every one of the five attributes sums to 278 liking points across its three groups, which puts the product's overall liking mean at 278 divided by 81, or 3.43 on a 9-point scale - below the scale midpoint, and the reason this product is a reformulation candidate at all.
Recompute the penalties. Converting the published percentages back to counts and applying the incidence-weighted formula reproduces every published penalty to within 0.02 - bitter taste at 3.01, sweet at 2.58 or 2.59, sour at 1.83, pulp at 0.96, color at 0.18. The arithmetic is not a black box, and if your own tool disagrees with a hand calculation on your data, the tool is wrong.
Why incidence alone ranks wrongly
Now put the two candidate rankings side by side.
| Attribute | Off-target incidence | Penalty | Expected liking recovered | Significant? |
|---|---|---|---|---|
| Bitter taste | 88.89% | 3.01 | 2.68 | yes (p = 0.001) |
| Sweet taste | 82.72% | 2.58 | 2.13 | yes (p = 0.000) |
| Sour taste | 79.01% | 1.83 | 1.45 | yes (p = 0.008) |
| Amount of pulp | 74.07% | 0.96 | 0.71 | no (p = 0.143) |
| Color | 41.98% | 0.18 | 0.08 | no (p = 0.749) |
The pulp row is the lesson. Nearly three quarters of respondents said the pulp level was wrong, and the evidence does not support spending anything on it. The people who complained about pulp liked the juice about as much as the people who did not, the difference is not distinguishable from zero at conventional thresholds, and the expected recovery is under a point of liking. Color is the same story in a milder form: 42% off-target, effectively no cost.
A prioritization built on how many people mentioned it would have put pulp fourth out of five and treated it as a real problem. A prioritization built on penalty puts it below the action line. This is the same trap in a different costume as ranking feature requests by mention count.
The conventional action line here is incidence-based rather than penalty-based, and it comes from the mean-drop plot: the paper divides the plot with "a vertical line representing 20% of the consumers", and treats the upper-right region - high incidence and high penalty - as the attributes "which have to be emphasized during the product development". Both conditions, not either.
The ceiling on any fix, and where it comes from
There is a clean upper bound hiding in this table, and it is worth stating because teams routinely over-promise on the back of a penalty analysis.
If you fixed bitterness perfectly, every off-target respondent would move into the just-right group and, at best, would then like the product as much as that group already does. The product mean would rise from 3.43 to 3.43 plus 2.68, which is 6.11 - and 6.11 is precisely the mean liking of the people who already said the bitterness was just right. That identity is not a coincidence; it is what the arithmetic says. The ceiling on any penalty fix is the liking score of the people who are already happy with that attribute.
Which means the ceiling inherits all the fragility of that group's mean - and here that group is nine people out of 81. One respondent in a nine-person cell moves its mean by 0.111, so three respondents rating one point differently would move the entire projected ceiling by a third of a point. The mean drop is a difference between two group means, and a difference between two means is far less stable than either mean on its own; when one of the two cells is small, the instability lands squarely on your headline number. That amplification is worked through in detail in error propagation in derived metrics and catastrophic cancellation in metric differences.
The practical rule: report the size of the JAR cell next to every penalty. A penalty computed against a just-right group of fewer than about 30 respondents is a direction, not an estimate.
A reporting template that survives scrutiny
For each attribute, five fields:
- Percent just about right - the plain-language health number
- Percent too little / percent too much - never netted
- Mean drop in each direction - on the overall liking scale
- Weighted penalty and its p-value - the cost, with its confidence
- n in the JAR cell - the fragility disclosure
Then one ranked list by expected recovery, with an explicit line under the attributes that clear both the 20% incidence bar and statistical significance. Everything below the line is documented, not scheduled.
Running penalty analysis in Koji
Penalty analysis has historically been a two-tool workflow: a survey platform to collect JAR and liking data, then a statistics package to collapse, compute and plot. The collection half was never the hard part - the hard part is that the output tells you which attribute costs you liking and never why it does.
- Collect both halves in one study. A
scalequestion carries overall liking; asingle_choicequestion carries each JAR item;multiple_choicecaptures the usage contexts you will want to cut by;rankingforces respondents to prioritize among the off-target attributes in their own words;yes_noscreens for exposure so you do not compute a penalty from people who never encountered the attribute. The six types are documented in the structured questions guide. - Attach an
open_endedprobe to each JAR item. Penalty analysis tells you which attribute costs you liking; the open channel is where the reason lives. Koji collects both at once instead of leaving the second half to a follow-up study. - Get the cause with the cost. When a respondent lands in a non-JAR group, Koji's AI interviewer follows up on that specific answer. So the report does not just say bitterness carries a 3.01 penalty; it carries the verbatim reasons from the 63 people who said too much, thematically grouped. A survey tool gives you the first number and leaves the second half of the job to a round of follow-up interviews you probably will not schedule.
- Let the analysis run itself. Koji aggregates structured answers into distributions automatically and generates the report, so the collapse-and-compare step is not a manual export. Refreshing a report costs 5 credits; a text interview costs 1 and a voice interview 3, which makes an 80-respondent diagnostic study genuinely routine rather than a quarterly event.
- Watch the small cells. Because Koji analyzes every transcript rather than a sample of them, a thin just-right group shows up as a thin group in the report rather than as an over-confident average.
- Re-run it as the product changes. A penalty table is a snapshot of one build. Because a Koji study can be re-fielded without re-recruiting a moderated panel, the same instrument can be run after each reformulation to confirm the penalty actually fell.
Against SurveyMonkey, Typeform or Qualtrics the difference is not the JAR widget or the crosstab. It is that penalty analysis identifies the attribute to fix and is structurally incapable of telling you what to change about it - and an AI-moderated interview closes that gap in the same pass, on every respondent, without a moderator.
Frequently asked questions
How many respondents does penalty analysis need?
Enough that the smallest cell you will act on is stable. The binding constraint is not total sample but the just-right group, which can be small when a product is badly off target - in the worked example above it was nine people for bitterness. As a working floor, aim for 100 or more respondents per product and treat any attribute whose JAR cell falls below 30 as directional only. The survey sample size guide covers the general case.
Can I use penalty analysis without an overall liking question?
No. The penalty is measured in units of overall liking, so without that question there is nothing to take the difference of. If you only have JAR items you can still report percent just-right and direction, which is genuinely useful, but you cannot rank by cost.
What if both directions carry a large penalty?
That is the polarized case, and it is a signal to segment rather than to compromise. When too much and too little both cost you real liking, moving the attribute toward the middle makes one group happier and the other unhappier, and the net may be close to zero. Look for a segmenting variable that separates the two camps, and consider making the attribute configurable instead of choosing a single value.
How does this differ from key driver analysis?
Key driver analysis regresses attribute ratings on overall satisfaction to find which attributes move the outcome. Penalty analysis works on signed distance from a target and answers a narrower, more actionable question: what does being off-target on this attribute cost, and in which direction. They are complementary, and key driver analysis is the better tool when your attributes run from low to high rather than around an optimum.
Is the 20% line a real threshold or a convention?
It is a convention, and a sensible one, not a statistical test. It exists because a large mean drop among 3% of respondents is arithmetically real and commercially irrelevant. Treat it as a default you can move with a stated reason - a 10% group may well be worth acting on if it is a high-value segment.
Should I trust a penalty that is not statistically significant?
Treat it as unproven rather than absent. The pulp attribute in the example returned p = 0.143 with a 0.96 mean drop, which means the data neither establishes the effect nor rules it out. The honest readout is not supported by this study, and if the attribute is cheap to fix you may fix it anyway - just do not present it as evidence-backed.
Related Resources
- Structured Questions in AI Interviews - collecting liking and JAR items in one study
- Just-About-Right Scales - designing the input data that penalty analysis consumes
- Key Driver Analysis - the complementary method for attributes with a maximum
- Importance-Performance Analysis - a different priority matrix and when to prefer it
- Error Propagation in Derived Research Metrics - why a difference between two group means is fragile
- Survey Sample Size - sizing the study so the just-right cell is not the weak link
Related Articles
Error Propagation in Research Metrics: What Happens to Uncertainty When You Combine Numbers (2026)
Averaging four sub-scores makes your number more precise. Subtracting two averages can make it meaningless. Both follow the same rule. Here is the rule, with worked examples.
Importance-Performance Analysis (IPA): The Priority Matrix Guide (2026)
How to run an importance-performance analysis: plot attribute importance against performance to find your fix-first priorities, avoid over-investing, and turn survey data into a decision.
Just-About-Right Scales: The Rating Question Whose Average Means Nothing (2026)
JAR scales measure signed distance from an optimum, so their mean is not a summary. How to field, report and interpret them without averaging away the answer.
Key Driver Analysis: How to Find What Actually Drives Customer Satisfaction
A complete guide to key driver analysis (KDA) — how to use correlation and regression to identify which factors most influence satisfaction, loyalty, and NPS, how to read an importance-performance matrix, and how AI shortens the path from data to decision.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)
A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.