{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-25T08:47:20.174Z"},"content":[{"type":"documentation","id":"49e36aef-797f-4f6f-bb00-3bc286e2735a","slug":"penalty-analysis-jar-fix-list","title":"Penalty Analysis: Turning Just-About-Right Data Into a Ranked Fix List (2026)","url":"https://www.koji.so/docs/penalty-analysis-jar-fix-list","summary":"Penalty analysis multiplies the mean drop in overall liking among off-target respondents by the share of respondents who are off target, producing the liking each attribute costs. Ranking by penalty disagrees with ranking by complaint volume: in the worked example, 74 percent of respondents flagged an attribute whose penalty was not statistically distinguishable from zero.","content":"You have just-about-right data on eight attributes. Six of them have a substantial group of users saying *too much* or *too little*. Your roadmap has room for one. Which do you fix?\n\nThe instinct is to fix the loudest complaint - the attribute with the biggest off-target group. That instinct is wrong often enough to be dangerous, because the number of people who noticed something is not the same as the amount of damage it did. Penalty analysis is the arithmetic that separates the two, and it produces a ranked fix list with a cost attached to each row.\n\n## The answer, up front\n\nPenalty analysis combines two numbers per attribute: the **mean drop** (how much lower overall liking runs among people who said the attribute was off-target, compared with people who said it was just right) and the **incidence** (what share of respondents said it was off-target). Multiply the direction-weighted mean drop by the incidence and you have the expected liking you are leaving on the table. Rank by that, not by incidence. In practice the two rankings disagree, and the attribute with 74% of users complaining can be worth less than the attribute with 22%.\n\nYou need two things to run it: a JAR item per attribute, and one overall liking score on a normal scale. If you have only the JAR items you can see direction but never cost.\n\n## The method, in three steps\n\nThe procedure is standard in sensory product development, where it is used to decide what to change in a reformulation. Iserliyska, Dzhivoderova and Nikovska set it out in three steps in *Current Trends in Natural Sciences* (volume 6, issue 11, 2017):\n\n1. **Collapse.** \"the JAR values are amalgamated into three groups\" - the too-little points, just-right, and the too-much points.\n2. **Compute the drops.** \"The mean overall liking (rating) is calculated for each group. The penalties (or mean drops) are calculated as the differences between the means of the two non-JAR categories and the mean of the JAR category.\"\n3. **Plot against incidence.** \"These values are plotted versus the percentage giving each response in a so called mean drop plot.\"\n\nStep 2 is where the cost appears. The mean drop is a difference between two group means on the *liking* scale, not on the JAR scale - which is exactly why you need the overall liking question. The overall penalty per attribute is then the incidence-weighted average of the two directional drops.\n\n## A worked example you can check\n\nThe same paper fields six commercial orange juices with 81 consumers, overall liking on a 9-point hedonic scale and JAR items on color, sweet taste, sour taste, bitter taste and amount of pulp. This is its penalty table for one product:\n\n| Attribute | Level | % | Mean overall liking | Mean drop | Penalty | p-value |\n| --- | --- | --- | --- | --- | --- | --- |\n| Color | not enough | 11.11% | 3.11 | 0.40 | 0.18 | 0.749 |\n| | JAR | 58.02% | 3.51 | | | |\n| | too much | 30.86% | 3.40 | 0.11 | | |\n| Sweet taste | not enough | 67.90% | 2.96 | 2.60 | 2.58 | 0.000 |\n| | JAR | 17.28% | 5.57 | | | |\n| | too much | 14.81% | 3.08 | 2.48 | | |\n| Sour taste | not enough | 25.93% | 2.85 | 2.02 | 1.83 | 0.008 |\n| | JAR | 20.99% | 4.88 | | | |\n| | too much | 53.09% | 3.14 | 1.74 | | |\n| Bitter taste | not enough | 11.11% | 3.44 | 2.66 | 3.01 | 0.001 |\n| | JAR | 11.11% | 6.11 | | | |\n| | too much | 77.78% | 3.04 | 3.06 | | |\n| Amount of pulp | not enough | 62.96% | 3.21 | 0.92 | 0.96 | 0.143 |\n| | JAR | 25.93% | 4.14 | | | |\n| | too much | 11.11% | 3.00 | 1.14 | | |\n\nTwo things are worth doing with a table like this before you trust it.\n\n**Check that it closes.** Each attribute partitions the same 81 people, so the liking totals must agree across attributes. They do: every one of the five attributes sums to 278 liking points across its three groups, which puts the product's overall liking mean at 278 divided by 81, or **3.43 on a 9-point scale** - below the scale midpoint, and the reason this product is a reformulation candidate at all.\n\n**Recompute the penalties.** Converting the published percentages back to counts and applying the incidence-weighted formula reproduces every published penalty to within 0.02 - bitter taste at 3.01, sweet at 2.58 or 2.59, sour at 1.83, pulp at 0.96, color at 0.18. The arithmetic is not a black box, and if your own tool disagrees with a hand calculation on your data, the tool is wrong.\n\n## Why incidence alone ranks wrongly\n\nNow put the two candidate rankings side by side.\n\n| Attribute | Off-target incidence | Penalty | Expected liking recovered | Significant? |\n| --- | --- | --- | --- | --- |\n| Bitter taste | 88.89% | 3.01 | 2.68 | yes (p = 0.001) |\n| Sweet taste | 82.72% | 2.58 | 2.13 | yes (p = 0.000) |\n| Sour taste | 79.01% | 1.83 | 1.45 | yes (p = 0.008) |\n| Amount of pulp | 74.07% | 0.96 | 0.71 | no (p = 0.143) |\n| Color | 41.98% | 0.18 | 0.08 | no (p = 0.749) |\n\nThe pulp row is the lesson. **Nearly three quarters of respondents said the pulp level was wrong, and the evidence does not support spending anything on it.** The people who complained about pulp liked the juice about as much as the people who did not, the difference is not distinguishable from zero at conventional thresholds, and the expected recovery is under a point of liking. Color is the same story in a milder form: 42% off-target, effectively no cost.\n\nA prioritization built on *how many people mentioned it* would have put pulp fourth out of five and treated it as a real problem. A prioritization built on penalty puts it below the action line. This is the same trap in a different costume as ranking feature requests by mention count.\n\nThe conventional action line here is incidence-based rather than penalty-based, and it comes from the mean-drop plot: the paper divides the plot with \"a vertical line representing 20% of the consumers\", and treats the upper-right region - high incidence *and* high penalty - as the attributes \"which have to be emphasized during the product development\". Both conditions, not either.\n\n## The ceiling on any fix, and where it comes from\n\nThere is a clean upper bound hiding in this table, and it is worth stating because teams routinely over-promise on the back of a penalty analysis.\n\nIf you fixed bitterness perfectly, every off-target respondent would move into the just-right group and, at best, would then like the product as much as that group already does. The product mean would rise from 3.43 to 3.43 plus 2.68, which is **6.11** - and 6.11 is precisely the mean liking of the people who already said the bitterness was just right. That identity is not a coincidence; it is what the arithmetic says. **The ceiling on any penalty fix is the liking score of the people who are already happy with that attribute.**\n\nWhich means the ceiling inherits all the fragility of that group's mean - and here that group is nine people out of 81. One respondent in a nine-person cell moves its mean by 0.111, so three respondents rating one point differently would move the entire projected ceiling by a third of a point. The mean drop is a difference between two group means, and a difference between two means is far less stable than either mean on its own; when one of the two cells is small, the instability lands squarely on your headline number. That amplification is worked through in detail in [error propagation in derived metrics](/docs/error-propagation-derived-research-metrics) and [catastrophic cancellation in metric differences](/docs/catastrophic-cancellation-metric-differences).\n\nThe practical rule: report the size of the JAR cell next to every penalty. A penalty computed against a just-right group of fewer than about 30 respondents is a direction, not an estimate.\n\n## A reporting template that survives scrutiny\n\nFor each attribute, five fields:\n\n- **Percent just about right** - the plain-language health number\n- **Percent too little / percent too much** - never netted\n- **Mean drop in each direction** - on the overall liking scale\n- **Weighted penalty and its p-value** - the cost, with its confidence\n- **n in the JAR cell** - the fragility disclosure\n\nThen one ranked list by expected recovery, with an explicit line under the attributes that clear both the 20% incidence bar and statistical significance. Everything below the line is documented, not scheduled.\n\n## Running penalty analysis in Koji\n\nPenalty analysis has historically been a two-tool workflow: a survey platform to collect JAR and liking data, then a statistics package to collapse, compute and plot. The collection half was never the hard part - the hard part is that the output tells you *which* attribute costs you liking and never *why* it does.\n\n- **Collect both halves in one study.** A `scale` question carries overall liking; a `single_choice` question carries each JAR item; `multiple_choice` captures the usage contexts you will want to cut by; `ranking` forces respondents to prioritize among the off-target attributes in their own words; `yes_no` screens for exposure so you do not compute a penalty from people who never encountered the attribute. The six types are documented in the [structured questions guide](/docs/structured-questions-guide).\n- **Attach an `open_ended` probe to each JAR item.** Penalty analysis tells you which attribute costs you liking; the open channel is where the reason lives. Koji collects both at once instead of leaving the second half to a follow-up study.\n- **Get the cause with the cost.** When a respondent lands in a non-JAR group, Koji's AI interviewer follows up on that specific answer. So the report does not just say bitterness carries a 3.01 penalty; it carries the verbatim reasons from the 63 people who said *too much*, thematically grouped. A survey tool gives you the first number and leaves the second half of the job to a round of follow-up interviews you probably will not schedule.\n- **Let the analysis run itself.** Koji aggregates structured answers into distributions automatically and generates the report, so the collapse-and-compare step is not a manual export. Refreshing a report costs 5 credits; a text interview costs 1 and a voice interview 3, which makes an 80-respondent diagnostic study genuinely routine rather than a quarterly event.\n- **Watch the small cells.** Because Koji analyzes every transcript rather than a sample of them, a thin just-right group shows up as a thin group in the report rather than as an over-confident average.\n- **Re-run it as the product changes.** A penalty table is a snapshot of one build. Because a Koji study can be re-fielded without re-recruiting a moderated panel, the same instrument can be run after each reformulation to confirm the penalty actually fell.\n\nAgainst SurveyMonkey, Typeform or Qualtrics the difference is not the JAR widget or the crosstab. It is that penalty analysis identifies the attribute to fix and is structurally incapable of telling you what to change about it - and an AI-moderated interview closes that gap in the same pass, on every respondent, without a moderator.\n\n## Frequently asked questions\n\n### How many respondents does penalty analysis need?\n\nEnough that the smallest cell you will act on is stable. The binding constraint is not total sample but the just-right group, which can be small when a product is badly off target - in the worked example above it was nine people for bitterness. As a working floor, aim for 100 or more respondents per product and treat any attribute whose JAR cell falls below 30 as directional only. The [survey sample size guide](/docs/survey-sample-size-guide) covers the general case.\n\n### Can I use penalty analysis without an overall liking question?\n\nNo. The penalty is measured in units of overall liking, so without that question there is nothing to take the difference of. If you only have JAR items you can still report percent just-right and direction, which is genuinely useful, but you cannot rank by cost.\n\n### What if both directions carry a large penalty?\n\nThat is the polarized case, and it is a signal to segment rather than to compromise. When *too much* and *too little* both cost you real liking, moving the attribute toward the middle makes one group happier and the other unhappier, and the net may be close to zero. Look for a segmenting variable that separates the two camps, and consider making the attribute configurable instead of choosing a single value.\n\n### How does this differ from key driver analysis?\n\nKey driver analysis regresses attribute *ratings* on overall satisfaction to find which attributes move the outcome. Penalty analysis works on *signed distance from a target* and answers a narrower, more actionable question: what does being off-target on this attribute cost, and in which direction. They are complementary, and [key driver analysis](/docs/key-driver-analysis-guide) is the better tool when your attributes run from low to high rather than around an optimum.\n\n### Is the 20% line a real threshold or a convention?\n\nIt is a convention, and a sensible one, not a statistical test. It exists because a large mean drop among 3% of respondents is arithmetically real and commercially irrelevant. Treat it as a default you can move with a stated reason - a 10% group may well be worth acting on if it is a high-value segment.\n\n### Should I trust a penalty that is not statistically significant?\n\nTreat it as unproven rather than absent. The pulp attribute in the example returned p = 0.143 with a 0.96 mean drop, which means the data neither establishes the effect nor rules it out. The honest readout is *not supported by this study*, and if the attribute is cheap to fix you may fix it anyway - just do not present it as evidence-backed.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - collecting liking and JAR items in one study\n- [Just-About-Right Scales](/docs/just-about-right-scale-product-research) - designing the input data that penalty analysis consumes\n- [Key Driver Analysis](/docs/key-driver-analysis-guide) - the complementary method for attributes with a maximum\n- [Importance-Performance Analysis](/docs/importance-performance-analysis-guide) - a different priority matrix and when to prefer it\n- [Error Propagation in Derived Research Metrics](/docs/error-propagation-derived-research-metrics) - why a difference between two group means is fragile\n- [Survey Sample Size](/docs/survey-sample-size-guide) - sizing the study so the just-right cell is not the weak link","category":"Analysis & Synthesis","lastModified":"2026-08-25T03:29:32.906017+00:00","metaTitle":"Penalty Analysis: Rank Product Fixes by What They Cost You (2026) | Koji","metaDescription":"Penalty analysis ranks attributes by the liking they cost, not by complaint volume. A worked example, the 20 percent action line, and the ceiling on every fix.","keywords":["penalty analysis","mean drop analysis","JAR penalty analysis","prioritize product fixes","mean drop plot","product reformulation research","attribute prioritization"],"aiSummary":"Penalty analysis multiplies the mean drop in overall liking among off-target respondents by the share of respondents who are off target, producing the liking each attribute costs. Ranking by penalty disagrees with ranking by complaint volume: in the worked example, 74 percent of respondents flagged an attribute whose penalty was not statistically distinguishable from zero.","aiPrerequisites":["Just-about-right data on at least one attribute","An overall liking or satisfaction score from the same respondents"],"aiLearningOutcomes":["Collapse JAR responses into three groups and compute directional mean drops","Weight mean drops by incidence to produce a comparable penalty","Apply the 20 percent incidence line alongside statistical significance","Recognize when a small just-right cell makes a penalty unstable"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}