{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-26T10:01:19.503Z"},"content":[{"type":"documentation","id":"5d78498e-7f40-459e-8cd1-3829fae640d3","slug":"attribute-lexicon-reference-anchors-research","title":"Attribute Lexicons and Reference Anchors: Getting Every Rater to Mean the Same Thing (2026)","url":"https://www.koji.so/docs/attribute-lexicon-reference-anchors-research","summary":"An attribute lexicon gives every rated attribute three parts: a term, a one-sentence definition, and an external reference anchor. Without the anchor, each rater scores against a private prototype and averages combine incompatible measurements. Build the lexicon, test convergence with a coefficient-of-variation threshold under 30 percent, then field the study.","content":"Two people rate the same screen a 4 out of 7 for responsiveness. You have no idea whether they agree. One of them is comparing it to the slowest tool they use at work; the other is comparing it to their phone camera. The number is identical and the measurement is not, because the word was never pinned to anything outside their heads.\n\nAn attribute lexicon fixes that before you collect a single response. It is a closed list of attributes, each with a written definition and a reference anchor: something real, nameable and re-checkable that fixes what the scale is measured against. Sensory scientists have used this method for decades because they had no choice, and product teams almost never use it because it feels like overkill until the day two researchers report opposite results from the same study.\n\n## The answer, up front\n\nIf you plan to average, compare or trend an attribute rating, you need three things per attribute and not one: a **term**, a one-sentence **definition**, and an external **reference anchor**. Without the third, each rater scores against a private prototype, and your average silently combines measurements taken on different instruments. Build the lexicon first, test that raters converge on it, and only then field the study. In Koji you attach the definition and the anchor to the question itself, so every participant reads the same standard before they answer, and the AI interviewer can probe when someone appears to be using a different one.\n\n## What a lexicon entry actually contains\n\nThe discipline that formalized this is sensory science, where panels have to rate things like astringency and aftertaste in numbers that hold up across sessions, laboratories and years. A 2026 lexicon-development study by Han and Tsai in *Foods* (volume 15, issue 12, article 2158) is a clean worked example. The authors built what they call the BQ Lexicon v.0 for a hybrid grape varietal: 21 defined descriptors, each one carrying a category, a code, and a column the paper labels Sensory Reference Standards.\n\nThe contents of that column are the whole point. Sweetness is not defined as *how sweet it seems*; it is anchored to a sucrose solution at 24.0 g/L, prepared to a published international method. Sourness is anchored to a citric acid solution. The visual attribute is anchored to two named Pantone chips, 19-1629 TCX and 19-1522 TCX, so that the phrase *red to purple* stops being a matter of opinion and becomes a comparison against a physical card.\n\nThree parts, then:\n\n1. **The term.** A single agreed word or short phrase. Not a synonym cluster.\n2. **The definition.** One sentence, written in terms of what is being judged, not how much of it is good.\n3. **The reference anchor.** A thing outside the rater that the term points at, which any rater can go and check.\n\nThe third part is the one product teams drop, and it is the only one that makes the other two enforceable.\n\n## What happens when the anchor is missing\n\nThe same study is unusually honest about its failures, which makes it more useful than a paper where everything worked. Several attributes did not reach agreement, and the authors name the reason directly: \"Without a single, universal reference standard, panelists may rely on different internal prototypes.\"\n\nThey give a specific case. The attribute for herb notes ran into trouble because, as the paper puts it, \"the herb (Her.E) attribute faced linguistic ambiguity in the Chinese context, where its semantic boundaries often overlap with vanilla, leading to divergence in scoring.\" Two raters used one word for two different things and the disagreement showed up as noise in a number.\n\nThe paper also lists one attribute whose reference column reads, plainly, \"Diverse profile; no standards yet.\" That is the honest state of most product-research attribute lists: a word, a scale, and nothing behind it.\n\nLook at what the missing anchor does to the data. For bitterness, the study reports a mean of 2.31 with a standard deviation of 3.07 - a coefficient of variation of 132.60%. The spread is larger than the average. A mean like that is not a measurement of the product; it is a record of the fact that the panel had not agreed what the word meant.\n\n## Building a lexicon for a product team\n\nYou are not rating wine, but the mechanics port directly.\n\n**Step 1: harvest the vocabulary from participants, not from the team.** Run an open_ended round and collect the words real users reach for. Teams that skip this end up measuring internal jargon. Koji's AI interviewer probes free-form answers automatically, so a single study can surface the natural vocabulary at a scale that manual moderation cannot reach - see the guide to [analyzing open-ended survey responses](/docs/ai-analyze-open-ended-survey-responses) for how to work the raw material.\n\n**Step 2: collapse synonyms deliberately.** Snappy, quick, fast, responsive and smooth are not five attributes. Decide which one survives and write down which words it absorbs, so a later reader knows what was folded in.\n\n**Step 3: define each surviving term in one sentence, with the evaluation stripped out.** *Perceived delay between tapping and visible response* is a definition. *How fast it feels, which matters a lot to users* is a definition plus a conclusion, and the conclusion will leak into the ratings.\n\n**Step 4: anchor it.** This is the step that does the work, and in software it is easier than in food. Anchors that hold up: a named build number, a specific competitor flow the rater is asked to run first, a recorded screen capture at a fixed frame rate, a shipped feature everyone has used. *Rate responsiveness from 1 to 5, where 1 is the export flow in build 4.2 and 5 is opening a new tab in your browser* is a scale two strangers can use the same way.\n\n**Step 5: check convergence before you field.** Give a small group the same stimulus and the same lexicon, and look at how much they disagree.\n\n## The convergence test, with a real threshold\n\nMost teams stop at step 4 and hope. The sensory literature gives you a cheap numeric gate instead. The Han and Tsai study classifies each attribute by coefficient of variation - the standard deviation divided by the mean, expressed as a percentage - with consensus at CV under 30%, low consensus at 30% to 70%, and a third category the authors call baseline noise for attributes with a mean under 3.0 and a CV above 70%.\n\nAdopt those bands as they stand and you have a defensible rule for shipping a lexicon:\n\n| CV across raters on the same stimulus | Reading | What to do |\n| --- | --- | --- |\n| Under 30% | Raters agree on the word | Field it |\n| 30% to 70% | Partial agreement | Rewrite the definition or find a harder anchor |\n| Over 70%, low mean | The word is not measuring anything | Cut the attribute |\n\nThat third row matters more than it looks. An attribute that nobody can rate consistently is not a weak signal to be averaged over more people; it is a broken instrument, and adding respondents makes the average tighter without making it truer. That is a different failure from ordinary sampling error, and it is the failure that [measurement system analysis](/docs/measurement-system-analysis-research-metrics) is built to detect.\n\n## Anchoring the scale, not just the word\n\nA defined term still needs defined scale points. Two rules carry most of the weight.\n\n**Label the endpoints with referents, not intensifiers.** *Extremely responsive* means whatever the rater wants. *As responsive as opening a new browser tab* does not.\n\n**Do not put the good end on the left every time.** A rater who learns that the right-hand side is always the flattering answer stops reading. The related question of how many points to offer is covered in [5-point vs 7-point Likert scales](/docs/5-point-vs-7-point-likert-scale); the lexicon question is orthogonal to it and comes first.\n\nOne more borrowing from the sensory protocol, and it is nearly free. In the orange juice study by Iserliyska, Dzhivoderova and Nikovska (*Current Trends in Natural Sciences*, volume 6, issue 11, 2017), samples were presented with \"Packaging was separated from samples in order to avoid the effect of brand knowledge\" labeled with three-digit random codes, and served under a balanced block design so that no product was always tasted first. The product-research equivalents are obvious once stated: strip the logo from the prototype, give each variant a neutral code, and rotate the order across participants.\n\n## Running this in Koji\n\nThe reason lexicons stay theoretical in most teams is that enforcing one used to require a moderator in the room. It does not any more.\n\n- **Attach the standard to the question.** A Koji `scale` question carries its own scale labels, so the anchor text sits in front of the participant at the moment they answer rather than in a document nobody opened.\n- **Constrain the vocabulary where it should be constrained.** Use `single_choice` or `multiple_choice` when the lexicon is settled, `ranking` when you need relative order across attributes, and `yes_no` for the gate questions that decide whether an attribute even applies. The full set of six question types is covered in the [structured questions guide](/docs/structured-questions-guide).\n- **Keep an unconstrained channel open anyway.** Pair every closed attribute with an `open_ended` follow-up. This is not politeness; it is the only defense against a perception with nowhere to go, which is a measurable distortion in its own right and the subject of [the dumping effect](/docs/dumping-effect-attribute-scales-research).\n- **Let the AI catch a rater using a private prototype.** When a participant gives an extreme rating, Koji's follow-up asks what they were comparing it to. A traditional survey collects the 4 and moves on; that single probe is the difference between a number and a measurement.\n- **Voice or text, same lexicon.** Koji runs both modes against the same study definition, so a spoken interview and a typed one produce comparable attribute data instead of two incompatible datasets.\n\nAgainst a form builder like SurveyMonkey, Typeform or Qualtrics, the gap is not the scale widget - everyone has scale widgets. It is that a static form cannot notice that a respondent has redefined your attribute mid-study, and an AI interviewer can.\n\n## What a lexicon does not fix\n\nIt does not fix a rater whose standard drifts over time, which is a separate problem covered in [panel conditioning](/docs/panel-conditioning-repeat-participants). It does not tell you which attributes matter to satisfaction - that is [key driver analysis](/docs/key-driver-analysis-guide). And it does not make a directional attribute averageable: for anything with an optimum rather than a maximum, you need a different scale shape entirely, which is why [just-about-right scales](/docs/just-about-right-scale-product-research) exist.\n\nWhat it does fix is the thing that quietly invalidates everything downstream: two numbers that look comparable and are not.\n\n## Frequently asked questions\n\n### How many attributes should a lexicon contain?\n\nFewer than you want. The Han and Tsai study settled on 21 descriptors for a product category with a famously rich vocabulary, and then reduced further for the quantitative phase. For a software feature, 6 to 12 well-anchored attributes will outperform 30 vague ones, because every attribute you add competes for the respondent's attention and increases the chance that two of them overlap semantically.\n\n### Is this not just writing good survey questions?\n\nIt overlaps, but the anchor is the difference. Guidance on [unbiased survey question wording](/docs/survey-question-wording-guide) tells you how to avoid leading or double-barreled phrasing, which is about the question. A lexicon is about the measurement standard behind the question - the external referent that makes a 4 from one person and a 4 from another the same quantity.\n\n### Can I build the lexicon from existing feedback instead of a new study?\n\nYes, and it is often the fastest route. Support tickets, reviews and past interview transcripts are a legitimate harvest source for step 1. Run them through thematic extraction to find the recurring vocabulary, then still run a small convergence check before fielding, because a word that appears often is not necessarily a word people use the same way.\n\n### What if raters disagree even with a reference anchor?\n\nThen you have learned something real rather than something noisy. Genuine disagreement after anchoring usually means the attribute is composite - it is bundling two perceptions that different people weight differently - and the fix is to split it. The wine study hit exactly this with aftertaste, an attribute the authors describe as composite and temporal.\n\n### Do I need trained raters for this to work?\n\nNo, and for consumer-facing work you actively should not use them. A trained panel is more precise and less representative at the same time; their agreement is bought with experience your customers do not have. Use anchors to raise agreement among ordinary users instead of raising ordinary users into experts.\n\n### How often should a lexicon be revised?\n\nVersion it and revise it when the product category changes, not on a calendar. Note that the study above labels its output v.0 rather than v.1 - the authors treat a lexicon as a living artifact with a version number, which is the right instinct. Every revision breaks comparability with earlier waves, so record the date and the change alongside the data.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types and when each one is the right instrument\n- [Just-About-Right Scales](/docs/just-about-right-scale-product-research) - what to do when an attribute has an optimum rather than a maximum\n- [The Dumping Effect](/docs/dumping-effect-attribute-scales-research) - why the attributes you leave out change the scores of the ones you keep\n- [5-Point vs 7-Point Likert Scale](/docs/5-point-vs-7-point-likert-scale) - choosing the number of scale points once the wording is settled\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - measuring agreement after the fact, the complement to building it beforehand\n- [Measurement System Analysis](/docs/measurement-system-analysis-research-metrics) - how much of an observed difference is the instrument rather than the product","category":"Research Methods","lastModified":"2026-08-25T03:29:32.906017+00:00","metaTitle":"Attribute Lexicons and Reference Anchors for Rating Scales (2026) | Koji","metaDescription":"Build an attribute lexicon with reference anchors so every rater means the same thing: definitions, anchors, and a convergence test you can ship against.","keywords":["attribute lexicon","reference anchors rating scale","sensory lexicon","anchored rating scale","rating scale definitions","descriptive analysis attributes","rater agreement"],"aiSummary":"An attribute lexicon gives every rated attribute three parts: a term, a one-sentence definition, and an external reference anchor. Without the anchor, each rater scores against a private prototype and averages combine incompatible measurements. Build the lexicon, test convergence with a coefficient-of-variation threshold under 30 percent, then field the study.","aiPrerequisites":["Familiarity with rating scales in surveys or interviews","Basic understanding of means and standard deviations"],"aiLearningOutcomes":["Write a lexicon entry with a term, definition and reference anchor","Choose reference anchors that work for software and digital products","Test rater convergence using coefficient-of-variation bands","Attach lexicon standards to Koji questions so participants see them at answer time"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}