Back to docs
Research Methods

Attribute Lexicons and Reference Anchors: Getting Every Rater to Mean the Same Thing (2026)

How to pin rating-scale attributes to external reference anchors so two raters produce comparable numbers, with a convergence test and thresholds you can ship against.

Two people rate the same screen a 4 out of 7 for responsiveness. You have no idea whether they agree. One of them is comparing it to the slowest tool they use at work; the other is comparing it to their phone camera. The number is identical and the measurement is not, because the word was never pinned to anything outside their heads.

An attribute lexicon fixes that before you collect a single response. It is a closed list of attributes, each with a written definition and a reference anchor: something real, nameable and re-checkable that fixes what the scale is measured against. Sensory scientists have used this method for decades because they had no choice, and product teams almost never use it because it feels like overkill until the day two researchers report opposite results from the same study.

The answer, up front

If you plan to average, compare or trend an attribute rating, you need three things per attribute and not one: a term, a one-sentence definition, and an external reference anchor. Without the third, each rater scores against a private prototype, and your average silently combines measurements taken on different instruments. Build the lexicon first, test that raters converge on it, and only then field the study. In Koji you attach the definition and the anchor to the question itself, so every participant reads the same standard before they answer, and the AI interviewer can probe when someone appears to be using a different one.

What a lexicon entry actually contains

The discipline that formalized this is sensory science, where panels have to rate things like astringency and aftertaste in numbers that hold up across sessions, laboratories and years. A 2026 lexicon-development study by Han and Tsai in Foods (volume 15, issue 12, article 2158) is a clean worked example. The authors built what they call the BQ Lexicon v.0 for a hybrid grape varietal: 21 defined descriptors, each one carrying a category, a code, and a column the paper labels Sensory Reference Standards.

The contents of that column are the whole point. Sweetness is not defined as how sweet it seems; it is anchored to a sucrose solution at 24.0 g/L, prepared to a published international method. Sourness is anchored to a citric acid solution. The visual attribute is anchored to two named Pantone chips, 19-1629 TCX and 19-1522 TCX, so that the phrase red to purple stops being a matter of opinion and becomes a comparison against a physical card.

Three parts, then:

  1. The term. A single agreed word or short phrase. Not a synonym cluster.
  2. The definition. One sentence, written in terms of what is being judged, not how much of it is good.
  3. The reference anchor. A thing outside the rater that the term points at, which any rater can go and check.

The third part is the one product teams drop, and it is the only one that makes the other two enforceable.

What happens when the anchor is missing

The same study is unusually honest about its failures, which makes it more useful than a paper where everything worked. Several attributes did not reach agreement, and the authors name the reason directly: "Without a single, universal reference standard, panelists may rely on different internal prototypes."

They give a specific case. The attribute for herb notes ran into trouble because, as the paper puts it, "the herb (Her.E) attribute faced linguistic ambiguity in the Chinese context, where its semantic boundaries often overlap with vanilla, leading to divergence in scoring." Two raters used one word for two different things and the disagreement showed up as noise in a number.

The paper also lists one attribute whose reference column reads, plainly, "Diverse profile; no standards yet." That is the honest state of most product-research attribute lists: a word, a scale, and nothing behind it.

Look at what the missing anchor does to the data. For bitterness, the study reports a mean of 2.31 with a standard deviation of 3.07 - a coefficient of variation of 132.60%. The spread is larger than the average. A mean like that is not a measurement of the product; it is a record of the fact that the panel had not agreed what the word meant.

Building a lexicon for a product team

You are not rating wine, but the mechanics port directly.

Step 1: harvest the vocabulary from participants, not from the team. Run an open_ended round and collect the words real users reach for. Teams that skip this end up measuring internal jargon. Koji's AI interviewer probes free-form answers automatically, so a single study can surface the natural vocabulary at a scale that manual moderation cannot reach - see the guide to analyzing open-ended survey responses for how to work the raw material.

Step 2: collapse synonyms deliberately. Snappy, quick, fast, responsive and smooth are not five attributes. Decide which one survives and write down which words it absorbs, so a later reader knows what was folded in.

Step 3: define each surviving term in one sentence, with the evaluation stripped out. Perceived delay between tapping and visible response is a definition. How fast it feels, which matters a lot to users is a definition plus a conclusion, and the conclusion will leak into the ratings.

Step 4: anchor it. This is the step that does the work, and in software it is easier than in food. Anchors that hold up: a named build number, a specific competitor flow the rater is asked to run first, a recorded screen capture at a fixed frame rate, a shipped feature everyone has used. Rate responsiveness from 1 to 5, where 1 is the export flow in build 4.2 and 5 is opening a new tab in your browser is a scale two strangers can use the same way.

Step 5: check convergence before you field. Give a small group the same stimulus and the same lexicon, and look at how much they disagree.

The convergence test, with a real threshold

Most teams stop at step 4 and hope. The sensory literature gives you a cheap numeric gate instead. The Han and Tsai study classifies each attribute by coefficient of variation - the standard deviation divided by the mean, expressed as a percentage - with consensus at CV under 30%, low consensus at 30% to 70%, and a third category the authors call baseline noise for attributes with a mean under 3.0 and a CV above 70%.

Adopt those bands as they stand and you have a defensible rule for shipping a lexicon:

CV across raters on the same stimulusReadingWhat to do
Under 30%Raters agree on the wordField it
30% to 70%Partial agreementRewrite the definition or find a harder anchor
Over 70%, low meanThe word is not measuring anythingCut the attribute

That third row matters more than it looks. An attribute that nobody can rate consistently is not a weak signal to be averaged over more people; it is a broken instrument, and adding respondents makes the average tighter without making it truer. That is a different failure from ordinary sampling error, and it is the failure that measurement system analysis is built to detect.

Anchoring the scale, not just the word

A defined term still needs defined scale points. Two rules carry most of the weight.

Label the endpoints with referents, not intensifiers. Extremely responsive means whatever the rater wants. As responsive as opening a new browser tab does not.

Do not put the good end on the left every time. A rater who learns that the right-hand side is always the flattering answer stops reading. The related question of how many points to offer is covered in 5-point vs 7-point Likert scales; the lexicon question is orthogonal to it and comes first.

One more borrowing from the sensory protocol, and it is nearly free. In the orange juice study by Iserliyska, Dzhivoderova and Nikovska (Current Trends in Natural Sciences, volume 6, issue 11, 2017), samples were presented with "Packaging was separated from samples in order to avoid the effect of brand knowledge" labeled with three-digit random codes, and served under a balanced block design so that no product was always tasted first. The product-research equivalents are obvious once stated: strip the logo from the prototype, give each variant a neutral code, and rotate the order across participants.

Running this in Koji

The reason lexicons stay theoretical in most teams is that enforcing one used to require a moderator in the room. It does not any more.

  • Attach the standard to the question. A Koji scale question carries its own scale labels, so the anchor text sits in front of the participant at the moment they answer rather than in a document nobody opened.
  • Constrain the vocabulary where it should be constrained. Use single_choice or multiple_choice when the lexicon is settled, ranking when you need relative order across attributes, and yes_no for the gate questions that decide whether an attribute even applies. The full set of six question types is covered in the structured questions guide.
  • Keep an unconstrained channel open anyway. Pair every closed attribute with an open_ended follow-up. This is not politeness; it is the only defense against a perception with nowhere to go, which is a measurable distortion in its own right and the subject of the dumping effect.
  • Let the AI catch a rater using a private prototype. When a participant gives an extreme rating, Koji's follow-up asks what they were comparing it to. A traditional survey collects the 4 and moves on; that single probe is the difference between a number and a measurement.
  • Voice or text, same lexicon. Koji runs both modes against the same study definition, so a spoken interview and a typed one produce comparable attribute data instead of two incompatible datasets.

Against a form builder like SurveyMonkey, Typeform or Qualtrics, the gap is not the scale widget - everyone has scale widgets. It is that a static form cannot notice that a respondent has redefined your attribute mid-study, and an AI interviewer can.

What a lexicon does not fix

It does not fix a rater whose standard drifts over time, which is a separate problem covered in panel conditioning. It does not tell you which attributes matter to satisfaction - that is key driver analysis. And it does not make a directional attribute averageable: for anything with an optimum rather than a maximum, you need a different scale shape entirely, which is why just-about-right scales exist.

What it does fix is the thing that quietly invalidates everything downstream: two numbers that look comparable and are not.

Frequently asked questions

How many attributes should a lexicon contain?

Fewer than you want. The Han and Tsai study settled on 21 descriptors for a product category with a famously rich vocabulary, and then reduced further for the quantitative phase. For a software feature, 6 to 12 well-anchored attributes will outperform 30 vague ones, because every attribute you add competes for the respondent's attention and increases the chance that two of them overlap semantically.

Is this not just writing good survey questions?

It overlaps, but the anchor is the difference. Guidance on unbiased survey question wording tells you how to avoid leading or double-barreled phrasing, which is about the question. A lexicon is about the measurement standard behind the question - the external referent that makes a 4 from one person and a 4 from another the same quantity.

Can I build the lexicon from existing feedback instead of a new study?

Yes, and it is often the fastest route. Support tickets, reviews and past interview transcripts are a legitimate harvest source for step 1. Run them through thematic extraction to find the recurring vocabulary, then still run a small convergence check before fielding, because a word that appears often is not necessarily a word people use the same way.

What if raters disagree even with a reference anchor?

Then you have learned something real rather than something noisy. Genuine disagreement after anchoring usually means the attribute is composite - it is bundling two perceptions that different people weight differently - and the fix is to split it. The wine study hit exactly this with aftertaste, an attribute the authors describe as composite and temporal.

Do I need trained raters for this to work?

No, and for consumer-facing work you actively should not use them. A trained panel is more precise and less representative at the same time; their agreement is bought with experience your customers do not have. Use anchors to raise agreement among ordinary users instead of raising ordinary users into experts.

How often should a lexicon be revised?

Version it and revise it when the product category changes, not on a calendar. Note that the study above labels its output v.0 rather than v.1 - the authors treat a lexicon as a living artifact with a version number, which is the right instinct. Every revision breaks comparability with earlier waves, so record the date and the change alongside the data.

Related Resources

Related Articles

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.

Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement

Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.

Measurement System Analysis: How Much of Your Segment Difference Is the Instrument? (2026)

How to separate real variation between customers from variation created by measuring them. The intraclass correlation, the four classes of monitor, probable error, and how to run an honest R&R study on a research metric.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)

A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.

A/B Testing vs. User Research: When to Use Each (And When to Use Both)

Understand when A/B testing and qualitative user research each shine, and how to combine them for better product decisions. Includes framework for choosing methods, real case studies, and how AI interviews make mixed methods accessible.