Back to docs
Research Operations

Authority Control for Research Repositories: One Entity, One Name (2026)

Why one competitor tagged five ways inverts your rankings, and how library-style authority control -- authorized names, see references, justification -- keeps repository counts honest.

Authority control is the practice of giving every entity in your research repository one authorized name, recording every variant that should point to it, and writing down which name won and why. Libraries have done this for over a century because a catalog that files the same author under five spellings is not five times as complete -- it is five times as wrong. Research repositories have exactly the same problem with competitors, features, integrations, personas and customer segments, and most teams only find out when a count that looked small turns out to be split across four tags.

This is not the same job as building a taxonomy. A taxonomy decides what categories exist. Authority control decides what a single thing is called, and which strings refer to it. You can have a perfect taxonomy and still undercount your biggest competitor by 60%.

What libraries mean by authority control

In library cataloging, authority control is the process of organizing records by using a single, distinct form of a name (the heading) or an identifier. Each name gets an authority record containing four things:

  • The authorized access point. The one form everything is filed under. The Library of Congress Name Authority File, for example, uses Diana, Princess of Wales, 1961-1997, not Princess Diana, Lady Di or Diana Spencer.
  • See references. Variant forms that are not used as headings but redirect to the authorized one. A catalog user who searches a variant is sent to the right place instead of an empty result.
  • See also references. Links between different authorized headings that are related -- a person and the organization they founded, a company and its former name.
  • Justification. The sources that support the choice of form. This is the field teams forget, and it is the one that stops the same argument happening again in six months.

At a larger scale, the Virtual International Authority File (VIAF) links national authority files so that one entity has one cluster across countries, and identifier schemes like ISNI and ORCID give people a number that survives any spelling. The principle is identical at every scale: the name is a pointer, and pointers must resolve to exactly one thing.

The two failures authority control prevents

Scattering: one entity, many strings

This is the one that bends your numbers. Suppose a year of win-loss interviews mentions a competitor in five forms:

Tag in the repositoryMentions
Salesforce14
SFDC9
SF6
salesforce.com4
Sales Cloud2
Actual total35

A second competitor, always called by one name, has 20 mentions. The dashboard shows it at the top with 20, and Salesforce second with 14. The ranking is inverted, and nothing in the repository is technically wrong -- every tag is a faithful record of what someone typed. The error only exists at the level of the entity, and no one is looking at that level.

Scattering compounds with time. Every new researcher, every new AI tagging run and every new integration adds a variant. A repository that is two years old without authority control does not have a few stray duplicates; it has a systematic bias toward entities with short, unambiguous names.

Conflation: one string, many entities

The reverse failure. SF in the table above might mean Salesforce in a CRM interview and San Francisco in a location question. Mercury might be a bank, a design tool or a planet in a space-tech study. Pipeline might be a sales pipeline or a data pipeline in the same B2B account. Merging on the string silently adds unrelated evidence to the count.

Libraries handle this by qualifying the heading (dates, places, roles) until it is unique. You need the same: Mercury (banking), Pipeline (sales), Pipeline (data).

Scattering undercounts. Conflation overcounts. They can happen to the same entity at once, which is why a total that happens to look right is not evidence that either problem is absent.

What an authority record looks like for a research team

You do not need MARC fields. You need a small table with six columns:

FieldExample
Authorized nameSalesforce (CRM)
Entity typeCompetitor
Variants (see from)SFDC; SF (CRM context only); salesforce.com; Sales Cloud; Sales Force
Related (see also)Slack (acquired); Tableau (acquired)
Scope noteUse for the CRM product and the company. Use Slack separately for Slack-specific evidence.
JustificationCompany's own product naming; chosen 2026-03 by research ops

Choose the authorized form by a rule, not by vote

Libraries pick the form most commonly found in the entity's own publications. For research, a workable rule is: the name the entity uses for itself, qualified only when needed for uniqueness. Writing the rule down matters more than which rule you pick, because the rule is what lets a new team member decide a new case without a meeting.

Establish on first use, not in a cleanup project

The cheapest moment to create an authority record is the first time an entity is tagged. The most expensive is the annual cleanup, when you are reconciling hundreds of variants with no memory of which ones meant what. Catalogers establish the heading when the first item arrives; treat new competitor or feature names the same way.

Keep the variants; do not delete them

It is tempting to merge variants by renaming every tag to the authorized form. Do that for the display, but keep the variant list. The variant list is what lets search find the entity when someone types SFDC next year, and what lets an automated tagger map a new string to an old entity. A catalog with see references deleted is a catalog where every variant search returns nothing.

Where authority control lives in the research workflow

At collection: closed options are authorized headings

The strongest form of authority control is preventing the variant at the point of collection. When you ask which tools did you evaluate? as a multiple-choice question with a fixed list, the options are the authorized headings. Every response arrives already controlled. Keep an Other field for the entities you did not anticipate, and treat every Other answer as a candidate for a new authority record.

At tagging: the variant list is the tagger's lookup

Whether tagging is done by a person or an AI model, give it the authority list and have it map to authorized forms. Language models are consistent within a run and inconsistent across runs; without a list, one batch produces Salesforce and the next produces SFDC for the same speech.

At synthesis: count entities, not strings

Before any finding with a number in it -- most-mentioned competitor, top-requested integration -- run the count against the authority list. If the ranking changes, the old ranking was a spelling statistic.

Common mistakes

  • Treating authority control as a taxonomy. Categories and names are separate problems. You need both.
  • Merging by string match. You fix scattering and create conflation.
  • Deleting variants after a merge. Search and future tagging lose their map.
  • No justification field. The same is it Sales Cloud or Salesforce? debate recurs every quarter.
  • One-time cleanup. Variants arrive continuously; control has to be continuous too.
  • Letting each researcher coin names. A repository with five contributors and no authority file has five dialects.

How Koji supports controlled naming

Authority control is cheaper when fewer variants are created in the first place, and that is where the interview platform matters.

  • Controlled options at the source. Koji's structured questions include six types -- open_ended, scale, single_choice, multiple_choice, ranking and yes_no. A single_choice or multiple_choice question with your authorized competitor or integration list collects clean entity names from every participant, with an open_ended follow-up for anything new.
  • Company context for the AI interviewer. Giving Koji your product, feature and competitor names up front helps the AI interviewer recognize your domain vocabulary and reflect it back consistently in follow-up questions, rather than introducing its own paraphrases.
  • One analysis pass per study. Koji's automatic analysis groups answers into themes across the whole study at once, which avoids the batch-to-batch drift you get when different people tag different weeks of interviews.
  • Voice and text in one corpus. Koji's voice and text interviews land in the same study and the same analysis, so a competitor named aloud and one typed into a chat box are counted together rather than in two separate tools with two naming habits.
  • Export into your repository of record. Koji's exports and integrations move themes and quotes into the tool where your authority list lives, so the naming rule is applied in one place.

Traditional survey tools like Qualtrics or SurveyMonkey can enforce a closed list, but they stop there: the open-text answers are left for someone to tag by hand weeks later. With Koji, the closed list, the AI follow-up and the analysis happen in the same study, which is where naming drift is cheapest to stop.

Frequently asked questions

What is authority control in a research repository?

It is the practice of giving each entity -- a competitor, feature, integration or segment -- one authorized name, recording the variant strings that should point to it, and documenting why that form was chosen. It prevents the same entity from being counted under several tags, and different entities from being merged under one.

How is authority control different from a taxonomy?

A taxonomy decides which categories exist and how they nest. Authority control decides what a single entity is called and which strings refer to it. A repository can have a well-designed taxonomy and still split one competitor across five tags, because that is a naming problem, not a categorization problem.

What are see and see also references?

A see reference sends a user from a variant form to the authorized one -- from SFDC to Salesforce (CRM). A see also reference links two different authorized headings that are related, such as a company and a product it acquired. Research teams need both: the first fixes scattering, the second preserves real relationships without merging entities.

Should I merge duplicate tags in my research repository?

Merge the display, but keep the variant list. Deleting variants means future searches for the old strings return nothing and automated taggers lose the map from new strings to old entities. Also check each merge for conflation: two tags with the same string may refer to different things.

How does AI tagging affect tag duplication?

AI taggers are usually consistent within a single run and inconsistent across runs, so they can create variants faster than people do. Supply the authority list as context and ask for authorized forms. Running analysis once across a whole study, as Koji does, also reduces drift between batches.

When should authority records be created?

On first use. The first time a new competitor or feature is tagged is the cheapest moment to choose its authorized name, write the scope note and record the justification. Cleanups done a year later have to reconstruct decisions no one remembers. Closed options in Koji structured questions make that first-use decision at the moment of collection.

Related Resources

Related Articles

AI Auto-Tagging for Customer Interviews: Code 100 Interviews in Minutes

How AI auto-tagging compresses 40+ hours of manual qualitative coding into minutes. Covers the two-cycle coding approach Koji uses (descriptive cycle-1 + axial cycle-2), the difference between auto-tagging and thematic analysis, building a codebook the AI respects, and how to validate AI-generated tags against your standards.

Why a Tagged Quote Is Not Evidence: The Archival Bond in Research Repositories

A quote filed under a theme is information; the interview it came from is evidence. The five bonds repositories delete at intake, and how to keep them.

Customer Feedback Categorization: How to Build a Feedback Taxonomy That Scales (2026)

A practical guide to categorizing customer feedback: designing a feedback taxonomy, choosing flat vs. hierarchical tags, avoiding tag sprawl, and using AI auto-tagging to turn thousands of unstructured comments into quantified themes.

How to Build a UX Research Repository: The Complete Guide

A research repository transforms scattered insights into a searchable organizational asset. Learn how to build one that teams actually use.

Is Your Research Repository Working? The Retrieval Metrics Nobody Collects

Four cheap metrics that measure whether colleagues can actually reach your research, and the failure mode that no governance audit can detect.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.