{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-22T10:57:23.787Z"},"content":[{"type":"documentation","id":"bad10f9a-8e0c-4afc-92d7-eda7c75b3657","slug":"authority-control-research-repository","title":"Authority Control for Research Repositories: One Entity, One Name (2026)","url":"https://www.koji.so/docs/authority-control-research-repository","summary":"Authority control gives each entity in a research repository one authorized name, a list of variant strings that point to it, and a recorded justification. It prevents scattering (one entity split across tags, which undercounts) and conflation (different entities merged under one string, which overcounts). It differs from a taxonomy, which defines categories. Closed-option questions at collection are the strongest form of control.","content":"**Authority control is the practice of giving every entity in your research repository one authorized name, recording every variant that should point to it, and writing down which name won and why.** Libraries have done this for over a century because a catalog that files the same author under five spellings is not five times as complete -- it is five times as wrong. Research repositories have exactly the same problem with competitors, features, integrations, personas and customer segments, and most teams only find out when a count that looked small turns out to be split across four tags.\n\nThis is not the same job as building a taxonomy. A taxonomy decides *what categories exist*. Authority control decides *what a single thing is called*, and which strings refer to it. You can have a perfect taxonomy and still undercount your biggest competitor by 60%.\n\n## What libraries mean by authority control\n\nIn library cataloging, authority control is the process of organizing records by using a single, distinct form of a name (the *heading*) or an identifier. Each name gets an **authority record** containing four things:\n\n- **The authorized access point.** The one form everything is filed under. The Library of Congress Name Authority File, for example, uses *Diana, Princess of Wales, 1961-1997*, not *Princess Diana*, *Lady Di* or *Diana Spencer*.\n- **See references.** Variant forms that are *not* used as headings but redirect to the authorized one. A catalog user who searches a variant is sent to the right place instead of an empty result.\n- **See also references.** Links between *different* authorized headings that are related -- a person and the organization they founded, a company and its former name.\n- **Justification.** The sources that support the choice of form. This is the field teams forget, and it is the one that stops the same argument happening again in six months.\n\nAt a larger scale, the Virtual International Authority File (VIAF) links national authority files so that one entity has one cluster across countries, and identifier schemes like ISNI and ORCID give people a number that survives any spelling. The principle is identical at every scale: **the name is a pointer, and pointers must resolve to exactly one thing.**\n\n## The two failures authority control prevents\n\n### Scattering: one entity, many strings\n\nThis is the one that bends your numbers. Suppose a year of win-loss interviews mentions a competitor in five forms:\n\n| Tag in the repository | Mentions |\n|---|---|\n| Salesforce | 14 |\n| SFDC | 9 |\n| SF | 6 |\n| salesforce.com | 4 |\n| Sales Cloud | 2 |\n| **Actual total** | **35** |\n\nA second competitor, always called by one name, has 20 mentions. The dashboard shows it at the top with 20, and Salesforce second with 14. **The ranking is inverted, and nothing in the repository is technically wrong** -- every tag is a faithful record of what someone typed. The error only exists at the level of the entity, and no one is looking at that level.\n\nScattering compounds with time. Every new researcher, every new AI tagging run and every new integration adds a variant. A repository that is two years old without authority control does not have a few stray duplicates; it has a systematic bias toward entities with short, unambiguous names.\n\n### Conflation: one string, many entities\n\nThe reverse failure. *SF* in the table above might mean Salesforce in a CRM interview and San Francisco in a location question. *Mercury* might be a bank, a design tool or a planet in a space-tech study. *Pipeline* might be a sales pipeline or a data pipeline in the same B2B account. Merging on the string silently adds unrelated evidence to the count.\n\nLibraries handle this by qualifying the heading (dates, places, roles) until it is unique. You need the same: *Mercury (banking)*, *Pipeline (sales)*, *Pipeline (data)*.\n\nScattering undercounts. Conflation overcounts. **They can happen to the same entity at once**, which is why a total that happens to look right is not evidence that either problem is absent.\n\n## What an authority record looks like for a research team\n\nYou do not need MARC fields. You need a small table with six columns:\n\n| Field | Example |\n|---|---|\n| Authorized name | Salesforce (CRM) |\n| Entity type | Competitor |\n| Variants (see from) | SFDC; SF (CRM context only); salesforce.com; Sales Cloud; Sales Force |\n| Related (see also) | Slack (acquired); Tableau (acquired) |\n| Scope note | Use for the CRM product and the company. Use *Slack* separately for Slack-specific evidence. |\n| Justification | Company's own product naming; chosen 2026-03 by research ops |\n\n### Choose the authorized form by a rule, not by vote\n\nLibraries pick the form most commonly found in the entity's own publications. For research, a workable rule is: **the name the entity uses for itself, qualified only when needed for uniqueness.** Writing the rule down matters more than which rule you pick, because the rule is what lets a new team member decide a new case without a meeting.\n\n### Establish on first use, not in a cleanup project\n\nThe cheapest moment to create an authority record is the first time an entity is tagged. The most expensive is the annual cleanup, when you are reconciling hundreds of variants with no memory of which ones meant what. Catalogers establish the heading when the first item arrives; treat new competitor or feature names the same way.\n\n### Keep the variants; do not delete them\n\nIt is tempting to merge variants by renaming every tag to the authorized form. Do that for the *display*, but keep the variant list. The variant list is what lets search find the entity when someone types *SFDC* next year, and what lets an automated tagger map a new string to an old entity. A catalog with see references deleted is a catalog where every variant search returns nothing.\n\n## Where authority control lives in the research workflow\n\n### At collection: closed options are authorized headings\n\nThe strongest form of authority control is preventing the variant at the point of collection. When you ask *which tools did you evaluate?* as a multiple-choice question with a fixed list, the options **are** the authorized headings. Every response arrives already controlled. Keep an *Other* field for the entities you did not anticipate, and treat every *Other* answer as a candidate for a new authority record.\n\n### At tagging: the variant list is the tagger's lookup\n\nWhether tagging is done by a person or an AI model, give it the authority list and have it map to authorized forms. Language models are consistent within a run and inconsistent across runs; without a list, one batch produces *Salesforce* and the next produces *SFDC* for the same speech.\n\n### At synthesis: count entities, not strings\n\nBefore any finding with a number in it -- *most-mentioned competitor*, *top-requested integration* -- run the count against the authority list. If the ranking changes, the old ranking was a spelling statistic.\n\n## Common mistakes\n\n- **Treating authority control as a taxonomy.** Categories and names are separate problems. You need both.\n- **Merging by string match.** You fix scattering and create conflation.\n- **Deleting variants after a merge.** Search and future tagging lose their map.\n- **No justification field.** The same *is it Sales Cloud or Salesforce?* debate recurs every quarter.\n- **One-time cleanup.** Variants arrive continuously; control has to be continuous too.\n- **Letting each researcher coin names.** A repository with five contributors and no authority file has five dialects.\n\n## How Koji supports controlled naming\n\nAuthority control is cheaper when fewer variants are created in the first place, and that is where the interview platform matters.\n\n- **Controlled options at the source.** Koji's [structured questions](/docs/structured-questions-guide) include six types -- open_ended, scale, single_choice, multiple_choice, ranking and yes_no. A single_choice or multiple_choice question with your authorized competitor or integration list collects clean entity names from every participant, with an open_ended follow-up for anything new.\n- **Company context for the AI interviewer.** Giving Koji your product, feature and competitor names up front helps the AI interviewer recognize your domain vocabulary and reflect it back consistently in follow-up questions, rather than introducing its own paraphrases.\n- **One analysis pass per study.** Koji's automatic analysis groups answers into themes across the whole study at once, which avoids the batch-to-batch drift you get when different people tag different weeks of interviews.\n- **Voice and text in one corpus.** Koji's voice and text interviews land in the same study and the same analysis, so a competitor named aloud and one typed into a chat box are counted together rather than in two separate tools with two naming habits.\n- **Export into your repository of record.** Koji's exports and integrations move themes and quotes into the tool where your authority list lives, so the naming rule is applied in one place.\n\nTraditional survey tools like Qualtrics or SurveyMonkey can enforce a closed list, but they stop there: the open-text answers are left for someone to tag by hand weeks later. With Koji, the closed list, the AI follow-up and the analysis happen in the same study, which is where naming drift is cheapest to stop.\n\n## Frequently asked questions\n\n### What is authority control in a research repository?\n\nIt is the practice of giving each entity -- a competitor, feature, integration or segment -- one authorized name, recording the variant strings that should point to it, and documenting why that form was chosen. It prevents the same entity from being counted under several tags, and different entities from being merged under one.\n\n### How is authority control different from a taxonomy?\n\nA taxonomy decides which categories exist and how they nest. Authority control decides what a single entity is called and which strings refer to it. A repository can have a well-designed taxonomy and still split one competitor across five tags, because that is a naming problem, not a categorization problem.\n\n### What are see and see also references?\n\nA see reference sends a user from a variant form to the authorized one -- from *SFDC* to *Salesforce (CRM)*. A see also reference links two different authorized headings that are related, such as a company and a product it acquired. Research teams need both: the first fixes scattering, the second preserves real relationships without merging entities.\n\n### Should I merge duplicate tags in my research repository?\n\nMerge the display, but keep the variant list. Deleting variants means future searches for the old strings return nothing and automated taggers lose the map from new strings to old entities. Also check each merge for conflation: two tags with the same string may refer to different things.\n\n### How does AI tagging affect tag duplication?\n\nAI taggers are usually consistent within a single run and inconsistent across runs, so they can create variants faster than people do. Supply the authority list as context and ask for authorized forms. Running analysis once across a whole study, as Koji does, also reduces drift between batches.\n\n### When should authority records be created?\n\nOn first use. The first time a new competitor or feature is tagged is the cheapest moment to choose its authorized name, write the scope note and record the justification. Cleanups done a year later have to reconstruct decisions no one remembers. Closed options in Koji structured questions make that first-use decision at the moment of collection.\n\n## Related Resources\n\n- [Structured questions in AI interviews](/docs/structured-questions-guide) -- closed options as authorized headings, with open_ended follow-ups for new entities.\n- [How to build a UX research repository](/docs/research-repository-guide) -- the repository structure authority control sits inside.\n- [Customer feedback categorization](/docs/customer-feedback-categorization) -- the taxonomy half of the problem.\n- [The archival bond in research repositories](/docs/archival-bond-research-quotes-context) -- keeping a quote attached to its context.\n- [Retrieval metrics for research repositories](/docs/research-repository-retrieval-metrics-known-item-test) -- testing whether people can find what was filed.\n- [AI auto-tagging for customer interviews](/docs/ai-auto-tagging-customer-interviews) -- where tagging drift starts.\n","category":"Research Operations","lastModified":"2026-09-22T03:26:50.848907+00:00","metaTitle":"Authority Control for Research Repositories","metaDescription":"One competitor tagged five ways can invert your rankings. How library-style authority control keeps research repository counts honest.","keywords":["authority control","research repository tags","tag normalization","controlled vocabulary","duplicate tags","research taxonomy","entity names"],"aiSummary":"Authority control gives each entity in a research repository one authorized name, a list of variant strings that point to it, and a recorded justification. It prevents scattering (one entity split across tags, which undercounts) and conflation (different entities merged under one string, which overcounts). It differs from a taxonomy, which defines categories. Closed-option questions at collection are the strongest form of control.","aiPrerequisites":["An existing research repository or tagging system"],"aiLearningOutcomes":["Distinguish authority control from taxonomy","Detect scattering and conflation in tag counts","Write a six-field authority record","Apply controlled names at collection, tagging and synthesis"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}