Back to blog
Research

The User Research Glossary (2026): 120 Terms for Buyers, PMs and Researchers

Every term you will meet in a research plan, a vendor quote or a sample provider invoice — methods, sampling, bias, statistics, metrics, commercial jargon like CPI, LOI and IR, and the new AI research vocabulary. Plain definitions, one line each.

Koji

Koji Team

Research · · 14 min read

Research has two vocabularies that rarely appear in the same glossary: the methods language researchers use, and the commercial language vendors and sample providers use. You need both. A PM who can define thematic analysis but not incidence rate will still be surprised by the invoice.

This glossary covers both, plus the AI research terms that only entered the language in the last two years. Definitions are one line each; where we have a full guide, the term links to it.

Methods and study types

Generative research — Research that explores an open problem space to discover needs, before a solution exists. Also called discovery or exploratory research.

Evaluative research — Research that tests a specific design, concept or product against a defined task or question.

Moderated research — A study where a human (or an AI moderator) is present, able to probe and follow up in real time.

Unmoderated research — A study participants complete alone, with no live moderator. Cheaper and faster; no follow-up questions. See moderated vs unmoderated research.

Usability test — A study where participants attempt realistic tasks so you can observe where the interface fails. See the usability testing guide.

Diary study — A longitudinal method where participants log experiences over days or weeks in their own context. See the diary study guide.

Card sorting — Participants group and label content items so you can learn their mental model of your information architecture. See the card sorting guide.

Tree testing — The inverse of card sorting: participants navigate a stripped-down hierarchy to find items, testing findability. See the tree testing guide.

Concept test — Exposing a proposition or design to target users before build, to measure appeal, comprehension and differentiation.

Jobs to Be Done (JTBD) — A framework that defines demand by the progress a customer is trying to make, not by demographics. See jobs to be done interviews.

Contextual inquiry — Observing and interviewing people while they work, in their own environment.

Focus group — A moderated group discussion, historically 6-10 people. Efficient for range of opinion, vulnerable to groupthink and dominant voices.

Guerrilla research — Fast, low-cost, low-ceremony research with convenient participants. See guerrilla user research.

Desk research — Synthesising existing sources — literature, analytics, support tickets, reviews — rather than collecting new data. See review mining.

Longitudinal study — Any design that measures the same people over time.

Wave — One measurement round in a tracking study. Wave-over-wave comparison needs far more sample than a single wave; see statistical power and MDE.

Triangulation — Corroborating a finding across multiple methods, sources or researchers to raise confidence. See triangulation in research.

Double Diamond — The discover-define-develop-deliver framing of a design process. See the Double Diamond guide.

Sampling and recruitment

Population — Everyone you want to draw conclusions about.

Sampling frame — The actual list you can draw from. The gap between frame and population is where coverage error lives.

Probability sample — Every member of the frame has a known, non-zero chance of selection. The only design that supports a formal margin of error.

Opt-in (non-probability) panel — Participants who volunteered. Cheap and fast; requires weighting and caution. See river sampling vs opt-in panels vs probability panels.

River sampling — Recruiting respondents in the moment from web traffic rather than from a standing panel.

Quota sampling — Filling predefined demographic cells until each is full. See the quota sampling guide.

Stratified sampling — Dividing the population into strata and sampling within each, to guarantee representation.

Purposive sampling — Deliberately selecting participants who fit criteria relevant to the research question. See the purposive sampling guide.

Snowball sampling — Existing participants refer others; used for hard-to-reach populations. See the snowball sampling guide.

Screener — The qualifying questionnaire that decides who enters your study. The single highest-leverage document in recruitment; see research screener questions.

Incidence rate (IR) — The percentage of people screened who qualify. A 10% IR means screening 10 people to get 1, and sample providers price accordingly.

Length of interview (LOI) — Expected survey or session duration in minutes. The second major input to sample pricing.

Cost per interview / cost per complete (CPI) — What a sample provider charges per qualified completed response. Driven mainly by IR and LOI. See how much survey sample costs.

Overquota — A respondent who qualifies but is turned away because their demographic cell is already full.

Professional respondent — Someone who takes surveys frequently enough that their answers stop resembling the market. See professional respondents and panel conditioning.

Panel conditioning — The distortion caused by repeated participation changing how someone answers.

ESOMAR 37 — The standard question set buyers use to audit an online sample provider. See how to choose a sample provider.

Incentive — Compensation for participation. See research participant incentives.

Saturation — The point where new interviews stop producing new themes. See data saturation in qualitative research.

Question design and scales

Open-ended question — A question inviting a free-text or spoken answer. The only question type that can surface something you did not think to ask.

Closed question — A question with a fixed answer set.

Likert scale — An agree-disagree rating scale, typically 5 or 7 points. See the Likert scale guide and 5-point vs 7-point.

Semantic differential — A scale anchored by opposing adjectives rather than agreement levels.

Matrix / grid question — Multiple items sharing one scale in a table. Efficient, and a well-documented driver of straightlining.

Straightlining — Selecting the same answer down a whole grid without reading. A core data-quality flag.

Double-barrelled question — One question asking two things at once, so the answer is uninterpretable.

Leading question — A question whose wording signals the expected answer.

Forced choice vs select-all — Asking each item individually rather than offering a check-all list. Pew found forced-choice formats produced endorsement rates 8 percentage points higher on average across 12 items, and up to 16 points higher — a larger effect than most order effects.

Question order effect — When answering one question changes how the next is answered. See question order bias.

Randomisation / rotation — Varying the order of items or options across respondents so position effects cancel out.

Attention check — An item designed to catch inattentive responding.

Structured questions — Fixed-format questions that produce chartable data. Koji supports six types — open_ended, scale, single_choice, multiple_choice, ranking and yes_no — inside a conversational interview; see the structured questions guide.

Bias and validity

Validity — Whether you measured what you meant to measure.

Reliability — Whether the same measurement repeats consistently.

Sampling bias — Systematic difference between who you studied and who you meant to study. See sampling bias.

Nonresponse bias — Distortion caused by who declines to answer. See nonresponse bias.

Survivorship bias — Studying only those who stayed, and missing everyone who left. See survivorship bias in customer research.

Social desirability bias — Answering to look good rather than truthfully. See social desirability bias.

Acquiescence bias — The tendency to agree regardless of content. See acquiescence bias.

Courtesy bias — Softening criticism to be polite to the researcher. See courtesy bias.

Demand characteristics — Participants inferring the study hypothesis and performing it. See demand characteristics.

Recall bias — Faulty memory distorting reported behaviour. See recall bias.

Central tendency bias — Clustering on the middle of a scale. See central tendency bias.

Extreme response bias — Habitually choosing scale endpoints. See extreme response bias.

Anchoring — An early number shaping later judgements. See anchoring bias.

Confirmation bias — Seeking and weighting evidence that supports what you already believe. See confirmation bias in user research.

Observer bias — The researcher expectations shaping what gets recorded. See observer bias.

Interviewer bias / moderator variance — Different moderators producing different answers to the same question. See interviewer bias.

Stated vs revealed preference — The gap between what people say they will do and what they do. See stated vs revealed preferences.

A fuller catalogue lives in the research bias guide and cognitive biases in user interviews.

Analysis and synthesis

Coding — Tagging segments of qualitative data with labels so patterns become countable. See how to code qualitative data.

Open, axial and selective coding — The three-phase grounded-theory coding progression. See open, axial and selective coding.

Thematic analysis — Identifying, refining and reporting patterns of meaning across a dataset. See the thematic analysis guide.

Affinity mapping — Physically or digitally clustering observations until themes emerge. See affinity mapping.

Inter-rater reliability — The degree to which two analysts independently assign the same codes.

Verbatim — A participant quote reproduced exactly.

Research repository — The searchable store of studies, transcripts, tags and findings. See the research repository guide.

Atomic research / nugget — Breaking findings into reusable evidence units linked to raw data.

Insight vs finding — A finding is what the data says; an insight is the non-obvious implication that changes a decision.

Metrics

NPS (Net Promoter Score) — Percentage of promoters (9-10) minus percentage of detractors (0-6) on an 11-point recommendation scale; ranges from -100 to +100. See the NPS survey guide and NPS benchmarks by industry.

Transactional vs relational NPS — Measured after a specific interaction versus about the relationship overall. See transactional vs relational NPS.

CSAT — Customer satisfaction, usually the percentage selecting the top scale points. See the CSAT survey guide.

CES — Customer Effort Score, measuring how easy an interaction felt. Comparison of all three: CSAT vs NPS vs CES.

SUS (System Usability Scale) — A 10-item questionnaire producing a 0-100 usability score, widely benchmarked against an average near 68. See the SUS guide.

eNPS — NPS applied to employees. See the eNPS guide.

Task success rate — The proportion of participants completing a task unaided.

Statistics you will be quoted

Margin of error (MoE) — The half-width of a confidence interval around a single estimate. At n=400 and 95% confidence it is about ±4.9 percentage points. See the margin of error guide.

Confidence level — How often the interval would contain the true value across repeated samples. Conventionally 95%.

Statistical power — The probability of detecting a real effect if one exists. Conventionally 80%.

Minimum detectable effect (MDE) — The smallest difference your design can reliably detect. Crucially, MDE is roughly 2.02 times the margin of error when comparing two independent equal waves at 95% confidence and 80% power — so n=400 per wave has a ±4.9pp MoE but a 9.9pp MDE. Detecting a 5-point change needs roughly 1,570 per wave. See statistical power and MDE.

Statistical significance — Whether an observed difference is unlikely under the null hypothesis. Not the same as importance.

Multiple comparisons — Running many tests inflates false positives: 20 segment tests at the 5% level give roughly a 64% chance of at least one false positive.

Weighting / raking — Adjusting a skewed sample toward known population benchmarks. See the survey weighting guide.

Design effect and effective sample size — Weighting reduces the precision your raw n implies; effective n is what actually determines your margin of error.

Sample size — How many responses you need. For qualitative work the question is different; see how many user interviews and the survey sample size guide.

Commercial and buying terms

This is the section most research glossaries skip, and the one that shows up on invoices.

Seat / licence — One named user. Zylo put average SaaS licence utilization at 54% in 2026 — roughly two seats bought per active user.

Credit — A prepaid unit consumed per action. Koji charges 1 credit per text response, 3 per voice interview, 5 per report refresh.

Overage — What you pay past your included allowance. Ask whether it is priced at parity with the plan rate or above it.

Shelfware — Licensed software nobody uses. The average organization wasted $19.8M a year on unused licences in Zylo 2026 data.

Metering unit — What the vendor actually counts: seats, studies, sessions, participant-minutes, responses, or a percentage of participant pay. It matters more than the sticker price — see per-seat vs usage-based pricing for research tools.

Platform fee — A percentage charged on top of participant rewards by recruitment marketplaces.

TCO (total cost of ownership) — Subscription plus incentives, recruitment, integration, and loaded researcher hours. See the user research cost calculator and researcher salary benchmarks.

MSA / order form — The master agreement and the document that actually defines your units and price. Definitions belong in the order form, not the marketing page.

DPA (Data Processing Agreement) — The GDPR Article 28 contract governing a vendor processing personal data on your behalf. See reviewing a research vendor DPA.

Sub-processor — A vendor your vendor uses. Every one is a party to your participant data.

Data residency — Where participant data is physically stored, and often a gating requirement in enterprise reviews. See enterprise security for AI research platforms.

Informed consent — The participant agreement to take part on a documented basis. See research consent form templates and interview recording consent laws.

AI research terms

AI-moderated interview — A conversational interview run by an AI that asks planned questions and generates its own follow-up probes, with no human moderator present. This removes moderator variance entirely rather than reducing it.

Synthetic users / AI personas — Model-generated respondents standing in for real people. Useful for rehearsal, not evidence; see AI personas vs real customer interviews.

Automatic thematic analysis — Model-driven clustering of open-ended responses into themes with linked verbatims.

Hallucination — A fluent, confident model output that is not grounded in the source data. In research tooling the mitigation is quote-level traceability back to transcripts.

Golden set / evaluation dataset — A curated set of examples with known correct answers, used to measure whether an AI feature is good enough to ship.

Model card — A structured disclosure document describing a model intended use, limitations and evaluation results. See AI model cards and user disclosure.

EU AI Act Article 50 — The transparency obligation, applicable since 2 August 2026, requiring people to be told they are interacting with an AI system, disclosed clearly at the latest at first interaction. An AI-moderated interview falls in scope, so the research instrument itself must disclose.

Automation bias / over-reliance — Accepting an AI output because it is confident rather than because it is right. See AI over-reliance and automation bias.

AI governance framework — The internal policy set covering how AI systems are approved, monitored and retired. See AI governance frameworks for research.

Frequently Asked Questions

What is the difference between incidence rate and response rate?

Incidence rate (IR) is the percentage of people screened who qualify for your study. Response rate is the percentage of people invited who respond at all. IR drives sample cost because a low IR means the provider screens many people per completed interview; response rate drives nonresponse bias.

What does CPI mean in market research?

CPI is cost per interview or cost per complete — what a sample provider charges for one qualified, completed response. It is driven mainly by incidence rate and length of interview, so the same study costs several times more with a 5% IR than a 50% IR.

How is minimum detectable effect different from margin of error?

Margin of error describes the precision of a single estimate; MDE describes the smallest change you could reliably detect between two measurements. For two independent equal waves at 95% confidence and 80% power, MDE is about 2.02 times the margin of error — so a n=400 wave with a ±4.9pp MoE can only reliably detect a change of about 9.9 points.

What is saturation in qualitative research?

Saturation is the point at which additional interviews stop surfacing new themes. It is a property of the data and the question, not a fixed number, which is why sample size in qualitative research is justified by observed saturation rather than a power calculation.

What does per-seat pricing mean for research tools?

Per-seat pricing charges by named user rather than by research activity. It suits small, continuously active teams and wastes money on occasional users — average SaaS licence utilization was 54% in 2026, meaning roughly half of seat spend goes unused.

Does an AI-moderated interview need to disclose that it is AI?

Yes, under EU AI Act Article 50, which has applied since 2 August 2026. Systems interacting directly with people must inform them that they are interacting with AI, clearly and at the latest at the time of the first interaction — a settings page or privacy policy does not satisfy this.

Put the vocabulary to work

Knowing the words is the cheap part. The expensive part is running enough conversations to have something worth naming.

Start with 10 free credits — three complete AI-moderated voice interviews that probe their own follow-ups, transcribe themselves and arrive as an automatic thematic analysis. From question to insight in hours, not weeks, with no research expertise required.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading