Back to docs
Research Methods

Construct Validity: How to Tell Whether You Are Measuring the Thing You Named (2026)

Construct validity is the question of whether your engagement score measures engagement. A guide to operationalization, convergent and discriminant evidence, jingle-jangle fallacies, and the discriminant test that kills most product metrics.

Answer first: construct validity is the question of whether the number you called "engagement" measures engagement. It is not established by a high Cronbach alpha, a large sample, or a clean-looking dashboard. It requires evidence that your measure correlates with things the construct should correlate with (convergent) and, far more importantly, that it does NOT correlate too highly with things it is supposed to be distinct from (discriminant). The discriminant test is the one nobody runs, and it is the one that reveals that a company tracking "trust," "satisfaction," and "loyalty" on three dashboards is usually tracking one thing three times. Validity is a property of the interpretation you draw, not of the questionnaire - which is why the same five questions can be valid evidence for one claim and worthless for another.

Every product organisation names things. Engagement. Activation. Delight. Trust. Intent. Each name eventually acquires a number, the number acquires a dashboard, and the dashboard acquires a quarterly target. What almost never happens in between is anyone asking whether the number is actually about the thing the name refers to.

That question is construct validity, and it is the oldest and most consequential question in measurement. This guide covers what it means, the two kinds of evidence that establish it, the two naming fallacies that destroy it, and what the metascience says about how often anybody bothers.

What a construct is, and why naming one is a measurement decision

A construct is something you cannot observe directly: satisfaction, trust, effort, motivation, engagement. You can only observe indicators - what people say, click, or do - and infer the construct from them. Operationalization is the step where you decide which observable things will stand in for the unobservable one.

The term entered the literature with Cronbach and Meehl's "Construct validity in psychological tests" (Psychological Bulletin 52(4), 281-302, 1955), which introduced the idea of a nomological network: a construct is defined by its lawful relationships to other constructs and to observable behaviour. You do not validate a measure by staring at it. You validate it by checking whether it behaves the way the construct is supposed to behave.

Here is the part that gets skipped in commercial work. The moment you write a name on a dashboard, you have made a measurement claim, and nobody reviewed it. "Engagement: 62" asserts that a specific arithmetic operation on specific observed data corresponds to a concept the whole company is now optimising. That assertion arrived with a chart rather than an argument.

The two kinds of evidence

Campbell and Fiske set the standard in "Convergent and discriminant validation by the multitrait-multimethod matrix" (Psychological Bulletin 56(2), 81-105, 1959). Their insight was that validation requires two moves, not one.

Convergent evidence: your measure correlates with other measures of the same construct, and with outcomes it should predict. If your satisfaction score does not relate to renewal, something is wrong.

Discriminant evidence: your measure does not correlate too highly with measures of constructs it is supposed to be different from.

Almost every commercial "validation" exercise runs only the first test. That is because convergent evidence is flattering and easy to find - correlate your new score with anything vaguely related and you will get a satisfying number. Discriminant evidence is the one that hurts, and it is therefore the one that carries information.

The discriminant test nobody runs, in one paragraph. Take the three or four headline constructs your company tracks separately - say trust, satisfaction, and likelihood to recommend. Correlate them with each other using the same respondents. If they correlate at 0.85, you do not have three constructs. You have one construct, three names, three dashboards, three owners, and three roadmap workstreams competing to move the same underlying number. This test takes an afternoon and is almost never performed, because the organisational cost of the answer is high and nobody is incentivised to ask.

Jingle and jangle

Two naming fallacies, both over a century old, both endemic to product work.

The jingle fallacy: two different things share a name, so people assume they are the same thing. Your growth team's "activation" is a behavioural event in a funnel. Your research team's "activation" is a felt sense of having got value. Same word in the same all-hands, two unrelated measures, and a debate that cannot resolve because the participants are not disagreeing about the world.

The jangle fallacy: one thing has two names, so people assume there are two things. The canonical modern case is grit and conscientiousness, which a large meta-analysis found correlated at .84 (Crede, Tynan & Harms, 2017) - a correlation so high that treating them as distinct constructs is difficult to defend. In product terms: "stickiness," "engagement," and "adoption" are frequently one construct wearing three badges.

The scale of this problem in the scientific literature is startling once quantified. There are at least 280 published scales for measuring depression (Santor et al., 2006), and at least 65 different scales for measuring emotions, 19 of which are devoted specifically to anger (Weidman et al., 2017). In another domain, results from a single behavioural task have been quantified in over 150 different ways across 130 publications (Elson, 2019). These are fields with peer review, and they still cannot agree what the thing is or how to count it.

How often is any of this checked? The metascience

Jessica Flake and Eiko Fried assembled the evidence in "Measurement Schmeasurement: Questionable Measurement Practices and How to Avoid Them" (Advances in Methods and Practices in Psychological Science, 2020). Their summary of the state of affairs is worth reading carefully, because these numbers describe peer-reviewed research, and commercial research is not held to a higher standard:

FindingSource
40 to 93 percent of measures across seven educational-behaviour journals lacked validity evidenceBarry et al. (2014)
Of 356 measurement instances in emotion research, 69 percent included no reference to prior research or any systematic development processWeidman, Steckler & Tracy (2017)
Of 433 scales in a random sample of Journal of Personality and Social Psychology articles from 2014, 40 percent were reported without their source, 19 percent without the number of items, and 9 percent without the response scaleFlake, Pek & Hehman (2017)
Of the item-based scales used across the Many Labs 2 replication project, 34 of 43 (79 percent) appeared to be ad hoc - created by the authors, used without supporting validity informationShaw et al. (2020)

Flake and Fried define questionable measurement practices as decisions that raise doubts about the validity of the measures and therefore of the study's conclusions. Their framing of the underlying incentive, quoting Orben and Przybylski's review of technology-use research, is the sharpest line in the literature: researchers pick and choose within and between questionnaires, "making the pre-specified constructs more of an accessory for publication than a guide for analyses."

Substitute "for the quarterly readout" and you have described the modal product metric.

Validity belongs to the interpretation, not the instrument

This is the single most useful correction, and it reverses how most teams think.

There is no such thing as a valid questionnaire. There is only a valid interpretation of scores from a questionnaire, for a specific use. The same five items can be strong evidence for "users found this flow easier than the old one" and worthless evidence for "users trust our AI features," even though the scores are identical, because validity is about the inference you are drawing.

The practical consequence: borrowing a validated instrument does not transfer its validity to your use of it. The System Usability Scale is genuinely validated - for measuring perceived usability of a system a respondent has just used. Administering it to people who have not used the product, or using it as a proxy for satisfaction, or dropping four of the ten items to fit a shorter survey, all break the chain of evidence that made it validated in the first place. Flake, Pek and Hehman found 18 cases in their review where two separate measures had been merged into one on the basis of the reliability coefficient alpha - alpha being used to justify collapsing constructs the authors had themselves defined as distinct. See Cronbach alpha and internal consistency for why that reasoning does not work.

The six questions to answer before a metric ships

Adapted from Flake and Fried's framework for the commercial case. If you cannot answer all six about a number on your dashboard, that number is a hypothesis wearing the costume of a fact.

  1. What is your construct? Define it in one sentence, without using the word itself and without the word "and." If you need "and," you have two constructs.
  2. Why did you select this measure? Name the alternatives you considered and why this one fits your definition better.
  3. What did you use to operationalize it? List the exact items, verbatim, with their response options.
  4. How did you quantify it? State the scoring rule - which items, what weighting, how missing data is handled.
  5. Did you modify an existing measure? If yes, say which items you changed or dropped, and accept that published validity evidence no longer applies unchanged.
  6. Did you create the measure on the fly? If yes, say so, and treat conclusions as provisional until validity evidence exists.
Evidence typeThe question it answersA cheap version you can actually run
Face validityDoes it look like it measures the construct?Show the items to five colleagues and ask what they think is being measured
Content validityDoes it cover the whole construct?List the facets of your definition; check each has at least one item
ConvergentDoes it relate to what it should?Correlate with a behavioural outcome you already log
DiscriminantIs it distinct from adjacent constructs?Correlate your headline metrics with each other. This is the test that matters
Criterion / predictiveDoes it predict what it should?Check whether last quarter scores predicted this quarter behaviour
Response processDo respondents interpret items as intended?Cognitive interviews - ask people what they thought the question meant

The last row is where product research has an advantage over academic psychometrics, and it is under-exploited. You do not need a multitrait-multimethod matrix to discover that half your respondents thought "the platform" meant the mobile app. You need to ask them.

The modern approach: the definition and the evidence come from the same session

The structural reason construct validity is skipped commercially is not that teams are careless. It is that the check requires qualitative evidence and the measure produces quantitative data, so the check lands in a different budget, a different quarter, and usually a different team. When validating a construct means commissioning a separate study, it does not happen.

AI-moderated research collapses that gap. When Koji runs a study, a scale question can be immediately followed - in the same session, in the respondent's own words - by "what were you thinking about when you answered that?" That is response-process evidence, the hardest and most informative kind, produced as a by-product of collecting the number rather than as a separate project.

Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - map onto this directly. The scale and single_choice items give you the quantitative indicator. The open_ended probe gives you the evidence about what respondents believed they were rating. A ranking item is frequently the better operationalization outright: if your construct is "which problem matters most," ranking measures it directly, whereas five separate ratings measure five things and then require you to defend an averaging rule.

Three specific advantages for construct work:

  • Discriminant evidence becomes affordable. The discriminant test needs your headline constructs measured on the same respondents. That is one Koji study, not four survey programmes owned by four teams.
  • The brief is the construct definition, timestamped. Because a Koji research brief states what you set out to measure before fielding, question 1 above has a documented answer rather than a reconstructed one - the same evidentiary property that defends against researcher degrees of freedom.
  • Cheap iteration makes real instrument development possible. Proper scale development means drafting, testing comprehension, revising, retesting. At six weeks a round nobody does it. When a pilot study runs in a day, the 79-percent ad-hoc rate stops being an economic inevitability.

Legacy survey platforms will let you name a metric anything you like and will never ask what it means, because they collect only the response and not the reasoning behind it. The gap between "we have a number" and "we know what the number is about" is exactly the gap an AI interviewer closes.

The audit: three questions, one afternoon

  • Take your top three metrics and write each definition in one sentence. If two definitions are hard to tell apart, run the discriminant correlation before your next planning cycle.
  • Pull the verbatim item wording for each. If you cannot find it, or it is not what you assumed, you have found your problem.
  • Ask ten users what they had in mind when they answered. If their answers do not match your one-sentence definition, the metric is measuring something real - just not the thing on the label.

Most teams fail the first step. The failure is free to discover and expensive to keep.

Frequently asked questions

What is construct validity in simple terms?

Construct validity is whether a measure actually captures the abstract concept it claims to. If you build an "engagement score," construct validity is the question of whether that number is about engagement rather than about, say, how many notifications you sent. It is established through evidence that the measure relates to things the construct should relate to and stays distinct from things it should not.

What is the difference between construct validity and reliability?

Reliability is consistency: whether the measure produces stable, internally coherent results. Construct validity is aboutness: whether those consistent results are about the right thing. A measure can be highly reliable and completely invalid, the way a mis-calibrated scale reliably gives the same wrong weight. Reliability is a precondition for validity but never evidence of it.

What is the difference between convergent and discriminant validity?

Convergent validity is evidence that your measure correlates with other measures of the same construct and with outcomes it should predict. Discriminant validity is evidence that it does not correlate too highly with constructs it should be distinct from. Convergent evidence is easy to obtain and flattering; discriminant evidence is the harder test and the one that reveals when several separately-tracked metrics are really one construct.

What are the jingle and jangle fallacies?

The jingle fallacy is assuming two things are the same because they share a name - two teams both saying "activation" while measuring unrelated things. The jangle fallacy is assuming two things are different because they have different names, as with grit and conscientiousness, which a large meta-analysis found correlated at .84. Both are naming errors that create real and expensive organisational confusion.

Can I just use a validated scale and skip construct validation?

Not entirely. Validity attaches to an interpretation for a specific use, not to the questionnaire itself. A validated instrument used with a different population, for a different inference, or with items removed no longer carries its original validity evidence. Using a well-developed scale as designed is a strong starting position, but any modification or repurposing puts the burden back on you.

How do I check construct validity without a statistics team?

Run the cheap versions. Show the items to colleagues and ask what they think is being measured. Check that your items cover every facet of your written definition. Correlate the measure against a behavioural outcome you already log. Correlate your headline metrics against each other to test discriminance. And ask ten respondents what they had in mind when answering - response-process evidence needs no model at all.

Related Resources

Related Articles

Cognitive Interviews: How to Test Your Survey Questions Before You Launch

A practical guide to cognitive interviewing — the pretesting technique that reveals whether your survey questions and interview guides are understood as intended. Covers think-aloud, verbal probing, sample sizing, and AI-powered approaches.

Factor Analysis in Survey Research: EFA, PCA & How to Read It (2026)

A practical guide to factor analysis for surveys: what EFA and PCA do, how to check KMO and eigenvalues, how many responses you need, how to name factors, and the mistakes to avoid.

Qualitative Research Validity and Reliability: How to Build Studies You Can Trust

A practical guide to Lincoln and Guba's trustworthiness framework — credibility, transferability, dependability, and confirmability — and how to build each into your qualitative research studies.

Reliability vs. Validity in Research: What They Mean and How to Get Both

A clear guide to reliability versus validity in research: precise definitions, the dartboard analogy, the types of each, how to improve them, and how AI-moderated interviews deliver consistent, accurate insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions

A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.