{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-10T16:07:02.238Z"},"content":[{"type":"documentation","id":"d8eba4cc-f466-4bc4-b56c-573f451ea27e","slug":"construct-validity-operationalization","title":"Construct Validity: How to Tell Whether You Are Measuring the Thing You Named (2026)","url":"https://www.koji.so/docs/construct-validity-operationalization","summary":"Construct validity asks whether a measure captures the concept it names. It requires convergent evidence (relates to what it should) and discriminant evidence (stays distinct from adjacent constructs) - the latter being the test rarely run, which often reveals that separately-tracked metrics are one construct. Metascience shows 79 percent of Many Labs 2 scales were ad hoc and 40 percent of JPSP scales were reported without a source. Validity attaches to an interpretation for a specific use, not to the questionnaire.","content":"**Answer first: construct validity is the question of whether the number you called \"engagement\" measures engagement. It is not established by a high Cronbach alpha, a large sample, or a clean-looking dashboard. It requires evidence that your measure correlates with things the construct should correlate with (convergent) and, far more importantly, that it does NOT correlate too highly with things it is supposed to be distinct from (discriminant). The discriminant test is the one nobody runs, and it is the one that reveals that a company tracking \"trust,\" \"satisfaction,\" and \"loyalty\" on three dashboards is usually tracking one thing three times. Validity is a property of the interpretation you draw, not of the questionnaire - which is why the same five questions can be valid evidence for one claim and worthless for another.**\n\nEvery product organisation names things. Engagement. Activation. Delight. Trust. Intent. Each name eventually acquires a number, the number acquires a dashboard, and the dashboard acquires a quarterly target. What almost never happens in between is anyone asking whether the number is actually about the thing the name refers to.\n\nThat question is construct validity, and it is the oldest and most consequential question in measurement. This guide covers what it means, the two kinds of evidence that establish it, the two naming fallacies that destroy it, and what the metascience says about how often anybody bothers.\n\n## What a construct is, and why naming one is a measurement decision\n\nA construct is something you cannot observe directly: satisfaction, trust, effort, motivation, engagement. You can only observe indicators - what people say, click, or do - and infer the construct from them. **Operationalization** is the step where you decide which observable things will stand in for the unobservable one.\n\nThe term entered the literature with Cronbach and Meehl's \"Construct validity in psychological tests\" (*Psychological Bulletin* 52(4), 281-302, 1955), which introduced the idea of a **nomological network**: a construct is defined by its lawful relationships to other constructs and to observable behaviour. You do not validate a measure by staring at it. You validate it by checking whether it behaves the way the construct is supposed to behave.\n\nHere is the part that gets skipped in commercial work. **The moment you write a name on a dashboard, you have made a measurement claim, and nobody reviewed it.** \"Engagement: 62\" asserts that a specific arithmetic operation on specific observed data corresponds to a concept the whole company is now optimising. That assertion arrived with a chart rather than an argument.\n\n## The two kinds of evidence\n\nCampbell and Fiske set the standard in \"Convergent and discriminant validation by the multitrait-multimethod matrix\" (*Psychological Bulletin* 56(2), 81-105, 1959). Their insight was that validation requires two moves, not one.\n\n**Convergent evidence:** your measure correlates with other measures of the same construct, and with outcomes it should predict. If your satisfaction score does not relate to renewal, something is wrong.\n\n**Discriminant evidence:** your measure does *not* correlate too highly with measures of constructs it is supposed to be different from.\n\nAlmost every commercial \"validation\" exercise runs only the first test. That is because convergent evidence is flattering and easy to find - correlate your new score with anything vaguely related and you will get a satisfying number. Discriminant evidence is the one that hurts, and it is therefore the one that carries information.\n\n**The discriminant test nobody runs, in one paragraph.** Take the three or four headline constructs your company tracks separately - say trust, satisfaction, and likelihood to recommend. Correlate them with each other using the same respondents. If they correlate at 0.85, you do not have three constructs. You have one construct, three names, three dashboards, three owners, and three roadmap workstreams competing to move the same underlying number. This test takes an afternoon and is almost never performed, because the organisational cost of the answer is high and nobody is incentivised to ask.\n\n## Jingle and jangle\n\nTwo naming fallacies, both over a century old, both endemic to product work.\n\n**The jingle fallacy:** two different things share a name, so people assume they are the same thing. Your growth team's \"activation\" is a behavioural event in a funnel. Your research team's \"activation\" is a felt sense of having got value. Same word in the same all-hands, two unrelated measures, and a debate that cannot resolve because the participants are not disagreeing about the world.\n\n**The jangle fallacy:** one thing has two names, so people assume there are two things. The canonical modern case is grit and conscientiousness, which a large meta-analysis found correlated at **.84** (Crede, Tynan & Harms, 2017) - a correlation so high that treating them as distinct constructs is difficult to defend. In product terms: \"stickiness,\" \"engagement,\" and \"adoption\" are frequently one construct wearing three badges.\n\nThe scale of this problem in the scientific literature is startling once quantified. There are at least **280 published scales for measuring depression** (Santor et al., 2006), and at least **65 different scales for measuring emotions, 19 of which are devoted specifically to anger** (Weidman et al., 2017). In another domain, results from a single behavioural task have been **quantified in over 150 different ways across 130 publications** (Elson, 2019). These are fields with peer review, and they still cannot agree what the thing is or how to count it.\n\n## How often is any of this checked? The metascience\n\nJessica Flake and Eiko Fried assembled the evidence in \"Measurement Schmeasurement: Questionable Measurement Practices and How to Avoid Them\" (*Advances in Methods and Practices in Psychological Science*, 2020). Their summary of the state of affairs is worth reading carefully, because these numbers describe peer-reviewed research, and commercial research is not held to a higher standard:\n\n| Finding | Source |\n| --- | --- |\n| **40 to 93 percent** of measures across seven educational-behaviour journals lacked validity evidence | Barry et al. (2014) |\n| Of 356 measurement instances in emotion research, **69 percent** included no reference to prior research or any systematic development process | Weidman, Steckler & Tracy (2017) |\n| Of 433 scales in a random sample of *Journal of Personality and Social Psychology* articles from 2014, **40 percent** were reported without their source, **19 percent** without the number of items, and **9 percent** without the response scale | Flake, Pek & Hehman (2017) |\n| Of the item-based scales used across the Many Labs 2 replication project, **34 of 43 (79 percent)** appeared to be ad hoc - created by the authors, used without supporting validity information | Shaw et al. (2020) |\n\nFlake and Fried define questionable measurement practices as decisions that raise doubts about the validity of the measures and therefore of the study's conclusions. Their framing of the underlying incentive, quoting Orben and Przybylski's review of technology-use research, is the sharpest line in the literature: researchers pick and choose within and between questionnaires, **\"making the pre-specified constructs more of an accessory for publication than a guide for analyses.\"**\n\nSubstitute \"for the quarterly readout\" and you have described the modal product metric.\n\n## Validity belongs to the interpretation, not the instrument\n\nThis is the single most useful correction, and it reverses how most teams think.\n\nThere is no such thing as a valid questionnaire. There is only a valid *interpretation* of scores from a questionnaire, for a specific use. The same five items can be strong evidence for \"users found this flow easier than the old one\" and worthless evidence for \"users trust our AI features,\" even though the scores are identical, because validity is about the inference you are drawing.\n\nThe practical consequence: **borrowing a validated instrument does not transfer its validity to your use of it.** The [System Usability Scale](/docs/system-usability-scale-guide) is genuinely validated - for measuring perceived usability of a system a respondent has just used. Administering it to people who have not used the product, or using it as a proxy for satisfaction, or dropping four of the ten items to fit a shorter survey, all break the chain of evidence that made it validated in the first place. Flake, Pek and Hehman found **18 cases** in their review where two separate measures had been merged into one *on the basis of the reliability coefficient alpha* - alpha being used to justify collapsing constructs the authors had themselves defined as distinct. See [Cronbach alpha and internal consistency](/docs/cronbachs-alpha-internal-consistency) for why that reasoning does not work.\n\n## The six questions to answer before a metric ships\n\nAdapted from Flake and Fried's framework for the commercial case. If you cannot answer all six about a number on your dashboard, that number is a hypothesis wearing the costume of a fact.\n\n1. **What is your construct?** Define it in one sentence, without using the word itself and without the word \"and.\" If you need \"and,\" you have two constructs.\n2. **Why did you select this measure?** Name the alternatives you considered and why this one fits your definition better.\n3. **What did you use to operationalize it?** List the exact items, verbatim, with their response options.\n4. **How did you quantify it?** State the scoring rule - which items, what weighting, how missing data is handled.\n5. **Did you modify an existing measure?** If yes, say which items you changed or dropped, and accept that published validity evidence no longer applies unchanged.\n6. **Did you create the measure on the fly?** If yes, say so, and treat conclusions as provisional until validity evidence exists.\n\n| Evidence type | The question it answers | A cheap version you can actually run |\n| --- | --- | --- |\n| Face validity | Does it look like it measures the construct? | Show the items to five colleagues and ask what they think is being measured |\n| Content validity | Does it cover the whole construct? | List the facets of your definition; check each has at least one item |\n| Convergent | Does it relate to what it should? | Correlate with a behavioural outcome you already log |\n| **Discriminant** | Is it distinct from adjacent constructs? | **Correlate your headline metrics with each other. This is the test that matters** |\n| Criterion / predictive | Does it predict what it should? | Check whether last quarter scores predicted this quarter behaviour |\n| Response process | Do respondents interpret items as intended? | [Cognitive interviews](/docs/cognitive-interview-guide) - ask people what they thought the question meant |\n\nThe last row is where product research has an advantage over academic psychometrics, and it is under-exploited. You do not need a multitrait-multimethod matrix to discover that half your respondents thought \"the platform\" meant the mobile app. You need to ask them.\n\n## The modern approach: the definition and the evidence come from the same session\n\nThe structural reason construct validity is skipped commercially is not that teams are careless. It is that the check requires qualitative evidence and the measure produces quantitative data, so the check lands in a different budget, a different quarter, and usually a different team. When validating a construct means commissioning a separate study, it does not happen.\n\nAI-moderated research collapses that gap. When Koji runs a study, a [scale question](/docs/scale-questions-guide) can be immediately followed - in the same session, in the respondent's own words - by \"what were you thinking about when you answered that?\" That is response-process evidence, the hardest and most informative kind, produced as a by-product of collecting the number rather than as a separate project.\n\nKoji's [six structured question types](/docs/structured-questions-guide) - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - map onto this directly. The `scale` and `single_choice` items give you the quantitative indicator. The `open_ended` probe gives you the evidence about what respondents believed they were rating. A `ranking` item is frequently the better operationalization outright: if your construct is \"which problem matters most,\" ranking measures it directly, whereas five separate ratings measure five things and then require you to defend an averaging rule.\n\nThree specific advantages for construct work:\n\n- **Discriminant evidence becomes affordable.** The discriminant test needs your headline constructs measured on the same respondents. That is one Koji study, not four survey programmes owned by four teams.\n- **The brief is the construct definition, timestamped.** Because a Koji research brief states what you set out to measure before fielding, question 1 above has a documented answer rather than a reconstructed one - the same evidentiary property that defends against [researcher degrees of freedom](/docs/p-hacking-researcher-degrees-of-freedom).\n- **Cheap iteration makes real instrument development possible.** Proper scale development means drafting, testing comprehension, revising, retesting. At six weeks a round nobody does it. When a [pilot study](/docs/pilot-study-user-research-guide) runs in a day, the 79-percent ad-hoc rate stops being an economic inevitability.\n\nLegacy survey platforms will let you name a metric anything you like and will never ask what it means, because they collect only the response and not the reasoning behind it. The gap between \"we have a number\" and \"we know what the number is about\" is exactly the gap an AI interviewer closes.\n\n## The audit: three questions, one afternoon\n\n- **Take your top three metrics and write each definition in one sentence.** If two definitions are hard to tell apart, run the discriminant correlation before your next planning cycle.\n- **Pull the verbatim item wording for each.** If you cannot find it, or it is not what you assumed, you have found your problem.\n- **Ask ten users what they had in mind when they answered.** If their answers do not match your one-sentence definition, the metric is measuring something real - just not the thing on the label.\n\nMost teams fail the first step. The failure is free to discover and expensive to keep.\n\n## Frequently asked questions\n\n### What is construct validity in simple terms?\n\nConstruct validity is whether a measure actually captures the abstract concept it claims to. If you build an \"engagement score,\" construct validity is the question of whether that number is about engagement rather than about, say, how many notifications you sent. It is established through evidence that the measure relates to things the construct should relate to and stays distinct from things it should not.\n\n### What is the difference between construct validity and reliability?\n\nReliability is consistency: whether the measure produces stable, internally coherent results. Construct validity is aboutness: whether those consistent results are about the right thing. A measure can be highly reliable and completely invalid, the way a mis-calibrated scale reliably gives the same wrong weight. Reliability is a precondition for validity but never evidence of it.\n\n### What is the difference between convergent and discriminant validity?\n\nConvergent validity is evidence that your measure correlates with other measures of the same construct and with outcomes it should predict. Discriminant validity is evidence that it does not correlate too highly with constructs it should be distinct from. Convergent evidence is easy to obtain and flattering; discriminant evidence is the harder test and the one that reveals when several separately-tracked metrics are really one construct.\n\n### What are the jingle and jangle fallacies?\n\nThe jingle fallacy is assuming two things are the same because they share a name - two teams both saying \"activation\" while measuring unrelated things. The jangle fallacy is assuming two things are different because they have different names, as with grit and conscientiousness, which a large meta-analysis found correlated at .84. Both are naming errors that create real and expensive organisational confusion.\n\n### Can I just use a validated scale and skip construct validation?\n\nNot entirely. Validity attaches to an interpretation for a specific use, not to the questionnaire itself. A validated instrument used with a different population, for a different inference, or with items removed no longer carries its original validity evidence. Using a well-developed scale as designed is a strong starting position, but any modification or repurposing puts the burden back on you.\n\n### How do I check construct validity without a statistics team?\n\nRun the cheap versions. Show the items to colleagues and ask what they think is being measured. Check that your items cover every facet of your written definition. Correlate the measure against a behavioural outcome you already log. Correlate your headline metrics against each other to test discriminance. And ask ten respondents what they had in mind when answering - response-process evidence needs no model at all.\n\n## Related Resources\n\n- [Reliability vs. Validity in Research](/docs/reliability-vs-validity-research) - how the two properties differ and why both are required\n- [Cronbach Alpha and Internal Consistency](/docs/cronbachs-alpha-internal-consistency) - why a high alpha is not validity evidence\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - whether your valid measure is comparable across groups\n- [Cognitive Interviews](/docs/cognitive-interview-guide) - testing how respondents actually interpret your items\n- [Structured Questions in Koji](/docs/structured-questions-guide) - the six question types and choosing the right operationalization\n- [How to Write Unbiased Survey Questions](/docs/survey-question-wording-guide) - wording problems that break construct validity\n- [Qualitative Research Validity](/docs/qualitative-research-validity) - the parallel standards for qualitative work\n- [Factor Analysis in Survey Research](/docs/factor-analysis-survey-research-guide) - testing whether your items form the structure you assumed","category":"Research Methods","lastModified":"2026-08-10T03:27:22.68992+00:00","metaTitle":"Construct Validity: Are You Measuring the Thing You Named? (2026 Guide)","metaDescription":"Construct validity explained for product teams: operationalization, convergent and discriminant evidence, jingle-jangle fallacies, and the discriminant test that reveals your three metrics are one construct.","keywords":["construct validity","operationalization","convergent validity","discriminant validity","nomological network","jingle jangle fallacy","measurement validity"],"aiSummary":"Construct validity asks whether a measure captures the concept it names. It requires convergent evidence (relates to what it should) and discriminant evidence (stays distinct from adjacent constructs) - the latter being the test rarely run, which often reveals that separately-tracked metrics are one construct. Metascience shows 79 percent of Many Labs 2 scales were ad hoc and 40 percent of JPSP scales were reported without a source. Validity attaches to an interpretation for a specific use, not to the questionnaire.","aiPrerequisites":["Familiarity with survey measures and metrics","Understanding of correlation"],"aiLearningOutcomes":["Define a construct and operationalize it defensibly","Distinguish convergent from discriminant validity evidence","Run the discriminant correlation test on your existing headline metrics","Identify jingle and jangle fallacies in your organisation metric names","Answer the six questions before a metric ships"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}