Construct Validity
Six frameworks for understanding whether a test measures what it claims to measure
Section 1
Six Frameworks for Construct Validity
Construct validity is the central question in psychological measurement: does a test actually measure the theoretical construct it's intended to measure? Over 70 years, several influential frameworks have emerged — each asking this question differently. Explore each below, applied to the same running example: a depression screening questionnaire (the PHQ-9).
Cronbach & Meehl (1955) · The Nomological Network
A construct derives its scientific meaning from its place in a nomological network — a system of lawful relationships connecting theoretical constructs to each other and to observable indicators. A construct with no such connections is not scientifically admissible: there is nothing to test, confirm, or refute.
Click on any connection to toggle it. Watch what happens to the construct's scientific status as you strip away its nomological embedding.
Click on any connection to toggle it. Watch what happens to the construct's scientific status as you strip away its nomological embedding.
Target construct
Related construct
Observable indicator
click lines to toggle
Nomological embedding: 100%
Key idea: Validity is not a property of a test alone — it is a property of a construct's place in theory. Without a nomological network, there is nothing to validate against. The construct's meaning is its web of lawful relationships.
Embretson (published as Whitely, 1983) later refined this approach by adding a second component alongside the nomological network: construct representation (the cognitive processes underlying test performance) sits next to nomothetic span (the web of external correlates). This distinguished the internal "how does the test work?" question from the external "what does it predict?" question — a distinction the original framework left implicit.
Campbell & Fiske (1959) · Multitrait-Multimethod Matrix
Campbell and Fiske operationalized validation by measuring multiple traits using multiple methods. This established two core principles: convergent validity (different methods measuring the same trait should agree) and discriminant validity (measures of different traits should diverge, even when sharing a method). Method variance — systematic variance coming from the measuring procedure rather than from the trait, which inflates the correlation between any two measures sharing a method — is the key threat.
The three sliders decide where each measure's variance comes from. Trait variance is the share reflecting the trait a measure targets, and method variance the share coming from the procedure used; between them they set how much of the measure is signal and how much is an artefact of asking. Trait correlation is different in kind — it is how much the underlying traits genuinely overlap, how related depression and anxiety really are before any measurement enters. Every cell is generated from a single trait-and-method decomposition, so the matrix is always one real data could produce. Watch what happens when method variance grows large relative to trait variance.
The three sliders decide where each measure's variance comes from. Trait variance is the share reflecting the trait a measure targets, and method variance the share coming from the procedure used; between them they set how much of the measure is signal and how much is an artefact of asking. Trait correlation is different in kind — it is how much the underlying traits genuinely overlap, how related depression and anxiety really are before any measurement enters. Every cell is generated from a single trait-and-method decomposition, so the matrix is always one real data could produce. Watch what happens when method variance grows large relative to trait variance.
Reliability (diagonal)
Convergent (same trait, diff method)
Heterotrait-monomethod
Heterotrait-heteromethod
The three checks above are Campbell and Fiske's first three. Their fourth asks that the same pattern of trait interrelationships appear in every heterotrait triangle — if depression and anxiety are the most correlated pair under self-report, they should be the most correlated pair under clinician rating too. Here all three traits share a single correlation, so each triangle is internally constant and there is no pattern left to violate.
Each measure is trait variance + method variance + error, so the sliders always produce a matrix real data could yield. Trait and method variance share one budget and cannot sum past 1 — which is why the method variance slider loses range as trait variance climbs. The diagonal shows each measure's reliable variance — the values Campbell and Fiske expected to be largest in the matrix, and under this model they always are.
Each measure is trait variance + method variance + error, so the sliders always produce a matrix real data could yield. Trait and method variance share one budget and cannot sum past 1 — which is why the method variance slider loses range as trait variance climbs. The diagonal shows each measure's reliable variance — the values Campbell and Fiske expected to be largest in the matrix, and under this model they always are.
Key idea: A high correlation is not evidence of validity on its own. What matters is which cell of the matrix it sits in.
Convergent validity is the PHQ-9 against a clinician's rating of the same trait: 0.65. Two quite different procedures — a patient ticking boxes, a clinician forming a judgment — agreeing that closely about who is depressed is what tells you the PHQ-9 is tracking depression, and not just a quirk of how people fill in questionnaires.
Discriminant validity asks whether that number beats what the PHQ-9 shares with a different trait. Against a self-report anxiety scale it correlates 0.30; against a clinician's anxiety rating, only 0.20. Depression and anxiety genuinely overlap by that second figure — the difference between the two, 0.10, is method variance and nothing else, bought purely by sharing a questionnaire format.
That gap is why crossing traits with methods is the whole point. Push method variance far enough and the anxiety questionnaire correlation overtakes the clinician's depression rating — at which point the PHQ-9 is telling you more about how people answer questionnaires than about who is depressed.
Convergent validity is the PHQ-9 against a clinician's rating of the same trait: 0.65. Two quite different procedures — a patient ticking boxes, a clinician forming a judgment — agreeing that closely about who is depressed is what tells you the PHQ-9 is tracking depression, and not just a quirk of how people fill in questionnaires.
Discriminant validity asks whether that number beats what the PHQ-9 shares with a different trait. Against a self-report anxiety scale it correlates 0.30; against a clinician's anxiety rating, only 0.20. Depression and anxiety genuinely overlap by that second figure — the difference between the two, 0.10, is method variance and nothing else, bought purely by sharing a questionnaire format.
That gap is why crossing traits with methods is the whole point. Push method variance far enough and the anxiety questionnaire correlation overtakes the clinician's depression rating — at which point the PHQ-9 is telling you more about how people answer questionnaires than about who is depressed.
Messick (1989; 1995) · Unified Validity
Messick argued that validity is a single, unified concept — not divisible into separate "types" (content validity, criterion validity, construct validity). Instead, construct validity subsumes all other forms, and it is evaluated through six distinguishable but interrelated facets of evidence. Crucially, Messick insisted that the consequences of test use are part of the validity argument — a position that remains controversial.
Click each facet to toggle whether that evidence has been gathered. Try the presets to see what typical versus comprehensive validation looks like.
Click each facet to toggle whether that evidence has been gathered. Try the presets to see what typical versus comprehensive validation looks like.
Validity argument completeness: 0/6
Key idea: Most validation studies address only content and external (criterion-related) evidence, ignoring consequences entirely. For Messick, this is an incomplete argument — you cannot claim a test is valid without considering how its use affects people.
Loevinger (1957) anticipated much of Messick's framework, arguing that construct validity subsumes all other forms. Her three components — substantive, structural, and external — directly prefigure Messick's more elaborated six facets. — Lissitz & Samuelsen (2007) pushed back in the other direction, arguing that consequences belong to a test's utility, not its validity. In their view, construct validity should rest strictly on internal evidence: content, reliability, and the cognitive processes items elicit.
Borsboom, Mellenbergh & van Heerden (2004) · Causal Realism
Borsboom and colleagues proposed a radically simple definition: a test is valid if and only if (1) the attribute exists and (2) variation in the attribute causally produces variation in test scores. This strips away the epistemic complexity of earlier frameworks and anchors validity in ontology. A test can predict outcomes perfectly and still fail to be valid if the causal link is missing.
Select a scenario to see how Borsboom's two conditions apply.
Select a scenario to see how Borsboom's two conditions apply.
Depression
(attribute)
(attribute)
PHQ-9
Scores
(test)
Scores
(test)
VALID
Key idea: Validity is about the world, not about evidence. A test is valid because the attribute it measures exists and causally drives the scores — regardless of whether we have gathered evidence for this. Conversely, no amount of correlational evidence can make a test valid if the causal link is absent.
Kane reframed validity as the evaluation of an interpretive argument (renamed the interpretation/use argument, or IUA, in later work, to give equal weight to how scores are used): a chain of inferences connecting observed test performance to real-world decisions. Each inference must be explicitly stated and supported with evidence and warrants. The overall validity of the argument equals the strength of its weakest link.
Adjust the warrant strength for each inference. Notice how a single weak link undermines the entire argument.
Adjust the warrant strength for each inference. Notice how a single weak link undermines the entire argument.
Overall argument strength
Weakest link: Strong
Key idea: Validity is not about a test — it is about an argument. A strong scoring inference cannot compensate for a weak extrapolation inference. This shifts the burden: instead of asking "is this test valid?", we ask "is this interpretation and use of test scores justified?"
Borsboom & Cramer (2013; Borsboom, 2017) · Network Psychometrics
Network psychometrics challenges a foundational assumption of all the frameworks above: that a psychological construct is a single latent variable that causes observed scores. Instead, constructs like depression may be networks of mutually reinforcing symptoms — sadness causes insomnia, insomnia causes fatigue, fatigue causes loss of interest, and so on. The "construct" is an emergent property of the network, not a hidden common cause.
Toggle between the latent variable view and the network view to see the same symptoms under radically different ontologies. In network view, click symptoms to activate them and trace the causal pathways.
Toggle between the latent variable view and the network view to see the same symptoms under radically different ontologies. In network view, click symptoms to activate them and trace the causal pathways.
Hub-and-spoke: a single latent variable "Depression" causes all observed symptoms.
Active symptoms: 0 / 7
View: Latent Variable
In the latent variable model, depression is a single hidden cause. All symptom correlations arise because they share this common cause.
Key idea: If constructs are networks rather than latent variables, then "construct validity" in the traditional sense may not apply. There is no single attribute to measure — instead, the question becomes whether the test adequately captures the network's structure and dynamics.
This complicates what Borsboom's first condition is asking for. If depression is a network rather than a unitary attribute, does “the attribute exists” simply fail — or does the network itself count as the real thing being measured? Borsboom wrote both frameworks and takes the second view: a network state can be perfectly real and can causally drive item responses. Either way, the test is measuring something quite different from what latent variable models assume.
Section 2
Where Does Validity Live?
Cronbach & Meehl
In the theory. Validity is about the construct's place in a nomological network of lawful relationships.
Campbell & Fiske
In the method. Validity requires disentangling trait variance from method variance through systematic design.
Messick
In the evidence and its consequences. Validity is a judgment based on the completeness of the evidentiary argument.
Borsboom
In the world. Validity is an ontological property — does the attribute exist and causally produce scores?
Kane
In the argument. Validity is the degree to which the interpretive/use argument is coherent and well-supported.
Network
In the structure. If constructs are emergent networks, validity is about capturing the system's topology, not measuring a hidden cause.
These are not just academic distinctions. Applied to the same test, these frameworks can reach different conclusions. A test can be ontologically valid (Borsboom) but lack consequential evidence (Messick). It can have a strong interpretive argument (Kane) while the construct itself is poorly embedded in theory (Cronbach & Meehl). The network view may question whether the construct is even the right unit of analysis. Choosing a framework shapes which questions you ask — and which evidence you gather. That these disagreements persist after decades goes to the heart of why psychology is difficult. It is not that our constructs are invisible — physics measures plenty of things nobody can see directly, from temperature to field strength. The difference is that physics has independent ways to calibrate its instruments and theory strong enough to specify what a good measurement would look like. For depression, intelligence, or personality we have neither, so there is no outside check on whether we have carved nature at its joints.
Citation
Persson, B. N. (2026). Construct Validity [Interactive visualization]. https://bjorn-persson.github.io/visualizations/construct-validity/
@misc{Persson2026constructvalidity,
author = {Björn N. Persson},
year = {2026},
title = {Construct Validity},
note = {Interactive visualization},
url = {https://bjorn-persson.github.io/visualizations/construct-validity/}}