Reliability, Validity & Measurement Error
An interactive archery-target metaphor for random error (reliability) and systematic error (validity)
Classical Test Theory: Observed Score = True Score + Error. Reliability concerns random error — how consistently a measure performs. Validity concerns systematic error — whether a measure hits its intended target.
Reliability (ρ): —
Validity: —
Diagnosis: —
Target (N = 40 arrows)
Observed Score Distribution
Each "arrow" represents one measurement attempt. The bullseye is the true score.
Random error scatters arrows around the cluster center (flattens and widens the distribution).
Systematic bias shifts the entire cluster away from the bullseye (displaces the mean without changing spread); use negative values to shift in the opposite direction.
High reliability + high validity = tight cluster on the bullseye. High reliability + low validity = tight cluster, wrong place.
Note: maximum validity = √reliability (CTT upper bound). A measure cannot correlate with any criterion more than the square root of its own reliability, so validity is bounded here accordingly. That bound assumes a perfectly measured criterion; when the criterion is itself fallible the tighter limit is $r_{xy} \le \sqrt{r_{xx'} \cdot r_{yy'}}$ — both measures' unreliability drags the ceiling down.
Citation
Persson, B. N. (2026). Reliability, Validity & Measurement Error [Interactive visualization]. https://bjorn-persson.github.io/visualizations/reliability-validity/
@misc{Persson2026reliabilityvalidity,
author = {Björn N. Persson},
year = {2026},
title = {Reliability, Validity \& Measurement Error},
note = {Interactive visualization},
url = {https://bjorn-persson.github.io/visualizations/reliability-validity/}}