Research Foundation

Validation Theory

The process of establishing that a measurement or inference captures what it is intended to capture — in the populations and contexts where it is used.

Overview

Validity is not a property of a test or an instrument. It is a property of the interpretations made from measurement results and the uses to which those interpretations are put. Validation is the process of accumulating and evaluating evidence for those interpretations.

The central move in 20th-century validation theory was a shift in what validity is a property of. Early frameworks treated validity as a feature of the measurement instrument — a test was valid if it measured what it said it measured. Samuel Messick's 1989 essay "Validity" — the definitive treatment in the field — argued instead that validity is a property of the interpretations made from scores and the uses to which those interpretations are put. A score that validly supports one inference may not support another. A test valid in one population may not be valid in another. Validation is therefore never finished: it is the ongoing process of accumulating and evaluating evidence for the interpretations and uses a test is put to.

Lee Cronbach and Paul Meehl's 1955 paper introduced the concept of construct validity: to claim that a measurement instrument is valid for a particular construct requires demonstrating that the instrument is embedded in a nomological network — a web of theoretical and empirical relationships connecting the construct to other constructs and to observable variables. The instrument must behave as theory predicts: correlating with what it should correlate with, not correlating with what it should not, and responding to interventions in the predicted direction. This is not a one-time demonstration but an ongoing programme of empirical inquiry.

Messick (1989) unified the previously fragmented validity concepts — content validity, criterion validity, construct validity — into a single construct validity framework with six aspects: content (does the instrument sample the domain appropriately?), substantive (does it engage the processes it claims to?), structural (does its internal structure reflect the construct's structure?), generalisability (do results hold across groups, settings, and tasks?), external (does it relate to other variables as expected?), and consequential (are the social consequences of its use acceptable?). The consequential aspect was controversial when introduced but is now recognised as essential — a test that produces valid scores but has discriminatory consequences fails on validity grounds.

Michael Kane's argument-based approach (2006, 2013) provides the most operationally useful framework for validation in practice. The interpretive argument — the chain of inferences from observed performance to a construct claim, through scoring, generalisation, extrapolation, and decision — should be stated explicitly and then subjected to systematic scrutiny. Each inference link is a claim that can be challenged, and validation research is organised around gathering evidence for or against each link. This makes validation a structured programme of inquiry rather than a property to be asserted once.

KANE'S INTERPRETIVE ARGUMENT Perfor- mance scoring Scoring generalising Generali- sation extrapolating Extra- polation deciding Decision each link is a claim — each claim requires evidence

Key Texts

Foundational works in this research tradition.

Cronbach & Meehl · 1955 · Psychological Bulletin
Construct Validity in Psychological Tests

The founding paper of construct validity: validity requires a theory — the nomological network — that specifies how the construct relates to other constructs and observable variables. Validity is a property of the whole inferential system, not just the instrument.

Messick · 1989 · in Educational Measurement (3rd ed.), ed. Linn
Validity

The definitive account: construct validity as the unified framework encompassing content, substance, structure, generalisability, external relations, and consequences. Validity attaches to interpretations and uses, not to instruments. The consequences of testing are part of validity evidence. Still the touchstone for measurement professionals.

Kane · 2006 · in Educational Measurement (4th ed.), ed. Brennan
Validation

The argument-based approach: state the interpretive argument — the chain of inferences from performance to construct claim — explicitly, then test each inference link. Validation is organised scrutiny of a stated argument, not a property to assert. The most actionable framework for validation programme design.

Kane · 2013 · Journal of Educational Measurement
Validating the Interpretations and Uses of Test Scores

The mature argument-based framework: extended and clarified. The interpretive argument as a structure for identifying what needs to be demonstrated and what evidence is required. Addresses common objections and connects the approach to broader validation theory.

AERA, APA & NCME · 2014
Standards for Educational and Psychological Testing

The professional standards that operationalise validation theory for practice: what evidence is required, how it should be gathered and reported, and the obligations of test developers and users. The normative baseline for any assessment programme claiming validity.

Related Research

Connected areas of inquiry.