Research Foundation

Psychometrics

The theory and practice of measuring psychological attributes through observable responses — the foundational discipline behind evidence-based assessment.

Overview

A psychological attribute — ability, knowledge, disposition, skill — cannot be observed directly. Psychometrics provides the theory and methods for inferring it from what can be observed: responses to tasks, questions, or situations.

Psychometrics is the scientific discipline concerned with the theory and practice of measuring psychological attributes through the systematic analysis of observable responses. The core inferential problem is that the things we most want to know about people — their cognitive ability, their level of expertise, their characteristic ways of responding — cannot be observed directly. They must be inferred from behaviour. Psychometrics provides the frameworks for doing this rigorously: specifying what is being inferred, designing tasks that provide evidence for the inference, and building statistical models that connect observed responses to latent attribute estimates.

The field's origins lie in Francis Galton's late-19th-century work on individual differences and Karl Pearson's statistical methods for analysing them. Charles Spearman's 1904 papers introduced factor analysis and the concept of g — general cognitive ability — establishing the idea that latent variables could be estimated from observed correlations among scores. This established the central logic of psychometrics: observable behaviour is a fallible indicator of an unobservable construct, and the goal is to estimate the construct with appropriate precision and uncertainty.

Classical Test Theory (CTT) — formalised by Lord and Novick in their landmark 1968 work — treats observed scores as composed of a true score and measurement error. CTT enabled systematic analysis of test reliability, provided the framework for estimating true score from observed score, and set standards for measurement quality that remain in use. Item Response Theory (IRT), pioneered by Georg Rasch and developed systematically by Frederic Lord, replaced total scores with probabilistic models of the individual response process: item difficulty, person ability, and item discrimination are estimated separately, and person-level estimates are produced that are theoretically independent of the specific items used. This is a significant advance — Rasch models in particular produce measures that sit on an interval scale rather than merely providing ordinal rankings.

Paul Meehl's Clinical versus Statistical Prediction (1954) addressed a question that persists in applied measurement: when do actuarial (mechanical, algorithmic) predictions outperform clinical (human, holistic) judgment? His systematic review found that actuarial methods equal or outperform clinical judgment in virtually every domain studied, for well-defined constructs with adequate base rates. This has direct implications for where AI-derived inferences can be trusted — and where they require different kinds of validation.

Key Texts

Foundational works in this research tradition.

Spearman · 1904 · American Journal of Psychology
"General Intelligence," Objectively Determined and Measured

The founding paper of quantitative psychology: factor analysis and the g factor. Establishes that correlated performance across different tasks implies an underlying latent variable — the logic that grounds all latent variable modelling in psychometrics.

Lord & Novick · 1968 · Addison-Wesley
Statistical Theories of Mental Test Scores

The formal treatise on Classical Test Theory: true scores, measurement error, reliability, and the conditions under which score-based inferences are warranted. The foundational text for a generation of measurement practitioners; still the reference for CTT.

Rasch · 1960 · Danish Institute of Educational Research
Probabilistic Models for Some Intelligence and Attainment Tests

The Rasch model: a probabilistic account of the response process that separates person ability from item difficulty. Person estimates are theoretically invariant across items, and item estimates across persons — enabling principled comparison across different test forms.

Lord · 1980 · Educational Testing Service
Applications of Item Response Theory to Practical Testing Problems

IRT in practice: item characteristic curves, parameter estimation, test information functions. The shift from total scores to latent variable measurement — and what that shift enables in terms of precision, comparability, and adaptive testing.

Meehl · 1954 · University of Minnesota Press
Clinical versus Statistical Prediction

Systematic review: actuarial (mechanical) prediction equals or outperforms clinical (human) judgment for well-defined constructs in virtually every domain studied. One of the most replicated findings in psychology. Has direct implications for where algorithmic inference can be trusted.

Related Research

Connected areas of inquiry.