Reliability and Validity in Measurement
Before you analyse a construct, you have to measure it well. Reliability and validity are the two properties that decide whether your measures are worth analysing at all—and they are the foundation on which every survey, scale, and structural model stands.
Much of management and social-science research studies things that cannot be observed directly—satisfaction, trust, engagement, perceived quality. We measure these latent constructs indirectly, through questionnaire items or other indicators, and the whole edifice of the analysis rests on how well that measurement is done. If the measures are poor, no analytical sophistication can rescue the study: a beautifully specified model built on a badly measured construct is a precise account of the wrong thing. Reliability and validity are the two properties that tell you whether your measurement is sound—and getting them right comes before, not after, the modelling.
This guide explains what reliability and validity are, how they differ, the specific types of each you will encounter, and how they are assessed. It underpins our survey design and SEM guides, and reflects the measurement-first discipline in our SEM & Psychometrics practice.
The two properties, and the target analogy
Reliability is consistency: the extent to which a measure produces stable, repeatable results. If the same respondent, in the same state, would answer similarly on repeated measurement—and if items meant to capture the same construct agree with one another—the measure is reliable. Validity is accuracy: the extent to which a measure actually captures the concept it is supposed to capture, rather than something else.
The classic way to see the difference is a target. Reliable-but-not-valid is a tight cluster of shots landing consistently in the wrong place—precise, but systematically off. Valid-but-not-reliable is shots scattered around the bullseye—centred on average, but too erratic to trust any single one. Reliable-and-valid is a tight cluster in the bullseye: consistent and on-target. The two are independent properties, and a measure needs both. Crucially, reliability is a precondition for validity—a measure that is not even consistent cannot be accurately capturing anything—but reliability alone is never enough, because you can measure the wrong thing with perfect consistency.
Reliability is necessary but not sufficient for validity. A consistent measure of the wrong construct is still wrong. This is why a high reliability coefficient, on its own, never establishes that a scale measures what you claim.
Types of reliability
Reliability is assessed in several complementary ways, depending on what kind of consistency matters. Internal consistency asks whether the items within a scale agree with one another—the most commonly reported form, often summarised by coefficients such as Cronbach’s alpha or composite reliability. Test–retest reliability asks whether the measure gives stable results when administered to the same respondents at different times. Inter-rater reliability asks whether different observers or coders, rating the same thing, agree—essential in qualitative coding and any observational measurement.
A caution on internal consistency: a widely used coefficient like Cronbach’s alpha is sensitive to the number of items, so a high value can reflect a long scale rather than a genuinely coherent one, and it rests on assumptions that are not always met. It is a useful indicator, not a guarantee—which is why modern practice increasingly reports composite reliability alongside it and does not treat a single number as the last word.
Types of validity
Validity, too, comes in several forms, and a strong study addresses the relevant ones rather than asserting validity in general. Content validity asks whether the measure covers the full conceptual domain of the construct—are the items a comprehensive, representative sample of what the concept includes? It is judged largely through expert review and careful conceptual work. Construct validity—the central concern—asks whether the measure behaves as the underlying construct should, and it has two important sub-parts. Convergent validity is whether items meant to measure the same construct actually correlate strongly with one another. Discriminant validity is whether the construct is genuinely distinct from other constructs it should differ from—whether, in effect, you are measuring one thing and not accidentally capturing a neighbour. Criterion validity asks whether the measure relates to an external outcome it should predict or correspond to.
These are not boxes to tick but evidence to assemble: establishing validity means building a case, from multiple angles, that the measure captures the intended construct and only that construct. In latent-variable work, convergent and discriminant validity are assessed formally as part of the measurement model—the step that must pass before any structural relationship is interpreted.
Why it comes first
The order is not negotiable: measurement is assessed before the substantive analysis, because everything downstream inherits its quality. A structural equation model, a regression on scale scores, a comparison of groups—all of it assumes the measures mean what they claim. If reliability is weak, relationships are attenuated and true effects can be masked; if validity is weak, you may find strong, significant relationships between things that are not what you think they are. Neither problem is visible in the final results, which look just as tidy either way—so the only protection is to establish measurement quality up front and report the evidence.
This is also what reviewers in measurement-heavy fields look for first. A paper that reports reliability coefficients, demonstrates convergent and discriminant validity, and justifies its scales is credible; one that jumps straight to the structural model invites the question that undoes many studies: how do you know you measured what you say you measured? Answering that question convincingly, before the analysis, is the mark of sound quantitative research.
The bottom line
Reliability and validity are the twin foundations of measurement: reliability is consistency, validity is accuracy, and a defensible measure needs both. Reliability is necessary but never sufficient—you can be consistently wrong—so validity, especially construct validity with its convergent and discriminant components, is where the real work lies. Assess both before the substantive analysis, using the types appropriate to your measures, and report the evidence. Do that, and the constructs you go on to model actually stand for what you claim; skip it, and every result that follows rests on a foundation no one has checked.
Frequently asked questions
Need to validate a scale or measurement model?
From reliability analysis and confirmatory factor analysis to convergent and discriminant validity, our team can build the measurement evidence your study—and its reviewers—require.