Causal Inference

Instrumental Variables & 2SLS Services

When a predictor is entangled with unobserved factors—so ordinary regression cannot separate cause from correlation—instrumental variables can recover a causal effect. Using a variable that shifts the predictor but affects the outcome only through it, IV (typically estimated by two-stage least squares) is the standard tool for endogeneity.

An instrumental variable (IV) is a variable that influences a treatment or predictor but affects the outcome only through that predictor and is unrelated to the unobserved confounders. It is used to estimate causal effects when the predictor is endogenous—correlated with the error term—so ordinary least squares is biased. Two-stage least squares (2SLS) is the standard IV estimator.

Addresses endogeneity 2SLS & related estimators Weak-instrument diagnostics Reproducible, journal-ready
The instrumental-variable causal diagram An instrument Z affects treatment X, which affects outcome Y. An unobserved confounder U affects both X and Y, but Z is unrelated to U and reaches Y only through X. instrumental_variables · identification Z X Y U relevance instrument unobserved confounder Z affects Y only through X, and is unrelated to U
IV identification causal path confounding

What instrumental variables do

Ordinary regression estimates a causal effect only if the predictor is unrelated to everything else in the error term. That fails constantly in observational management, economics, and social research, because of endogeneity: omitted variables that drive both the predictor and the outcome, simultaneity (the predictor and outcome influence each other), or measurement error. When a predictor is endogenous, its ordinary-least-squares coefficient conflates the causal effect with the confounding, and no amount of the usual controls fixes it if the confounder is unobserved.

An instrumental variable offers a way through. The instrument, Z, is a variable that moves the endogenous predictor X but influences the outcome Y only through X—not directly, and not via the unobserved confounders. Intuitively, IV uses only the part of the variation in X that is driven by the instrument (and is therefore “clean” of confounding) to estimate the effect of X on Y. In practice this is done by two-stage least squares (2SLS): the first stage predicts X from the instrument (and controls), and the second stage regresses Y on that predicted, confounding-free portion of X.

When to use it

Reach for IV when you have good reason to believe a key predictor is endogenous—and you can identify a credible instrument. Classic instruments in applied research include policy rules, lottery or random allocation, geographic or institutional features, and other sources of variation in the predictor that plausibly have no direct path to the outcome. IV also underpins the fuzzy regression discontinuity design, where the cutoff serves as the instrument.

The difficulty—and it is a real one—is that a valid instrument must satisfy strong conditions, and finding one is often the hardest part of a study. IV is not a routine fix to be applied whenever a coefficient looks suspicious; a weak or invalid instrument can make the cure worse than the disease. We assess instrument credibility rigorously before recommending the approach, because a defensible IV analysis lives or dies on the instrument.

At a glance

What a valid instrument must satisfy

The conditions an instrument has to meet
ConditionWhat it requiresHow it is assessed
RelevanceThe instrument genuinely predicts the endogenous predictorFirst-stage strength (e.g. F-statistic); testable
ExclusionThe instrument affects the outcome only through the predictorArgued from theory/design; not directly testable
IndependenceThe instrument is unrelated to unobserved confoundersArgued from theory/design; not directly testable
(With extra instruments)Over-identifying restrictionsOver-identification tests provide partial evidence
Methodology

Weak instruments, exclusion, and what IV estimates

Two problems determine whether an IV analysis is credible. The first is weak instruments: if the instrument only weakly predicts the endogenous predictor, 2SLS becomes badly biased (toward the very OLS estimate it was meant to correct) and its inference unreliable. This is checkable—first-stage strength is reported and assessed against established guidance—and weak-instrument-robust methods are used where needed. The second is the exclusion restriction: the assumption that the instrument affects the outcome only through the predictor. This is the crux of any IV study, it generally cannot be tested statistically, and it must be argued from theory and the design. Over-identification tests help only when there are more instruments than endogenous predictors, and even then provide partial evidence.

It also matters what IV actually estimates. When effects differ across units, 2SLS identifies a local average treatment effect (LATE)—the effect for those whose behaviour is shifted by the instrument (the “compliers”), not the average effect across everyone. That is a meaningful, well-defined quantity, but it is not always the population effect a reader assumes, so we state which effect is being estimated rather than implying a universal one.

The exclusion restriction is an assumption, not a result. An instrument’s validity rests on it affecting the outcome only through the predictor—which usually cannot be tested and must be justified by theory and design. A weak instrument can also make IV more biased than OLS, so first-stage strength is always checked.

Software

We deliver IV/2SLS in established, reproducible tools—R (ivreg, fixest) and Stata—reporting the first stage and its strength, weak-instrument-robust inference where appropriate, over-identification tests when applicable, and an explicit statement of the estimand, with versioned code.

How we work

How we deliver an IV / 2SLS study

IV sits within our wider causal-inference practice—so the endogeneity problem is diagnosed, the instrument is scrutinised, and the estimand is stated honestly.

We start by establishing whether endogeneity is genuinely the problem and, if so, assessing the credibility of any candidate instrument against the relevance, exclusion, and independence conditions—before committing to the approach. We then estimate the model by 2SLS (or a suitable alternative), report and evaluate first-stage strength, apply weak-instrument-robust inference where needed, and run the available specification and over-identification checks.

Reporting sets out the endogeneity argument, the instrument and the case for its validity, the first stage, the estimator, the diagnostics, and—crucially—which effect (LATE or otherwise) is being identified.

You receive the IV estimate with appropriate inference, the first-stage results and strength diagnostics, weak-instrument-robust results where relevant, over-identification and specification checks, a clear statement of the estimand and the exclusion argument, and reproducible analytical code and analysis-ready files (where appropriate and permitted). Where no credible instrument exists, we say so and recommend a more defensible design.

Where we apply it

Instrumental variables across Management & Allied Studies

Endogeneity is pervasive in observational social science—so IV is a core tool wherever a key predictor is entangled with unobservables, across the disciplines we serve.

Economics & Public Policy

The classic home of IV—using policy rules, random allocation, or institutional features to estimate effects of endogenous choices and exposures.

Finance & Accounting

Addressing endogeneity in relationships such as governance and performance, or financing and investment, with credible instruments.

Management & Organizational Research

Estimating effects of strategic or managerial choices that firms make non-randomly, where OLS conflates choice with confounders.

Marketing & Consumer Research

Handling endogenous marketing decisions—such as price or advertising set in response to demand—with instruments for the endogenous variable.

Strategy & Entrepreneurship

Effects of endogenous firm decisions (entry, investment, alliances) where selection makes naive estimates unreliable.

Operations & Information Systems

Effects of endogenously adopted technologies or practices, instrumented by exogenous sources of adoption variation.

FAQ

Instrumental variables: common questions

An instrumental variable (IV) is a variable that influences a treatment or predictor but affects the outcome only through that predictor, and is unrelated to the unobserved confounders. It is used to estimate causal effects when the predictor is endogenous—correlated with the error term—so ordinary least squares is biased. IV uses only the confounding-free variation in the predictor, typically via two-stage least squares.
2SLS is the standard IV estimator. In the first stage, the endogenous predictor is regressed on the instrument (and any controls) to obtain its predicted, confounding-free portion. In the second stage, the outcome is regressed on that predicted portion. This isolates the part of the predictor’s variation driven by the instrument, yielding a causal estimate when the instrument is valid.
Three: relevance (the instrument genuinely predicts the endogenous predictor, which is testable via first-stage strength), the exclusion restriction (the instrument affects the outcome only through the predictor), and independence (the instrument is unrelated to unobserved confounders). Relevance is testable; exclusion and independence generally are not and must be argued from theory and the research design.
A weak instrument is one that predicts the endogenous predictor only weakly. Weak instruments make 2SLS badly biased—often toward the OLS estimate it was meant to correct—and its standard inference unreliable. First-stage strength (for example, the first-stage F-statistic) is reported and assessed against established guidance, and weak-instrument-robust methods are used when strength is in doubt.
When treatment effects differ across units, IV/2SLS estimates a local average treatment effect (LATE)—the effect for the “compliers,” whose treatment status is shifted by the instrument—rather than the average effect for the whole population. This is a well-defined and meaningful quantity, but it is not always the population effect a reader might assume, so the estimand should be stated explicitly.

Facing an endogeneity problem?

If a key predictor is entangled with unobserved factors and you have a credible instrument, IV/2SLS can recover the causal effect—with first-stage strength checked, the exclusion argument made explicit, and the estimand stated plainly.