Causal Inference

Difference-in-Differences Services

When a policy, reform, or intervention affects some units but not others, difference-in-differences (DiD) estimates its causal effect by comparing how outcomes change over time in the treated group against a comparison group. It is one of the most widely used quasi-experimental designs—and one where recent methodological advances have changed best practice.

Difference-in-differences (DiD) estimates the causal effect of a treatment by comparing the before-to-after change in outcomes for a treated group with the change over the same period for an untreated comparison group. Subtracting the comparison group’s change removes trends common to both, isolating the treatment effect—provided the parallel-trends assumption holds.

Parallel-trends tested Staggered & modern estimators Event-study & robustness Reproducible, journal-ready
The difference-in-differences design Treated and control outcomes over pre and post periods, with a dashed counterfactual parallel to the control trend and the treatment effect shown as the gap. difference_in_differences · parallel trends pre post intervention effect treated control counterfactual
Parallel trends treated control

What difference-in-differences does

Many important questions in management, economics, and policy concern the effect of something that happened to some units but not others—a new regulation adopted by some states, a minimum-wage change in some regions, a programme rolled out to some firms, a reform some organizations underwent. A simple before-and-after comparison of the treated group is confounded by everything else changing over time; a simple treated-versus-untreated comparison is confounded by pre-existing differences between the groups. Difference-in-differences addresses both at once.

It takes two differences. The first is the before-to-after change in the treated group. The second is the before-to-after change in a comparison group that was not treated. Subtracting the second from the first removes any trend common to both groups—the general drift the treated group would have followed anyway—leaving the effect attributable to the treatment. The comparison group, in effect, supplies the counterfactual: what would have happened to the treated group without the intervention.

When to use it

DiD suits natural-experiment settings: a policy, reform, or intervention affects some units at a known time while comparable units are unaffected, and you observe outcomes for both groups before and after. It is a workhorse of policy evaluation and empirical management and economics precisely because such staggered, partial roll-outs are common and randomized experiments are often impossible.

Its credibility rests on a specific, testable-in-part assumption: parallel trends—that, absent the treatment, the treated and comparison groups would have followed the same trend over time. If the groups were already diverging before the intervention, the design attributes that divergence to the treatment and misleads. This is the assumption we scrutinise first, before any estimate is interpreted, and it is why a defensible DiD is far more than a regression with an interaction term.

At a glance

What DiD handles that simpler comparisons don’t

Isolating the treatment effect
ComparisonWhat it estimatesMain confound
Treated: before vs afterChange over time in the treated groupEverything else changing over time
After: treated vs controlGroup difference after treatmentPre-existing differences between groups
Difference-in-differencesTreatment effect, net of common trendsRelies on parallel trends holding
Methodology

Parallel trends—and the staggered-timing problem

The parallel-trends assumption cannot be proven, but it can be examined. The standard evidence is an event-study specification, which estimates the treated–control difference in each period before and after the intervention: if the pre-treatment differences are flat and close to zero, the trends were plausibly parallel before treatment, which supports (though does not guarantee) the assumption afterwards. We report this evidence, consider whether anything other than the treatment changed at the same time, and treat parallel trends as something to defend rather than assert.

A major development concerns staggered adoption—when units are treated at different times. It has been shown that the conventional two-way fixed-effects regression, long used for this case, can produce badly biased or even wrong-signed estimates, because already-treated units get used as comparisons for later-treated ones. A family of modern estimators addresses this, and using them where timing is staggered is now expected in careful work. We select the estimator to match the treatment-timing structure rather than defaulting to the traditional specification.

With staggered timing, the classic two-way fixed-effects DiD can mislead. When units are treated at different times, the traditional regression can be biased—so modern staggered-DiD estimators are used, and parallel trends is examined with an event study rather than assumed.

The wider DiD family & software

Related designs extend DiD when its assumptions are strained: triple-difference adds a second comparison to net out confounding trends; synthetic control and synthetic DiD construct a weighted comparison when no single clean control exists. We deliver DiD and its extensions in reproducible tools (R and Stata, including modern staggered-DiD packages) with event-study plots, robustness checks, and appropriately clustered inference.

How we work

How we deliver a difference-in-differences study

DiD sits within our wider causal-inference practice—so the design is scrutinised, the estimator matches the timing structure, and the parallel-trends evidence is presented honestly.

We start from your setting—what was treated, when, and which units form a credible comparison—and assess whether a DiD design is defensible before estimating anything. We examine parallel trends with an event-study specification, select an estimator appropriate to the treatment timing (including modern staggered-adoption methods), and run the robustness and placebo checks that a careful DiD requires.

Reporting sets out the design and its identifying assumption, the parallel-trends and event-study evidence, the estimator and why it was chosen, and the inference approach—transparently, so a reviewer can judge the causal claim.

You receive the treatment-effect estimate with appropriate (typically clustered) inference, the event-study plot and parallel-trends assessment, robustness and placebo checks, and reproducible analytical code and analysis-ready files (where appropriate and permitted). The result is a causal estimate whose assumptions are stated and defended, not a coefficient presented without its design.

Where we apply it

Difference-in-differences across Management & Allied Studies

Staggered, partial roll-outs of policies and interventions are everywhere in applied social science—so DiD is among the most-used causal designs in the disciplines we serve.

Economics & Public Policy

Evaluating reforms, regulations, taxes, and programmes that some regions or groups adopt while others do not—a core DiD setting.

Management & Organizational Research

Estimating the effect of a practice, policy, or event adopted by some firms or units at a known time.

Finance & Accounting

Effects of regulatory changes, disclosure rules, or listing events on firms exposed versus comparable firms that are not.

Marketing & Consumer Research

Impact of a campaign, price change, or platform change rolled out to some markets or segments and not others.

Strategy & Entrepreneurship

Effects of policy shocks, entry events, or ecosystem changes on exposed versus unexposed firms.

Operations & Information Systems

Effects of a technology, process, or system change adopted by some sites or units at staggered times.

FAQ

Difference-in-differences: common questions

Difference-in-differences (DiD) estimates the causal effect of a treatment by comparing the before-to-after change in outcomes for a treated group with the change over the same period for an untreated comparison group. Subtracting the comparison group’s change removes trends common to both groups, isolating the treatment effect—provided the parallel-trends assumption holds.
Parallel trends is the assumption that, absent the treatment, the treated and comparison groups would have followed the same trend over time. It is the key identifying assumption of DiD. It cannot be proven, but it can be examined—most commonly with an event-study specification showing that pre-treatment differences between the groups were flat. If the groups were already diverging before treatment, DiD can misattribute that divergence to the treatment.
When units are treated at different times, the conventional two-way fixed-effects regression can produce biased or even wrong-signed estimates, because already-treated units are used as comparison units for later-treated ones. Recent econometric work established this problem and developed estimators that avoid it. Where treatment timing is staggered, using a modern staggered-DiD estimator rather than the traditional specification is now expected in careful work.
A simple before-after comparison of the treated group attributes all change to the treatment, even though other things change over time. DiD adds an untreated comparison group and subtracts its change, removing trends common to both—so the estimate reflects the treatment rather than the general drift the treated group would have experienced anyway. This is what makes DiD a causal design rather than a descriptive one.
Where no single untreated group is a credible counterfactual, related designs can help: a triple-difference adds a further comparison to net out confounding trends, and synthetic control or synthetic DiD constructs a weighted combination of untreated units to serve as the comparison. Which is appropriate depends on the data and setting; we assess whether a credible counterfactual can be constructed before proceeding.

Evaluating a policy, reform, or roll-out?

If some units were affected and comparable units were not, difference-in-differences can estimate the causal effect—with parallel trends examined, the right estimator for your timing, and robustness reported in full.