Causal Inference

Causal Machine Learning Services

Machine learning excels at prediction, but a good predictor is not a causal estimate. Causal machine learning brings the two together—using flexible algorithms to control for many covariates and to uncover how an effect varies across units, while preserving valid causal inference. It answers not just “what is the effect?” but “for whom is it larger?”

Causal machine learning combines machine-learning algorithms with causal-inference methods to estimate treatment effects. It uses flexible models to adjust for high-dimensional confounders and to estimate heterogeneous (subgroup and individual-level) treatment effects—via methods such as double/debiased machine learning and causal forests—while retaining valid confidence intervals, which naive prediction models do not provide.

Double/debiased ML Causal forests & CATE Heterogeneous effects, valid inference Reproducible, journal-ready
Heterogeneous treatment effects Estimated treatment effect varying across a covariate, with an average effect line and individual estimates spread above and below it. causal_ml · treatment effect varies by unit treatment effect covariate (e.g. firm size, tenure) 0 average effect effect is larger for some units than others
Effect heterogeneity estimated effect (CATE) average

What causal machine learning does

Machine-learning models are built to predict, and they do it well—but a model that predicts an outcome accurately does not tell you the causal effect of changing a treatment, and its coefficients cannot be read as effects. Causal machine learning bridges this gap: it deploys the flexibility of ML for the parts of a causal problem where flexibility helps—adjusting for confounders and modelling how effects vary—while using inference procedures that keep the resulting treatment-effect estimates valid.

Two contributions matter most. First, double (debiased) machine learning uses ML to flexibly adjust for a large number of potential confounders—too many, or too non-linear, for conventional regression—without the regularisation bias that would otherwise distort the effect estimate, and it delivers valid confidence intervals through a specific orthogonalisation and cross-fitting procedure. Second, methods such as causal forests estimate heterogeneous treatment effects: instead of one average effect, they estimate how the effect varies across units and subgroups (the conditional average treatment effect, or CATE), and can identify who benefits most. This moves the question from “does it work on average?” to “for whom, and by how much?”

When to use it

Causal ML is valuable when you have many potential confounders or complex, non-linear relationships that make manual specification difficult; when you want to estimate effect heterogeneity rigorously rather than through arbitrary subgroup splits; and when you are conducting policy or programme evaluation and need both a credible average effect and an understanding of how it is distributed. It also strengthens mediation and moderation questions where the relationships are high-dimensional.

One point deserves emphasis, because it is often misunderstood: causal ML is not a substitute for a research design. It relaxes the functional-form assumptions of traditional methods—how confounders enter the model—but it does not manufacture identification. Double machine learning for confounder adjustment still assumes the confounders are observed (as in selection-on-observables); it cannot fix unobserved confounding any more than matching can. The design assumptions still have to hold; causal ML makes the estimation more flexible and honest within them, not the identification stronger.

At a glance

Predictive ML vs causal ML

Two different goals for machine learning
Predictive MLCausal ML
GoalPredict the outcome accuratelyEstimate the effect of a treatment
AnswersWhat will happen?What happens if we intervene?
InferencePredictive accuracy metricsValid confidence intervals for effects
HeterogeneityNot the focusEstimates who benefits (CATE)
Still needsGood data & validationA valid identification strategy
Methodology

Valid inference, and honest heterogeneity

The reason causal ML is more than “ML applied to a treatment variable” is that naive approaches produce biased effect estimates and invalid inference. Double machine learning addresses this deliberately: it separates the prediction tasks (modelling the outcome and the treatment as functions of confounders) from the effect estimation, uses an orthogonal (Neyman) formulation that is insensitive to small errors in those predictions, and applies cross-fitting to avoid overfitting bias. The payoff is a treatment-effect estimate with valid standard errors even when flexible ML is used for the nuisance components—something a single black-box model cannot deliver.

For heterogeneous effects, honesty is the operative word—quite literally. Estimating how effects vary and then testing significance on the same data invites false discoveries, so methods like the causal forest use sample-splitting (“honest” estimation) so that the subgroups are found on different data from the one used to estimate their effects. Even so, discovered heterogeneity should be validated and interpreted cautiously: a subgroup with an apparently large effect needs confirmation, not just a striking chart. And because the models are flexible, we pair them with the transparency and sensitivity analysis that keep results interpretable and defensible rather than opaque.

Flexible estimation is not a shortcut around identification. Causal ML relaxes functional-form assumptions and can estimate who benefits, but it does not create causal identification—double ML for confounder adjustment still assumes confounders are observed. The design assumptions must hold; the method makes estimation within them more flexible and inference valid.

Software

We deliver causal ML in established, reproducible tools—R (grf, DoubleML) and Python (EconML, DoubleML)—with cross-fitting, valid confidence intervals, honest heterogeneity estimation, validation of discovered subgroups, and sensitivity analysis, all with versioned code.

How we work

How we deliver a causal ML study

Causal ML sits within our wider causal-inference practice—so the identification strategy comes first, and the machine learning serves valid effect estimation rather than replacing the design.

We start with the identification strategy—what makes a causal claim credible in your setting—and only then bring in ML for what it does well: flexibly adjusting for high-dimensional confounders (double ML) and estimating how effects vary across units (causal forests). We use cross-fitting for valid inference and honest estimation for heterogeneity, and validate any discovered subgroups rather than reporting them at face value.

Reporting sets out the identification assumptions, the ML methods and why they were used, the cross-fitting and honesty procedures, the average and heterogeneous effects with valid inference, and a sensitivity analysis—so the results are both flexible and defensible.

You receive the average treatment effect with valid confidence intervals, the heterogeneous (CATE) estimates identifying where effects are larger or smaller, the validation of any subgroups, a sensitivity analysis, and reproducible analytical code and analysis-ready files (where appropriate and permitted)—with the identification assumptions and the observed-confounding caveat stated plainly.

Where we apply it

Causal machine learning across Management & Allied Studies

Rich, high-dimensional data and the question of who benefits are increasingly central to applied research—so causal ML is a fast-growing tool across the disciplines we serve.

Economics & Public Policy

Programme and policy evaluation with many covariates, and estimating which groups a policy helps most—a core causal-ML setting.

Marketing & Consumer Research

Estimating heterogeneous responses to interventions—which customers respond most to a campaign, offer, or change.

Management & Organizational Research

Effects of practices adjusted for rich firm and employee covariates, and where those effects are strongest.

Finance & Accounting

High-dimensional confounder adjustment and heterogeneous effects across firms, markets, or conditions.

Operations & Information Systems

Effects of technology or process changes with many controls, and targeting where interventions pay off.

Education & Behavioural Science

Which learners or participants benefit most from an intervention, estimated rigorously rather than by ad hoc splits.

FAQ

Causal machine learning: common questions

Causal machine learning combines machine-learning algorithms with causal-inference methods to estimate treatment effects. It uses flexible models to adjust for high-dimensional confounders and to estimate heterogeneous (subgroup and individual-level) treatment effects—via methods such as double/debiased machine learning and causal forests—while retaining valid confidence intervals, which naive predictive models do not provide.
Predictive ML aims to forecast an outcome accurately and is judged on predictive accuracy; its coefficients are not causal effects. Causal ML aims to estimate the effect of intervening on a treatment, provides valid confidence intervals for that effect, and can estimate how it varies across units. Crucially, causal ML still requires a valid identification strategy—accurate prediction alone does not establish a causal effect.
Double machine learning uses ML to flexibly adjust for many potential confounders while keeping the treatment-effect estimate valid. It separates the prediction tasks from effect estimation, uses an orthogonal formulation that is insensitive to small prediction errors, and applies cross-fitting to avoid overfitting bias—delivering valid confidence intervals even when flexible ML is used for the nuisance components. It assumes the confounders are observed.
Heterogeneous treatment effects describe how an effect varies across units or subgroups (the conditional average treatment effect, or CATE) rather than a single average. A causal forest is a method for estimating them, using sample-splitting (“honest” estimation) so subgroups are found on different data than is used to estimate their effects, which guards against false discoveries. Discovered heterogeneity should still be validated and interpreted cautiously.
No. Causal ML relaxes functional-form assumptions—how confounders enter the model—but it does not create causal identification. Double ML for confounder adjustment still assumes the confounders are observed and cannot fix unobserved confounding, just as matching cannot. The design assumptions still have to hold; causal ML makes estimation more flexible and inference valid within them, not the identification stronger.

Many confounders, or need to know who benefits?

When your problem has high-dimensional controls or the distribution of the effect matters, causal machine learning estimates it with valid inference—double ML for adjustment, causal forests for heterogeneity, and the identification assumptions kept front and centre.