Propensity Score Matching & Weighting Services
When treated and untreated groups differ on observed characteristics, a raw comparison confuses the treatment with those differences. Propensity score methods—matching, weighting, and entropy balancing—construct a comparison that is balanced on the observed covariates, so the estimate reflects the treatment rather than pre-existing group differences.
Propensity score methods estimate treatment effects from observational data by balancing the treated and untreated groups on observed covariates. The propensity score is each unit’s estimated probability of being treated given its characteristics; matching, weighting, or balancing on it makes the groups comparable—so under the assumption of no unobserved confounding, the adjusted comparison estimates the treatment effect.
What propensity score methods do
In observational data, units are not randomly assigned to treatment—firms choose to adopt a practice, employees select into a programme, customers opt into a scheme. As a result, the treated and untreated groups usually differ systematically on their characteristics, and a raw outcome comparison mixes the treatment effect with those pre-existing differences (selection bias). Propensity score methods address this by making the groups comparable on observed covariates before the outcomes are compared.
The propensity score is a unit’s estimated probability of receiving the treatment given its observed characteristics. Its usefulness comes from a key property: comparing treated and untreated units with the same propensity score is, on the observed covariates, like comparing like with like. That score can then be used in several ways—matching treated units to similar untreated ones, weighting units (inverse-probability weighting) to construct a balanced pseudo-population, or as one input alongside outcome modelling. Entropy balancing is a related weighting approach that reweights the comparison group to match the treated group’s covariate moments directly, achieving balance by construction. The goal in every case is the same: a comparison in which the groups differ (on observables) only in whether they were treated.
When to use them—and the crucial caveat
These methods suit observational studies where treatment was not randomized, you can measure the characteristics that drive selection, and there is enough overlap between the groups to find comparable units. They are widely used across management, economics, and policy for exactly the settings where an experiment is impossible but rich covariate data exist. Related treatment-effect models extend the toolkit to specific selection structures.
The crucial caveat must be stated plainly: propensity score methods only adjust for observed covariates. Unlike difference-in-differences or instrumental variables, they do not address unobserved confounding—they rest on the strong assumption of no hidden bias (“selection on observables”). If an unmeasured factor drives both treatment and outcome, the estimate remains biased no matter how good the balance on observables looks. This is why matching alone is not a substitute for a design-based approach, and why credible practice always accompanies it with a sensitivity analysis for hidden bias.
Ways to use the propensity score
| Approach | How it creates balance | Notes |
|---|---|---|
| Matching | Pairs treated units with similar untreated units | Intuitive; may discard unmatched units |
| Inverse-probability weighting | Weights units by the inverse of their treatment probability | Uses all units; sensitive to extreme weights |
| Entropy balancing | Reweights to match covariate moments directly | Achieves specified balance by construction |
| Doubly robust | Combines weighting with outcome modelling | Consistent if either model is correct |
Balance, overlap, and hidden bias
The right way to judge a propensity score analysis is not the propensity model’s fit but the balance it achieves. After matching or weighting, the treated and comparison groups should be similar on every covariate—assessed with standardized mean differences (commonly against a small threshold) and distributional checks, and reported in a balance table or love plot. If balance is not achieved, the specification is revised until it is; a significant treatment estimate on top of poor balance is not credible. Overlap (common support) matters too: there must be treated and untreated units across the range of propensity scores, or some units have no comparable counterparts and the estimate relies on extrapolation.
The decisive issue, however, is unobserved confounding. Because these methods assume selection on observables, the honest complement to any propensity score estimate is a sensitivity analysis that asks how strong an unmeasured confounder would have to be to overturn the result. A finding that is robust to plausible hidden bias is more credible than one that a modest unobserved factor could erase. We treat this sensitivity assessment as part of the analysis, not an optional extra—and we are clear that a well-balanced matched sample is still an observational comparison, not a randomized experiment.
Balancing observables does not remove unobserved confounding. Propensity score methods make groups comparable only on measured covariates; a hidden confounder still biases the estimate. Credible practice reports balance and a sensitivity analysis for hidden bias—and treats matching as an observational, not experimental, comparison.
Software
We deliver propensity score matching, weighting, entropy balancing, and doubly robust estimation in established, reproducible tools—R (MatchIt, WeightIt, cobalt, ebal) and Stata—with balance diagnostics, overlap checks, appropriate inference, and hidden-bias sensitivity analysis, all with versioned code.
How we deliver a propensity score analysis
Propensity score methods sit within our wider causal-inference practice—so balance is demonstrated, overlap is checked, and the selection-on-observables assumption is tested rather than assumed.
We start by identifying the covariates that drive selection into treatment and confirming there is enough overlap to support a credible comparison. We estimate the propensity score, apply the appropriate approach (matching, weighting, entropy balancing, or a doubly robust combination), and—before interpreting any effect—demonstrate covariate balance and assess common support.
Reporting sets out the covariates and selection argument, the balance achieved (table or love plot), the overlap assessment, the estimator, and a sensitivity analysis for unobserved confounding—so the strength and the limits of the causal claim are both clear.
You receive the treatment-effect estimate with appropriate inference, the balance and overlap diagnostics, the hidden-bias sensitivity analysis, robustness checks across specifications, and reproducible analytical code and analysis-ready files (where appropriate and permitted)—with the selection-on-observables assumption and its implications stated plainly.
Propensity score methods across Management & Allied Studies
Non-random selection into a practice, programme, or choice is the norm in observational social science—so propensity score methods are widely used across the disciplines we serve, where rich covariate data exist.
Management & Organizational Research
Effects of practices, certifications, or programmes that firms or employees adopt non-randomly, balanced on observed characteristics.
Economics & Public Policy
Programme and policy evaluation where participants self-select and rich covariates are available—a core matching setting.
Finance & Accounting
Effects of corporate choices (adoptions, listings, financing) on firms matched to comparable non-adopters.
Marketing & Consumer Research
Effects of programme enrolment or channel adoption where customers opt in, using matched or weighted comparisons.
Education & Learning Sciences
Effects of participation in a programme or intervention where enrolment is non-random but observable characteristics are rich.
Operations & Information Systems
Effects of voluntarily adopted technologies or practices, comparing adopters to balanced non-adopters.
Propensity score methods: common questions
Comparing groups that were not randomly assigned?
If treated and untreated units differ on measurable characteristics, propensity score matching, weighting, or entropy balancing can construct a comparable comparison—with balance demonstrated, overlap checked, and hidden-bias sensitivity reported.