Research Methods

Causal Inference & Policy Evaluation

Correlation is easy; a defensible causal claim is not. We build the identification strategy your question needs—the right source of exogenous variation, the right comparison group, and the validity tests that decide whether the design holds—so your effect estimate survives the scrutiny that causal work attracts.

Quasi-experimental & experimental R · Stata · Python Identification tested, not assumed Reproducible, journal-ready
Sample difference-in-differences plot A difference-in-differences chart: treatment and control groups follow parallel trends before an intervention, after which the treated group diverges upward, and the gap between the two lines represents the estimated treatment effect. did_estimate · treatment effect Time → intervention ATT
Sample output treated control
Overview

Identification is the whole game

In causal work, the estimator is rarely the hard part—the identification is. A regression coefficient becomes a causal effect only when something in the data creates variation in the treatment that is unrelated to the outcome's other determinants. Naming that source of variation, and defending it, is what separates a credible causal paper from a correlational one dressed up in causal language.

We start there. Given how your treatment or policy actually varied—a threshold, a staggered rollout, an eligibility rule, an instrument, a natural experiment—we choose the design that variation can support, then subject its assumptions to the tests a careful referee will run: parallel trends before a difference-in-differences, density and continuity around a discontinuity, first-stage strength and exclusion for an instrument.

The result is an effect estimate that comes with its own defense: the identifying assumption stated plainly, the evidence for it presented, and a sensitivity analysis showing how large a violation would have to be to change the conclusion. That is what makes a causal claim publishable—and usable for a real decision.

Who We Work With

For anyone who needs to prove an effect

If your question is "did X actually cause Y," this is the right desk to write to.

PhD Researchers & Doctoral Candidates

A credible identification strategy for a dissertation chapter—built to survive a committee and a referee, and explained so you can defend it yourself.

Faculty & Academic Researchers

Design and estimation for causal papers—modern DiD, RDD, IV, and matching done to the current methodological standard.

Policy Organizations & Agencies

Program and policy impact evaluation that stands up to external review—and reports honestly what the evidence does and doesn't support.

Research Institutes & Think Tanks

Applied causal work for evidence programs where the finding will be quoted, contested, and acted upon.

Development & Applied Economists

Field experiments, natural experiments, and quasi-experimental designs on program and administrative data.

Corporates & Industry R&D

Causal measurement of interventions—pricing, policy, and program changes—translating academic identification into commercial decisions.

Capabilities

The modern causal toolkit, matched to your variation

Organized by where the identifying variation comes from. If your setting suggests a design not listed here, ask—this is the core, not the boundary.

Panel & Time-Based Designs

Difference-in-differences & event studies

For treatments that switch on over time—including the modern estimators that handle staggered adoption correctly.

  • Difference-in-differences
  • Staggered DiD
  • Event-study models
  • Triple difference
  • Interrupted time series
  • Synthetic control
  • Synthetic DiD
Quasi-Experimental Designs

Discontinuities & instruments

For treatments assigned by a rule, a threshold, or an as-good-as-random instrument—with the validity tests each design requires.

  • Regression discontinuity
  • Regression kink
  • Instrumental variables
  • 2SLS
  • Propensity score matching
  • Entropy balancing
  • Treatment-effect models
Causal Machine Learning

Heterogeneity & high-dimensional controls

For estimating who is affected and by how much—modern methods used inside a credible design, not as a substitute for one.

  • Causal forests
  • Double machine learning
  • Heterogeneous treatment effects
  • Mediation analysis
  • Sensitivity analysis
  • Policy impact evaluation
How the Analysis Works

Six steps from question to credible effect

A transparent, best-practice sequence—the design chosen for your setting and every identifying assumption documented and tested. Nothing is a black box.

Steps are adapted to your source of variation: a threshold, a staggered rollout, an instrument, or a randomized intervention. We confirm the identification strategy with you before estimation begins.

  1. 1

    Frame

    Define the treatment, the outcome, the population, and—above all—the source of exogenous variation the claim will rest on.

    Inputs: treatment · outcome · source of variation

  2. 2

    Design

    Select the identification strategy that the variation can support, and specify the comparison group and estimand precisely.

    Designs: DiD · RDD · IV · matching · synthetic control

  3. 3

    Validate

    Test the identifying assumptions the design depends on—the exact checks a careful referee will run.

    Tests: parallel trends · McCrary · first-stage · balance

  4. 4

    Estimate

    Estimate the treatment effect with appropriate estimators and inference, using modern methods where staggered timing or heterogeneity require them.

    Inference: clustered · robust · randomization · bootstrap

  5. 5

    Stress-test

    Probe how far the result holds—placebo tests, alternative comparison groups, and sensitivity to assumption violations.

    Checks: placebo · alternative controls · sensitivity bounds

  6. 6

    Report

    Produce interpretable effect estimates, event-study and diagnostic figures, methodology, and reproducible code you keep.

    Output: estimates · figures · methods text · R/Stata/Python code

Rigor by default

The checks that decide whether a causal claim survives

An effect estimate is only as good as the assumption behind it. Testing that assumption—and showing how sensitive the result is to it—is standard on every causal engagement.

Included on every project

  • Explicit statement of the identifying assumption
  • Parallel-trends, pre-trend, and event-study evidence
  • Instrument strength and exclusion, or RDD validity tests
  • Placebo tests and sensitivity to assumption violations
  • Reproducible, versioned code you keep
What You Receive

Every engagement, delivered in full

Not a black-box result and a number, but a complete, documented package you can submit, defend, and reproduce.

  • Clean, documented datasets and analysis files
  • A clearly stated identification strategy
  • Validity and assumption tests for the design
  • Placebo, robustness, and sensitivity analysis
  • Event-study, effect, and diagnostic figures
  • Interpretation of the treatment effect and its scope
  • Reproducible R, Stata, or Python code
  • Journal-ready methodology and results sections
  • Technical responses to methodological reviewer comments, where required
Where this fits

Part of a larger arc

Causal inference is strongest when the design ahead of it is deliberate and the reporting after it is precise—each handled with the same care.

Stage 02 · Design

Research Design & Planning

Identification strategy, power, and specification decided before estimation begins.

Explore methods
Stage 06 · Validate

Statistical & Methodological Audit

An independent check of assumptions, specification, and reproducibility before submission.

Explore audit
Stage 08 · Publish

Publication & Research Support

Methods and results reporting, journal selection, and reviewer-response support.

Explore support
FAQ

Common questions

Answers to what most researchers and project leads ask before we begin a causal inference engagement.

By the source of variation in your data. A policy that switched on at a threshold suggests regression discontinuity; a staggered rollout across units and time suggests modern difference-in-differences; an as-good-as-random instrument suggests IV. We start from where the exogenous variation actually comes from, then choose the design its assumptions can support.
Yes. Classic two-way fixed-effects difference-in-differences is biased under staggered adoption with heterogeneous effects. We use modern estimators (Callaway & Sant'Anna, Sun & Abraham, and related approaches) and report event-study evidence on pre-trends and dynamic effects.
That is the core of the work. We run and report the checks each design relies on—parallel-trends and pre-trend tests for DiD, McCrary density and covariate-continuity tests for RDD, first-stage strength and over-identification for IV—plus sensitivity analysis for how far a violation would have to go to overturn the result.
Where they fit. Double machine learning, causal forests, and related methods are valuable for high-dimensional controls and for estimating heterogeneous treatment effects—but they are tools inside a credible design, not a substitute for one. We use them when the identification strategy justifies them.
Yes. Our Statistical & Methodological Audit reviews the identification strategy, the validity tests, and the robustness of an existing causal analysis—the exact points a referee scrutinizes—whether or not we ran it.
Yes. Policy impact evaluation for institutes, agencies, and organizations is a core use case. We apply the same identification rigor to applied evaluations, and report results in language decision-makers can act on without overstating what the evidence supports.
R, Stata, and Python, matched to the estimator and your environment. You receive versioned, commented code alongside the results and a methods section written to journal standards.
Yes, and it is the highest-leverage point to involve us. The credibility of a causal claim is mostly determined at the design stage—choosing the right source of variation, the comparison group, and the pre-registration and power strategy—long before estimation.

Trying to prove an effect?

Tell us how your treatment or policy varied—we'll tell you the design, the assumptions it rests on, and what it takes to defend it.