Meta-Analysis & Evidence Synthesis

Multilevel & Three-Level Meta-Analysis Services

When studies report more than one effect size—several outcomes, subgroups, or time points each—those effects are not independent, and a standard meta-analysis gets the statistics wrong. Multilevel (three-level) meta-analysis models the nesting of effects within studies directly, so your pooled estimate and its uncertainty are correct.

Multilevel meta-analysis—commonly a three-level meta-analysis—accounts for effect sizes nested within studies. Level 1 is the sampling variance of each effect, level 2 is variation between effects within a study, and level 3 is variation between studies. It correctly handles multiple, dependent effect sizes per study that standard meta-analysis assumes are independent.

Handles dependent effect sizes Within- & between-study variance R (metafor) · robust variance Reproducible, journal-ready
Three-level nesting of effect sizes within studies A diagram showing effect sizes nested within studies, which are nested within the overall meta-analysis, across three levels. three_level_meta_analysis · nesting L3: between studies Study A (L2) Study B (L2) each dot = one effect size (L1) effect (L1)
Nested structure effect size study

What multilevel meta-analysis does

Standard meta-analysis assumes each study contributes one independent effect size. In reality, a single study often reports many—several outcome measures, several subgroups, multiple time points, or several comparisons—and these effects from the same study are correlated. Treating them as if they were independent double-counts information: it understates the standard errors, narrows confidence intervals artificially, and produces “significant” findings that will not hold.

A multilevel (three-level) meta-analysis solves this by modelling the data’s nested structure directly. It partitions the variance into three levels: level 1, the sampling variance of each individual effect size; level 2, the variation among effect sizes within a study; and level 3, the variation between studies. By estimating the within-study and between-study variance separately, it uses all the effect sizes without pretending they are independent—so the pooled estimate and its uncertainty are correct.

When to use it

Reach for a three-level model whenever your studies contribute more than one effect size each and you do not want to throw information away. The common alternatives are worse: averaging each study’s effects into one discards real variation, and selecting a single effect per study is arbitrary and wasteful. A three-level model keeps every effect while respecting the dependence.

It is the natural choice when you want to use all reported effects, when you are interested in within-study variation in its own right, or when you plan to run moderator analyses that need the full set of effects. Where the correlation structure among effects is complex or unknown, it pairs naturally with robust variance estimation for extra protection.

At a glance

Handling multiple effects per study

Three ways to deal with dependent effect sizes
Average per studyPick one per studyThree-level model
Uses all effects?No (collapsed)No (discarded)Yes
Respects dependence?Yes, but crudelyAvoids it by dropping dataYes, modelled
Within-study variance?LostLostEstimated
RiskThrows away variationArbitrary, wastefulNeeds enough data per level
Methodology

Getting it right

A three-level model estimates two variance components—within-study (level 2) and between-study (level 3)—and interpreting them is part of the value: the split tells you whether heterogeneity sits mostly within studies or between them, which shapes how you read the evidence and design moderator analyses. Estimating both reliably needs enough studies and enough effects per study; with too little data at a level, that variance component is hard to pin down.

The three-level model assumes a particular correlation structure among within-study effects. When that structure is unknown or complex, robust variance estimation (RVE) is the standard companion—it produces valid standard errors without requiring the correlations to be known, and modern practice often combines a multilevel model with RVE for the most defensible inference.

Multiple effects per study are not free data—they are dependent data. Ignoring the nesting inflates significance; a three-level model (often with robust variance estimation) uses every effect while keeping the inference honest.

Software

We fit three-level models in established, reproducible tools—principally R’s metafor package, with robust variance estimation where appropriate—and deliver versioned code so the analysis can be checked and rerun.

How we work

How we deliver a three-level meta-analysis

Multilevel meta-analysis sits within our wider meta-analysis and evidence-synthesis service, run on a full systematic-review workflow—so the model is built on a sound, reproducible review.

We begin with a registered protocol, a comprehensive documented search, careful extraction of every effect size with its study identifier, and risk-of-bias assessment. We then fit the three-level model, add robust variance estimation where the dependence structure warrants it, and run the moderator and sensitivity analyses your question needs.

Reporting follows PRISMA standards, with the multilevel structure and variance components reported transparently.

You receive the pooled estimate with correct standard errors, the within- and between-study variance components and what they imply, any moderator results, a full heterogeneity and bias assessment, and reproducible code and data. The result is an analysis that uses all your evidence and stands up in peer review.

Where we apply it

Three-level meta-analysis across Management & Allied Studies

Multiple effect sizes per study are the norm in social-science synthesis—so three-level models are widely applicable across the disciplines we serve.

Management & Organisational Studies

Studies reporting effects across several outcomes, teams, or measures—pooled without double-counting the dependent effects.

Economics & Public Policy

Evaluations reporting multiple estimates per study or programme, kept in the analysis with honest standard errors.

Marketing & Consumer Research

Experiments with several conditions or outcome measures per paper, synthesised while respecting within-study correlation.

Finance & Accounting

Studies reporting effects across markets, periods, or specifications—modelled as effects nested within studies.

Education & Learning Sciences

A classic three-level setting: multiple effect sizes per study across outcomes, grades, or subgroups.

Psychology & Behavioural Science

Multi-outcome and multi-measure studies pooled with within- and between-study variance modelled separately.

FAQ

Three-level meta-analysis: common questions

A three-level meta-analysis models effect sizes nested within studies. It partitions variance into three levels: level 1 is the sampling variance of each effect, level 2 is variation among effects within a study, and level 3 is variation between studies. This correctly handles multiple, dependent effect sizes per study that a standard two-level meta-analysis wrongly assumes are independent.
Whenever your studies contribute more than one effect size each—multiple outcomes, subgroups, time points, or comparisons—and you want to keep all of them. The alternatives (averaging effects per study, or picking one) either discard information or are arbitrary. A three-level model uses every effect while modelling the dependence, giving correct standard errors.
Because effects from the same study are correlated. Treating them as independent double-counts information, understates the standard errors, narrows confidence intervals artificially, and produces significant-looking findings that will not replicate. A three-level model accounts for the within-study correlation so the inference is honest.
Robust variance estimation (RVE) produces valid standard errors for dependent effect sizes without needing to know the exact correlation structure among them. It is the standard companion to a three-level model when that structure is unknown or complex, and modern practice often combines the two for the most defensible inference.
Estimating both variance components reliably needs enough studies and enough effect sizes per study. With too little data at a level, that variance component is hard to pin down. We assess whether the structure your data supports is appropriate and, where it is thin, use robust variance estimation and report the limitation transparently.

Studies with several effect sizes each?

If your evidence base has multiple correlated effects per study, a three-level meta-analysis lets you use all of them without inflating significance. We design and deliver it—with robust variance estimation where it helps.