Meta-Analysis & Synthesis 10 min read

Understanding Heterogeneity in Meta-Analysis

A meta-analysis produces a single pooled estimate—but that number only means something if the studies behind it are estimating the same effect. Heterogeneity is the measure of whether they are, and interpreting it correctly is what separates a meaningful synthesis from a misleading average.

The headline output of a meta-analysis is a single pooled effect—a diamond at the bottom of a forest plot that seems to settle the question. But that number carries a hidden condition: it is only meaningful if the studies it combines are estimating broadly the same underlying effect. When they are not—when the true effect genuinely varies from study to study—the pooled average can obscure far more than it reveals, blending together effects that point in different directions into one number that describes none of them. Heterogeneity is the concept that captures this, and understanding it is essential to reading, and conducting, a meta-analysis honestly.

This guide explains what heterogeneity is, how it is measured, what to do about it, and—most importantly—why it should be treated as a finding to explain rather than a nuisance to suppress. It deepens our guide to conducting a systematic review and meta-analysis and reflects the synthesis practice in our Meta-Analysis & Evidence Synthesis work.

What heterogeneity is

In a meta-analysis, the individual studies never produce identical results—their effect estimates always differ somewhat. Some of that difference is simply sampling error: each study is based on a finite sample, so its estimate wobbles around the truth by chance. Heterogeneity refers to the variation beyond what chance alone would produce—genuine differences in the true effect across studies. When heterogeneity is low, the studies are essentially all estimating the same effect and differ only by chance; when it is high, the true effect itself varies across the studies, for real reasons.

The distinction matters enormously for interpretation. Low heterogeneity means a pooled average is a sensible summary—there is one effect, and the meta-analysis pins it down precisely. High heterogeneity means there is no single effect to summarise; the studies are answering the same question in different contexts and getting genuinely different answers, and averaging them produces a number that may not apply to any real situation.

Two forest plots: low heterogeneity with studies clustering, and high heterogeneity with studies scattered
Low heterogeneity means a pooled average is meaningful; high heterogeneity means the average hides more than it shows - a finding to explain.

How heterogeneity is measured

Several statistics quantify heterogeneity, and they answer slightly different questions. Cochran’s Q is a test of whether any heterogeneity is present beyond chance—but it is sensitive to the number of studies, often failing to detect real heterogeneity when studies are few and flagging trivial heterogeneity when they are many, so a significant or non-significant Q should not be over-interpreted. I-squared is the most widely reported: it expresses the proportion of the total variation across studies that is due to real heterogeneity rather than chance, on a 0–100% scale, which makes it intuitive and comparable across analyses. Tau-squared estimates the actual variance of the true effects across studies, in the units of the effect size, which is what random-effects models use directly.

A common mistake is to lean on rules of thumb for I-squared—treating, say, 25%, 50%, and 75% as fixed “low/moderate/high” cut-offs. These are rough guides, not laws: the same I-squared can mean different things depending on the effect sizes and the number of studies, and I-squared says nothing about the direction or practical importance of the variation. The statistics should inform judgement, not replace it—the forest plot, showing where each study actually falls, is often more revealing than any single number.

Heterogeneity is a finding, not a flaw. High heterogeneity is not a sign the meta-analysis failed—it is telling you that the effect depends on something. The right response is to investigate what, not to hide it behind a pooled average.

Fixed-effect vs random-effects models

How you handle heterogeneity is bound up with the choice of model. A fixed-effect model assumes there is one true effect that every study is estimating, with differences due only to sampling error—appropriate only when heterogeneity is genuinely negligible. A random-effects model assumes the true effect varies across studies and estimates the average of that distribution of effects, explicitly accounting for the between-study variance (tau-squared). Because genuine heterogeneity is the norm rather than the exception in social-science, management, and most applied research—studies differ in populations, settings, measures, and design—the random-effects model is usually the more realistic and defensible default.

It is worth being clear about what each pooled estimate then means. Under a fixed-effect model, the pooled number is an estimate of the one common effect. Under a random-effects model, it is an estimate of the average effect across a distribution—which is a different, and often more honest, quantity when effects genuinely vary. Reporting the model choice and its rationale is part of a transparent meta-analysis.

Investigating and explaining heterogeneity

When substantial heterogeneity is present, the valuable work is explaining it. Two tools do this. Subgroup analysis splits the studies into groups—by population, setting, method, or design—and asks whether the effect differs systematically between them; finding that an intervention works in one context but not another is often more useful than the overall average. Meta-regression generalises this, modelling how the effect size varies with study-level characteristics treated as continuous or multiple moderators. Both turn heterogeneity from a problem into a source of insight: they identify what the effect depends on, which is frequently the most important contribution a synthesis can make.

These analyses come with cautions—they are observational comparisons across studies, vulnerable to confounding at the study level, and prone to false positives if many subgroups are tested—so they are best specified in advance and interpreted as exploratory unless pre-planned. But approached carefully, investigating heterogeneity is where a meta-analysis earns its keep, moving beyond “on average, this works” to “this works, for these cases, for these reasons.”

When not to pool at all

Sometimes the honest conclusion is that heterogeneity is so severe, or the studies so diverse, that no pooled estimate is meaningful. If the studies measure genuinely different things, in incomparable populations, using incompatible designs, forcing them into one number produces a meaningless average dressed up as a precise finding. In such cases a structured narrative synthesis, or a synthesis that groups only comparable studies, is more truthful than a single diamond. Recognising when not to pool is itself part of the analysis—and resisting the pull to compute a headline number simply because the software will is a mark of a careful reviewer.

The bottom line

Heterogeneity is the question at the heart of every meta-analysis: are these studies estimating the same effect, or different ones? Measure it with I-squared and tau-squared alongside the forest plot, but judge it—don’t apply cut-offs mechanically. Where it is genuine, prefer a random-effects model, investigate the sources through subgroup analysis or meta-regression, and be willing not to pool when the studies are simply too different. Treated this way, heterogeneity stops being an inconvenience and becomes what it should be: the most informative part of the synthesis, telling you not just whether an effect exists, but where, and why, it varies.

Frequently asked questions

Heterogeneity is variation in the true effect across the included studies, over and above the variation expected from sampling error (chance). Low heterogeneity means the studies are essentially estimating the same effect; high heterogeneity means the true effect genuinely differs across studies, so a single pooled average may not describe any real situation well.
Common measures include Cochran’s Q (a test for whether heterogeneity exists beyond chance, sensitive to the number of studies), I-squared (the proportion of total variation due to real heterogeneity, on a 0–100% scale), and tau-squared (the estimated variance of true effects, in effect-size units, used by random-effects models). These inform judgement rather than replacing it—the forest plot is often more revealing.
A fixed-effect model assumes one true effect that all studies estimate, appropriate only when heterogeneity is negligible. A random-effects model assumes the true effect varies and estimates the average of that distribution. Because genuine heterogeneity is the norm in applied research—studies differ in populations, settings, and design—the random-effects model is usually the more realistic and defensible default.
Investigate it rather than hide it. Use subgroup analysis (splitting studies by population, setting, method, or design) and meta-regression (modelling how the effect varies with study characteristics) to explain what the effect depends on. If heterogeneity is so severe that the studies are not comparable, the honest choice is not to pool at all, but to present a structured narrative synthesis instead.

Making sense of heterogeneity in your synthesis?

From choosing the right model to subgroup analysis and meta-regression, our team can help you interpret heterogeneity correctly—and know when a pooled estimate is, and isn’t, appropriate.