Simpson's Paradox, Explained

After reading this you will be able to spot when a pooled trend lies, compute the two slopes that disagree, and decide whether to trust the within-group answer or the combined one.

What Simpson's paradox is

Simpson's paradox is a reversal of direction. Split your data into groups, and every group shows the same trend, say, more of X goes with more of Y. Pool the groups into one cloud, and the trend flips: more X now goes with less Y. Nothing about the individual relationships changed. Only the act of merging did.

Here is a compact numeric hook. Group A sits low and left, at points like (1, 2) and (2, 3). Group B sits high and right, at (5, 6) and (6, 7). Inside each group, when X goes up by 1, Y goes up by 1: a slope of +1. But look at the group centres. Group A averages (1.5, 2.5); group B averages (5.5, 6.5). Now push group B down. Move its Y values to 1 and 2, so its centre falls to (5.5, 1.5). Each group still slopes up by +1 internally, but the line connecting the two centres now runs from high-left to low-right. Fit one line to all four points and the slope is negative. That is the paradox in four points.

The visualizer plays this out continuously. Two coloured groups each keep an upward regression line, while the line through everything slopes down. The hidden driver is group membership: a variable correlated with both axes that the pooled fit ignores.

When it matters and when it does not

The paradox matters whenever you plan to act on a pooled statistic. A hospital reports a higher death rate than its rival, but treats sicker patients. A drug looks worse overall, but was given more often to severe cases. A university seems to admit men at a higher rate, but women applied to more competitive departments. In each case the pooled number points one way and the honest, within-group number points the other.

It does not matter when group membership is genuinely irrelevant to both axes. If the groups are just random halves of one population, their centres line up along the same trend, and pooling changes nothing but the sample size. The paradox needs a confounder: something that shifts a group along X and along Y at the same time.

A reversed pooled trend is a signal, not a verdict. It tells you a confounder is present. It does not by itself tell you the within-group answer is the correct one to report. That depends on what question you are asking.

The formula and the intuition

Fit an ordinary least squares line to a set of points. The slope is the covariance of X and Y divided by the variance of X.

b = \frac{\sum_i (x_i - \bar{x})(y_i - \bar{y})}{\sum_i (x_i - \bar{x})^2}

Here x_i and y_i are the coordinates of point i, \bar{x} and \bar{y} are the means over all points in the fit, and b is the slope. The numerator is the sum of cross products; the denominator is the spread in X.

Now split the points into two groups. The pooled numerator breaks into two pieces: variation within each group, plus variation between the group centres. Write the total cross-product sum as

S_{xy} = \sum_{g} S_{xy}^{(g)} + \sum_{g} n_g (\bar{x}_g - \bar{x})(\bar{y}_g - \bar{y})

The first term, S_{xy}^{(g)}, is the within-group cross product for group g: it carries the sign of each group's own slope. The second term is the between-group part: it depends on how the group centres (\bar{x}_g, \bar{y}_g) sit relative to the grand mean, weighted by group size n_g. When the centres are arranged so that high-X groups have low Y, this term is large and negative. If it outweighs the sum of positive within-group terms, the pooled slope flips sign. That single algebraic split is the whole paradox.

Reproducing the demo data

The demo loads two groups of five points each. Group A (blue): (1, 4), (2, 5), (3, 6), (4, 7), (5, 8). Group B (orange): (6, 2), (7, 3), (8, 4), (9, 5), (10, 6). Watch the two slopes disagree.

  1. Group A centre: \bar{x}_A = 3, \bar{y}_A = 6. Within A, X and Y both rise by 1 per step, so S_{xy}^{(A)} = 10 and S_{xx}^{(A)} = 10. Slope b_A = 10/10 = +1.
  2. Group B centre: \bar{x}_B = 8, \bar{y}_B = 4. Same internal spacing, so S_{xy}^{(B)} = 10, S_{xx}^{(B)} = 10. Slope b_B = +1. Both groups slope up.
  3. Grand means over all 10 points: \bar{x} = 5.5, \bar{y} = 5.
  4. Between-group cross product: group A contributes 5(3 - 5.5)(6 - 5) = 5 \cdot (-2.5)(1) = -12.5. Group B contributes 5(8 - 5.5)(4 - 5) = 5 \cdot (2.5)(-1) = -12.5. Between sum = -25.
  5. Pooled numerator: S_{xy} = (10 + 10) + (-25) = -5. Pooled denominator: within S_{xx} = 20 plus between 5(2.5)^2 + 5(2.5)^2 = 62.5, total 82.5.
  6. Pooled slope: b = -5 / 82.5 \approx -0.0606.

Both groups climb at +1. The pooled line drifts down at -0.0606. The between-group term of -25 swamped the within-group total of +20, and the sign flipped.

Ten demo points. Each colour group climbs by 1 per step, but the least squares line through all ten runs slightly downhill (slope -0.0606, intercept 5.333).

Watch the sign flip

The instructive move is the vertical gap between the two group centres. Slide it and the pooled slope crosses zero.

With the demo groups, dropping group B's centre lowers the pooled slope. At a vertical offset of 0 (both centres at the same height), the pooled slope equals +1, matching the groups. As group B's centre falls below group A's, the pooled slope decreases and passes through zero, then goes negative, even though each group's own slope stays fixed at +1.

Reading the result correctly

When the pooled and within-group slopes disagree, ask which comparison answers your question. Two rules help.

Within-group slope
The relationship you would see if you held group membership fixed. This is usually the causal question: for a given patient severity, does the treatment help?
Pooled slope
The relationship in the mixed population, confounder and all. It answers a descriptive question: across everyone as they actually are, how do X and Y move together?

Most of the time the within-group answer is what you want, because the confounder is a nuisance you did not intend to compare across. In the demo, if the colour is patient severity and X is drug dose, the honest answer is +1: more dose helps within each severity level. The pooled -0.0606 only reflects that sicker patients (group B) both got shifted along X and had worse outcomes.

But not always. If group membership is itself the thing you are studying, and cannot be intervened on, the pooled figure may be the fair summary. Deciding requires knowledge outside the numbers. The statistics cannot pick for you.

As group B's centre drops (offset going negative), the pooled slope falls linearly and crosses zero near an offset of -0.9. Each group's own slope stays at +1 the whole time. The marker at -3.667 shows the offset used by the demo, where the pooled slope is about -0.0606 after accounting for the demo's own +5 gap.

Common mistakes

The first mistake is trusting a single pooled line without checking for structure. Always look at the scatter coloured by any grouping you have: sex, site, batch, cohort, year. If the colours form separated clouds, the pooled slope is suspect.

The second mistake is the opposite: assuming a confounder must exist because someone can name one. A named variable only causes reversal if it is correlated with both axes. Test it. Compute the between-group term. In the demo it was -25, large enough to flip a within-group total of +20. If the between-group term is near zero, pooling is safe.

The third mistake is stopping at two groups. Real confounders often have many levels, and a variable can even be a continuous confounder. The same algebra applies: split the cross product into within and between parts. If the between part fights the within part hard enough, the sign flips.

Before reporting any slope or rate, recompute it inside each level of every plausible confounder. If the direction is stable across all of them, the pooled number is honest. If it flips in even one split, report the within-group figures and explain the split.

Related tools

Simpson's paradox is one of several places where intuition about aggregated data misleads. To build the underlying feel for slopes and correlation, try Guess the Correlation, which trains your eye to read r off a scatter. For the machinery of confidence in a slope estimate, see Bootstrap Resampling Visualizer. To see how easily a "significant" pattern appears in pure noise, spend time with the p-Hacking Simulator. And for two more probability results that defy first guesses, look at the Monty Hall Simulator and Nontransitive Dice, where A beats B beats C beats A.

Frequently asked questions

Is Simpson's paradox a real paradox?

No. It is a genuine arithmetic fact, not a contradiction. Pooled averages weight groups by size and position, so a weighted combination of upward pieces can point downward. It surprises people because they expect trends to survive aggregation, but the algebra above shows exactly why they need not.

How do I know which answer to trust, pooled or within-group?

Decide what you would do with the number. If you want the effect of X holding the confounder fixed, use the within-group slopes. If you want a plain description of the mixed population as it is, the pooled slope is fair. The data alone cannot choose; you need to know whether the grouping variable is a nuisance or the point.

Can it happen with more than two groups?

Yes. The cross-product split into within and between parts works for any number of groups and even for a continuous confounder. As long as the between-group term has enough magnitude and the opposite sign, the pooled slope reverses.

Does a bigger sample fix it?

No. More data makes both the within-group and pooled slopes more precise, but they still disagree. In the demo, copying each point a thousand times leaves both group slopes at +1 and the pooled slope at -0.0606. Sample size sharpens estimates; it does not remove confounding.

What is the difference between this and a spurious correlation?

A spurious correlation is a link with no causal path, often driven by a lurking variable. Simpson's paradox is a specific, stronger case: the lurking variable does not just create a link, it reverses the sign of one. Every Simpson case involves confounding, but not every confounded correlation reverses direction.