Berkson's Paradox, Explained

After reading this you will be able to explain why two independent traits look negatively correlated inside any group filtered by their sum, compute the effect yourself, and spot it in real datasets built through admission, hiring or publication filters.

What it is and why the handsome ones seem rude

Pick two traits that have nothing to do with each other in the general population. Call them attractiveness and kindness. In the whole species their correlation is roughly zero: knowing someone is attractive tells you nothing about whether they are kind.

Now filter. You only date people who clear some minimum standard, and that standard is a combination: you will overlook a bit of rudeness if the person is striking, and you will overlook plain looks if the person is very kind. In effect you accept anyone whose sum of the two traits clears a bar. Inside that dating pool a strange thing happens. The attractive people in your pool tend to be less kind, and the kind ones tend to be less attractive. A negative correlation appears out of nowhere.

Nothing changed about the people. The correlation is an artifact of the filter. This is Berkson's paradox: conditioning on the sum of two independent variables makes them negatively correlated. Joseph Berkson described it in 1946 for hospital data, where two unrelated diseases correlate among inpatients because either disease raises the chance of admission.

When the effect appears, and when it does not

You get Berkson's paradox whenever three things hold. First, you have two variables that combine to decide selection. Second, selection keeps the high combinations and drops the low ones (a threshold on the sum, or a top-k cut). Third, you then measure the correlation within the selected group only, forgetting the ones you dropped.

The classic settings are all selection filters. Hospital admission selects on illness burden. Hiring selects on a blend of skill and interview polish. Academic publication selects on novelty plus statistical significance. In each case the correlation you measure inside the filtered set says more about the filter than about the world.

The paradox is not the same as Simpson's paradox. Simpson's paradox is about a trend reversing when you merge groups. Berkson's paradox is about a correlation appearing when you select a subset. Both fool you, but by different mechanisms.

The effect is weak or absent when selection is random, when it depends on only one of the two traits, or when you keep almost everyone (a very low threshold barely cuts the cloud). It is strongest when the bar is high and the true correlation is near zero.

The formula and the intuition behind it

Let the two traits be X and Y, each with variance \sigma^2 and true correlation \rho. The selection keeps points where the sum S = X + Y is large. The intuition is arithmetic: if the sum is forced to be high, then whenever X is high, Y is free to be low, and whenever X is low, Y must be high to compensate. Conditioning on the sum ties the two together with opposite signs.

Take the clean case \rho = 0 and equal variances. Before selection the covariance of X and Y is zero. Now condition on a fixed value of the sum, X + Y = s. Then Y = s - X exactly, so the two are perfectly and negatively related along that line:

\text{Corr}(X, Y \mid X + Y = s) = -1

Here the left side is the correlation of X and Y once you fix the sum to a single number s. Fixing the sum forces a straight line of slope -1, so the correlation is exactly -1.

Real selection uses a threshold, not a single value: keep everyone with S \ge t. That admits a band of sums, not one line, so the induced correlation lands between 0 and -1. The higher the threshold t, the thinner the surviving band and the closer the measured correlation gets to -1. Sample correlation is the usual Pearson coefficient:

r = \frac{\sum_{i}(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_i (x_i - \bar{x})^2}\,\sqrt{\sum_i (y_i - \bar{y})^2}}

where x_i, y_i are the trait values of point i, and \bar{x}, \bar{y} are the means computed over whichever set you are measuring. That last phrase is the whole paradox: compute r over everyone and you get about 0; compute it over the selected set and you get a solid negative.

Reproducing the demo data

The demo scatters N = 500 points with two independent standard-normal traits (true correlation \rho = 0) and keeps everyone whose trait sum clears a threshold. Work a small version by hand first, then read off the full run.

  1. Take eight points with coordinates (x, y): (2.0, -1.0), (-1.5, 1.8), (0.4, 0.6), (1.2, 1.0), (-0.8, -0.5), (1.6, 0.2), (0.9, 1.4), (-1.1, 0.3). By construction the two columns are unrelated: their full-set correlation is about -0.05, near zero.
  2. Apply the rule "keep sum x + y \ge 1.5". The sums are 1.0, 0.3, 1.0, 2.2, -1.3, 1.8, 2.3, -0.8. Four points survive: (1.2, 1.0), (1.6, 0.2), (0.9, 1.4), and no others reach the bar except those two plus the sum-2.2 and sum-2.3 points, giving the kept set (1.2, 1.0), (1.6, 0.2), (0.9, 1.4).
  3. For the kept set the means are \bar{x} = 1.233 and \bar{y} = 0.867. The centered products sum to a negative number: the high-x point (1.6, 0.2) pairs high x with low y, dragging the correlation down.
  4. Computing Pearson r on those three kept points gives about -0.86. The traits went from unrelated to strongly negative purely because the filter kept only the high-sum corner.

The full N = 500 run at a threshold that keeps the top third of the cloud produces a selected correlation near -0.5, while the full-population correlation stays within 0.05 of zero. Toggle the dimming of rejected points and the negative slope is all you see, exactly as a hospital study or a dating pool would present it.

What the two regression lines are telling you

The tool draws two regression lines: one through all points, one through the selected points. The full-population line is nearly flat, slope close to 0, because the traits are independent. The selected-set line slopes downward. Read the gap between them as the size of the illusion the filter creates.

The full cloud shows no tilt (the flat line has slope near zero). The high-sum corner (upper right) is what a sum threshold keeps, and inside it the tilt runs downward.

The measured slope of the selected line depends on the threshold. Slide the bar up and the surviving band thins toward the anti-diagonal, so the slope steepens toward -1. Slide it down and you readmit the rounded cloud, so the slope relaxes back toward 0. The chart below traces that relationship.

Keep everyone (100%) and the correlation is zero. Cut harder and the induced negative correlation grows, approaching -1 as the kept fraction shrinks to a thin sliver.

A scatter of 500 points with two independent standard-normal traits, plus a slider for the sum threshold. As you raise the threshold from the low end (keeping all points, correlation near 0) to the high end (keeping a thin corner, correlation near -0.9), the kept points highlight, the two regression lines redraw, and the two correlation numbers update live.

Common mistakes when reading filtered data

The first mistake is trusting a correlation from data you did not sample yourself. Admissions records, customer lists, survey respondents who bothered to reply: all are filtered sets. A negative correlation inside them may be pure Berkson.

The second mistake is inventing a mechanism. Faced with "attractive people are ruder," the mind reaches for a causal story about ego or entitlement. The math needs no such story. A sum threshold produces the same pattern from two coins that never met.

The cure is knowing your denominator. Before you interpret a correlation, ask what filter produced the sample and whether that filter depends on both variables. If it does, expect a spurious negative and look for the rejected points you never saw.

The third mistake is assuming the effect is small. In the worked example a threshold keeping the top few points pushed the correlation to -0.86 from a true value of about -0.05. Selection can flip a correlation clean past zero and make it strong.

Related tools on this site

If selection artifacts interest you, three neighbors are worth a visit. The Simpson's paradox visualizer shows the merging version of a fooling correlation. The Base rate visualizer shows how a filter (a positive test) reshapes probabilities the same way. And Regression to the mean explains why extreme selected performers drift back, another artifact of picking the tail.

To build intuition for correlation itself, try Guess the correlation, and to see how easily filtering and repeated testing conjure false signals, run the p-hacking simulator.

Frequently asked questions

Why does selecting on a sum create a negative correlation?

Because a high sum can be reached by being high on one trait or the other. Inside the selected set, a point that is high on X did not need much Y to clear the bar, so kept high-X points tend to have lower Y. The trade-off is baked into the threshold.

Does it only work when the true correlation is zero?

No. Selection always pushes the measured correlation more negative than the true value. Starting from \rho = 0.4, a strong sum threshold can still drag the selected correlation to zero or below. The default of zero just makes the effect easiest to see.

What happens if I select on the difference instead of the sum?

Selecting on X - Y being large induces a positive correlation in the kept set, by the mirror-image argument. Any linear selection function bends the correlation; a sum bends it negative, a difference bends it positive.

Is a top-k cut different from a threshold?

Only in framing. Keeping the top k by total is the same as choosing whatever threshold happens to admit exactly k points. Both thin the cloud toward the high-sum corner and both produce the negative correlation.

How do I avoid being fooled in real data?

Identify the filter that created your sample and check whether it depends on both variables you are correlating. If you cannot recover the rejected cases, treat any correlation inside the filtered set as suspect and report the selection rule alongside the number.