The Central Limit Theorem, Explained

After reading this you can predict the shape, center and spread of the distribution of sample means, check the prediction against a simulation, and know when the theorem fails you.

What the theorem claims

Pick any population you like. It can be lopsided, spiky, or flat. Draw a sample of n independent values, average them, and record that one average. Do it again. Do it thousands of times. The histogram of those averages is not shaped like the population. It is shaped like a bell.

That is the central limit theorem (CLT). The distribution of the sample mean approaches a normal distribution as n grows, regardless of the population shape, as long as the population has a finite mean and variance.

Here is a hook you can check in the demo. Take a single fair die. Its six outcomes are equally likely, so its own histogram is flat, not bell-shaped at all. Now roll n = 30 dice, average them, and plot that average. Repeat. The flat die turns into a tight bell centered on 3.5. Nothing about one die is bell-shaped. The average of thirty is.

When to lean on it, and when not

The CLT is the reason so much of statistics assumes normality. Confidence intervals for a mean, t-tests, and standard error bars all rest on it. When you compute an average from a decent-sized sample, you can treat that average as normal even if the raw data is not.

It has limits. The theorem is about the mean, not about individual values. A skewed population stays skewed no matter how many values you observe. Only the averages straighten out.

The CLT needs a finite variance. Distributions with infinite or undefined variance, such as the Cauchy distribution, never converge to a normal shape. The average of a thousand Cauchy samples is just as wild as a single one. Finite variance is the fine print, not a formality.

It is also approximate for small n. A mildly skewed population might look normal by n = 10. A heavily skewed one might still be visibly lopsided at n = 30 and only settle down near n = 100. Bigger samples converge faster, and the more skewed the population, the larger the n you need.

The formula and why the two parts appear

Two numbers describe the distribution of the mean: where it centers and how wide it spreads.

\bar{X} \approx \mathcal{N}\left(\mu,\; \frac{\sigma^2}{n}\right)

Here \bar{X} is the sample mean, \mu is the population mean, \sigma^2 is the population variance, and n is the sample size. The mean of the averages equals the population mean exactly. The variance of the averages is the population variance divided by n.

The spread, called the standard error, is the square root of that variance.

\text{SE} = \frac{\sigma}{\sqrt{n}}

The \sqrt{n} in the denominator is the whole story of why averages are stable. Quadruple the sample size and the standard error halves, because \sqrt{4} = 2. To cut the spread by a factor of ten you need a hundred times as many samples. That diminishing return is why polls of a few thousand people are common but polls of a few million are rare: the extra precision is not worth the cost.

The mean centering on \mu follows from linearity: the average of unbiased draws is unbiased. The 1/n shrink in variance follows from independence: independent errors partly cancel, and the cancellation grows with n.

A worked example with a fair die

Rolling 30 dice, ten thousand times

Use the demo defaults: a single six-sided die as the population, sample size n = 30. First pin down the population numbers.

  1. The mean of one die is \mu = (1+2+3+4+5+6)/6 = 3.5.
  2. The variance of one die is \sigma^2 = \frac{1}{6}\sum (k-3.5)^2 = 35/12 \approx 2.917. So \sigma \approx 1.708.
  3. The predicted center of the sample means is \mu = 3.5.
  4. The predicted standard error is \text{SE} = 1.708 / \sqrt{30} = 1.708 / 5.477 \approx 0.312.

So the theorem predicts the averages of 30 dice pile up around 3.5 with a standard deviation of about 0.312. That means roughly 95% of the averages should land within 3.5 \pm 2(0.312), that is between 2.88 and 4.12. Run the demo with these defaults and the overlaid normal curve will sit almost exactly on the histogram.

The bell centered on 3.5 with standard error 0.312. Almost all of the mass sits between 2.88 and 4.12, matching the two-sigma estimate.

Watching the width shrink with n

The single most instructive thing to vary is n. Everything else stays fixed while the bell tightens. Using the die population, the standard error at four sample sizes is worth tabulating.

Standard error of the die average, \sigma = 1.708
Sample size n√nSE = σ/√n95% range around 3.5
11.0001.7080.08 to 6.92
52.2360.7641.97 to 5.03
305.4770.3122.88 to 4.12
12010.950.1563.19 to 3.81

Notice the pattern in the SE column. Going from n = 30 to n = 120 is a fourfold increase in samples, and the standard error drops from 0.312 to 0.156, exactly half. That is the \sqrt{n} law in action.

With a fair-die population (mean 3.5, standard deviation 1.708), the sample mean is approximately normal with standard error 1.708/√n. At n=1 the spread is 1.708 and the shape is flat and blocky; at n=5 the spread is 0.764; at n=30 it is 0.312; at n=120 it is 0.156 and sharply peaked. The center stays at 3.5 throughout.

Reading the output correctly

Three checks tell you whether the theorem is holding in front of you.

Center
The histogram of means should sit on \mu, the population mean, not on some shifted value. For the die that is 3.5.
Spread
The observed standard deviation of the means should match \sigma/\sqrt{n}. For 30 dice, expect about 0.312. If your simulation reports 0.31 or 0.32, that is a match.
Shape
The overlaid normal curve should trace the histogram. For a symmetric population like the die this happens fast. For a skewed one, watch a small residual lean at low n that fades as n climbs.

A useful contrast: raising the number of repeated samples (say from 1000 to 100000 draws) makes the histogram smoother and the center land more precisely, but it does not narrow the bell. Only raising n, the size of each sample, narrows the bell. Confusing these two knobs is the most common misreading.

Common mistakes

If the bell looks too wide, check whether you changed the number of repeats when you meant to change the sample size n. Repeats sharpen the picture. Sample size shrinks the spread. They are different controls with different effects.

The mistake worth naming first: claiming the raw data becomes normal. It does not. Ten thousand die rolls plotted individually stay uniform across the six faces. Only the averages of grouped rolls go bell-shaped.

Second, assuming n = 30 is a magic threshold. The number 30 is a rule of thumb for mild skew. A strongly skewed population, such as an exponential income distribution, can still show a right lean at n = 30. Push to n = 100 or beyond and check the shape again.

Third, forgetting independence. The theorem assumes independent draws. If your samples are correlated (consecutive days of the same stock, repeated measurements on the same person), the effective sample size is smaller than n, and the true standard error is larger than \sigma/\sqrt{n}.

Fourth, using it on a distribution with no finite variance. There the theorem simply does not apply, and no amount of averaging produces a bell.

Related tools on this site

The Galton Board builds the same bell physically, by dropping balls through pegs, and is the clearest companion to this demo. To see averages settle onto their expected value one trial at a time, try the Law of Large Numbers. For random sampling used to estimate quantities, the Monte Carlo Playground and Buffon's Needle both rely on the same convergence ideas. To turn one dataset into a confidence interval by resampling, see the Bootstrap Resampling Visualizer. And to watch how random increments accumulate over time rather than average out, explore the Random Walk Explorer.

Frequently asked questions

Does the central limit theorem make my data normal?

No. It makes the distribution of the sample mean normal. Your raw data keeps whatever shape it had. A skewed population stays skewed no matter how many points you collect. Only averages of samples straighten into a bell.

Why is n = 30 the number everyone quotes?

It is a rough guideline for populations that are not too skewed. For symmetric populations you often see a good bell by n = 10. For heavily skewed populations even n = 30 can look lopsided, and you may need n = 100 or more. Always check the shape rather than trusting the number.

What happens if the population has no finite variance?

The theorem does not apply. The Cauchy distribution is the classic example: the average of a thousand draws is just as spread out as a single draw, and no bell ever forms. Finite variance is a required condition.

How does the standard error change if I quadruple the sample size?

It halves. The standard error is \sigma/\sqrt{n}, so multiplying n by 4 multiplies the denominator by \sqrt{4} = 2. For the die at n = 30 the SE is about 0.312; at n = 120 it is about 0.156.

What is the difference between more repeats and a larger sample size?

More repeats give you more sample means to plot, which smooths the histogram but does not change its width. A larger sample size n shrinks the width of the histogram, because each mean is averaged over more values. To watch the bell tighten, raise n.