The Bootstrap, Explained
After reading this you can take one modest dataset, resample it thousands of times, and read a 95% confidence interval straight off the resampled statistics, without ever assuming a bell curve.
What the bootstrap does and why it feels like cheating
Suppose you measured the waiting time at a clinic for 12 patients and got a mean of 8.3 minutes. How sure are you of that 8.3? A different 12 patients would have given a different mean. The honest question is: how much does the mean wobble from sample to sample?
The classic answer uses a formula for the standard error and assumes the sampling distribution is roughly normal. The bootstrap answers the same question with almost no assumptions. It treats your sample as if it were the whole population, then draws new samples from it, with replacement, each the same size as the original. Recompute the statistic on each draw. The spread of those recomputed values estimates how much your statistic would wobble in reality.
It feels like getting something from nothing. It is not. The trick is that your sample already carries information about its own variability, and resampling extracts it.
The plug-in principle behind it
The bootstrap rests on one idea: replace the unknown true distribution F with the empirical distribution \hat{F} that puts equal weight 1/n on each observed data point. If you cannot sample from the real world again, sample from the copy you have.
Here n is your sample size, x_i is the i-th observation, and \mathbf{1}\{\cdot\} is 1 when the condition holds and 0 otherwise. So \hat{F} just counts what fraction of your data falls at or below x. Drawing a bootstrap sample means drawing n values from this staircase, which is the same as picking n of your points at random with replacement.
Because you draw with replacement, a single bootstrap sample repeats some points and omits others. On average each bootstrap sample leaves out a fixed fraction of the original. The probability a given point is never chosen in n draws is (1 - 1/n)^n, which approaches 1/e \approx 0.368 as n grows. So a typical bootstrap sample contains about 63.2% of the distinct original points.
Building the interval from the resamples
The full recipe has four steps.
- Compute the statistic \hat{\theta} on the original data (say the mean or the median).
- Draw a bootstrap sample of size n with replacement and compute \hat{\theta}^{*}_{b}.
- Repeat step 2 for B resamples, collecting \hat{\theta}^{*}_{1}, \dots, \hat{\theta}^{*}_{B}.
- Read percentiles off that collection.
The bootstrap standard error is just the standard deviation of the resampled statistics:
where \bar{\theta}^{*} is the mean of the resampled statistics. For a 95% percentile interval you sort the B values and take the 2.5th and 97.5th percentiles. That pair of numbers is your interval. No t-table, no normal assumption, no closed-form variance.
Use a large B. For a standard error, B = 200 is often enough. For the tails of a 95% interval you want at least 2000 resamples, because the 2.5th percentile is estimated from far fewer effective points than the center.
A worked example on the demo data
Ten numbers, one interval for the mean
Take the demo sample of ten values:
4, 7, 5, 9, 3, 8, 6, 12, 5, 7
- The sum is
66, so the original mean is \hat{\theta} = 66/10 = 6.6. - Draw a resample of size 10 with replacement. One draw might land on indices giving
7, 7, 5, 9, 12, 6, 5, 4, 8, 7. Its mean is 70/10 = 7.0. - Another draw gives
3, 5, 5, 4, 7, 6, 5, 9, 3, 8, mean 55/10 = 5.5. - A third gives
12, 8, 9, 7, 7, 6, 8, 5, 9, 4, mean 75/10 = 7.5.
Notice the three resample means (7.0, 5.5, 7.5) scatter around 6.6. Repeat this 5000 times and the collection of means clusters tightly, with a spread near 0.79. That number is the bootstrap standard error. Compare it to the textbook standard error of the mean, s/\sqrt{n}. The sample standard deviation here is s \approx 2.63, so s/\sqrt{10} \approx 0.83. The bootstrap lands close to the formula, which is exactly what should happen for a mean.
The percentile interval for the mean comes out roughly 5.1 to 8.2. Because the ratio between the bootstrap SE (0.79) and the classical SE (0.83) is near 1, you gained little for the mean here. The bootstrap earns its keep on harder statistics.
Watch the interval settle
The interval is not fixed the moment you start. Each new resample nudges the estimated percentiles. Early on, with a few hundred resamples, the endpoints jitter by a few tenths. As the count climbs into the thousands, they lock in.
When to reach for it, and when not to
The bootstrap shines when you have no clean formula for the standard error. The median is the classic case. There is a formula for the variance of the sample median, but it depends on the unknown density at the median, which is awkward to estimate. Resampling sidesteps that entirely. The same holds for trimmed means, correlation coefficients, ratios, and quantiles.
It is not magic, and it fails in predictable ways.
The bootstrap cannot invent data you never collected. With n = 5, there are only a handful of distinct values to reshuffle, so the resampled distribution is coarse and the interval is unreliable. It also breaks for statistics that depend on the extreme tails, such as the sample maximum, because the largest observed value caps every resample.
It also assumes your original sample was drawn independently from the population. If the data are correlated (a time series, clustered measurements), plain resampling destroys that structure and understates the uncertainty. Specialized block bootstraps exist for those cases.
Reading the output honestly
A 95% percentile interval of 5.1 to 8.2 means this: if you repeated the whole experiment many times and built an interval each time by this method, about 95% of those intervals would contain the true mean. It does not mean there is a 95% probability the true mean lies in this particular interval. The true mean is a fixed number; it is in or out.
- Bootstrap distribution
- The histogram of the statistic recomputed on every resample. Its spread estimates the standard error.
- Percentile interval
- The interval between the 2.5th and 97.5th percentiles of that histogram, for 95% coverage.
- Coverage
- The long-run fraction of built intervals that trap the true value. The percentile method aims for 95% but can undercover for skewed statistics or tiny samples.
One more caution: a bootstrap interval is only as good as the sample it came from. A biased sample of 10 patients gives a tight, confident, and wrong interval. Resampling propagates the data you have; it does not correct for how you got it.
Related tools on this site
To see randomness converge on a fixed answer, throw darts in the Monte Carlo Playground or watch running averages settle in the Law of Large Numbers. The bell shape your resampled means take is explained by the Central Limit Theorem Demo. For a physical bell curve building itself from chance, drop balls through the Galton Board. And to see how easy it is to fool yourself with resampled tests, run the p-Hacking Simulator. If you want to feel correlation before you estimate it, try Guess the Correlation.
Frequently asked questions
How many resamples do I need?
For a standard error, 200 is often enough. For a 95% percentile interval, use at least 2000, and 5000 gives stable endpoints. More resamples reduce the Monte Carlo noise in the interval, not the uncertainty from the original sample, which is set by n.
Does the bootstrap fix a small sample?
No. With n = 10 your interval reflects the wobble of a size-10 estimate, and that wobble is large. Resampling 10 points a million times does not add information; it only maps out the uncertainty already present.
Why does the median give a lumpy histogram?
The median of any resample is always one of your observed values (or an average of two of them). With 10 fixed numbers, only a handful of distinct medians can appear, so the bootstrap distribution has visible steps rather than a smooth curve. The mean can land anywhere, so its histogram looks continuous.
Is the percentile interval the best method?
It is the simplest and works well when the bootstrap distribution is roughly symmetric. For skewed statistics, refinements like the bias-corrected and accelerated (BCa) interval give better coverage. The percentile method is the right place to start and the one this tool draws.