The A/B Test Peeking Problem
After reading this you will know why checking an A/B test every day and stopping at the first p \lt 0.05 turns a 5% false-positive rate into 30% or more, and what to do instead.
What peeking is and why it hurts
You launch an A/B test. Every morning you refresh the dashboard, run a two-proportion z-test on the data so far, and promise yourself that you will stop and declare a winner the moment the p-value drops below 0.05. That habit is called peeking, and it quietly wrecks your error rate.
Here is the hook. Suppose variant A and variant B are identical. There is no real difference to find. A single honest test at \alpha = 0.05 should call a false winner about 5 times in 100. But if you peek daily for a month and stop at the first crossing, simulation shows you call a false winner about 30 times in 100. Same test, same threshold, six times the errors. The only thing that changed is how many times you looked.
This tool makes both variants identical on purpose. Every "significant" result it ever produces is a false positive by construction, so the fraction of experiments that stop with a "win" is the false-positive rate. That gives you a clean instrument for measuring the damage.
The p-value is a random walk
The core problem is that a p-value computed on accumulating data is not a fixed number. It moves as the sample grows. Under the null hypothesis (no real difference), the test statistic behaves like a random walk. Each new day nudges it up or down by a small amount driven by noise.
A single fixed threshold at 0.05 corresponds to the test statistic crossing about \pm 1.96. A random walk that runs long enough will cross any fixed line eventually, purely by chance. Statisticians call this "sampling to a foregone conclusion." The more looks you take, the more chances the walk gets to wander across the line, and stopping at the first crossing freezes that error in place.
The inflation has nothing to do with a broken test. Every single look in this simulator is an honest, correctly computed two-proportion z-test. The 5% guarantee applies to one look. It says nothing about the minimum of thirty correlated looks.
The two-proportion z-test behind each look
Each daily peek compares the conversion rate of A against B. With x_A conversions out of n_A visitors and x_B out of n_B, the test uses the pooled rate and this statistic:
Here \hat{p}_A = x_A / n_A and \hat{p}_B = x_B / n_B are the observed rates, and \hat{p} = (x_A + x_B)/(n_A + n_B) is the pooled rate under the assumption that both are equal. The two-sided p-value is the probability of a standard normal value at least as far from zero as z:
\Phi is the standard normal cumulative distribution function. You reject at |z| \ge 1.96, which is p \le 0.05. On any one look, under the null, that happens 5% of the time. The trouble is that you do not take one look.
How much peeking inflates the error
Imagine the extreme case where each of k looks were independent. The chance that at least one look crosses is:
Plug in numbers. For k = 5 looks that is 1 - 0.95^5 = 0.226, about 23%. For k = 30 daily looks it is 1 - 0.95^{30} = 0.785, about 79%.
Real peeking looks are not independent, because each day reuses yesterday's data plus a little more. That correlation pulls the realised rate down from the 79% ceiling into the 30% to 40% range you actually observe. The independent formula overstates the damage, but it explains the direction: more looks, more error. The table shows the honest simulated numbers next to the naive ceiling.
| Looks k | Naive ceiling 1-0.95^k | Simulated peeking rate |
|---|---|---|
| 1 | 0.050 | 0.050 |
| 5 | 0.226 | 0.14 |
| 10 | 0.401 | 0.21 |
| 20 | 0.642 | 0.28 |
| 30 | 0.785 | 0.33 |
The sequential boundary that fixes it
You do not have to give up early stopping. Group-sequential designs let you look many times while keeping the total error at 5%. They do it by demanding stronger evidence at early looks, when little data has arrived. The O'Brien-Fleming boundary is the common choice. It spends almost no error budget early and relaxes toward the ordinary 0.05 threshold at the final look.
A convenient approximation for the two-sided z-boundary at look j out of K looks is:
where z_K \approx 2.024 for K = 5 looks at overall \alpha = 0.05. At the first look (j = 1) this demands z_1 = 2.024 \cdot \sqrt{5} = 4.53, which corresponds to p \approx 0.0000059. That is a very high bar, so early false stops almost never happen. At the last look (j = 5) it relaxes to z_5 = 2.024, close to the familiar 1.96. Add up the error spent across all five looks and it totals 5%.
The sequential boundary still lets you win early. If B is genuinely far better, the z-statistic can clear even the tall early bar, and you stop on day 3 with honest evidence. You give up nothing except the ability to be fooled by noise.
Reproducing the demo run
Run the simulator with its default fields: baseline conversion rate 0.10, daily sample 1000 visitors per arm, test length 30 days, and both variants identical. Compare the three strategies over 1000 simulated experiments.
- Fixed horizon. Look only once, on day 30, at \alpha = 0.05. Across 1000 experiments about 50 stop with a false win. Realised rate:
0.050. This matches the guarantee. - Daily peeking. Look every day, stop at the first p \lt 0.05. Across 1000 experiments about 330 stop early with a false win. Realised rate:
0.33. - O'Brien-Fleming. Look every day against the shrinking boundary. Across 1000 experiments about 50 stop with a false win, spread across the looks. Realised rate:
0.051.
The single-experiment view shows one p-value trace wandering. Watch it dip under 0.05 around day 12, pop back above it by day 18, and settle near 0.4 at day 30. The fixed test that looks only on day 30 never sees the dip. The peeking strategy stops at day 12 and records a phantom win.
Common mistakes
The single worst move is running a fixed-horizon test and then stopping it early because the p-value looks good. That is exactly the peeking behaviour that inflates your error to 30% or more. Decide your stopping rule before you collect data, then follow it.
- Stopping only on wins
- If you would stop early for a good result but keep waiting after a bad one, you have a one-sided ratchet. The dips below
0.05get locked in, the recoveries do not. - Treating "not yet significant" as "keep going"
- Extending a fixed test because the result is "almost there" is peeking with extra steps. Each extension is another draw from the false-positive lottery.
- Confusing per-look and overall error
- The 5% is a promise about one look. Over 30 looks the overall error is a different, much larger number. Do not quote the per-look figure as if it covered the whole test.
- Ignoring correlation between looks
- Do not use the independent ceiling 1-0.95^k as your real error. It overstates the damage. The simulated rate is lower because consecutive looks share most of their data.
Related tools
Peeking is one flavour of a wider problem: testing until something looks significant. The p-Hacking Simulator shows the sibling failure of running many tests and reporting the winner. The Monte Carlo Playground covers the sampling machinery that all of these simulations rest on. For the underlying idea that a p-value under the null wanders like a random walk, see the Random Walk Explorer. To see why extreme early results tend to fade, the Regression to the Mean tool is a good companion, and the Base Rate Visualizer shows a different way that a "significant" test can mislead you.
Frequently asked questions
Is peeking always wrong?
No. Peeking is fine if your stopping rule accounts for it. The problem is peeking with a fixed 0.05 threshold and no correction. Use a group-sequential boundary or always-valid inference and you can look as often as you like.
Why is the real peeking rate only 0.33, not 0.79?
Because the looks are correlated. Day 12 and day 13 share almost all their data, so they are not two independent chances to fail. The independent formula 1-0.95^{30} = 0.785 is an upper bound. The correlation drags the realised rate down to about 0.33.
Does a bigger daily sample fix the problem?
No. Larger samples make each look more precise, but the null p-value is still a random walk, and it still crosses 0.05 eventually with enough looks. More data per day does not reduce the number of chances to cross.
What is always-valid inference?
It is a family of methods (sequential p-values and confidence sequences) whose guarantees hold no matter when you stop. They let you monitor continuously and stop at any moment while keeping the false-positive rate at 5%. They trade a little sensitivity for the freedom to peek.
How do I pick a fixed sample size in advance?
Run a power calculation. Choose the smallest effect worth detecting, the false-positive rate (5%) and the power you want (often 80%), then compute the sample size those imply. Collect exactly that much, look once, and stop.