Run an A/B test until it reaches the sample size you calculated before starting, and for at least one full week - ideally two - so every day of the week is represented. Never stop just because the dashboard shows 95% confidence on day three. The duration comes from three things: your baseline conversion rate, the smallest lift you want to detect, and how much traffic you have.
Step 1: decide the minimum lift worth detecting
Ask: what is the smallest improvement that would change a decision? For a pricing page that might be a 10% relative lift; for a button colour, nothing small is worth the effort. The smaller the lift, the more traffic you need - roughly four times the traffic to detect half the effect.
Step 2: look up the sample size
At 95% confidence and 80% power (the usual defaults), the visitors needed per variant are approximately:
| Baseline conversion | 10% relative lift | 20% relative lift | 30% relative lift |
|---|---|---|---|
| 1% | ~163,000 | ~42,700 | ~19,800 |
| 2% | ~80,700 | ~21,100 | ~9,800 |
| 3% | ~53,200 | ~13,900 | ~6,400 |
| 5% | ~31,200 | ~8,200 | ~3,800 |
| 10% | ~14,800 | ~3,800 | ~1,800 |
These come from the standard two-proportion sample size formula. Multiply by the number of variants (two for a classic A/B test) to get total traffic.
Step 3: convert sample size into days
Example: a product page converts at 3%, gets 2,000 visitors a day, and the team wants to detect a 20% lift. That needs about 13,900 per variant, or 27,800 total - about 14 days. If the result says 3 days, still run a full week; if it says 6 months, test something bolder.
| Situation | What to do |
|---|---|
| Under 7 days needed | Run for 7 to 14 days anyway to cover weekly cycles |
| 1 to 4 weeks needed | Ideal; run as planned |
| More than 6 to 8 weeks needed | Test a bigger change, test higher in the funnel, or use a higher-traffic page |
Why at least one full week
Visitors behave differently on Monday morning and Saturday night, and email sends, paydays and promotions create spikes. A test that runs Tuesday to Thursday only measures Tuesday-to-Thursday visitors. Running whole weeks averages these cycles out.
Mistakes that create false winners
- Peeking. Checking significance every day and stopping at the first 95% reading pushes the real false-positive rate well above 5%. Decide the sample size first and check significance at the end.
- Changing the test mid-way. Editing a variant or splitting traffic differently resets the experiment.
- Novelty effect. Returning visitors sometimes click a new design simply because it is new. Longer tests let this fade.
- Too many variants. Each extra variant needs its own full sample and raises the chance of a lucky winner.
- Ignoring segments after the fact. Slicing results by device or country until something "wins" is fishing. Plan segments before starting.
When the test ends
Enter the final visitors and conversions for each variant into the A/B test significance calculator. If you reach 95% confidence and the lift is big enough to matter, ship it. If not, the honest conclusion is "no detectable difference" - which is still useful, because it tells you this change is not where the growth is.
Frequently asked questions
What is the minimum time to run an A/B test?
At least one full week, and ideally two, even if the required sample size is reached sooner.
Can I stop an A/B test early if it reaches 95% confidence?
Not safely. Stopping at the first significant reading inflates false positives. Wait for the sample size you planned.
What if my site does not have enough traffic?
Test bigger changes that could produce 20 to 30% lifts, test on your highest-traffic pages, or measure an earlier step in the funnel such as add-to-cart.