How long should an A/B test run?

5 min read

Short answer

Long enough to find the smallest change worth knowing about, in whole weeks, fixed before the test starts. With 1,000 people a week and a 4% base rate, 2 weeks can find a change of about +70%, 4 weeks about +48% and 8 weeks about +33%.

Decide how long an A/B test runs before it starts, in whole weeks, from the smallest change you want to be able to find. Then read the result once, at the end. The length follows from two numbers you know today: how many people enter the test each week, and how often they reach the goal now.

Why not stop as soon as B looks better?

Because early differences are mostly noise. With a few hundred people per variant, one group can run well ahead for days and then fall back. A test at 95% confidence calls a false winner in about 5 of 100 tests where nothing changed, but only if you look once. Check every day and stop at the first result that clears the bar, and each look is another chance for noise to clear it: the share of false winners rises several times over.

Watching the numbers is fine. Deciding before the planned end is what goes wrong, which is why the length has to be fixed first.

What decides how long a test needs?

Two numbers: how many people take part, and the base rate, the share who reach the goal today. Together they set the smallest change a test can find. More people find smaller changes, and so does a higher base rate. Here is a test with 4,000 people, 2,000 per variant:

Base rate B must reach Smallest change it can find
1% 2.1% +107%
4% 5.9% +48%
10% 12.8% +28%
20% 23.7% +18%

A purchase goal at 1% needs a change that doubles it before a test of this size can see it. An action more people take, such as adding to cart or starting checkout, shows a far smaller change with the same people.

How is the smallest detectable change calculated?

With the arcsine method, a two-sided test at 95% confidence and 80% power. Take today’s rate p and the people per variant n, which is half the people in the test. The rate B has to reach is

p₂ = sin²(asin √p + 1.4 · √(2 / n))

and the smallest relative change is p₂ / p − 1. The 1.4 is half of 1.96 + 0.84, the values that stand for 95% two-sided and 80% power. In a spreadsheet, with the base rate in B1 as a decimal (0.04) and the people per variant in B2:

=SIN(ASIN(SQRT(B1))+1.4*SQRT(2/B2))^2/B1-1

80% power means a real change of exactly that size is found in about 8 of 10 tests. A smaller real change can still show up, less reliably, and a test that finds nothing has not shown that nothing changed.

How long at 1,000 people a week and a 4% base rate?

Say 1,000 people a week enter the test, each of them new to it, and 4% of them buy today. Each week adds 500 people to each variant:

Length People in the test Per variant B must reach Smallest change
2 weeks 2,000 1,000 6.8% +70%
4 weeks 4,000 2,000 5.9% +48%
8 weeks 8,000 4,000 5.3% +33%

Four times the people finds a change about half the size. Going smaller gets expensive: a change of +10%, from 4% to 4.4%, needs about 79,000 people, a year and a half at this traffic. On Free, where a test runs at most 3 weeks, the same shop can find about +56%.

So at this traffic, test changes you expect to move the rate by a third or more, such as a shorter checkout, a new product page or a clearer delivery promise, and leave small tweaks to shops with more visitors. The free A/B test calculator runs the same maths for your base rate and people a week.

Why run whole weeks?

Because weekdays differ. Some people browse on Sunday evening and buy on Monday, business visitors disappear at the weekend, and payday shifts the end of the month. A test that stops on a Thursday gives some weekdays more weight than others; whole weeks give each the same.

Then leave room for the goal. People who enter on the last day need time to buy, so a window after the planned weeks lets them reach the goal before the test is judged. If your buyers take days to decide, pick a window that fits: in the sample shop, the median buyer took 10 days from first visit to purchase.

How does MIRA FIVE plan and judge a test?

Under Experiments you choose the goal or action that decides the test, a length of 2, 3, 4, 6 or 8 weeks and a window of 1, 7, 14 or 30 days, and the planner shows the smallest change that length can detect (“What this experiment can find”), projected from the people of the last 28 days so returning people count once. In the browser, only people who consented are counted. While the status reads Counting there is no verdict; it is taken once, at the planned end plus the window, as a two-sided test at 95% whose 95% range must agree, with at least 100 people per variant. It reads Variant B wins, worse, no big difference (when the 90% range stays within ±10%) or no clear winner. On Free a test runs at most 3 weeks, with length plus window inside 28 days. In the sample shop, a checkout button test ran 4 weeks plus a 7-day window: 49 of 1,204 people reached “Order paid” with the original (4.1%) and 71 of 1,188 with Variant B (6.0%), and the result reads Variant B wins. The Experiments page shows the screens.

Questions and answers

Can I stop a test early when B is clearly ahead?
Better not. Early leads are often noise, and stopping at the first good-looking day turns noise into winners. Read the result once, at the planned end; MIRA FIVE gives no verdict while counting.
What if my shop has little traffic?
Test bigger changes, run longer, or decide the test by an action more people take, such as starting checkout. A higher base rate lets the same number of people find a smaller change.
Why should a test run in whole weeks?
Because people shop differently on a Monday and on a Sunday. Whole weeks give every weekday the same weight, and MIRA FIVE's lengths of 2, 3, 4, 6 or 8 weeks are all whole weeks.
How long can a test run on the free plan?
Up to 3 weeks, and its length plus window must fit in 28 days. Pro offers every length up to 8 weeks, and every plan includes Experiments.

Read next

How to set up a purchase goal

Set up conversion tracking with a purchase goal, revenue read from your events, one total per currency, and the time from first visit to paid plan.

All guides

See which channel brings buyers

Free for 25,000 events a month. No card needed.