Skip to content
Get CRO Audit

A/B Testing

Stop calling A/B tests early because the graph looks good

Every experimentation program eventually meets the same moment. A test has been live for six days, the variant line is comfortably above control, the dashboard says 94% chance to beat baseline, and someone asks the reasonable question: can we just ship it?

Usually the honest answer is not yet. Not because statistics are fussy, but because early results lie in predictable ways.

Why early winners fade

Three effects stack up in the first days of most tests:

  • Small samples swing. With a few hundred conversions per arm, a handful of large orders or a single good day can create a gap that disappears as data accumulates.
  • Peeking inflates false positives. If you check results daily and stop the moment significance appears, your real false positive rate is far higher than the 5% you think you are accepting.
  • Weekly cycles are real. Monday visitors and Saturday visitors behave differently. A test that has not run through full weeks is comparing different mixes of people.

Decide the stopping rule before launch

The fix is boring and effective: agree on how the test ends before it starts. That means three numbers written into the test plan:

  1. Minimum detectable effect. The smallest lift that would actually matter to the business. If a 2% relative lift would not change any decision, do not size the test to detect it.
  2. Required sample size per variant, calculated from your baseline conversion rate and that effect size.
  3. Minimum duration, rounded up to whole weeks, usually at least two.

If the test reaches both the sample and the duration, you read the result. If it does not, you keep waiting or you accept that the test cannot answer the question with the traffic you have.

Read more than one metric

A variant that lifts conversion rate while lowering average order value can lose money. Agree on one primary metric, usually conversion rate or revenue per visitor, and two or three guardrails such as order value, refund rate or lead quality. A winner has to improve the primary metric without breaking a guardrail.

What to do when the answer is “inconclusive”

Many well-run tests end flat. That is information, not failure. It tells you the change did not matter enough to detect, which frees you to test something bolder. The most common mistake after a flat test is to re-run a nearly identical variant. The better move is to go back to the research and ask whether the hypothesis targeted a real problem at all.

A test is finished when it reaches the sample size and duration you planned, not when the graph looks good.

Scroll to Top