Conversion Optimization
How to run A/B tests in ecommerce

In short
An ecommerce A/B test starts with a data backed hypothesis that states what you will change, which metric it should move and why. You then calculate the required sample size, run the test for at least two full weeks, read the result against a significance level chosen in advance, and only then ship the winner.
Contents
- How do you run an A/B test in ecommerce?
- How do you write a good A/B test hypothesis?
- Which metric should you measure?
- How do you calculate sample size?
- How long should a test run?
- How do you read statistical significance?
- What should you test in ecommerce?
- Which tools can you use for A/B testing?
- What are the most common A/B testing mistakes?
- Key takeaways
How do you run an A/B test in ecommerce?
An ecommerce A/B test randomly splits traffic into two groups, shows one the current page and the other a changed version, and measures which performs better. A trustworthy test needs a data backed hypothesis, a sample size calculated before launch, a run time of full weeks, and a result judged against a significance level you set in advance.
The value of A/B testing is that it settles opinion battles with real behavior. A badly designed test, however, is worse than no test at all, because it stamps "proven" on a wrong conclusion. This guide walks through setting one up properly.
How do you write a good A/B test hypothesis?
Every test starts with an observation. Analytics, session recordings, support tickets or usability tests reveal a problem. Nielsen Norman Group points out that A/B tests produce much stronger variations when they are informed by insights from user research.
Write the hypothesis using this pattern:
Because we observed [observation], if we [change], then [metric] will increase, because [reason].
Example: "Support logs show mobile shoppers frequently ask about shipping cost. If we move shipping cost and delivery date directly under the price, mobile add to cart rate will increase, because shoppers will see the total cost at the moment of decision."
A good hypothesis has one change, one primary metric and a testable reason. "Let's redesign the product page" is not a hypothesis.
Which metric should you measure?
The primary metric is the single metric that will decide the test, and you choose it before launch. In ecommerce it is usually one of these:
- Purchase conversion rate
- Revenue per visitor
- Add to cart rate (only as an intermediate metric for product page tests)
Secondary metrics such as average order value, return rate and page speed are there to catch side effects. A change that lifts add to cart but lowers purchases is not a winner.
How do you calculate sample size?
You need to know how many users a test requires before it starts. That takes three inputs: your baseline conversion rate, the minimum effect you want to detect, and your statistical power and significance level. A common practice is 95% confidence with 80% power.
For those settings, a practical approximation is: sample per variant ≈ 16 x p x (1 - p) / d², where p is the baseline conversion rate and d is the absolute difference you want to detect.
Example: a store with a 2% conversion rate wants to detect a 10% relative lift, from 2% to 2.2%.
- p = 0.02, d = 0.002
- Sample per variant ≈ 16 x 0.02 x 0.98 / 0.000004 ≈ 78,400 sessions
- Total for two variants ≈ 156,800 sessions
At 5,000 sessions a day, the test needs about 32 days. This math also explains why small stores cannot reliably test small tweaks.
| Baseline rate | Target relative lift | Approx. sample per variant |
|---|---|---|
| 2% | 10% | 78,400 |
| 2% | 20% | 19,600 |
| 4% | 10% | 38,400 |
| 4% | 20% | 9,600 |
These are example values calculated with the same formula. For real tests, use the calculator in your testing tool.
How long should a test run?
Reaching the sample size is not the only condition. Run tests for at least two full weeks so that weekday and weekend behavior is represented in both variants. Promotions, holidays and big sale days distort the traffic mix, so interpret tests that overlap them carefully. Fix the duration before launch and do not stop early because the result looks good.
How do you read statistical significance?
A 95% confidence level means that if there were truly no difference, you would see a result this extreme only about 5% of the time. It does not mean the winner is certainly better; it means the result is unlikely to be pure chance.
Read results in this order:
- Did you reach the planned sample size and duration?
- Is the traffic split as expected, for example 50/50?
- Is the difference in the primary metric significant?
- Does the confidence interval include an effect that matters to the business?
- Are there negative side effects in secondary metrics?
- Is the result consistent across mobile and desktop?
What should you test in ecommerce?
Your traffic cannot support every idea, so keep a backlog and score each idea from 1 to 10 on expected impact, strength of evidence and ease of implementation. The highest scores get tested first.
Areas with high potential in most stores include the placement of shipping and delivery information, the free shipping threshold and how it is shown, image order and in scale images, a sticky mobile add to cart button, guest checkout and field count, default sorting and filter layout on category pages, and bundles or complementary product suggestions. Button colors and one word copy tweaks rarely move anything measurable and cannot be tested reliably on most stores' traffic.
Document every test in one place: hypothesis, screenshots, dates, sample size, primary and secondary results and the decision. Losing tests are as valuable as winners, because they stop the same idea from being retested six months later.
Which tools can you use for A/B testing?
Since Google Optimize was discontinued in 2023, teams have moved to independent platforms such as VWO, Optimizely, AB Tasty and Convert, or to testing apps in their ecommerce platform's app store. When choosing, check whether the tool runs without visible flicker, how it integrates with GA4, whether it has a sample size calculator, and whether it supports server side tests for things like pricing or shipping rules.
What are the most common A/B testing mistakes?
- Stopping early: large early differences usually shrink over time.
- Changing too much at once: you will never know which change mattered.
- Switching metrics after seeing results: the primary metric must be chosen up front.
- Testing with too little traffic: small stores should test bold, larger changes.
- Hunting for winners in segments: look at dozens of segments and some will be significant by chance.
- Shipping the winner and stopping measurement: keep tracking performance after launch.
If you lack the traffic to test, usability testing, session recordings and controlled before and after analyses are better tools.
If you would like to build a testing roadmap together, you can book a free growth analysis with the Performetic team through our contact page. You can also see results from this approach in our case studies.
Key takeaways
- An A/B test starts with one data backed hypothesis and one primary metric chosen in advance.
- Calculate sample size before launch; small effects require a lot of traffic.
- Run tests for at least two full weeks and never stop early.
- Read significance together with traffic split, secondary metrics and device consistency.
- Without enough traffic, rely on qualitative research and controlled before and after analysis.
Frequently asked questions
How much traffic do you need for an A/B test?
It depends on your baseline conversion rate and the size of the effect you want to detect. Detecting a 10% relative lift on a 2% conversion rate needs roughly 78,000 sessions per variant. Smaller stores should test bigger, bolder changes that produce larger effects.
How long should an A/B test run?
Run it until it reaches the calculated sample size and at least two full weeks have passed, so weekday and weekend behavior appears in both variants. Decide the duration before launch and never stop a test early just because the result looks promising.
What can replace Google Optimize?
Google Optimize was discontinued in 2023. Alternatives include independent tools such as VWO, Optimizely, AB Tasty and Convert, as well as testing apps built for ecommerce platforms. Compare them on flicker, GA4 integration and support for server side tests.
Can you run several A/B tests at the same time?
Yes, on different pages and areas that do not influence each other. Running tests that can interact within the same journey, for example on the product page and checkout at once, can muddy results. In those cases, run the tests one after another.