Key takeaways
Shopify’s built-in A/B testing is the Experiment type in Rollouts, available on the Grow plan or higher. It splits visitors 50/50 by default and can test theme edits, a replacement theme, a checkout configuration or discounts. Decide the sample size before you start and don’t stop early: at a 2% conversion rate, detecting a 20% relative lift takes about 20,000 visitors per version. With less traffic, test bigger changes, measure a step closer to the change, such as add-to-cart rate, or use research instead of a test.
- Shopify’s built-in A/B testing is the Experiment type in Rollouts, available on the Grow plan or higher.
- Work out the sample size before you start: at a 2% conversion rate, a 20% lift needs about 20,000 visitors per version.
- Don’t stop a test early because it looks significant; repeated peeking inflates false positives.
- With low traffic, test bigger changes, measure a step closer to the change or use research instead.
- Start new baselines after Shopify’s September 2026 change to how sessions are measured.
Shopify’s built-in testing: Rollouts
Rollouts, under Markets > Rollouts in the Shopify admin, schedule and test changes to your store. There are three types. A Launch makes changes permanent, an Event applies them for a set period and then rolls them back, and an Experiment shows a treatment to part of your traffic and compares it with a control. Rollouts are available on the Basic plan or higher; Experiments need the Grow plan or higher.
An Experiment can change your main theme, by editing it or replacing it with another theme, your checkout and accounts configuration, which catalogues are active and which discounts apply. Since June 2026 it can compare two entirely different themes or checkout setups. New Experiments split traffic 50/50 and end after 90 days by default. They run on the online store only, vintage themes aren’t supported and a rollout can’t change Liquid templates.
Theme Experiments report conversion rate, add-to-cart rate, reached-checkout rate and bounce rate; checkout Experiments report checkout conversion rate. Shopify’s rollout documentation doesn’t set a significance threshold or say when a result is final, so decide your sample size and end date yourself.
When a testing app fits better
Rollouts can’t change Liquid templates, and its documentation doesn’t cover price or shipping-rate tests. Testing apps in the Shopify App Store go further: Shoplift’s documentation lists template, theme, URL-redirect and price tests, and Intelligems documents price, shipping and checkout tests. Check what each app supports on your plan, and how it assigns visitors, before you install it.
Google Optimize, which many stores used for free testing, shut down on September 30, 2023; Google said it would invest in third-party A/B testing integrations for Google Analytics instead. Run the test in Rollouts or an app, and use your analytics to look at segments and funnels.
Price tests come with a Canadian rule. If a variant shows a compare-at price, that “regular” price must meet one of the Competition Bureau’s ordinary-price tests: more than half of sales at that price or higher, or the product offered in good faith at that price or higher for a substantial period. The Bureau usually looks at the year before or after the promotion.
Work out the sample before you start
A test needs enough visitors to tell a real difference from noise, and the smaller the difference, the more visitors it needs. Evan Miller’s free sample-size calculator, a standard reference, asks for your current conversion rate and the smallest lift worth detecting. At a 5% significance level and 80% power, it gives these numbers:
| Current conversion rate | Lift to detect | Visitors per version | Time at 10,000 visitors a month |
|---|---|---|---|
| 2% | 20%, to 2.4% | About 19,800 | About 4 months |
| 2% | 10%, to 2.2% | About 78,000 | About 16 months |
| 3% | 20%, to 3.6% | About 13,000 | About 2½ months |
| 1% | 20%, to 1.2% | About 40,000 | About 8 months |
The time column assumes every visitor enters the test and traffic is split evenly between two versions.
Set the baseline and the length
Measure your baseline in the unit the test assigns. Shopify reports conversion rate per session, while testing tools usually assign visitors, who can have several sessions. And because Shopify changed how it measures sessions between September 21 and 23, 2026, filtering out identified bots and no longer ending sessions at midnight UTC, take the baseline from after that week.
Run every test for whole weeks, so weekday and weekend shoppers are both represented, and avoid launching during a sale or a big campaign unless that’s what you’re testing.
Don’t stop when it looks good
The significance a dashboard reports assumes you fixed the sample size in advance. Evan Miller showed that checking a running test repeatedly and stopping at the first significant result inflates false positives: in his worst case, checking after every visitor, a nominal 5% false-positive rate becomes 26.1%. Even ten looks make an apparent 1% significance level worth about 5%.
Set the sample size and end date before launch, and don’t act on interim results. If you need the option to stop early for a clear winner, use a method built for it, such as Miller’s sequential test, which allows early stopping without the peeking problem.
What to do when traffic is low
Most independent stores can’t detect a 10% lift in a reasonable time. That doesn’t make testing pointless; it changes what to test and how.
- Test bigger changes. A new product-page layout or a different offer can move results far more than a button colour, and large effects need far fewer visitors.
- Measure a step closer to the change. Add-to-cart rate happens more often than a purchase, so a product-page test reaches its sample sooner; check completed orders before you roll out the winner.
- Test where the traffic is. A change to the template every product page uses collects visitors faster than a change to one product.
- Use research instead. Session recordings, heatmaps and post-purchase surveys show why people hesitate without needing a statistical result. Microsoft Clarity is free with no traffic limits, though on Shopify it records checkout pages only for Plus stores. Check consent requirements first, especially for visitors in Quebec.
- Ship and watch. When a test isn’t practical, release the change and compare it with the same period before, allowing for campaigns, seasonality and stock. It’s weaker evidence, so keep the change easy to undo.
Write the hypothesis first
Before building a variant, write down the change, the metric it should move and why. For example: “Putting the size guide next to the size selector will raise add-to-cart rate and reduce fit-related returns, because return reasons point to sizing.” Choose one primary metric and a few guardrails, such as completed orders, average order value and returns, so that a test that raises add-to-cart rate while lowering purchases doesn’t count as a win.
Keep a log of every test: the hypothesis, dates, sample, result and what you did next. Record inconclusive tests too; they show which ideas made no difference you could measure.
A testing checklist
- Write the hypothesis, primary metric and guardrails before building anything.
- Calculate the sample size from your own baseline, measured after September 23, 2026.
- Set the end date in advance and run for whole weeks.
- Don’t stop early because a dashboard shows a winner.
- Check completed orders and average order value before rolling out a winner.
- Log every result, including the inconclusive ones.
Frequently asked questions
Does Shopify have built-in A/B testing?
Yes. The Experiment type in Shopify Rollouts compares a treatment with a control on your online store. Rollouts are available on the Basic plan or higher, but Experiments need the Grow plan or higher.
How much traffic do I need to A/B test a Shopify store?
It depends on your conversion rate and the size of change you want to detect. At a 2% conversion rate, a 20% relative lift takes about 20,000 visitors per version at a 5% significance level and 80% power, and a 10% lift takes about 78,000.
Can I stop a test as soon as it shows a significant result?
Not if it was planned as a fixed-sample test. Checking repeatedly and stopping at the first significant result inflates the false-positive rate. Decide the sample size in advance, or use a sequential method designed for early stopping.
Can I still use Google Optimize?
No. Google Optimize and Optimize 360 shut down on September 30, 2023. Use Shopify Rollouts or a testing app, and analyse the results in Shopify’s reports or your analytics tool.
Sources and further reading
- Shopify Help Center: Rollouts
- Shopify Help Center: Requirements and considerations for rollouts
- Shopify Help Center: Types of rollouts and changes
- Shopify Help Center: Creating a rollout
- Shopify Help Center: Rollout analytics
- Shopify Changelog: A/B test new themes and checkout configurations
- Shopify Help Center: Changes to sessions and conversion rate
- Evan Miller: How Not To Run an A/B Test
- Evan Miller: Sample size calculator
- Evan Miller: Simple Sequential A/B Testing
- Google: Google Optimize sunset
- Shoplift: Choose the right test type
- Intelligems: What you can test
- Microsoft Learn: Clarity FAQ
- Microsoft Learn: Clarity on Shopify
- Competition Bureau: Ordinary selling price
Keep exploring



