Testing

How to A/B Test Changes on Your Shopify Store

How to run Shopify tests that give real answers: what to test, how much traffic you need, and why stopping early is the most common mistake.

11 min readTesting · CRO

Most Shopify A/B tests produce a confident conclusion that is wrong. Not because the tooling is bad, but because ecommerce testing is statistically harder than it looks and the failure modes are invisible from inside the dashboard.

This guide covers the parts that actually determine whether a test tells you the truth: what to test, how much traffic you need, how long to run, when to stop, and what to do instead when your store is too small to test properly. Because for a lot of Shopify stores, the honest answer is that A/B testing is the wrong tool — and knowing that is worth more than running tests that can never resolve.

What A/B testing actually gives you#

It gives you one thing: a causal answer. Version B caused more orders than version A, rather than merely coinciding with more orders.

That matters because everything else about a store changes at the same time as your change. Traffic mix shifts, a campaign launches, a competitor discounts, the season turns. A before-and-after comparison confuses all of that with your change. A properly run split test does not, because both versions experience the same conditions simultaneously.

That is the whole value proposition, and it is a big one. It is also why cutting statistical corners defeats the purpose entirely — a badly run test gives you a before-and-after comparison wearing a lab coat.

Do you have the traffic?#

Answer this before anything else.

The uncomfortable arithmetic: the smaller your baseline conversion rate and the smaller the effect you want to detect, the more traffic you need — and the relationship is steep, not linear. Detecting a small relative improvement on a low base can require a volume of conversions that many stores do not generate in a quarter.

Run the numbers before you commit. Use any reputable sample size calculator and give it four inputs:

  1. Baseline conversion rate for the specific page and audience you are testing, not your blended store rate.
  2. Minimum detectable effect — the smallest improvement worth shipping. Be honest. A 1% relative lift is not worth a two-month test.
  3. Statistical significance level, conventionally 95%.
  4. Statistical power, conventionally 80%. This is the one people omit, and omitting it is why so many tests are underpowered and inconclusive.

The calculator returns the visitors per variation you need. If that number is more than your page gets in six to eight weeks, do not run the test. You will either stop early and be wrong, or run for a season and have seasonality contaminate the result.

Two ways to make the arithmetic work:

  • Test bigger changes. A wholesale rethink of a product page produces a larger effect than a button colour, and larger effects need less traffic to detect. Small stores testing small changes is the classic dead end.
  • Test earlier funnel metrics. Add-to-cart rate has a much higher base rate than purchase rate, so it needs far less traffic. Use it as your primary metric with purchase rate as a guardrail — while remembering that a lift in add-to-carts that does not reach checkout is not a win.

What to test#

Test things where you have a real hypothesis and a mechanism. "I wonder if green converts better" is not a hypothesis. "Shoppers are abandoning because shipping cost appears late — showing it on the product page should raise checkout completion" is.

Get hypotheses from evidence, in roughly this order of yield:

  • Friction signals. Rage clicks and dead clicks point at things that are actually broken. Fix these; do not test them. A broken thing does not need a control group.
  • Scroll depth. If most visitors never reach your reviews, testing a version with reviews moved up has an obvious mechanism. See scroll depth analytics.
  • Session replay. Watch ten failed sessions on the page. The recurring hesitation is your hypothesis.
  • Support inbox. Recurring pre-purchase questions are objections your page fails to answer.
  • Funnel drop-offs. The largest drop weighted by traffic reaching it is where testing pays best.

High-yield test candidates on Shopify, in rough order:

  1. Product page buy-box content — delivery estimate, shipping cost, returns line placement.
  2. Cart page structure — shipping estimate, free-shipping progress, express wallet placement.
  3. Collection page filtering and sort defaults.
  4. Homepage value proposition and primary action.
  5. Popup timing and trigger, which is often net-negative and rarely tested.

Low-yield candidates that consume most testing budgets: button colours, minor copy tweaks, trust badge shuffling, hero image swaps with no accompanying message change.

How to run it without breaking it#

Randomise properly and keep assignment sticky. A returning visitor must see the same variation every time. If assignment is per-session rather than per-visitor, a shopper can see version A on Monday and version B on Wednesday, which contaminates the result and produces a genuinely confusing experience.

Run for full weeks. Weekday and weekend traffic behave differently. A test that runs Tuesday to Friday oversamples weekday behaviour. Always run in whole seven-day multiples, minimum two full weeks even when the sample size is reached sooner.

Do not overlap tests on the same page unless your tool supports proper multivariate handling. Two concurrent tests on the same template will interact and neither result will be interpretable.

Avoid flicker. Client-side testing tools that repaint the page after load create a visible flash of the original version. This annoys users, hurts perceived speed, and biases results toward the control. If your tool flickers, fix it before you trust anything it tells you.

Watch the speed cost. Testing scripts are render-blocking by nature. A test that slows the page can lose on speed while winning on design, and you will draw the wrong conclusion.

Do not change anything else mid-test. No price changes, no new apps, no campaign launches targeted at the tested page. If something changes, note it and consider the test compromised.

Stopping: the mistake that ruins most tests#

This is the single most common and most damaging error in ecommerce testing.

Peeking and stopping when significance appears is not a valid procedure. If you check daily and stop the moment the tool says 95%, your real false-positive rate is far above 5% — you have effectively run many tests and kept the one that looked good. Evan Miller's breakdown of repeated significance testing puts a number on this: under continuous monitoring, a supposed 5% error rate can actually land above 25%. Every conversion tool encourages this by showing a live significance readout.

Do it properly:

  1. Fix the sample size and duration before the test starts. Write them down.
  2. Run to that number. Do not stop early because the trend looks good.
  3. Only then read the result.

If you genuinely need to monitor for a disaster mid-test, define a stopping rule for harm only — for example, stop if the variation is dramatically worse — and never use the same latitude to declare a winner.

Two related traps:

  • Novelty effect. Returning customers react to any change simply because it is new. Early results skew positive and then decay. Another argument for full-duration runs.
  • Segment mining. Finding that a change won for mobile visitors in one country when the overall test was flat is not a finding. It is what randomness looks like when you slice it enough times. If you want to test a segment, test that segment deliberately from the start.

Reading the result honestly#

A flat result is a real result. It means the change was not worth the effort, which is genuinely useful information. Most tests do not win; that is normal and not a sign you are testing badly.

A win needs a mechanism. If a variation wins and you cannot explain why in one sentence, be suspicious. Explainable wins generalise to other pages; inexplicable ones often do not replicate.

Check the guardrails. A change that raises add-to-carts while lowering checkout completion has moved a problem downstream. A change that raises conversion rate while lowering average order value may lose money. Always look at revenue per visitor alongside conversion rate.

Value it in money. Translate the lift into expected revenue over a period, so you can rank the next round of work — that is what impact and revenue attribution exist for. A percentage with no monetary value attached does not help you prioritise.

What to do if you cannot test#

Most Shopify stores cannot run a properly powered test on purchase rate. That is not a failure; it is arithmetic. Here is the honest alternative hierarchy.

Fix broken things without testing. If a variant picker fails silently or a button does not respond on mobile, ship the fix. You do not need a control group to justify repairing something.

Test on higher-base-rate metrics. Add-to-cart rate and checkout-started rate need far less traffic than purchase rate.

Use sequential testing with discipline. Change one thing, log the exact date, and compare a clean period after against the same period last year — not against last week, which is dominated by traffic mix. This is weaker evidence than a split test. It is much better than nothing, and it is what most small stores should actually do.

Use qualitative evidence more heavily. Ten session replays and a read of your support inbox will identify problems that no amount of testing volume would surface anyway. For low-traffic stores, qualitative work has a better return than statistics.

Batch changes deliberately. If you cannot isolate individual effects, ship a coherent set of improvements to one page, then evaluate the page as a whole against last year. You lose attribution to individual changes and keep a defensible read on direction.

The short version#

  1. Calculate the sample size before you start. If you cannot reach it in six to eight weeks, do not run the test.
  2. Test big changes with a real hypothesis, sourced from behaviour, not from opinion.
  3. Run full weeks, minimum two, and do not stop early.
  4. Read the guardrail metrics, not just the headline.
  5. If you cannot test, fix what is broken, change one thing at a time, log it, and compare year over year.

The discipline is what produces trustworthy answers, not the tool. A store that ships three well-evidenced fixes a quarter and measures them honestly will beat one that runs a dozen underpowered tests and believes all of them.

Next: the Shopify CRO checklist for what to work on, and 12 reasons your store isn't converting for where to look first.

Try it on your store

Stop guessing. Fix it and see the money.

DynoWeb watches how shoppers actually behave on your storefront, points at the step that is costing you orders, and shows the revenue each fix moved.