// Ecommerce and DTC
Conversion Rate Optimization for Ecommerce and DTC Stores
Sample size is rarely what stops a store from testing. The promotional calendar moves the baseline every few weeks. A result measured in peak season does not carry to February. The return that cancels the sale lands a month after the test was called.
What changes in Ecommerce and DTC
- The promotional calendar moves the conversion baseline every few weeks, so a year-round holdout says more than a long run time does
- A win measured in peak season was measured on visitors who had already decided to buy something, and February asks a different question
- Returns land weeks after the order, so a checkout win counted in orders can become a loss counted in kept revenue
- A free shipping threshold moves basket size and gross margin in the same test, so one number cannot settle it
- Repeat buyers with a saved card and first-time visitors typing one are doing different tasks on the same checkout page
Traffic is usually the thing that stops a testing programme. On a store it usually does not.
A site converting at 2% with 200,000 sessions a month can finish a two-arm test at a 10% detection threshold inside a month. The arithmetic is checkable. Separating 2% from 2.2% at 95% confidence and 80% power takes roughly 78,000 sessions per arm, so about 157,000 for the test. You have that. What you do not have is a stable population to spend it on.
Your problem is the calendar
Four weeks of traffic on a store spans a promotion, at least one payday, and possibly a holiday. Someone browsing in the quiet week and someone browsing on the second day of a sale are different people with different intent. The baseline conversion rate moves underneath you as the mix changes.
For the comparison between the two arms this is fine, and it is the part teams talk themselves out of. A concurrent split shows both arms the same weeks, so a calendar shift hits both equally. The test is still valid.
What breaks is the part nobody writes down. A checkout change that wins during a sale won on people who had already decided to buy something. February is a different question asked by a different visitor, and the win may be smaller, absent or reversed. So a result measured in peak gets recorded as a result measured in peak, and it gets re-run.
The other thing we put in is a holdout. A small slice of traffic that sees nothing all year, so the programme has something to be measured against at the end of it. Individual tests answer individual questions. A holdout is how you find out whether a year of them added up to anything.
The metric has to survive the return window
An order is not the end state. A return arrives weeks later and takes the revenue back out, and in some categories it takes shipping both ways with it.
That matters because the changes with the largest effect on orders are often the ones with the largest effect on returns. A size guide that pushes a hesitant buyer over the line. A free shipping threshold that adds a second item to the basket. A product page leaning on the flattering photograph. Each of those can win on orders and lose on what the customer kept.
So the readout comes twice. The call at the stopping rule, on orders, which is what the pre-registered plan fixed. Then a revision once the return window has closed, on kept revenue. Where the two disagree, the second one is the answer. We say that before the first test runs, because a programme reporting only the earlier number will always look better than it is.
Two populations in one funnel
A repeat buyer with a saved card and a first-time visitor typing a card number are doing different tasks on the same checkout page. A change helping one can hurt the other, and a blended conversion rate averages them into a flat result that hides both.
The split has to be pre-registered. Deciding afterwards which segment your test worked on is how a losing test gets written up as a win. We name the segments before traffic starts and report each one. The sample split and the loss of precision that comes with it are stated in the same table.
Tests we will not run in peak
We do not touch checkout during your highest revenue weeks. A defect in a variant costs more in those five weeks than the test could win back in a year. The population you would be measuring is also the one that generalises worst.
We do not build invented urgency either. A countdown that resets on refresh, or a stock count nobody ever checked against the warehouse, is a false statement about your inventory. On a store that is a different kind of exposure than it is on a software landing page, because someone can verify it. Where the scarcity is real, show it and show where the number came from. Where it is not, the widget does not get built.
This is the Ecommerce and DTC view of Conversion Rate Optimization. That page covers how the work runs whatever the sector.