Most small sites do not have the traffic to run a valid A/B test, run them anyway, and act on the results. This is how to tell whether testing is available to you, and what to do when it is not.
Key takeaways
- A reliable test needs around 300 conversions per variant to detect a 25% lift. Below that you are measuring randomness.
- Stopping a test the moment it looks significant is the most common way to manufacture a fake winner.
- Test one substantial change at a time. Button colours are why testing has a reputation for wasting time.
- Under a few hundred conversions a month, qualitative research and obvious-fix work beat testing outright.
On this page
Whether you can test at all
The traffic you need depends on your current conversion rate and the size of the improvement you want to detect. Smaller improvements need dramatically more data, and this relationship surprises almost everyone.
| Improvement to detect | Conversions needed per variant | Realistic for |
|---|---|---|
| 50% relative | Around 95 | Almost any site |
| 25% relative | Around 310 | Steady mid-size sites |
| 10% relative | Around 1,700 | High-volume sites only |
| 5% relative | Around 6,400 | Large ecommerce and SaaS |
Click to download this table as an image
A site with sixty conversions a month cannot detect anything smaller than a transformation, and transformations do not come from testing button copy. That is not a reason for despair. It is a reason to spend the effort somewhere with a better return.
How tests produce false winners
Watching a running test and stopping when it reaches significance guarantees false positives. Significance fluctuates constantly during a test; if you keep checking and stop at the first favourable crossing, you will find one eventually in almost any test, including one between two identical pages.
Peeking
Decide the sample size and the end date before you start, then do not stop early no matter how the numbers look. If you want to monitor the test, monitor it for breakage, not for results.
Running too briefly
Run for whole weeks, always. Tuesday traffic behaves differently from Saturday traffic, and a test that runs Monday to Thursday measures a slice of your audience rather than your audience. Two full weeks is a sensible floor even when the sample arrives sooner.
Testing too many things
Every extra variant and every extra metric is another chance for coincidence to look like a discovery. Four variants measured on five metrics gives twenty opportunities for a random result to reach conventional significance, and roughly one will.
What is worth testing
Test changes big enough that you would be willing to ship them permanently on judgement alone. If the change is too trivial to defend without a test, it is too trivial to detect with one.
- The core offer or its framing: what you are actually promising, and at what price. That is a positioning question first.
- The number and nature of form fields. Removing fields is among the most reliably positive changes available, and form completion is one of the numbers worth watching.
- Page structure: what appears before the fold, and in what order. This is where landing page design decisions are won or lost.
- Whether a step in the funnel needs to exist at all.
- The headline, when it changes the proposition rather than the wording.
Notice that these are strategic questions wearing test clothing. That is the point: the value is in resolving a genuine disagreement about direction, not in finding a two per cent lift.
What to do when you cannot test
Below testing volume, the fastest gains come from watching real people use the site and fixing what obviously breaks. This is less rigorous and considerably more productive.
- Watch five session recordings of people who abandoned the funnel. Five is usually enough to find something embarrassing. Microsoft Clarity records sessions for free, and we can add it during an analytics and dashboard setup.
- Complete your own checkout or enquiry on a phone, on mobile data, as a stranger would. If it feels slow, a speed and Core Web Vitals fix often beats any test.
- Read your support enquiries for questions the site should have answered.
- Fix everything that is plainly broken before optimising anything that works.
- Ship the improvement and compare month to month, accepting that you will not have proof, only direction.
Our UX and conversion review is built for exactly this situation, a structured pass over the funnel that finds the obvious losses without needing statistical power you do not have.
On a small site, the first ten fixes are not close calls. You do not need a test to tell you the form is broken on mobile.
Serhii Yelbaiev, Well Web Marketing
Frequently asked questions
Until it reaches the sample size you calculated in advance, and always for whole weeks. Two full weeks is a practical minimum regardless of how quickly the sample accumulates.
That is a real and useful result. It means the change did not matter at a scale you can detect, so ship whichever version is simpler to maintain and move on.
Not meaningfully. Direct that effort to qualitative research instead: recordings, interviews and your own support inbox will find more in a week than an underpowered test will in a quarter.
It is a convention, not a law. For a low-risk, cheap-to-reverse change, a lower bar is defensible. For anything expensive or hard to undo, keep the conventional threshold.
Sources and method
- Sample size figures are standard two-proportion calculations at 95% confidence and 80% power, from a 3% baseline conversion rate, recalculated in October 2026. Check your own numbers with Evan Miller’s sample size calculator.
- Evan Miller, How Not To Run an A/B Test (why stopping early creates false winners)
- Practical observations come from our own conversion work with small and medium clients. Last reviewed October 2026.
