Blog · Craft
— Craft··8 min read

A/B testing for small sites: how to test with under 20k monthly visitors

Joona Heinonen· Choco Media · Rovaniemi

A/B testing on a small site can feel like trying to hear a whisper in a crowd. You set up an experiment, wait two weeks, and your stats tool tells you the result is “not significant” — so you either abandon the test or, worse, act on noise. At Choco Media, we run A/B testing for clients whose sites receive under 20,000 monthly visitors, and the thing we hear most often is: “we tried it, it didn’t work.” Usually the test worked fine. The method was wrong for the traffic level.

This post is for founders, marketing leads, and in-house teams who want to make data-driven decisions without waiting until they have enterprise-scale traffic. You will leave with a testing method that fits your volume, a prioritisation lens for which pages to touch first, and a realistic picture of what statistical confidence actually requires at this scale.

Low-traffic testing is not a second-class version of proper CRO. It just requires different tools and different expectations — and once you know both, it is genuinely useful.

Why standard A/B testing breaks down below 20k visitors

The standard frequentist A/B test — the kind built into most testing tools by default — was designed for high-volume environments. It calculates statistical significance based on the assumption that you will accumulate enough conversions in both variants to reliably distinguish a real effect from random variation. At 20,000 monthly visitors with a 2% conversion rate, you might collect 400 conversions per month. Split 50/50 across two variants, that is 200 conversions each. To detect a 15% relative improvement with 80% statistical power at 95% confidence, you typically need 3,000+ conversions per variant. You are looking at 15+ months per test.

That is not CRO. That is archaeology.

The two problems this creates

Neither outcome is inevitable. The fix is choosing methods built for your traffic level — not squeezing your traffic into a tool designed for ten times your volume.

Sequential testing: the method that fits small sites

Sequential testing (also called “always-valid inference” or “anytime-valid testing”) solves the core problem: it lets you look at results continuously without inflating false positives. Traditional frequentist tests fix a sample size in advance and penalise you for peeking. Sequential tests build the sample size check into the statistical model itself.

Tools like Optimizely Stats Engine and VWO SmartStats use variants of this approach. You can check results daily, stop when significance is reached, and trust the number — because the method accounts for the sequential checking.

How to set it up

Bayesian testing: reasoning under uncertainty

Bayesian testing approaches the question differently: instead of asking “is this result statistically significant?”, it asks “what is the probability that variant B is better than variant A, given the evidence so far?”

A 90% probability that variant B is better is not a guarantee — but it is a usable signal. On a small site, it might be the most honest answer the data can give you.

Bayesian methods work well for small sites because they incorporate prior knowledge (what conversion rates typically look like on similar pages) and update continuously as data arrives. Tools like AB Tasty and some Google Analytics-adjacent stacks offer Bayesian modes, and lightweight Bayesian calculations can also be run directly in Google Sheets with published templates.

What Bayesian tests give you that frequentist tests do not

The trade-off: Bayesian results require interpretive judgement. A 75% probability that B is better might be enough to ship in one context and not in another. Define your decision threshold before you see the result — not after.

Where to focus: high-intent pages first

On a small site, you cannot test everywhere. Traffic is a finite resource, and spreading it across many simultaneous experiments means none of them accumulate signal fast enough to be useful. The prioritisation rule we apply in client work is simple: test where intent is highest and the page has the most direct influence on conversion.

This usually means:

Our conversion rate optimisation service starts every engagement with a page priority matrix — because the method matters less than testing the right thing first.

Micro-conversions: faster signal without waiting for the final event

If your primary conversion happens too infrequently to accumulate test data at a useful pace, move up the funnel. Micro-conversions — events that precede and predict the final conversion — give you faster signal while remaining directionally useful.

Micro-conversions worth tracking

The caveat: micro-conversions can diverge from macro-conversions. A variant that drives more CTA clicks might not drive more actual sign-ups if the post-click experience is weaker. Use micro-conversions as a directional filter to short-list variants — not as a replacement for the final conversion metric. Session replay tools like Microsoft Clarity (free) or Hotjar can help you understand which on-page behaviours actually precede conversions in your current traffic before you decide what to test.

Traffic segmentation: test on the right audience

One underused tactic on small sites: instead of splitting all visitors between variants, test only the segment most relevant to your hypothesis. If you are testing a headline aimed at first-time organic visitors, there is no reason to include returning direct visitors — they will dilute the signal and may react differently, polluting results.

Segments worth isolating

A tighter, more relevant segment means your test accumulates conversions from the people the change actually matters to. The results are more likely to hold when you roll out the winner to everyone.

The flip side: a narrower segment means fewer users per variant. This is exactly where sequential or Bayesian methods become essential — they tolerate lower absolute conversion counts better than classical significance tests do. The CRO playbook we published covers how segmentation fits into a broader optimisation programme.

The testing calendar: structure that prevents the common mistakes

Most small-site testing programmes fail not because of the statistics but because of the process. Tests get interrupted, run only over weekends, or get called early when a team meeting happens to fall mid-experiment. A lightweight calendar discipline prevents this.

Rules we apply to every test

What to do when a test is inconclusive

An inconclusive test is not a failed test. It is information: either the change does not have a meaningful effect on this page, or the effect is smaller than your site can detect in a reasonable timeframe. Both outcomes are useful. If a major hero section redesign produces no detectable improvement, the bottleneck is probably elsewhere — in your offer, your pricing, your traffic quality, or a post-click experience issue.

Next steps after an inconclusive result

One pattern we see repeatedly: teams test button colour while the actual bottleneck is a confusing pricing structure or a form that asks for too much information too early. A CRO audit — qualitative research first, then quantitative testing — almost always reveals a clearer priority than intuition alone. If you would like to talk through your current setup, get in touch and we will take a look together.

— Work with Choco Media

Want posts like this working for your business?

10–40 SEO + AI-optimised blog posts a month, researched, senior-edited and published straight to your site. Built to rank on Google and get cited by ChatGPT, Claude and Gemini.

See plans — from €199/mo →
No start-up fee · Price locked for 12 months · Cancel any time after
← All storiesNext story →
— Free tips, monthly

Get the playbook, for free.

One short letter a month — the prompts we use, the campaigns that worked, the AI tools worth the time. No sales pitch, just field notes.

— Want us to do it for you?

Hire the agency.

AI-accelerated content, paid media, brand and web — delivered by one small team that talks to itself. Currently taking on a handful of clients each quarter.

Book a call