Analytics and Tracking

Sample Size

Also called required sample, test size

The number of visitors and conversions a test needs before a difference can be told apart from ordinary noise.

Quick facts: Sample Size

Category
Analytics and Tracking
Also called
required sample, test size
Level
Intermediate
Affects
Test reliability, run length, whether a test is worth starting
Where to see it
Free sample size calculators, testing platforms such as VWO or Optimizely, GA4
In this article4
  1. How sample size is calculated
  2. Why sample size matters
  3. Where sample size goes wrong
  4. How to act on it

How sample size is calculated

Any result contains noise. Two identical pages shown to two random groups will produce slightly different conversion rates simply because different people arrived. Sample size is the amount of traffic needed before a real difference can be told apart from that everyday variation with reasonable confidence.

A sample size calculator asks for a small set of inputs, and each one moves the answer:

  • Your current conversion rate on the page being tested — lower rates need more traffic.
  • The smallest improvement that would be worth acting on — small improvements need far more traffic than large ones.
  • How confident you want to be that a declared winner is genuine.
  • How willing you are to miss a real improvement that the test fails to detect.

The output is a number of visitors per version, which you then divide into your weekly traffic to get a run length. That run length is the honest answer to whether the test is worth starting at all.

Why sample size matters

It decides whether a test can answer its question before it begins, which saves weeks of work that would have produced nothing. Calculating it first often reveals that a page cannot support the test at all, and that is a useful finding, not a disappointment.

It also protects against the most expensive testing mistake: rolling out a change that never actually worked. A difference read too early is usually noise, and building the rest of the site around it spreads the error everywhere.

Where sample size goes wrong

Watching a test daily and stopping it the moment the variant is ahead is the classic error. Early in a test the two lines cross repeatedly, so a leader will always appear if you look often enough. Deciding the sample size in advance and then ignoring the dashboard until it is reached is the only defence.

Counting the wrong thing is the second error. Sample size is about visitors and conversions, not days, so a test run for a fixed fortnight on thin traffic is not finished simply because the fortnight ended. Counting sessions rather than people, or including traffic the test never applied to, inflates the count and finishes the test early on paper only.

How to act on it

Calculate before you build. If the required run length runs to months, do not run that test: pick a bolder change, test on a page with more traffic, or test earlier in the funnel where volumes are higher. A large difference is detectable with far less traffic than a subtle one.

Where traffic is genuinely limited, which is common for smaller businesses and for markets like Nepal where a niche audience is small, test a measure that occurs more often, such as clicks to the enquiry step rather than completed sales. Set the required sample alongside your hypothesis, run whole weeks so weekday and weekend behaviour is included, and judge the result with statistical significance rather than by eye.

Do and do not

Do

  • Calculate the required sample before building the test
  • Divide it by weekly traffic to get a run length
  • Run whole weeks so weekends are included

Do not

  • Stop the test the moment a version pulls ahead
  • Count days instead of visitors and conversions
  • Test a subtle change on a low-traffic page

Questions people ask about this

How do I work out the sample size I need?

Use a free sample size calculator and give it your current conversion rate, the smallest improvement that would be worth acting on, and how confident you want to be. It returns the visitors needed per version. Divide that by your weekly traffic to that page, and you have the run length.

My site does not get enough traffic. What can I test?

Test things that happen more often and change more sharply. Measure clicks through to the enquiry step rather than completed sales, test on your highest-traffic page rather than a niche one, and make the change bold rather than subtle. Large differences need far fewer visitors before they can be trusted.

Can I stop a test early if one version is clearly winning?

Only if the planned sample has been reached. Early in any test the two versions trade places repeatedly, so an apparent winner appears simply because you kept looking. Stopping at that moment is how teams roll out changes that never worked. Set the sample and the end date first, then leave it alone.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.