How sample size is calculated
Any result contains noise. Two identical pages shown to two random groups will produce slightly different conversion rates simply because different people arrived. Sample size is the amount of traffic needed before a real difference can be told apart from that everyday variation with reasonable confidence.
A sample size calculator asks for a small set of inputs, and each one moves the answer:
- Your current conversion rate on the page being tested — lower rates need more traffic.
- The smallest improvement that would be worth acting on — small improvements need far more traffic than large ones.
- How confident you want to be that a declared winner is genuine.
- How willing you are to miss a real improvement that the test fails to detect.
The output is a number of visitors per version, which you then divide into your weekly traffic to get a run length. That run length is the honest answer to whether the test is worth starting at all.
Why sample size matters
It decides whether a test can answer its question before it begins, which saves weeks of work that would have produced nothing. Calculating it first often reveals that a page cannot support the test at all, and that is a useful finding, not a disappointment.
It also protects against the most expensive testing mistake: rolling out a change that never actually worked. A difference read too early is usually noise, and building the rest of the site around it spreads the error everywhere.
Where sample size goes wrong
Watching a test daily and stopping it the moment the variant is ahead is the classic error. Early in a test the two lines cross repeatedly, so a leader will always appear if you look often enough. Deciding the sample size in advance and then ignoring the dashboard until it is reached is the only defence.
Counting the wrong thing is the second error. Sample size is about visitors and conversions, not days, so a test run for a fixed fortnight on thin traffic is not finished simply because the fortnight ended. Counting sessions rather than people, or including traffic the test never applied to, inflates the count and finishes the test early on paper only.
How to act on it
Calculate before you build. If the required run length runs to months, do not run that test: pick a bolder change, test on a page with more traffic, or test earlier in the funnel where volumes are higher. A large difference is detectable with far less traffic than a subtle one.
Where traffic is genuinely limited, which is common for smaller businesses and for markets like Nepal where a niche audience is small, test a measure that occurs more often, such as clicks to the enquiry step rather than completed sales. Set the required sample alongside your hypothesis, run whole weeks so weekday and weekend behaviour is included, and judge the result with statistical significance rather than by eye.