How an A/B test works
Two versions of something — a landing page, a headline, an email subject line, an ad — run at the same time, and arriving traffic is split between them at random. Because the split is random and the timing is identical, the only systematic difference between the two groups is the change you made, so a difference in outcome can be attributed to it rather than to the day of the week or the campaign that sent the visitors.
The test needs a single decided metric before it starts, usually enquiries or sales rather than clicks. It also needs enough traffic and enough time for the result to mean something. With small numbers, one unusually good week can make the losing version look like the winner, which is why a test has to run to a planned stopping point rather than until the result looks pleasing.
Why A/B testing matters
It replaces opinion with evidence at the exact point where opinions are strongest and least reliable — what a page should say, which offer works, whether the form is too long. Teams argue about these endlessly, and a test settles the argument for that audience, on that page, at that time.
It also protects you from confident redesigns. Changing everything at once and watching the total move tells you nothing about which change helped, and a redesign that performs worse is usually discovered too late to unpick. Testing one change at a time is slower but leaves you knowing something afterwards.
Where A/B tests go wrong
Stopping early is the classic error. Results swing wildly at the start, and a test watched daily will always show a winner at some point. Deciding the sample size and duration in advance, then leaving it alone, removes the temptation.
Testing on thin traffic is the more common problem for small businesses. If a page receives a handful of enquiries a month, no split test will reach a trustworthy answer within a useful timeframe, and running one anyway produces confident nonsense. Other faults: changing several things at once so the winner cannot be explained, running a test across a festival or sale that distorts behaviour, and measuring clicks when what you needed was revenue. Traffic sent to only one version by an ad or an email also breaks the randomisation and invalidates the whole comparison.
How to run one properly
Start with a written prediction: what you are changing, what you expect to happen, and why. Choose one outcome metric, work out roughly how much traffic and how long you will need, and commit to that before switching it on. Change one meaningful thing — a genuinely different offer or structure, not a button colour — because small cosmetic changes need far more traffic to detect than they are worth.
If your traffic is too thin to test, do not fake it. Make changes based on evidence you can gather instead, such as recordings, enquiry conversations and obvious usability faults, and reserve testing for the pages that receive enough volume. Where testing is viable, it belongs inside a wider conversion rate optimisation process, and the result only counts once it clears statistical significance.