What a confidence interval measures
The result you see at the end of a test is one sample of reality, not reality itself. Run the same test on a different week, with a different set of visitors, and the number would move. A confidence interval puts a bracket around the measured result and says that, given the data actually collected, the true effect plausibly sits somewhere inside that bracket.
The width of the bracket is the honest part. A small sample produces a wide interval, because a handful of extra conversions either way would swing the outcome. As traffic accumulates the interval tightens around a single figure. Nothing about the change being tested has altered; you simply know more about it.
The bracket is normally reported as a lower and an upper bound on the difference between the variant and the control. When both bounds sit above zero, the variant is plausibly better across the whole range. When the bracket straddles zero, the same data are consistent with a gain, with no change at all, and with a loss.
Why confidence intervals matter
A single headline number invites a decision the evidence cannot support. Told that a redesigned checkout button improved conversions, nobody in the room asks whether the improvement might be nothing. The interval forces that question into the open, which is exactly what you want before rebuilding a site around a result.
It also improves the business conversation. A range lets you plan against the pessimistic end: if the worst plausible outcome still pays back the development work, the change is worth shipping even if the optimistic end never materialises. That is far more useful than arguing about whether a test won, and it is a large part of what disciplined conversion rate optimisation actually buys you.
Common mistakes with confidence intervals
The commonest error is reading the interval as a promise. It is not a guarantee that the true effect lies inside it; it is a statement about how the method behaves across many repeated tests. Treat it as a measure of what you still do not know, not as a boundary reality has agreed to respect.
The second is quoting only the midpoint. A wide interval with a flattering midpoint is a weak result dressed as a strong one, and stakeholders will remember the midpoint long after the caveat is forgotten. The third is testing several variants at once and reporting whichever bracket suits the preferred option — more comparisons mean more chances for noise to look like a finding.
How to act on it
Ask for the interval, not just the winner, on every test result you are shown, and check first whether zero falls inside it. If it does, the honest summary is that the test was inconclusive. That is a legitimate outcome, and usually a cheap one compared with rolling out a change that does nothing.
Where the interval is too wide to act on, the answer is more data or a bolder change, not a more generous reading. Decide before launch how small an effect would still be worth having, size the test around it, and hold the run to its planned length so statistical significance and the interval are read once, at the end, rather than watched daily.