How peeking works
A test result is not a fixed number waiting to be revealed. It wanders. Early in a run, when only a handful of conversions have landed on each side, one extra sale can swing the comparison dramatically, and the two variants trade places repeatedly before the picture settles.
Peeking is checking that wandering line repeatedly and stopping when it happens to be in a pleasing position. The problem is not curiosity — it is that every extra look is another chance for random variation to cross the threshold you are watching for. Look often enough at a test between two identical pages and, sooner or later, one of them will appear to be winning.
The statistics that testing tools report assume a single check at a planned end point. Read them repeatedly and stop on the first favourable reading, and the confidence figure on screen no longer describes what you actually did. It flatters you, and the amount by which it flatters you grows with the number of looks.
Why peeking matters
Peeking is the main reason a testing programme can be full of winners and still produce no change in the annual conversion rate. Each early stop banks a result that was mostly luck, the change ships, and the effect never appears in the accounts. Over a year that adds up to real development time spent on nothing.
It also damages trust. Once a team notices that tested changes do not show up in revenue, testing gets written off as theatre, and the genuinely good ideas lose their evidence base along with the weak ones. Protecting the discipline is worth more than any individual result.
Where peeking goes wrong in practice
It rarely looks like cheating. It looks like a manager asking how the test is doing on Monday, a dashboard emailing a daily summary, or a client meeting that happens to fall in week one. The pressure to answer with a verdict rather than “too early to say” is what does the damage.
A subtler version is stopping a test because it is losing. Killing a variant early on weak evidence is the same error pointed the other way, and it can bury a change that would have worked. The mirror image is extending a test past its planned end because the result has not yet turned favourable — waiting for a better number is peeking with extra steps.
What to do about it
Decide the end point before launch, using the effect you need to detect and the traffic that implies, and record it where everyone can see it. Then agree what may be looked at during the run: delivery, spend, errors and obviously broken pages, but not the conversion comparison.
Where stakeholders genuinely need progress updates, report progress towards the planned sample size rather than the current score. If early reads are unavoidable, use a testing tool that supports sequential or always-valid methods, which are built for repeated checking and adjust the threshold accordingly. Finally, read the result once, with its range, and let the planned test duration rather than the mood in the room decide when the experiment is over.