Split testing is treated as a universally applicable practice, but its reliability depends entirely on volume. Below a certain level of traffic the method produces results indistinguishable from chance.

Confidence depends on conversions, not visitors

The relevant sample is the number of conversions observed, and a site with modest traffic and a low conversion rate accumulates those very slowly.

Detecting a small improvement requires far more data than detecting a large one, and the improvements available on a competent page are usually small.

A test that would need months of traffic to resolve is not a test the business can run, whatever the tooling suggests.

Early results swing widely and look convincing

In the first days of a test the difference between variants moves substantially, because each additional conversion shifts a small denominator.

Watching a dashboard during that period produces a strong impression of a winner, which is why tests are so often stopped early.

Stopping when a result looks good systematically selects for random fluctuations, so the recorded win rate of a testing programme can be high while the revenue effect is absent.

Running many tests multiplies false positives

Testing many variants raises the probability that at least one shows a difference by chance alone, independently of whether any real effect exists.

Programmes that report a long list of small wins with no corresponding movement in overall conversion are usually observing this.

The aggregate metric is the check, because it does not benefit from the selection that produces the individual results.

Larger changes are testable when small ones are not

A low-traffic site can detect a substantial effect, which points toward testing meaningful changes — a different offer, a restructured page, a changed pricing model — rather than variations in wording.

These tests are riskier and slower to build, and they are the only ones the available sample can actually resolve.

The alternative is to accept that many decisions will be made on judgement and qualitative evidence, which is a more honest position than a test that cannot conclude.

Qualitative methods do not need volume

Watching a handful of people attempt a task reveals obstacles clearly, because a confused user is confused visibly rather than statistically.

That evidence identifies problems but not magnitudes, so it is best used to generate the large changes worth testing rather than to replace measurement.

Used together the sequence is workable: qualitative work finds the problem, a substantial change addresses it, and the aggregate metric confirms whether anything moved.

Sessions of this kind also cost a fraction of what a testing programme consumes in engineering time, which matters most at exactly the traffic levels where the testing programme cannot conclude anything.