What is A/B testing? A practical guide

A/B testing (split testing) is the practice of showing two or more versions of a page, message or experience to comparable slices of live traffic and measuring which performs better against a defined goal. It replaces opinion with evidence: instead of debating the headline, you run both and let visitor behaviour decide.

How an A/B test actually runs

A well-formed test starts with a hypothesis ("a benefit-led headline will lift demo requests, because visitors currently bounce before understanding the offer"), a single primary metric, and a variant that isolates the change. Traffic is split between control and variant; the test runs until it reaches sufficient sample size; the outcome — win, loss or inconclusive — is recorded with the learning either way.

Losses are not failures. A recorded loss is knowledge the next brief inherits; an unrecorded one is a test someone will run again next year.

Why testing programs stall

The most common killer of testing programs is not statistics — it is the developer queue. When every variant needs a ticket, a sprint and a release window, "we should test that" quietly becomes "we shipped it and hoped". The second killer is broken variants: a test that renders wrong on mobile burns traffic and trust. The third is amnesia: without a durable record of hypotheses and outcomes, teams repeat old tests and re-litigate settled questions.

Fixing throughput matters more than fixing sophistication. A program that ships four honest tests a month compounds; a perfect test that ships quarterly does not.

What makes a trustworthy variant

Before a variant sees traffic it should be verified on the real page: layout intact, elements present, behaviour correct on mobile. And when the underlying site changes mid-test, variants need re-checking — silent breakage mid-flight is worse than not testing, because it produces confident wrong answers.

From testing to an experimentation program

The step-change comes when tests stop being events and become a system: a backlog of hypotheses, a steady shipping cadence, and a searchable record connecting each outcome to the campaigns it informed. That record is the asset — it is what lets a new team member inherit three years of learning instead of a folder of screenshots.

Evolve, Cresia’s experimentation product, attacks the throughput problem directly: describe the change in plain language, get a variant built and verified on your live page, ship it through your existing testing platform, and keep every outcome in one compounding record.

Common questions

How long should an A/B test run?

Until it reaches the sample size your metric needs — typically at least one full business cycle (a week or two) so weekday and weekend behaviour are both represented. Stopping early because a variant is "clearly winning" is the classic way to ship noise.

What should we test first?

The highest-traffic page with the clearest goal — usually a landing page headline, hero or form. Early wins fund the program politically; early losses on obscure pages do not.

Do we need a data scientist to run A/B tests?

No. Modern testing platforms handle the statistics. What programs actually lack is throughput (variants built and verified quickly) and memory (outcomes recorded where the next person can find them).

A/B test or multivariate test?

Start with A/B. Multivariate tests need much more traffic and answer subtler questions. Most teams have bigger, simpler questions waiting.

Keep learning

See the platform behind the practice.

Cresia puts creative, experiments, media and measurement in one workspace.

Request a demo