What is A/B testing? A practical guide
A/B testing (split testing) is the practice of showing two or more versions of a page, message or experience to comparable slices of live traffic and measuring which performs better against a defined goal. It replaces opinion with evidence: instead of debating the headline, you run both and let visitor behaviour decide.
How an A/B test actually runs
A well-formed test starts with a hypothesis ("a benefit-led headline will lift demo requests, because visitors currently bounce before understanding the offer"), a single primary metric, and a variant that isolates the change. Traffic is split between control and variant; the test runs until it reaches sufficient sample size; the outcome — win, loss or inconclusive — is recorded with the learning either way.
Losses are not failures. A recorded loss is knowledge the next brief inherits; an unrecorded one is a test someone will run again next year.
Why testing programs stall
The most common killer of testing programs is not statistics — it is the developer queue. When every variant needs a ticket, a sprint and a release window, "we should test that" quietly becomes "we shipped it and hoped". The second killer is broken variants: a test that renders wrong on mobile burns traffic and trust. The third is amnesia: without a durable record of hypotheses and outcomes, teams repeat old tests and re-litigate settled questions.
Fixing throughput matters more than fixing sophistication. A program that ships four honest tests a month compounds; a perfect test that ships quarterly does not.
What makes a trustworthy variant
Before a variant sees traffic it should be verified on the real page: layout intact, elements present, behaviour correct on mobile. And when the underlying site changes mid-test, variants need re-checking — silent breakage mid-flight is worse than not testing, because it produces confident wrong answers.
From testing to an experimentation program
The step-change comes when tests stop being events and become a system: a backlog of hypotheses, a steady shipping cadence, and a searchable record connecting each outcome to the campaigns it informed. That record is the asset — it is what lets a new team member inherit three years of learning instead of a folder of screenshots.
Evolve, Cresia’s experimentation product, attacks the throughput problem directly: describe the change in plain language, get a variant built and verified on your live page, ship it through your existing testing platform, and keep every outcome in one compounding record.
Common questions
How long should an A/B test run?
Until it reaches the sample size your metric needs — typically at least one full business cycle (a week or two) so weekday and weekend behaviour are both represented. Stopping early because a variant is "clearly winning" is the classic way to ship noise.
What should we test first?
The highest-traffic page with the clearest goal — usually a landing page headline, hero or form. Early wins fund the program politically; early losses on obscure pages do not.
Do we need a data scientist to run A/B tests?
No. Modern testing platforms handle the statistics. What programs actually lack is throughput (variants built and verified quickly) and memory (outcomes recorded where the next person can find them).
A/B test or multivariate test?
Start with A/B. Multivariate tests need much more traffic and answer subtler questions. Most teams have bigger, simpler questions waiting.