A/B Testing
A/B testing compares two versions of a page or flow against a single metric to determine which performs better. Valid results require enough conversions to reach statistical significance and changes substantial enough to produce a detectable difference.
Most A/B tests fail before they start, because the change is too small to detect or the traffic too low to reach significance. Testing a button colour on 300 monthly conversions will never produce a trustworthy answer.
We calculate the required sample before running anything, test changes big enough to move the number, and stop tests at the planned point rather than when the result looks good, which is how most teams fool themselves.
What's included
- Hypothesis development from funnel data
- Sample size calculation before launch
- Test implementation and QA
- Significance monitoring with a pre-set stopping rule
- Result analysis including inconclusive outcomes
- Winner implementation
- Test log so learnings accumulate
Questions people actually ask
How much traffic do I need to A/B test?
Roughly 1,000 conversions per month as a working floor. Fewer than that and tests run for months or produce false positives. With low traffic, make well-reasoned changes and track the trend rather than running underpowered tests.
How long should a test run?
Until it reaches the sample size you calculated, and at least two full weeks to cover weekly cycles. Stopping early because a variant is ahead is the most common way teams ship changes that do nothing.
What if a test is inconclusive?
That is a legitimate and common result, and it usually means the change was too small to matter. Keep the simpler version and test something more substantial.

