A/B Testing Without Lying to Your Boss

· 3 min read · Syed Omar Faruk Towaha
A/B Testing Without Lying to Your Boss

Your team changed the "Buy" button from blue to green. After three days, the green version has a 4% higher conversion rate and a p-value of 0.048. The product manager is drafting a celebratory Slack message.

Please, wait.

Mistake 1: peeking and stopping when it looks good

Peeking
Even with no real difference, the p-value wanders. Stop at the lucky moment and you 'win.'

A p-value bounces around as data arrives. If you check every day and stop the first time it dips below 0.05, you'll find "significant" results far more often than 5% of the time, even when the two versions are identical. Check often enough and you can make anything win.

Fix: decide the sample size before the test, and only make the decision when you reach it. If you need to monitor continuously, use methods designed for it (sequential testing), not a plain t-test checked daily.

Mistake 2: no sample size calculation

Small effects need big samples. If your conversion rate is 3% and you want to detect a 5% relative improvement, you may need tens of thousands of users per variant.

from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

effect = proportion_effectsize(0.0315, 0.03)   # 3.0% -> 3.15%
n = NormalIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.8)
print(round(n))   # users needed per group, roughly 100k+

If you can't get that many users, either test something bolder or accept that you can't detect small effects.

Mistake 3: testing twenty metrics

Look at enough metrics and one will move by chance: "conversion didn't change, but tablet users in Ontario on Tuesdays spent 11% more!" That's not a finding; that's a lottery ticket.

Fix: choose one primary metric before starting. Track a few guardrail metrics (page load time, refunds, complaints) to make sure you didn't break anything. Treat everything else as ideas for future tests, not conclusions.

Mistake 4: ignoring the novelty effect

Users click the shiny new thing because it's new. Run tests for at least one or two full weeks so you cover weekday and weekend behaviour and let the novelty wear off.

Mistake 5: broken randomisation

If version B only shows on new browsers, or the assignment changes when users log in, your groups differ in ways that have nothing to do with the button. Run an A/A test (same version in both groups) occasionally. If it shows a "significant" difference often, your setup is broken.

How to report results honestly

Instead of "green wins!", report:

Conversion: 3.10% (blue) vs 3.18% (green). Difference: +0.08 percentage points, 95% confidence interval −0.03 to +0.19. Not statistically significant at our planned sample size. Guardrails unchanged.

It's less exciting. It's also true, and it protects your credibility for the day when you do find a real winner.

The boss conversation

Bosses don't actually want false winners. They want good decisions. A test that honestly says "no detectable difference" saves the company from building on a fake improvement. Say that clearly, and you'll be trusted much longer than the analyst who finds a winner every week.

// related

// prefer the terminal?

Open the terminal blog and type read ab-testing-without-lying-to-your-boss.