Back to Blog
Skills

A/B Testing for PMs: A Practical Guide

You don't need a stats degree to run good experiments. You do need to avoid the handful of mistakes that make most A/B test results meaningless.

PM Job BoardJuly 27, 20267 min read
Share:

A/B testing has a credibility problem, and it's not the math's fault.

The math is fine. The problem is how tests get run: peeked at daily, stopped the moment they look good, launched without a hypothesis, and interpreted by whoever wants the feature to win. Then teams conclude "experimentation doesn't work here" and go back to shipping on vibes.

You don't need to be a statistician to run trustworthy experiments. You need discipline about a small number of things. Here they are.

What A/B Testing Is Actually For

An A/B test answers one question: did this change cause that outcome? Not "is this correlated," not "do users say they like it." Caused.

That's a genuinely rare thing to know in product work. Most of our evidence is correlational or anecdotal, which is fine for forming hypotheses but dangerous for confirming them. Experiments are how you check whether your beautiful theory survives contact with random assignment.

Which also tells you what A/B testing is not for:

  • Deciding strategy. No experiment will tell you whether to enter a new market.
  • Detecting tiny effects with small traffic. More on this below.
  • Replacing judgment. Tests evaluate options. Someone still has to generate good options, which is what product discovery is about.

Before You Test: The Hypothesis

Every real experiment starts with a falsifiable sentence:

"We believe [change] will cause [outcome] for [audience] because [reasoning]."

For example: "We believe reducing signup fields from 9 to 4 will increase signup completion by at least 10% for organic visitors, because session recordings show heavy drop-off on the address fields."

Notice what this forces:

  • A primary metric, chosen before you see any data (signup completion)
  • An expected effect size (at least 10%), which determines whether the test is even feasible
  • Reasoning, so that whatever happens, you learn something about your model of the user

A test without a written hypothesis isn't an experiment. It's a slot machine with dashboards.

The Statistics You Actually Need

You can skip the derivations. You cannot skip these four concepts:

Sample size comes first

Before launching, use any sample size calculator to answer: given my baseline rate and the smallest effect I care about, how many users do I need? This is the step that kills most bad tests early, and that's a feature. If the calculator says you need 400,000 users per arm and you get 5,000 signups a month, you cannot run this test. Ship it based on judgment instead, and be honest that you did.

Don't peek and stop

Checking results daily is fine. Stopping the test the first day it crosses significance is not. Random fluctuation crosses the significance line all the time on the way to nowhere. Decide the duration (based on your sample size calculation), run the full duration, then read the result. If you want the freedom to stop early, you need sequential testing methods; most experimentation platforms offer them. Use the platform's version, don't improvise.

Significance isn't importance

"Statistically significant" means "probably not noise." It doesn't mean "big" or "worth it." With enough traffic, a 0.1% lift is significant and still might not justify the complexity you added. Look at the confidence interval, not just the verdict.

One primary metric

If you check twenty metrics, one of them will look significant by pure chance. Declare a single primary metric in your hypothesis. Everything else is secondary: useful for context and generating new hypotheses, not for declaring victory.

Guardrails: The Metrics That Protect You

Every test should carry guardrail metrics: things that must not get worse. Testing an aggressive upsell prompt? Guardrails are churn, support tickets, and task completion. Testing a signup simplification? Guardrail is downstream activation, because you might be signing up people who never intended to use the product.

A "winning" test that quietly damages a guardrail is a losing test with good PR. Our product metrics guide covers picking metrics that reflect real value; the same logic applies inside experiments.

Running the Test: Practical Details

  • Run full weeks. Weekday and weekend users behave differently. A Tuesday-to-Friday test oversamples one population. Two full weeks is a sensible default minimum.
  • Watch for novelty effects. Anything visually new gets clicked more for a while. If your lift decays across the test period, that's novelty, not value.
  • Check the split is actually random. A sample ratio mismatch (you expected 50/50, you got 53/47) invalidates the test. Good platforms flag this. Take the flag seriously.
  • Don't ship mid-test changes. Changing the variant halfway through means you tested two things and learned about neither.

Interpreting Results Like an Adult

Three outcomes, three responses:

It won. Check guardrails, check segments for anyone harmed, then ship. Write down the result where future PMs can find it, including effect size. Institutional memory of past experiments is absurdly valuable and almost nobody maintains it.

It lost. This is information, not failure. Your model of the user was wrong somewhere. The teams that get compounding value from experimentation are the ones that treat losses as updates: what did we believe that turned out false? Roughly half of well-run tests at mature companies fail to beat control. If everything you test wins, your tests are broken or your bar is too low.

It's flat. The most common and most awkward outcome. Either the change genuinely doesn't matter (ship the simpler version, usually control) or your test was underpowered to detect the true effect. This is why the sample size step matters: it's the difference between "no effect" and "we couldn't see."

The Politics of Experimentation

Here's the part no stats course covers. Experiments produce results people don't want. The feature the VP sponsored loses. The redesign everyone loved is flat. What happens next determines whether your experimentation culture is real.

Your job as PM:

  • Pre-commit publicly. Before the test, state the hypothesis, the primary metric, and what result would mean ship vs. kill. It's much harder to reinterpret a result you predicted in writing.
  • Never rerun until you get the answer you want. Rerunning a lost test with no changes until it wins is p-hacking with extra steps.
  • Protect the messenger. If a data scientist tells you your feature lost, thank them, visibly. How you react the first time determines what you get told the second time. More on that relationship in working with data scientists.

When Not to Test

Skip the experiment when:

  • Traffic can't support it (see sample size, above)
  • The change is strategically required regardless (compliance, rebrand, platform migration)
  • The cost of testing exceeds the decision's value. A two-week test to pick a button label nobody will notice is process cosplay.
  • You'd ignore a negative result anyway. Then don't pretend. Ship it and monitor.

Judgment doesn't disappear in an experimentation culture. It moves up a level: deciding what's worth testing at all.

The Bottom Line

Write the hypothesis. Size the test before launching. Run it the full duration. One primary metric, guardrails always, honest readouts regardless of politics. That's 90% of good experimentation, and none of it requires a statistics degree.

Experimentation skills are increasingly a line item in PM job descriptions, especially for growth roles. If you can describe a real test you ran, including one that lost, you'll stand out.

Find teams that take experimentation seriously at productmanagerjobboard.com.

Share:
P

PM Job Board

Helping product managers find their next great opportunity. Follow us for career tips, interview advice, and industry insights.

More in Skills

Ready to Find Your Next PM Role?

Browse hundreds of Product Manager jobs at top companies, from startups to FAANG.