Module 1 · Experimentation and A/B Testing
Why experiment
Why comparing users who did something with users who didn't can't tell you whether the thing worked, and how random assignment fixes that.
About 20 minutes
The problem
Paystream's product team redesigned the signup flow. A month after launch, someone runs the numbers: new users who signed up through referrals complete identity verification (KYC) far more often than those from social ads. Should the company spend more on referrals? Probably, but the data can't say how much a referral causes: people who arrive through a friend may simply be more committed in the first place.
The same trap catches almost every "before and after" or "users who did X versus users who didn't" analysis. Those who chose X, or who arrived after the change, differ in ways you can't see. The only reliable way to measure the effect of a change is to decide at random who gets it. That's an experiment, or A/B test, and it's how product, marketing and pricing decisions are made at companies that take evidence seriously.
The concept
Correlation isn't causation, for a specific reason
When people choose a treatment (a feature, a channel, a plan), the choosers differ from the non-choosers: more engaged, richer, more urban. Those differences, not the treatment, may explain the outcome. They're called confounders.
Randomisation breaks the link
If a coin decides who sees the new signup flow (B) and who sees the old one (A), then on average the two groups are the same in every way, seen and unseen. Any difference in outcomes beyond chance is caused by the flow. That's why A/B tests are the gold standard.
The pieces of an A/B test
| Piece | Paystream's onboarding test |
|---|---|
| Unit randomised | each new user, at signup |
| Variants | A: old flow (control); B: new flow (treatment) |
| Primary metric | completed KYC within 7 days |
| Secondary metrics | transactions and value in the first 14 days |
| Guardrails | metrics that mustn't get worse (support tickets, failed payments) |
Example
The onboarding experiment: 12,000 new users, randomly assigned at signup.
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/experiments/"
onboarding = pd.read_csv(base + "onboarding.csv")
print(onboarding["variant"].value_counts())
onboarding.groupby("acquisition_channel")["completed_kyc_7d"].mean().round(3)variant
B 6054
A 5946
Name: count, dtype: int64
acquisition_channel
Agent 0.406
Organic 0.397
Referral 0.471
Social ads 0.337
Name: completed_kyc_7d, dtype: float64Referral users verify most and social-ads users least, but users choose how they arrive, so that comparison is confounded. The A/B comparison isn't, because the coin decided the variant:
onboarding.groupby("variant")[["completed_kyc_7d", "txns_first_14d"]].mean().round(3)completed_kyc_7d txns_first_14d
variant
A 0.382 1.615
B 0.426 1.821The new flow's users verify more often and transact more. Because assignment was random, this difference is the effect of the flow plus chance, and the next lessons show how to separate the two. Check that randomisation worked by comparing the groups on things the flow can't affect:
pd.crosstab(onboarding["acquisition_channel"], onboarding["variant"], normalize="columns").round(3)variant A B
acquisition_channel
Agent 0.302 0.299
Organic 0.206 0.202
Referral 0.249 0.256
Social ads 0.242 0.243The channel mix is nearly identical in both variants, exactly what random assignment should produce.
Walkthrough
- Load the data and look at the columns:
onboarding.head(). - Compare KYC completion by platform. Is that comparison causal? Why not?
- Compare the variants on platform and region mix, as above. Are they balanced?
- Think of a recent decision at a company you know that was based on a "users who did X" comparison. What confounders might explain the result?
Practice
Practice
What share of onboarding users completed KYC within 7 days in variant A (the old flow)? As a percentage, one decimal place.
Task
6 minA bank finds that customers who use its budgeting tool save twice as much as those who don't, and wants to say the tool doubles savings. Write a short reply (40 to 120 words) explaining why that conclusion isn't safe, naming two possible confounders, and describing the experiment that would answer the question.
Your work is checked for
- Explains self-selection (choose, chose, already, different)
- Names at least two confounders (income, motivation, age, engagement, education, savers)
- Describes random assignment
- Between 40 and 120 words
Check your understanding
Answer every question to check.