Module 4 · Experimentation and A/B Testing
Analysing conversion tests
Test the difference between two conversion rates with a two-proportion z-test, report it with a confidence interval and relative lift, and say what it means in users and money.
About 25 minutes
The problem
The onboarding test has finished: 12,000 users, randomisation checked. Now the head of product wants one sentence: "Did the new flow work, and by how much?"
"B converted 42.6% and A 38.2%, so B wins" isn't enough. With random samples, some difference is always there by chance. The analysis has to say whether this difference is bigger than chance can explain, how big the true effect plausibly is, and what that means for the business.
The concept
The two-proportion z-test
For conversion rates p_A and p_B from n_A and n_B users:
- Pooled rate p = (conversions A + conversions B) ÷ (n_A + n_B).
- Standard error SE = √[p(1 − p)(1/n_A + 1/n_B)].
- z = (p_B − p_A) ÷ SE, and the two-sided p-value = 2 × P(Z > |z|).
A p-value below 0.05 means a difference this large would be rare if the flows were really the same.
Report the effect, not just the p-value
- Absolute difference: p_B − p_A, in percentage points.
- Confidence interval for it: difference ± 1.96 × √[p_A(1 − p_A)/n_A + p_B(1 − p_B)/n_B]. It gives the range of effects consistent with the data.
- Relative lift: (p_B − p_A) ÷ p_A.
Translate it
"+4.5 points" means little to a director. "About 45 more verified customers per 1,000 signups, around 5,800 more a year at current signup rates" means a lot.
Example
import pandas as pd
import numpy as np
from scipy import stats
onboarding = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/experiments/onboarding.csv")
summary = onboarding.groupby("variant")["completed_kyc_7d"].agg(conversions="sum", users="count")
summary["rate"] = summary["conversions"] / summary["users"]
summary.round(4)conversions users rate
variant
A 2270 5946 0.3818
B 2581 6054 0.4263pA, pB = summary.loc["A", "rate"], summary.loc["B", "rate"]
nA, nB = summary.loc["A", "users"], summary.loc["B", "users"]
pooled = summary["conversions"].sum() / summary["users"].sum()
z = (pB - pA) / np.sqrt(pooled * (1 - pooled) * (1 / nA + 1 / nB))
p_value = 2 * stats.norm.sf(abs(z))
diff = pB - pA
se = np.sqrt(pA * (1 - pA) / nA + pB * (1 - pB) / nB)
print(f"Difference: {diff:.2%} 95% CI: {diff - 1.96 * se:.2%} to {diff + 1.96 * se:.2%}")
print(f"Relative lift: {diff / pA:.1%} z = {z:.2f} p-value = {p_value:.2g}")Difference: 4.46% 95% CI: 2.70% to 6.21%
Relative lift: 11.7% z = 4.97 p-value = 6.6e-07The new flow raised 7-day KYC completion by about 4.5 percentage points, a relative lift of about 12%, and the confidence interval is well clear of zero. With about 430 signups a day, that's roughly 19 more verified customers every day, or around 7,000 a year.
Walkthrough
- Run the cells and check the z-test by hand for the first step (the pooled rate).
- Compare your result with
stats.chi2_contingencyon the 2×2 table of variant against KYC. The p-values should match closely. - Work out how many extra verified customers the flow would produce in a year.
- Write the result for the head of product (the task below).
Practice
Practice
What is the absolute difference in 7-day KYC completion (B − A), in percentage points? One decimal place.
Practice
What is the lower end of the 95% confidence interval for the difference, in percentage points? One decimal place.
Task
6 minWrite the result for the head of product in 40 to 110 words: the effect with its confidence interval, whether it's statistically significant, and what it means in customers per day or year.
Your work is checked for
- Gives the effect in points or percent
- Gives a confidence interval (CI, range, between … and)
- Says whether it's significant
- Translates into customers
- Between 40 and 110 words
Check your understanding
Answer every question to check.