Module 6 · Experimentation and A/B Testing
Segments, multiple tests and novelty
Look inside a result by segment without fooling yourself, correct for running many tests at once, and spot novelty effects that fade once users get used to a change.
About 25 minutes
The problem
The onboarding result is in, and everyone wants to slice it. Marketing asks for it by acquisition channel, the regional managers by state, the mobile team by platform. Ten regions later, four show a "significant" effect and six don't. The Kano manager wants credit; the Delta manager wants to know what went wrong.
Meanwhile, the banner test (setting aside its broken split) showed a click rate far above control in its first week, which the marketing team is already quoting. By week three, the gap had gone.
Both are classic ways to over-read an experiment. Segment results are noisy, and every extra comparison is another chance of a false alarm. And people often react to something new in ways that don't last.
The concept
Segments: look, but carefully
- Each segment has fewer users, so its estimate is much noisier than the overall one.
- "Significant here, not there" doesn't mean the effect differs between segments. Compare the segments' effects directly, or look at whether their confidence intervals overlap.
- Trust a segment difference when it was planned in advance, is large, and has a mechanism (a reason it should differ).
Multiple comparisons
Test 10 segments at α = 5% and, even with no real differences, you'd expect about one false alarm by chance. The simplest correction is Bonferroni: with m tests, use α ÷ m (here 0.05 ÷ 10 = 0.005) for each.
Novelty effects
Existing users often click on something just because it's new. The effect fades as they get used to it. Plot the effect by week: if it shrinks towards zero, judge the change on the later weeks, or run the test longer. (Its mirror image, a primacy effect, is when users first resist a change and then adapt.)
Example
The onboarding effect by platform, where there's a reason to expect a difference (the new flow fixed an Android camera step):
import pandas as pd
import numpy as np
from scipy import stats
base = "https://academy.cloudtechanalytics.com/datasets/experiments/"
onboarding = pd.read_csv(base + "onboarding.csv")
def effect(df):
a = df.loc[df["variant"] == "A", "completed_kyc_7d"]
b = df.loc[df["variant"] == "B", "completed_kyc_7d"]
pooled = (a.sum() + b.sum()) / (len(a) + len(b))
z = (b.mean() - a.mean()) / np.sqrt(pooled * (1 - pooled) * (1 / len(a) + 1 / len(b)))
return pd.Series({"users": len(df), "effect_pts": round((b.mean() - a.mean()) * 100, 1), "p_value": round(2 * stats.norm.sf(abs(z)), 4)})
onboarding.groupby("platform").apply(effect, include_groups=False)users effect_pts p_value
platform
Android 9288.0 5.2 0.0000
iOS 2712.0 1.7 0.3745The effect is much larger on Android, as the mechanism predicts. Now the regions, with no particular reason to expect differences:
regions = onboarding.groupby("region").apply(effect, include_groups=False).sort_values("p_value")
regions["significant_at_5pct"] = regions["p_value"] < 0.05
regions["significant_bonferroni"] = regions["p_value"] < 0.05 / len(regions)
regionsusers effect_pts p_value significant_at_5pct significant_bonferroni
region
Kano 1249.0 9.7 0.0005 True True
Kaduna 853.0 9.6 0.0038 True True
Enugu 862.0 7.2 0.0329 True False
Anambra 826.0 7.3 0.0347 True False
Lagos 3120.0 2.7 0.1240 False False
Oyo 1018.0 3.2 0.2851 False False
Rivers 1098.0 3.2 0.2872 False False
Delta 977.0 3.1 0.3251 False False
Ogun 793.0 2.5 0.4748 False False
FCT 1204.0 0.7 0.7907 False FalseThe new flow works everywhere on average; the regional estimates scatter around the overall effect because each region is small. After the Bonferroni correction fewer regions pass, and nothing here suggests the flow works differently in Kano from Delta. Now the banner's novelty effect, by week:
banner = pd.read_csv(base + "banner_daily.csv")
banner["week"] = (pd.to_datetime(banner["date"]) - pd.Timestamp("2026-06-01")).dt.days // 7 + 1
weekly = banner.groupby(["week", "variant"])[["users", "clicks"]].sum()
(weekly["clicks"] / weekly["users"]).unstack().round(4)variant A B
week
1 0.0308 0.0466
2 0.0304 0.0374
3 0.0316 0.0312
4 0.0305 0.0335B's first-week click rate is far above A's, but by weeks three and four the two are close. The first week was novelty. (And remember lesson 3: this test's split is broken anyway.)
Walkthrough
- Run the cells. Compare the Android and iOS confidence intervals for the effect.
- Check the effect by acquisition channel. Is there a mechanism that would make it differ?
- Count how many regions you'd expect to be "significant" by chance alone if the flow did nothing.
- Write what you'd tell the Kano and Delta managers (the task below).
Practice
Practice
What is the effect of the new flow on Android users, in percentage points? One decimal place.
Practice
How many regions are significant after the Bonferroni correction?
Task
6 minThe Delta regional manager asks why the new flow "didn't work" in Delta. Write a reply (40 to 110 words) explaining noise in small segments, multiple comparisons, and what the evidence actually says about Delta.
Your work is checked for
- Explains that segments are small or noisy
- Mentions multiple comparisons (many regions, by chance, false alarm)
- Says the overall effect is the best estimate for Delta
- Between 40 and 110 words
Check your understanding
Answer every question to check.