Module 8 · Experimentation and A/B Testing
When you can't randomise
Estimate an effect when a change was rolled out to some places and not others, with difference-in-differences, and check the parallel-trends assumption it depends on.
About 25 minutes
The problem
In week 14 of 2026, Paystream launched cash-out agents (shops where customers can withdraw cash from their wallet) in three northern and south-eastern states: Kano, Kaduna and Enugu. Nobody randomised anything: the operations team chose states where it already had partners. Three months later, the question is whether agents increased weekly active users.
Weekly active users in the three agent states rose by about 6%. But a before-and-after comparison can't separate the agents' effect from everything else that changed in those weeks: seasons, salaries, a competitor, a fuel price rise. You need a comparison group, and a method that uses it.
The concept
Difference-in-differences (DiD)
Compare the change in the treated group with the change in a comparison group over the same period:
effect ≈ (treated after − treated before) − (comparison after − comparison before)
or, in percentages, the treated group's growth minus the comparison group's growth. The comparison group's change stands in for what would have happened to the treated states without the agents.
The parallel-trends assumption
DiD only works if, without the treatment, both groups would have moved in parallel. You can't check that directly, but you can check that they did move in parallel before the change. If their pre-period trends differ, the estimate is suspect.
Other options when you can't randomise
- A staggered rollout: launch in different places at different times, which gives several before-and-after comparisons.
- Holdouts: keep a random set of places or users without the change for a while.
- If possible, randomise next time. Even randomising which states go first would have made this question easy.
Example
import pandas as pd
import numpy as np
rollout = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/experiments/rollout.csv")
rollout["group"] = np.where(rollout["state"].isin(["Kano", "Kaduna", "Enugu"]), "Agent states", "Other states")
rollout["period"] = np.where(rollout["week"] >= 14, "after", "before")
weekly = rollout.groupby(["group", "period", "week"])["weekly_active_users"].sum().groupby(["group", "period"]).mean()
table = weekly.unstack()[["before", "after"]]
table["change_pct"] = (table["after"] / table["before"] - 1) * 100
table.round(1)period before after change_pct
group
Agent states 5242.2 5545.5 5.8
Other states 15861.1 15538.2 -2.0The agent states grew while the other states shrank slightly over the same weeks. The difference-in-differences:
did = (table.loc["Agent states", "change_pct"] - table.loc["Other states", "change_pct"]) / 100
print(f"Before-and-after in agent states: {table.loc['Agent states', 'change_pct'] / 100:+.1%}")
print(f"Difference-in-differences estimate: {did:+.1%}")Before-and-after in agent states: +5.8%
Difference-in-differences estimate: +7.8%The before-and-after figure understates the effect, because the other states show that this period was a slightly weaker one. Now the parallel-trends check: the two groups' weekly totals, each indexed to their own average in the weeks before the launch:
pre = rollout[rollout["week"] < 14]
idx = rollout.groupby(["group", "week"])["weekly_active_users"].sum().unstack(0)
idx = idx / idx.loc[1:13].mean()
idx.loc[[1, 5, 9, 13, 14, 18, 22, 26]].round(3)group Agent states Other states
week
1 0.949 0.949
5 1.001 0.994
9 0.996 1.007
13 1.014 1.016
14 1.053 0.978
18 1.033 0.949
22 1.064 0.979
26 1.128 1.044Before week 14, the two groups move closely together; after it, the agent states pull ahead. That's the pattern that makes the DiD estimate believable.
Walkthrough
- Run the cells and plot the indexed weekly totals for both groups, with a line at week 14.
- Calculate the DiD for each agent state separately. Are they similar?
- Run a "placebo" DiD: pretend the launch was in week 7, using only weeks 1 to 13. The estimate should be close to zero.
- Write the result for the operations director (the task below).
Practice
Practice
What is the difference-in-differences estimate of the agents' effect on weekly active users? As a percentage, one decimal place.
Task
8 minWrite the result for the operations director (50 to 130 words): the estimate, why the before-and-after figure alone was misleading, the assumption behind it and the evidence for it, and how to make the next rollout easier to measure.
Your work is checked for
- Gives the estimate as a percentage
- Explains the before-and-after problem (other states, same period, would have happened)
- Names the parallel-trends assumption or its check
- Suggests randomising, staggering or a holdout next time
- Between 50 and 130 words
Check your understanding
Answer every question to check.