Module 7 · Experimentation and A/B Testing
Guardrails and decisions
Weigh a winning primary metric against guardrails that get worse, value both sides in money over a sensible horizon, and make a decision you can defend.
About 25 minutes
The problem
Paystream tested raising its transfer fee from ₦10 to ₦25. The finance director is delighted: fee revenue per user more than doubled in four weeks. The product team is worried: users made fewer transfers, and more of them had stopped using the app by day 28.
Both are right, and that's the point of guardrails. An experiment that only looks at its primary metric can recommend changes that win this month and lose the customer base. The decision needs both sides on the same scale, usually money, over a long enough time to see the cost.
The concept
Primary metric and guardrails, together
Report every pre-agreed metric with its effect and confidence interval. A change that improves the primary metric but significantly harms a guardrail isn't a win; it's a trade-off to decide.
Put both sides in naira
- The gain: extra fee revenue per user per month.
- The cost: users lost, each worth their future revenue (fees, interest on balances, other products) over a horizon such as a year.
Short tests measure the gain fully but the cost only partly: churn keeps accumulating long after four weeks. A decision based only on the test window favours changes that harvest revenue and pay for it later.
Decision options
Not just "ship" or "don't": a smaller increase (₦15 or ₦20), the higher fee only for large transfers, or a further test that measures retention over longer.
Example
import pandas as pd
import numpy as np
from scipy import stats
fee = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/experiments/fee_test.csv")
summary = fee.groupby("variant")[["transfers_28d", "fee_revenue_28d_ngn", "active_on_day_28"]].mean()
summary.round(3)transfers_28d fee_revenue_28d_ngn active_on_day_28
variant
Control 8.810 88.099 0.937
Higher fee 7.586 189.654 0.903Each difference, with its test:
c = fee[fee["variant"] == "Control"]
t = fee[fee["variant"] == "Higher fee"]
for metric in ["fee_revenue_28d_ngn", "transfers_28d", "active_on_day_28"]:
test = stats.ttest_ind(t[metric], c[metric], equal_var=False)
change = t[metric].mean() - c[metric].mean()
print(f"{metric}: change {change:+.3f} ({change / c[metric].mean():+.1%}) p = {test.pvalue:.2g}")fee_revenue_28d_ngn: change +101.555 (+115.3%) p = 6.5e-178
transfers_28d: change -1.224 (-13.9%) p = 2.2e-10
active_on_day_28: change -0.034 (-3.6%) p = 2.5e-08Revenue per user more than doubles, but transfers fall and more users are inactive by day 28: all three changes are far too large to be chance. Now value the trade-off over a year for 100,000 users, assuming a retained user is worth ₦3,000 a year across all products. The test can't say whether the churn gap was a one-off or will repeat every month, so try both:
users = 100_000 # active users who'd get the higher fee
extra_fee_per_user_year = (t["fee_revenue_28d_ngn"].mean() - c["fee_revenue_28d_ngn"].mean()) * 13
extra_churn = c["active_on_day_28"].mean() - t["active_on_day_28"].mean()
value_per_user_year = 3000
scenarios = {
"Gap stays at the 4-week level": extra_churn,
"Same gap again every month": 1 - (1 - extra_churn) ** 12,
}
for name, lost_share in scenarios.items():
gain = users * extra_fee_per_user_year * (1 - lost_share)
loss = users * lost_share * value_per_user_year
print(f"{name}: extra fees ₦{gain / 1e6:,.1f}m, users lost worth ₦{loss / 1e6:,.1f}m, net ₦{(gain - loss) / 1e6:,.1f}m")Gap stays at the 4-week level: extra fees ₦127.6m, users lost worth ₦10.1m, net ₦117.4m
Same gap again every month: extra fees ₦87.4m, users lost worth ₦101.3m, net ₦-13.8mIf the extra churn is a one-off, the fee increase pays handsomely. If the same gap opens up again every month, a third of the affected users are gone within a year and the higher fee loses money. A four-week test can't tell these apart, so the responsible decision is not "ship ₦25": it's to test smaller increases and measure retention for longer.
Walkthrough
- Run the cells. Add confidence intervals for each difference.
- Change the value of a retained user to ₦6,000. Does the decision flip?
- Find the monthly churn gap at which the two sides break even.
- Write the recommendation (the task below).
Practice
Practice
What is the average fee revenue per user over 28 days in the Higher fee group? (A rounded figure is fine.)
Practice
By how many percentage points is the share of users active on day 28 lower in the Higher fee group? One decimal place.
Task
8 minWrite the recommendation to the finance director and head of product (60 to 150 words): what the test showed on the primary metric and the guardrails, the value of each side, the key uncertainty, and your recommended next step.
Your work is checked for
- Reports the revenue gain
- Reports a guardrail (transfers or users active)
- Uses naira values
- Names the key uncertainty (longer term, beyond four weeks, value of a user)
- Recommends a next step (test, smaller, ₦15, ₦20, large transfers)
- Between 60 and 150 words
Check your understanding
Answer every question to check.