Module 9 · Experimentation and A/B Testing
"Final project: Paystream's experiment review"
Plan your final project, a full review of Paystream's four experiments and a design for the next one, and start by checking every test before trusting any result.
About 20 minutes
The problem
Paystream's leadership wants a review of the quarter's experiments before it makes three decisions: whether to roll out the new signup flow, whether to raise the transfer fee, and whether to expand cash-out agents to more states. It also wants to know whether the banner result the marketing team keeps quoting can be trusted.
Your final project is that review, plus a properly designed next experiment. The first step in any review is the one people skip: check that each test is sound before reading its results.
The concept
A review checklist for every experiment
| Check | Question |
|---|---|
| Design | Was there one primary metric, an MDE and a planned sample size? |
| Randomisation | Does the split match the plan (sample ratio mismatch)? Are the groups balanced? |
| Duration | Whole weeks? Long enough for the guardrails? Any novelty? |
| Analysis | Effect with a confidence interval, not just a p-value? Skewed metrics handled? |
| Segments | Planned, with a mechanism, and corrected for multiple comparisons? |
| Decision | Guardrails valued? Next step clear? |
What each test can support
- Onboarding: a clean randomised test with a clear effect. Ready to decide.
- Banner: broken split and a novelty effect. Not trustworthy; rerun.
- Fee: clean, but a trade-off with an uncertain long-run cost. Needs a follow-up test.
- Agents: not randomised; difference-in-differences with a plausible parallel-trends check. Reasonable evidence, to be confirmed by the next rollout.
Example
The quick health check across the user-level tests:
import pandas as pd
from scipy import stats
base = "https://academy.cloudtechanalytics.com/datasets/experiments/"
onboarding = pd.read_csv(base + "onboarding.csv")
fee = pd.read_csv(base + "fee_test.csv")
banner = pd.read_csv(base + "banner_daily.csv")
checks = {
"onboarding": onboarding["variant"].value_counts(),
"fee": fee["variant"].value_counts(),
"banner": banner.groupby("variant")["users"].sum(),
}
for name, counts in checks.items():
print(f"{name:11s} split {counts.to_dict()} SRM p-value {stats.chisquare(counts).pvalue:.3g}")onboarding split {'B': 6054, 'A': 5946} SRM p-value 0.324
fee split {'Control': 4025, 'Higher fee': 3975} SRM p-value 0.576
banner split {'A': 110094, 'B': 104210} SRM p-value 5.18e-37Only the banner test fails. Everything else in your review can build on the other three.
Walkthrough
- Run the health check, then work through the checklist for each experiment.
- Write a one-paragraph verdict per experiment: trustworthy or not, what it shows, and the decision.
- Design the follow-up fee test: variants, primary metric, guardrails, MDE, sample size and duration.
- Open the project brief on the course page and plan the write-up.
Practice
Practice
What is the SRM p-value for the fee test? Two decimal places.
Task
8 minWrite a verdict for each of the four experiments, one line each in the form Experiment | trustworthy? | what it shows | decision.
Your work is checked for
- Four lines in the form Experiment | trustworthy | shows | decision
- The banner test is marked untrustworthy
- Mentions the sample ratio mismatch or novelty for the banner
- The fee decision is a follow-up test, not a straight rollout
Check your understanding
Answer every question to check.