Module 8 · Feature Engineering and Model Evaluation
"Final project: Paystream's retention model"
Plan your final project, an end-to-end churn model with point-in-time features, time-based validation, calibration, targeting and monitoring, and start with a feature of your own.
About 20 minutes
The problem
Paystream's head of growth wants a churn model the retention team can use every month from June 2026, and a plan for keeping it honest as the competitor keeps pushing. You'll build it end to end: features as of each month-end, a time-based test, a fair model comparison, calibrated probabilities, a calling plan with its value, and monitoring.
The best projects also add something the course didn't: a new feature that captures behaviour the others miss. This lesson starts you on that.
The concept
The project, step by step
| Step | Deliverable | Lesson |
|---|---|---|
| Define | population, snapshot, horizon and target, written down | 1 |
| Engineer | the course's features plus at least two of your own, each checked against churn | 2 |
| Validate | training snapshots with complete labels, a later test snapshot, the leakage checks | 3 |
| Compare | logistic regression and gradient boosting, with reasons for the choice | 4 |
| Calibrate | a reliability table and the Brier score | 5 |
| Target | gains, lift and a calling plan with its value | 6 |
| Monitor | drift checks by segment, a retraining rule and triggers | 7 |
Ideas for new features
- Variety: the number of different transaction types in the last 90 days. Customers who use several services may be stickier.
- Large-value share: the share of value from transfers, where competitors usually compete.
- Weekday pattern: whether activity is concentrated on salary week.
- Failure streaks: the longest run of consecutive failed transactions.
Example
A first new feature, variety, built point-in-time like the others:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/wallet/"
customers = pd.read_csv(base + "customers.csv", parse_dates=["signup_date"])
tx = pd.read_csv(base + "transactions.csv", parse_dates=["transaction_date"])
s = pd.Timestamp("2026-03-31")
recent = tx[(tx["transaction_date"] <= s) & (tx["transaction_date"] > s - pd.Timedelta(days=90))]
future = tx[(tx["transaction_date"] > s) & (tx["transaction_date"] <= s + pd.Timedelta(days=60))]
march = pd.DataFrame({"variety": recent.groupby("customer_id")["type"].nunique()})
march["churned"] = (~march.index.isin(future["customer_id"])).astype(int)
march.groupby("variety")["churned"].agg(["size", "mean"]).round(3)size mean
variety
1 59 0.373
2 136 0.221
3 227 0.145
4 339 0.086
5 446 0.034Check the pattern before trusting it: does churn fall steadily as variety rises, or is the difference driven by a small group?
Walkthrough
- Write your definitions at the top of the notebook.
- Add the variety feature to
build_table, and design one more of your own. - Check each new feature's relationship with churn, and confirm neither uses data after the snapshot.
- Open the project brief on the course page and plan the remaining steps.
Practice
Practice
At the 31 March 2026 snapshot, what is the churn rate of active customers who used only one transaction type in the previous 90 days? As a percentage, one decimal place.
Task
6 minDescribe two new features you'll add to the churn model, one per line in the form Name | how it's calculated (with its time window) | why it might predict churn. Neither may use data after the snapshot.
Your work is checked for
- Two lines in the form Name | calculation | reason
- Each calculation names a time window
- No future data (after the snapshot, next, following)
Check your understanding
Answer every question to check.