Module 10 · Machine Learning Fundamentals
From model to decision
Package a model as a pipeline that takes raw applications, score new cases with reasons, save it for others to use, and document it with a model card.
About 30 minutes
The problem
Your notebook contains a good default model, but nobody at Ladder can use a notebook. The credit officers need to enter an application (business type, revenue, loan amount, and so on) and get back a probability, a recommendation and the main reasons, in a form they trust. The data team needs to run the same model next month without redoing your cleaning steps by hand.
Getting from "a model in a notebook" to "a model the business uses" is where many data science projects stall. The two tools that bridge the gap are a pipeline that does every step from raw data to prediction, and a model card that tells everyone what the model is for and where it shouldn't be trusted.
The concept
A full pipeline
Until now, you've one-hot encoded with pd.get_dummies before training. That's fragile: a new application has one row, so get_dummies can't create the same columns. A scikit-learn ColumnTransformer does the preparation inside the model:
OneHotEncoder(handle_unknown="ignore")for text columns, remembering the categories it saw in training;StandardScaler()for numbers;- then the model.
Fit the whole pipeline on the training data, and it takes raw rows, exactly like the CSV, from then on.
Scoring and reasons
For each new application, return the probability, the decision at the agreed threshold, and the top reasons. Reasons matter as much as the score: they're what a credit officer can act on and explain.
Saving and loading
joblib.dump(pipeline, "default_model.joblib") saves the fitted pipeline to a file; joblib.load brings it back, ready to score. Record the date, the data used and the scikit-learn version alongside it.
A model card
A one-page description of the model: purpose, intended users, data, features (and those deliberately excluded), performance, threshold and its business basis, known limitations, fairness checks, monitoring plan, owner and review date.
Example
The production pipeline, trained on raw columns without region:
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
loans = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/loans/loans.csv")
loans["amount_to_revenue"] = loans["loan_amount_ngn"] / loans["monthly_revenue_ngn"]
categorical = ["business_type", "group_loan", "has_guarantor"]
numeric = ["borrower_age", "years_in_business", "monthly_revenue_ngn", "loan_amount_ngn",
"amount_to_revenue", "term_months", "interest_rate_monthly_pct", "previous_loans",
"previous_late_payments", "mobile_money_txns_per_month"]
X = loans[categorical + numeric]
y = loans["defaulted"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
pipeline = Pipeline([
("prepare", ColumnTransformer([
("categories", OneHotEncoder(handle_unknown="ignore", drop="first"), categorical),
("numbers", StandardScaler(), numeric),
])),
("model", LogisticRegression(max_iter=1000)),
])
pipeline.fit(X_train, y_train)
print("Test AUC:", round(roc_auc_score(y_test, pipeline.predict_proba(X_test)[:, 1]), 3))Test AUC: 0.765Now score a new application, exactly as a credit officer would enter it:
application = pd.DataFrame([{
"business_type": "Food vendor", "group_loan": "No", "has_guarantor": "No",
"borrower_age": 31, "years_in_business": 1, "monthly_revenue_ngn": 250000,
"loan_amount_ngn": 600000, "term_months": 6, "interest_rate_monthly_pct": 4.5,
"previous_loans": 1, "previous_late_payments": 2, "mobile_money_txns_per_month": 8,
}])
application["amount_to_revenue"] = application["loan_amount_ngn"] / application["monthly_revenue_ngn"]
p = pipeline.predict_proba(application[categorical + numeric])[0, 1]
print(f"Probability of default: {p:.0%}")
print("Recommendation:", "refer for review" if p >= 0.25 else "approve")Probability of default: 59%
Recommendation: refer for reviewAnd the reasons, from each feature's contribution to the score:
prepare, model = pipeline.named_steps["prepare"], pipeline.named_steps["model"]
contrib = pd.Series(prepare.transform(application[categorical + numeric])[0] * model.coef_[0],
index=prepare.get_feature_names_out())
contrib.sort_values(ascending=False).head(3).round(2)numbers__amount_to_revenue 0.85
numbers__previous_late_payments 0.52
numbers__years_in_business 0.38
dtype: float64The biggest pushes towards default are the loan's size relative to revenue (2.4 times monthly revenue), two late payments on the previous loan and little time in business: three reasons a credit officer can explain, and that a guarantor or a smaller loan could address.
Walkthrough
- Run the cells, then save the pipeline:
import joblib; joblib.dump(pipeline, "default_model.joblib"), and load it back in a new cell to check it scores the same. - Change the application: add a guarantor, or halve the loan amount. How much does the probability fall?
- Score five real applications from the test set and compare the recommendations with what actually happened.
- Write the model card (the task below).
Practice
Practice
What probability of default does the pipeline give the example application? As a percentage, rounded to the nearest whole number.
Task
10 minWrite a model card for Ladder's default model, with one line each starting: Purpose:, Users:, Data:, Features excluded:, Performance:, Threshold:, Limitations:, Monitoring: and Owner:.
Your work is checked for
- Purpose, Users and Data lines
- Features excluded line, mentioning region
- Performance line with a number
- Threshold line with a number and its basis
- Limitations, Monitoring and Owner lines
Check your understanding
Answer every question to check.