Module 4 · Machine Learning Fundamentals
Linear regression
Fit your first model, measure it against the baseline, discover why predicting the log of the target works far better for prices, and read the model's coefficients as business facts.
About 30 minutes
The problem
The property company wants more than a number. Landlords ask: "How much more could I charge if I made the flat serviced? Is 24-hour power worth it?" A model that predicts rent well is useful; a model that also explains what drives rent is far more useful, because it answers those questions too.
Linear regression does both. It's the oldest and simplest machine learning model, and still one of the most used, precisely because you can read it. This lesson fits one, finds an important flaw, and fixes it with one line of code.
The concept
What linear regression learns
It predicts the target as a weighted sum of the features:
rent ≈ intercept + w₁ × bedrooms + w₂ × size + w₃ × serviced + w₄ × (is it in Ikoyi?) + …
Training finds the weights (the coefficients) that make the predictions as close as possible to the real rents in the training data.
The flaw: prices multiply, they don't add
A linear model says being serviced adds a fixed amount, say ₦1.5m, everywhere. But in reality it adds a percentage: about the same proportion in Kubwa as in Ikoyi, so many more naira in Ikoyi. Prices, salaries and sales usually behave like this.
The fix: model the log of the target
Train on np.log(rent) and convert predictions back with np.exp. On the log scale, percentage effects become additive, which is exactly what a linear model can learn. A coefficient c then means: this feature multiplies rent by exp(c), a change of (exp(c) − 1) × 100 percent.
Reading the error
Use the same test set and the same measure (MAE) as the baselines, so the comparison is fair. Add R² to see how much of the variation the model explains.
Example
The setup, then a plain linear regression:
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
rentals = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/rentals/listings.csv")
typical = rentals.groupby(["area", "property_type", "bedrooms"])["annual_rent_ngn"].transform("median")
rentals = rentals[rentals["annual_rent_ngn"] <= 4 * typical].copy()
rentals["size_sqm"] = rentals["size_sqm"].fillna(
rentals.groupby(["property_type", "bedrooms"])["size_sqm"].transform("median")
)
features = ["area", "property_type", "bedrooms", "bathrooms", "size_sqm", "serviced",
"furnished", "power", "parking_spaces", "year_built"]
X = pd.get_dummies(rentals[features], drop_first=True, dtype=int)
y = rentals["annual_rent_ngn"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
plain = LinearRegression().fit(X_train, y_train)
pred = plain.predict(X_test)
print("Plain linear MAE:", round(mean_absolute_error(y_test, pred)), " R²:", round(r2_score(y_test, pred), 3))Plain linear MAE: 2216317 R²: 0.759Better than the area baseline (₦2.87m), but only a little. Now the same model on the log of rent:
log_model = LinearRegression().fit(X_train, np.log(y_train))
pred_log = np.exp(log_model.predict(X_test))
print("Log-rent linear MAE:", round(mean_absolute_error(y_test, pred_log)), " R²:", round(r2_score(y_test, pred_log), 3))Log-rent linear MAE: 827735 R²: 0.934The error falls by almost two-thirds. Same features, same model, one transformation. And now the coefficients answer the landlords' questions:
effect = (np.exp(pd.Series(log_model.coef_, index=X.columns)) - 1) * 100
effect[["serviced_Yes", "furnished_Yes", "power_Prepaid meter", "bedrooms"]].round(1)serviced_Yes 29.4
furnished_Yes 19.3
power_Prepaid meter -8.9
bedrooms 24.3
dtype: float64Read them as: a serviced flat rents for about 29% more than a similar unserviced one; furnished adds about 19%; a flat with only a prepaid meter rents for about 9% less than one with 24-hour power (the reference category); each extra bedroom adds about 24%, holding the other features constant.
Walkthrough
- Run the cells. Compare both models' MAE with your baselines from lesson 3.
- Plot predicted against actual rent for the log model (
plt.scatter(y_test, pred_log)), with log scales on both axes. The points should lie close to a diagonal line. - Look at the area coefficients:
effect.filter(like="area_").sort_values(). They're relative to Ajah, the reference area. - Find a listing in the test set with a large error. Can you see why the model got it wrong?
- Write down the log model's MAE: it's the number to beat in lesson 5.
Practice
Practice
What is the test MAE of the log-rent linear model? (A rounded figure is fine.)
Practice
According to the log model, by what percentage does being serviced increase rent, all else equal? One decimal place.
Task
6 minA landlord in Gwarinpa asks: "If I make my 3-bedroom flat serviced and furnished, how much more could I charge?" Write a short answer (40 to 120 words) using the model's coefficients, with the percentages, a sensible caveat about what the model can't tell you, and no promise of an exact figure.
Your work is checked for
- Uses the serviced and furnished percentages
- Includes a caveat (on average, similar flats, doesn't account for, cost, depends)
- No exact promise (guarantee, definitely, will get)
- Between 40 and 120 words
Check your understanding
Answer every question to check.