Module 1 · Machine Learning Fundamentals
What machine learning is
What machine learning can and can't do, the difference between regression and classification, the workflow every project follows, and your first look at the two datasets this course uses.
About 25 minutes
The problem
A property company in Lagos wants to tell landlords what rent to ask for a flat, instantly, from a few details. A microfinance bank wants to know, before it lends, which small-business borrowers are likely to default. Both have years of records. Neither can write down the rules: rent depends on area, size, power, finishing and a dozen other things at once, and nobody can say exactly how much each one adds.
That's the kind of problem machine learning (ML) is for: when you have many examples with known answers, but the rule connecting the inputs to the answer is too complicated to write by hand. This course teaches you to build, test and explain such models in Python with scikit-learn, the library most data scientists start with, and to recognise when a model isn't the answer at all.
The concept
What a model learns
A machine learning model learns a function from features (the inputs, such as area, bedrooms, size) to a target (the answer, such as rent), from examples where both are known. It then predicts the target for new cases where only the features are known.
Two kinds of supervised learning
| Type | Target | Examples |
|---|---|---|
| Regression | a number | rent, delivery time, monthly sales |
| Classification | a category | default or not, fraud or not, will the customer leave? |
Both are supervised: the training data includes the right answers. Unsupervised learning (such as grouping customers into segments) has no target; this course focuses on supervised learning, which is where most business value is.
The workflow
- Frame the question: what will be predicted, for whom, and what decision will it change?
- Prepare the data: clean it, handle gaps, turn categories into numbers.
- Split it: train on some data, test on data the model hasn't seen.
- Start with a baseline: the simplest possible prediction, to beat.
- Train and compare models.
- Evaluate honestly, with the right measure for the decision.
- Explain and deploy responsibly, and keep monitoring.
When not to use machine learning
- When a simple rule works ("loans over ₦5m need a manager's approval").
- When you have too few examples, or none with known answers.
- When nobody will act differently because of the prediction.
- When mistakes would be harmful and you can't explain the model's decisions.
Example
The two datasets for this course. First, 2,400 rental listings in Lagos and Abuja:
import pandas as pd
rentals = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/rentals/listings.csv")
print(rentals.shape)
rentals.head(3)(2400, 14)
listing_id city area property_type bedrooms bathrooms size_sqm serviced furnished power parking_spaces year_built listed_date annual_rent_ngn
0 RL-00001 Lagos Yaba Mini flat 1 1 NaN No Yes 24-hour power 0 1995 2025-08-08 1550000
1 RL-00002 Lagos Ajah Flat 1 1 69.0 No No Prepaid meter and generator 1 2014 2026-05-17 1550000
2 RL-00003 Lagos Surulere Flat 1 1 90.0 No No Prepaid meter and generator 0 1990 2025-09-02 1600000And 5,000 microfinance loans, with whether each one defaulted:
loans = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/loans/loans.csv")
print(loans.shape)
loans["defaulted"].value_counts()(5000, 16)
defaulted
0 4407
1 593
Name: count, dtype: int64The rentals are a regression problem (predict annual_rent_ngn). The loans are a classification problem (predict defaulted, 1 or 0). Notice that only about 1 loan in 8 defaulted. That imbalance will matter a great deal in lesson 7.
Walkthrough
- Open a new notebook in Google Colab (colab.research.google.com). scikit-learn is already installed.
- Load both datasets with the code above.
- Look at the rentals with
rentals.describe()andrentals["area"].value_counts(). Which areas are most listed? - Look at the loans with
loans.describe(). What's the range of loan amounts and monthly revenue? - For each dataset, write down the target, three features you'd expect to matter, and the decision a prediction would change.
Practice
Practice
What is the median annual rent across all listings?
Practice
What percentage of loans defaulted? One decimal place.
Task
5 minFor each of these four problems, say whether it's regression, classification or not a machine learning problem, with a one-line reason. One per line in the form Problem | answer | reason.
- Estimating how many days a shipment will take
- Deciding whether an insurance claim looks fraudulent
- Applying a 10% discount to orders over ₦1m
- Predicting whether a customer will cancel their subscription next month
Your work is checked for
- Four lines in the form Problem | answer | reason
- Shipment days is regression
- The discount rule is not machine learning
- Fraud and cancellation are classification
Check your understanding
Answer every question to check.