Module 3 · Data Analytics Foundations
Types of data
Structured and unstructured data, numbers and categories, and the most important question about any table - what does one row mean?
About 25 minutes
The problem
You open Kolanut's orders.csv for the first time. Some columns hold numbers you can add up (quantity). Some hold numbers you must never add up (customer_id). Some hold dates. Before you calculate anything, you need to know what kind of data each column is, and what a single row stands for.
The concept
Structured vs unstructured
- Structured data fits in rows and columns with a fixed meaning: orders, invoices, attendance records. This course works with structured data.
- Unstructured data has no fixed layout: emails, WhatsApp messages, scanned contracts, photos. It holds value too, but needs other techniques to analyse.
Quantitative vs qualitative
| Type | What it is | Examples |
|---|---|---|
| Quantitative, discrete | Counts: whole numbers | Quantity ordered, number of staff |
| Quantitative, continuous | Measurements: can take any value | Weight in kg, revenue in naira |
| Qualitative, nominal | Categories with no order | Region, product category, payment method |
| Qualitative, ordinal | Categories with an order | Job level (Junior < Mid < Senior), rating (Poor, Fair, Good) |
Identifiers look like numbers but aren't. customer_id 42 is not "twice" customer 21. Adding or averaging IDs is meaningless. Treat them as labels.
Dates and times deserve their own type. They let you group by month, measure time between events, and compare periods.
Grain: what one row represents
The grain of a table is the answer to "one row = one what?". In Kolanut's data:
| File | One row is |
|---|---|
orders.csv | one product on one order (an order line) |
customers.csv | one customer |
products.csv | one product |
Getting the grain wrong causes real errors. If you think each row in orders.csv is a whole order and count rows to get "number of orders", you'll count some orders more than once whenever a shop buys several products on the same day.
Example
The first rows of orders.csv:
| order_id | order_date | customer_id | product_id | quantity | unit_price | discount_pct |
|---|---|---|---|---|---|---|
| 10001 | 2025-01-01 | 27 | 3 | 14 | 18600 | 0 |
| 10002 | 2025-01-01 | 56 | 1 | 7 | 13200 | 0 |
| 10003 | 2025-01-01 | 37 | 2 | 4 | 3600 | 0 |
order_id,customer_id,product_id: identifiers (labels, not quantities).order_date: a date.quantity: quantitative, discrete.unit_price: quantitative, the price actually charged for one unit (one pack).discount_pct: quantitative, the percentage taken off this line: 0, 5 or 10.
Reading the first row: order line 10001 was 14 packs of product 3 at ₦18,600 a pack, with no discount.
Walkthrough
Open orders.csv in Google Sheets (File → Import → Upload) or Excel (File → Open).
- Look at the header row. Each column name tells you what the column holds.
- Press Ctrl + ↓ (Cmd + ↓ on a Mac) in column A to jump to the last row. The row number tells you how many rows there are. Remember that row 1 is the header.
- For each column, decide: identifier, number, category or date?
- Ask the grain question: what does one row stand for?
Practice
Practice
How many order lines (data rows, not counting the header) are in orders.csv?
Practice
How many different products does Kolanut sell? (Look in products.csv.)
Check your understanding
Answer every question to check.