Module 3 · Observability and Site Reliability
Structured logs
Why logs should be structured (JSON) rather than free text, how to filter and count them like data, and how to use them to find the event that started an incident.
About 25 minutes
The problem
Metrics told you when the outage started and which resource ran out. They can't tell you why the database's connection pool suddenly filled at 09:40 on the dot. Something happened at 09:40. Logs record events, and one of them is the answer.
Tallybook's services write structured logs: one JSON object per line, with named fields. That makes them data you can filter, count and join, rather than text to squint at.
The concept
Free text versus structured
2026-08-31 09:58:12 ERROR db pool exhausted for /api/invoices after 5000ms (40/40)
{"ts": "2026-08-31T09:58:12.031Z", "level": "error", "service": "api", "message": "db pool exhausted: no connection within 5000 ms", "route": "GET /api/invoices", "pool_in_use": 40, "pool_size": 40, "trace_id": "4f1c..."}The second can be filtered by service, grouped by route, and linked to a trace by trace_id, without fragile text matching.
Log levels
debug (detail for developers), info (normal events), warn (something odd, still working), error (a request failed). Alert on rates of errors, not individual lines.
Good logging habits
- Log events with context (IDs, durations, counts), not prose.
- Include a trace ID so a log line leads to the full request.
- Never log secrets or personal data (passwords, card numbers, full phone numbers).
Example
Load the logs from 09:30 to 10:45 and count by level and service:
import pandas as pd
logs = pd.read_json("https://academy.cloudtechanalytics.com/datasets/observability/app_logs.jsonl", lines=True)
logs["ts"] = pd.to_datetime(logs["ts"])
print(len(logs), "log lines")
pd.crosstab(logs["service"], logs["level"])1080 log lines
level error info warn
service
api 189 604 118
web 140 0 0
worker 0 29 0Errors per 10 minutes, by message:
errors = logs[logs["level"] == "error"]
errors.groupby([errors["ts"].dt.floor("10min").dt.strftime("%H:%M"), "message"]).size().unstack(fill_value=0)message db pool exhausted: no connection within 5000 ms upstream timed out after 30000 ms
ts
09:40 25 29
09:50 32 30
10:00 37 26
10:10 43 24
10:20 37 23
10:30 15 8The errors start in the 09:40 window and stop by 10:40. Now the key question: what else happened at those moments? Look at every log line that isn't a routine request or error, from the worker and from configuration changes:
events = logs[~logs["message"].isin(["request completed", "slow db connection acquire", "db pool exhausted: no connection within 5000 ms",
"upstream timed out after 30000 ms", "bulk send batch sent"])]
print(events[["ts", "service", "message", "job", "invoices_queued", "concurrency", "changed_by"]].to_string(index=False))ts service message job invoices_queued concurrency changed_by
2026-08-31 09:40:02.114000+00:00 worker bulk send job started month-end-bulk-send 41250.0 24.0 NaN
2026-08-31 10:33:40.502000+00:00 api config reloaded: db pool size 40 -> 80 NaN NaN NaN ada
2026-08-31 10:34:05.871000+00:00 worker bulk send job throttled month-end-bulk-send NaN 4.0 adaThat's the story. At 09:40:02, the month-end bulk send job started, queueing tens of thousands of invoices to send with 24 at a time, each needing a database connection from the same pool the API uses. At 10:33 and 10:34, Ada doubled the pool and throttled the job. The errors stopped. The web servers had nothing to do with it.
Walkthrough
- Run the cells. How many "bulk send batch sent" lines are there, and how many invoices did they send?
- Which routes appear most in the error lines? Does any route escape?
- Pick one error line's trace ID. Lesson 4 shows what a trace ID leads to.
- Write three logging rules for Tallybook's developers (the task below).
Practice
Practice
How many error-level log lines are there in the file?
Task
5 minWrite logging rules for Tallybook's developers, one per line starting with a dash: at least four, covering structure, trace IDs, what never to log, and levels.
Your work is checked for
- At least four rules, each starting with -
- Structured (JSON, fields)
- Trace IDs
- Never log secrets or personal data
- Levels
Check your understanding
Answer every question to check.