Module 10 · Observability and Site Reliability
"Final project: Tallybook's reliability review"
Plan your final project, a reliability review that finds the real cause of the month-end outages with metrics, logs and traces, and sets the SLOs, alerts, capacity plan and practices that prevent them.
About 20 minutes
The problem
Tallybook's CTO has read three explanations of the month-end outages: too few web servers, a cryptominer, and now a database connection pool. The board wants one clear account and a plan. Your final project is the reliability review: what really happened, proved with telemetry, and what Tallybook will measure and change so it doesn't happen again.
The concept
The parts of the review
| Part | Built in |
|---|---|
| The cause, from metrics, logs and traces | lessons 1 to 4 |
| SLOs and August's performance against them | lesson 5 |
| Alerting: burn-rate rules and the alert clean-up | lessons 6 and 7 |
| Capacity: pool sizing and job isolation | lesson 8 |
| Toil and the error budget policy | lesson 9 |
Reconcile the explanations
Good reviews don't just give the right answer; they explain why the earlier ones were incomplete. The web fleet and the miner were real issues that made things worse, but the evidence shows the pool was the cause. Say so, with the data.
Example
The headline evidence in one table: the four golden signals for the API and database, before and during the incident.
import pandas as pd
metrics = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/observability/metrics.csv", parse_dates=["minute"]).set_index("minute")
periods = {"before (08:40-09:39)": ("2026-08-31 08:40", "2026-08-31 09:39"), "during (09:40-10:33)": ("2026-08-31 09:40", "2026-08-31 10:33")}
rows = []
for label, (start, end) in periods.items():
for service in ["web", "api", "db"]:
m = metrics[metrics["service"] == service].loc[start:end]
rows.append({"period": label, "service": service, "requests_per_min": round(m["requests"].mean()),
"error_rate_pct": round(m["errors"].sum() / m["requests"].sum() * 100, 2),
"p95_ms": round(m["p95_ms"].median()), "saturation_pct": round(m["saturation_pct"].median(), 1)})
pd.DataFrame(rows).set_index(["service", "period"]).sort_index()requests_per_min error_rate_pct p95_ms saturation_pct
service period
api before (08:40-09:39) 8293 0.05 581 70.8
during (09:40-10:33) 8272 37.64 14960 100.0
db before (08:40-09:39) 24878 0.00 173 81.8
during (09:40-10:33) 24815 0.00 284 100.0
web before (08:40-09:39) 11847 0.06 621 66.8
during (09:40-10:33) 11817 26.38 15000 66.8Traffic is almost unchanged; the database's saturation went to 100% and the API's errors and latency followed. The web tier's own saturation barely moved, which is why adding web servers wouldn't have helped.
Walkthrough
- Complete the evidence: the log events, the trace comparison and the Little's law calculation.
- Write the SLOs, August's results and the error budget used.
- Write the alerting changes and the capacity plan.
- Open the project brief on the course page and plan the write-up.
Practice
Practice
What was the API's error rate during the incident, in per cent? Two decimal places.
Task
10 minWrite the executive summary of your reliability review (100 to 200 words): the real cause with the evidence from metrics, logs and traces, why the earlier explanations were incomplete, the SLO result, and the changes (alerts, capacity, policy).
Your work is checked for
- Names the cause (connection pool, bulk job)
- Cites metrics, logs and traces
- Addresses earlier explanations (web servers, miner)
- An SLO result with a percentage
- Changes (alert, pool, isolate, policy)
- Between 100 and 200 words
Check your understanding
Answer every question to check.