Module 5 · Observability and Site Reliability
SLIs and SLOs
Define service level indicators from what users experience, set objectives and error budgets, and measure Tallybook's August against them, including how much of the month's budget two mornings used.
About 25 minutes
The problem
"Is the app reliable enough?" can't be answered with a feeling. Tallybook's customers can't send invoices when requests fail or crawl, and Tallybook's engineers can't ship features while firefighting. An agreed, measured target settles arguments in both directions: when to slow down and fix reliability, and when reliability is good enough to keep building.
The concept
SLI, SLO, error budget
- A service level indicator (SLI) measures what users experience, as a ratio of good events to all events. For example: the share of requests that succeed, or that finish in under one second.
- A service level objective (SLO) is the target for that ratio over a period: 99.9% of requests succeed, over 30 days.
- The error budget is what's left: 0.1% of requests may fail. Spending it on releases and experiments is fine; running out means reliability work comes first.
Choose SLIs users would recognise
| SLI | Good event |
|---|---|
| Availability | request didn't return a 5xx error |
| Latency | request finished within 1 second |
Measure them where users meet the service (the load balancer or web tier), not deep inside.
Pick realistic objectives
100% is the wrong target: it's impossible, and chasing it stops all change. Pick what users actually need, and what the system can achieve.
Example
August's daily totals, measured at the web tier:
import pandas as pd
daily = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/observability/daily_sli.csv", parse_dates=["date"])
daily["availability"] = 1 - daily["errors_5xx"] / daily["requests"]
daily["fast"] = 1 - daily["slow_requests"] / daily["requests"]
month = daily[["requests", "errors_5xx", "slow_requests"]].sum()
availability = 1 - month["errors_5xx"] / month["requests"]
fast = 1 - month["slow_requests"] / month["requests"]
print(f"August availability: {availability:.4%} (SLO 99.9%)")
print(f"August requests under 1 s: {fast:.4%} (SLO 99%)")
daily.sort_values("availability")[["date", "requests", "availability", "fast"]].head(4)August availability: 99.6917% (SLO 99.9%)
August requests under 1 s: 98.9508% (SLO 99%)
date requests availability fast
30 2026-08-31 8316694 0.979174 0.950193
27 2026-08-28 8198501 0.984194 0.958629
0 2026-08-01 1947225 0.999360 0.996685
28 2026-08-29 1997031 0.999361 0.996642Both SLOs were missed for August. How much of the error budget did each day use?
SLO = 0.999
budget = (1 - SLO) * month["requests"] # failed requests allowed in August
daily["budget_used"] = daily["errors_5xx"] / budget
print(f"Error budget: {budget:,.0f} failed requests; used: {month['errors_5xx'] / budget:.0%}")
print(daily.sort_values("budget_used", ascending=False)[["date", "errors_5xx", "budget_used"]].head(3).round(3).to_string(index=False))Error budget: 118,739 failed requests; used: 308%
date errors_5xx budget_used
2026-08-31 173205 1.459
2026-08-28 129583 1.091
2026-08-04 2802 0.024The other 29 days together used about half the month's budget. Two month-end mornings used about two and a half budgets on their own. That makes the decision obvious: the month-end bulk job and the connection pool are the reliability work to do before anything else, and an error budget policy should say so in advance (lesson 9).
Walkthrough
- Run the cells. Recalculate August without the 28th and 31st. Would Tallybook have met both SLOs?
- Compute a 28-day rolling availability. On which day did it first fall below 99.9%?
- Why measure at the web tier rather than at the database?
- Write Tallybook's SLOs (the task below).
Practice
Practice
What was August's availability? As a percentage, two decimal places.
Task
6 minWrite Tallybook's SLOs: for each of availability and latency, one line giving the SLI (what counts as good), where it's measured, the target and the window. Add a line saying what happens when the error budget runs out.
Your work is checked for
- An availability line
- A latency line with a time threshold
- Targets as percentages
- A window (30 days, 28 days, rolling, month)
- Where measured (web tier, load balancer)
- Error budget consequence
Check your understanding
Answer every question to check.