Module 2 · Observability and Site Reliability
Metrics and percentiles
Why averages hide what users experience, how percentiles (p50, p95, p99) show it, how to read latency over a day, and which metrics to put on a service's first dashboard.
About 25 minutes
The problem
Tallybook's old dashboard showed one latency number: the average. On a normal day it said about 200 milliseconds, and everyone relaxed. But users don't experience averages. One in twenty requests can be several times slower than the average, and the people making those requests are the ones who complain, refresh, and retry, adding more load.
The concept
Percentiles
The p95 is the time that 95% of requests are faster than; 5% are slower. p50 is the median; p99 is the slowest 1%.
| Measure | Tells you |
|---|---|
| p50 | the typical experience |
| p95 | what a meaningful minority experiences; good for SLOs |
| p99 | the tail: often a different problem (cold caches, lock waits, retries) |
| Average | dragged by outliers, matches nobody's experience |
Percentiles don't average
You can't average p95s across minutes or servers and get the p95 of the whole. Compute percentiles from the raw data, or use histogram metrics that can be combined.
A first dashboard
For each service: requests per minute, error rate, p50 and p95 (or p99) latency, and saturation of its tightest resource, on the same time axis.
Example
The API's latency percentiles before the incident, by hour:
import pandas as pd
metrics = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/observability/metrics.csv", parse_dates=["minute"])
api = metrics[metrics["service"] == "api"].set_index("minute")
normal = api.loc["2026-08-31 06:00":"2026-08-31 09:39"]
normal.groupby(normal.index.hour)[["requests", "p50_ms", "p95_ms", "p99_ms"]].mean().round(0)requests p50_ms p95_ms p99_ms
minute
6 3372.0 207.0 445.0 812.0
7 5856.0 243.0 515.0 925.0
8 7717.0 269.0 567.0 1008.0
9 8339.0 278.0 586.0 1037.0Even on a good morning, the slowest 1% of API requests take several times as long as the median, and latency rises with traffic. Now the same measures during the incident:
during = api.loc["2026-08-31 09:40":"2026-08-31 10:33"]
pd.DataFrame({"before (06:00-09:39)": normal[["p50_ms", "p95_ms", "p99_ms"]].median(),
"during (09:40-10:33)": during[["p50_ms", "p95_ms", "p99_ms"]].median()}).round(0)before (06:00-09:39) during (09:40-10:33)
p50_ms 254.0 6058.0
p95_ms 539.0 14960.0
p99_ms 962.0 15708.0During the incident even the median request took about six seconds, and p95 and p99 were around 15 seconds. A dashboard showing only the average would have shown "slow"; the percentiles show that nearly every request was badly delayed, and the error rate (lesson 1) shows more than a third of API requests failing outright.
Walkthrough
- Run the cells. Compute the ratio of p99 to p50 for each hour before the incident. Is it stable?
- Plot p50, p95 and p99 for the API across the day on a logarithmic axis.
- Why is averaging the three services' p95s meaningless?
- Sketch the four panels of the API's dashboard (the task below).
Practice
Practice
What was the API's median p95 latency during the incident, in milliseconds?
Task
6 minDesign the API's first dashboard: one line per panel, starting with a dash, naming the metric, how it's shown and why it's there. Cover all four golden signals.
Your work is checked for
- At least four panels, each starting with -
- Traffic (requests per minute)
- Errors (rate or percentage)
- Latency with percentiles
- Saturation (pool, CPU, connections)
- No averages for latency
Check your understanding
Answer every question to check.