Module 7 · Cloud Fundamentals: Cost, Scaling and Reliability
Availability and SLOs
Measure availability against a target, calculate how components in series and in parallel combine, and find which single change would remove most of the downtime.
About 25 minutes
The problem
Tallybook promises customers that the app is available "99.9% of the time". Over June, July and August, there were six outages. Nobody had added them up against the promise, and nobody could say which change would prevent the most downtime: more web servers, a better deployment process, or a different database setup.
The concept
SLOs and error budgets
A service level objective (SLO) is an availability target, such as 99.9% a month. The allowed downtime is the error budget: 0.1% of a 30-day month is 43.2 minutes. When the budget is spent, reliability work takes priority over new features.
Components in series
If the app needs the load balancer and the web tier and the API and the database, its availability is the product of theirs. Every component in the chain lowers the total.
Components in parallel
If any one of several redundant copies is enough, the chance they're all down at once is the product of their unavailabilities. Two copies at 99.5% each give 1 − 0.005² = 99.9975%.
Single points of failure
A component with no redundancy (one database in one zone) often dominates downtime, no matter how many web servers you add.
Example
Downtime by month against a 99.9% SLO:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/cloud/"
outages = pd.read_csv(base + "outages.csv", parse_dates=["start"])
outages["month"] = outages["start"].dt.strftime("%Y-%m")
minutes_in_month = {"2026-06": 30 * 1440, "2026-07": 31 * 1440, "2026-08": 31 * 1440}
monthly = outages.groupby("month")["minutes"].sum().to_frame("downtime_min")
monthly["budget_min"] = [round(minutes_in_month[m] * 0.001, 1) for m in monthly.index]
monthly["availability_pct"] = [round(100 * (1 - d / minutes_in_month[m]), 3) for m, d in zip(monthly.index, monthly["downtime_min"])]
print(monthly)
outages.groupby("component")["minutes"].sum().sort_values(ascending=False)downtime_min budget_min availability_pct
month
2026-06 69 43.2 99.840
2026-07 86 44.6 99.807
2026-08 102 44.6 99.772
component
Database 102
Web 102
Load balancer 31
API 22
Name: minutes, dtype: int64Tallybook missed its target every month. The database and the web tier at month-end account for most of the downtime. Now the design arithmetic, with illustrative availabilities for each component:
def series(*parts):
total = 1.0
for p in parts:
total *= p
return total
def parallel(p, copies):
return 1 - (1 - p) ** copies
LB, SERVER, DB_SINGLE_ZONE, DB_MULTI_ZONE = 0.9999, 0.995, 0.995, 0.9995 # illustrative
designs = {
"now: 6 web, 4 API, single-zone DB": series(LB, parallel(SERVER, 6), parallel(SERVER, 4), DB_SINGLE_ZONE),
"multi-zone DB": series(LB, parallel(SERVER, 6), parallel(SERVER, 4), DB_MULTI_ZONE),
"multi-zone DB, 2 web, 2 API": series(LB, parallel(SERVER, 2), parallel(SERVER, 2), DB_MULTI_ZONE),
}
for name, a in designs.items():
print(f"{name}: {a:.5%} (about {(1 - a) * 30 * 1440:.0f} minutes down in a 30-day month)")now: 6 web, 4 API, single-zone DB: 99.49005% (about 220 minutes down in a 30-day month)
multi-zone DB: 99.94000% (about 26 minutes down in a 30-day month)
multi-zone DB, 2 web, 2 API: 99.93501% (about 28 minutes down in a 30-day month)With a single-zone database, the design can't reach 99.9% however many servers it has: the database alone allows more downtime than the whole budget. A multi-zone database (a standby copy in a second zone that takes over automatically) moves the design comfortably past the target. Notice too that the month-end outages weren't a design-availability problem at all: every server was up, there just weren't enough of them (lesson 6).
Walkthrough
- Run the cells. Which outages would a multi-zone database have prevented?
- Which outage would a gradual deployment (one server at a time) have prevented? Which would certificate expiry alerts have prevented?
- Calculate the availability with 3 API servers instead of 4. Does it matter?
- Write the reliability plan (the task below).
Practice
Practice
How many minutes of downtime did Tallybook have in August?
Practice
What is the error budget for a 30-day month at 99.9%, in minutes? One decimal place.
Task
8 minWrite the reliability plan: one line per outage cause starting with a dash, each naming the change that prevents it and how much downtime it would have saved over the three months.
Your work is checked for
- At least four lines, each starting with -
- Multi-zone database
- Autoscaling or capacity for month-end
- Safer deployment (gradual, one server, canary, rollback)
- Certificate renewal or expiry alert
- Minutes saved
Check your understanding
Answer every question to check.