Across this track, Tallybook's month-end outage has been blamed on too few web servers and on a cryptominer. In this course you find the real cause. With a day of per-minute metrics you'll read the four golden signals and see that the database connection pool, not the web tier, ran out; with structured logs you'll find the bulk invoice-sending job that started at 09:40:02; with distributed traces you'll see requests spending seconds waiting for a connection while their queries barely slowed. Then you'll build the reliability practice that prevents a repeat: SLIs and SLOs with error budgets (two mornings used about two and a half months' worth), multi-window burn-rate alerts, a clean-up of pages that were nine-tenths noise, connection pool sizing with Little's law and job isolation, and an error budget policy with toil measured and cut. Every number comes from running the code on the incident's data.