Module 8 · Cloud & DevOps Engineer Capstone: Ready for the Sale
Game day and the go/no-go
Test the platform by breaking it on purpose, read the game-day results against their targets, turn the whole review into a readiness checklist, and give leadership an honest go/no-go for the sale.
About 25 minutes
The problem
Three weeks before the sale, the team ran a game day: a planned morning of breaking things in a controlled way, to see whether the system and the people respond as the plans say. Plans that have never been tested are hopes. Now the head of engineering wants the answer to the question in the brief: are we ready?
The concept
A game day
Each drill has a scenario, a success criterion and, where it matters, a target time, such as a recovery time objective (RTO) for restoring the database. Record what actually happened, including the surprises.
Pass, fail and what it teaches
A failed drill is a success for the game day: you found the problem before the sale did. Each failure becomes an action with an owner and a date, then a re-test.
Go, no-go, or go with conditions
A readiness decision lists what's done, what's open, and the conditions that must be met by a date. "Go if the backup restore passes a re-test by 20 November" is more useful than a vague "mostly ready".
Example
The game-day results:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/platform/"
gameday = pd.read_csv(base + "gameday.csv")
gameday["overrun_minutes"] = (gameday["actual_minutes"] - gameday["target_minutes"]).clip(lower=0)
print(gameday[["drill_id", "scenario", "target_minutes", "actual_minutes", "passed"]].to_string(index=False))
print(f"\nPassed {gameday['passed'].sum()} of {len(gameday)}")drill_id scenario target_minutes actual_minutes passed
G-1 Kill two checkout-api instances under load 5.0 3.0 1
G-2 Fail over orders-db to the replica 5.0 4.0 1
G-3 Restore orders-db from last night's backup 60.0 155.0 0
G-4 Roll back a bad checkout-api deploy 10.0 7.0 1
G-5 Payment gateway returns errors for 10 minutes 2.0 18.0 0
G-6 Burn-rate alert fires and reaches on-call 5.0 9.0 0
G-7 Traffic at 1.6x last year's peak in staging NaN NaN 1
Passed 4 of 7The failures, with what happened:
for _, d in gameday[gameday["passed"] == 0].iterrows():
print(f"{d['drill_id']} {d['scenario']} ({d['actual_minutes']:.0f} min against {d['target_minutes']:.0f}):\n {d['notes']}\n")G-3 Restore orders-db from last night's backup (155 min against 60):
Backup bucket unencrypted and unlabelled; restore steps were not written down; took 2.5 hours.
G-5 Payment gateway returns errors for 10 minutes (18 min against 2):
No fallback: checkout showed a blank error page until the gateway recovered.
G-6 Burn-rate alert fires and reaches on-call (9 min against 5):
Alert went to an email list; on-call saw it 9 minutes later.The capacity work held: the platform took 1.6 times last year's peak, and losing instances or failing over the database caused little harm. But three things failed, and each is serious on sale day. Restoring the database took over two and a half hours against a one-hour target. The bucket was the unencrypted, unowned one from lesson 1, and nobody had written the steps down. A payment gateway outage showed customers a blank page. And the page went to an email list, so on-call saw it after nine minutes, against a five-minute target.
The readiness checklist, from the whole review:
checklist = pd.DataFrame([
("Pooler, larger database, autoscaling to 30 (PR 214 fixed)", "Done", "Capacity drill passed at 1.6x"),
("Policy checks block dangerous plans in the pipeline", "Done", "Caught the database replacement"),
("Manual production resources imported into Terraform", "Open", "Replica, worker, backups bucket, bastion"),
("Backups encrypted, owned, and restore runbook written", "Open", "Restore drill failed: 155 min against 60"),
("Payment gateway fallback: retry message and queued orders", "Open", "Drill failed: blank page for 18 min"),
("Fast burn-rate page routed to on-call phones", "Open", "Alert reached on-call after 9 min by email"),
("Noisy alerts removed (CPU, heartbeat)", "Done", "On-call load cut by most of the volume"),
("Deploy rules and sale-week freeze agreed", "Done", "Tests required; no late-Friday deploys"),
("Idle machines deleted and staging scheduled", "Done", "Saves about $2,930 a month"),
], columns=["item", "status", "evidence"])
print(checklist.to_string(index=False))
print(f"\n{(checklist['status'] == 'Done').sum()} done, {(checklist['status'] == 'Open').sum()} open")item status evidence
Pooler, larger database, autoscaling to 30 (PR 214 fixed) Done Capacity drill passed at 1.6x
Policy checks block dangerous plans in the pipeline Done Caught the database replacement
Manual production resources imported into Terraform Open Replica, worker, backups bucket, bastion
Backups encrypted, owned, and restore runbook written Open Restore drill failed: 155 min against 60
Payment gateway fallback: retry message and queued orders Open Drill failed: blank page for 18 min
Fast burn-rate page routed to on-call phones Open Alert reached on-call after 9 min by email
Noisy alerts removed (CPU, heartbeat) Done On-call load cut by most of the volume
Deploy rules and sale-week freeze agreed Done Tests required; no late-Friday deploys
Idle machines deleted and staging scheduled Done Saves about $2,930 a month
5 done, 4 openThe honest answer is go, with conditions. The capacity problem that caused last year's outage is fixed and tested. But four items are open, and three of them failed a drill. Each needs an owner, a date before the freeze, and a re-test.
Walkthrough
- Run the cells.
- Give each open item an owner and a date, at least a week before the sale.
- Plan the re-tests: which drills run again, and when?
- Write the sale-day runbook's first page: who's on call, how to reach them, and the first three things to check.
- Write the executive summary (the task below), then open the project brief on the course page.
Practice
Practice
How many game-day drills failed?
Practice
By how many minutes did the database restore overrun its target?
Task
12 minWrite the executive summary for the head of engineering (120 to 230 words): what went wrong last year, what you've fixed and the evidence it works, what's still open, the cost effect, and your go/no-go with conditions and dates.
Your work is checked for
- Last year's cause (connections, pool)
- Uses numbers
- Evidence (load test, drill, game day)
- Open items (restore, backup, payment, paging)
- Cost effect
- A go/no-go decision
- Conditions with dates or deadlines
- Between 120 and 230 words
Check your understanding
Answer every question to check.