Module 9 · CI/CD and Containers
Rollbacks and recovery
Compare rolling back with fixing forward, measure time to restore under the old and new processes, and design releases (feature flags, backwards-compatible changes) that make recovery fast.
About 20 minutes
The problem
Even with good tests and canaries, some bad releases reach customers. What matters then is how fast service is restored. Under the old process, Tallybook's engineers usually tried to fix the problem and push a new release ("fix forward"), which meant working out the cause first, under pressure. Under the new one, the first move is always to go back to the last good image.
The concept
Roll back or fix forward?
| Roll back | Fix forward | |
|---|---|---|
| How | redeploy the previous image | write, test and release a fix |
| Speed | minutes, without understanding the cause | as long as diagnosis and the fix take |
| When | the default | when rolling back isn't possible (for example, data was already changed) |
Restore service first; understand the cause afterwards.
Making rollbacks possible
- Keep previous images in the registry, tagged by version.
- Make database changes backwards-compatible (add a column first, start using it in a later release, remove the old one later still), so the previous app version still works.
- Use feature flags: ship new code switched off, turn it on separately, and switch it off instantly if it misbehaves, with no deployment at all.
Example
Incidents and their restore times under each process:
import pandas as pd
deploys = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/cicd/deployments.csv")
incidents = deploys[deploys["caused_incident"] == 1]
print(incidents.groupby(["pipeline", "result"])["minutes_to_restore"].agg(["count", "median", "max"]))
print()
print("Total customer-facing minutes of failed releases:")
print(incidents.groupby("pipeline")["minutes_to_restore"].sum())count median max
pipeline result
new rolled_back 2 21.5 25.0
old fixed_forward 3 216.0 278.0
rolled_back 5 194.0 306.0
Total customer-facing minutes of failed releases:
pipeline
new 43.0
old 1600.0
Name: minutes_to_restore, dtype: float64Under the old process, fixing forward took much longer than rolling back, and even the rollbacks took about three hours at the median, because rolling back meant reverting the code and re-running the whole manual process. Under the new pipeline, the two releases that slipped past the canary were rolled back in minutes.
Walkthrough
- Run the cells. How many hours of failed-release time did the old process cause in three months, and the new one?
- Design a backwards-compatible change for renaming a database column used by the app.
- Which kinds of Tallybook features would you put behind a feature flag?
- Write the recovery runbook (the task below).
Practice
Practice
What was the total number of minutes to restore across old-process incidents?
Task
6 minWrite the recovery runbook for a bad release that got past the canary, one numbered step per line: at least five steps covering detection, the rollback (how, and who decides), communication, finding the cause, and preventing a repeat.
Your work is checked for
- At least five numbered steps
- Detection (alert, monitor)
- Rollback to the previous image or version
- Communication (status page, customers, channel)
- Finding the cause after restoring
- Prevention (test, canary rule, check)
Check your understanding
Answer every question to check.