Module 9 · LLM Evaluation and Safety in Production
Incident response
Measure how long incidents took to detect and fix, decide severity and first actions in advance, roll back safely, and write a blameless postmortem that leaves the system better than before.
About 20 minutes
The problem
Paystream's three incidents took very different times to detect: the outage was found the same day; the provider update after four days; the index rebuild after twelve. Each was handled by whoever happened to be around, with no agreed steps, and none of them led to lasting changes.
When an AI feature misbehaves, the first hour matters: who decides, what gets switched off, how customers are told. And after it's over, the most valuable thing is a clear, honest account of what happened, so it doesn't happen again.
The concept
Measure incidents
- Time to detect: from start to detection.
- Time to resolve: from detection to fix.
- How detected: your own alerts, or someone else (customers, social media)? Incidents found by others are the ones your monitoring missed.
Prepare before it happens
| Prepared in advance | Example |
|---|---|
| Severity levels | High: wrong answers about money, or safety failures. Medium: slow or unavailable |
| Owner on call | one named person with the authority to act |
| Safe fallbacks | pin the model version; roll back to the last good release or index; switch the assistant to "hand over everything" |
| Customer message | a pre-written notice that a person will help |
Blameless postmortems
After each incident, write: timeline, impact, root cause, why it wasn't caught earlier, and actions with owners and dates. Focus on systems, not people. "Nobody checked the index" becomes "the index rebuild had no automated check that all articles were present."
Example
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
incidents = pd.read_csv(base + "incidents.csv", parse_dates=["started", "detected", "resolved"])
daily = pd.read_csv(base + "daily_metrics.csv", parse_dates=["date"])
incidents["days_to_detect"] = (incidents["detected"] - incidents["started"]).dt.days
incidents["days_to_resolve"] = (incidents["resolved"] - incidents["detected"]).dt.days
def affected(row):
window = daily[daily["date"].between(row["started"], row["resolved"])]
return int(window["conversations"].sum())
incidents["conversations_affected"] = incidents.apply(affected, axis=1)
incidents[["incident_id", "severity", "how_detected", "days_to_detect", "days_to_resolve", "conversations_affected"]]incident_id severity how_detected days_to_detect days_to_resolve conversations_affected
0 INC-01 Medium Latency alert 0 0 1835
1 INC-02 High Customer complaints to support leads 4 1 11478
2 INC-03 High A customer's social media post 12 1 28395Most of the impact sits in detection time, not repair time. Once found, each incident was fixed within a day or so. The index incident ran for two weeks and touched tens of thousands of conversations, almost all of them before anyone knew. Faster detection (lessons 7 and 8) is the biggest improvement available.
Walkthrough
- Run the cell. Using lessons 7 and 8, work out how many days sooner each incident could have been detected, and how many fewer conversations it would have affected.
- For the index incident, write the timeline in five lines.
- Decide the safe fallback for each incident type.
- Write the postmortem (the task below).
Practice
Practice
How many conversations took place during incident INC-03 (from start to resolution)?
Task
10 minWrite a blameless postmortem for INC-03 with lines starting Summary:, Impact:, Root cause:, Why it wasn't caught:, and at least two lines starting Action: that each name an owner and a date.
Your work is checked for
- A Summary line
- An Impact line with a number
- A Root cause line
- A Why it wasn't caught line
- At least two Action lines with an owner and a date
- Blames no individual by name
Check your understanding
Answer every question to check.