Module 10 · LLM Evaluation and Safety in Production
"Final project: Paystream's AI quality and safety plan"
Plan your final project, the evaluation and safety programme for a live AI assistant, with a release decision, red-team fixes, a fair guardrail, monitoring that would have caught every incident, and an incident playbook.
About 20 minutes
The problem
Paystream's head of support and its security lead want one document: how the assistant will be kept accurate and safe from now on. It must answer three immediate questions (should r2 or r3 ship? what threshold should the guardrail use? why did the incidents take so long to find?) and set up the process that answers them next time without a crisis.
Your final project is that programme, built from the data in this course.
The concept
What the plan contains
| Part | Built in |
|---|---|
| Regression suite review and new cases | lesson 2 |
| Release comparison and decision | lessons 3 and 4 |
| Red-team results and fix list | lesson 5 |
| Guardrail threshold and fairness | lesson 6 |
| Alerts with control limits | lessons 7 and 8 |
| Incident playbook and a postmortem | lesson 9 |
A backtest of your monitoring
The strongest evidence for a monitoring plan is to run it on the past: for each incident, the date your alerts would have fired, against the date it was actually found.
Example
The start of the backtest: the first date each alert would have fired, for each incident.
import numpy as np
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
daily = pd.read_csv(base + "daily_metrics.csv", parse_dates=["date"]).set_index("date")
incidents = pd.read_csv(base + "incidents.csv", parse_dates=["started", "detected", "resolved"])
def upper_limit(series):
past = series.shift(1).rolling(28, min_periods=14)
return past.mean() + 3 * past.std()
refusal_rate = daily["refusals"] / daily["conversations"]
alerts = pd.DataFrame({
"refusal": refusal_rate > upper_limit(refusal_rate),
"latency": daily["p95_latency_ms"] > upper_limit(daily["p95_latency_ms"]),
})
graded = daily[["graded_correct", "graded_sample"]].rolling(7).sum()
p = 0.9
alerts["accuracy_7d"] = graded["graded_correct"] / graded["graded_sample"] < p - 3 * np.sqrt(p * (1 - p) / graded["graded_sample"])
def first_alert(row):
window = alerts.loc[row["started"]:row["resolved"]]
fired = window[window.any(axis=1)]
return fired.index.min()
incidents["would_detect"] = incidents.apply(first_alert, axis=1)
incidents["days_saved"] = (incidents["detected"] - incidents["would_detect"]).dt.days
incidents[["incident_id", "started", "detected", "would_detect", "days_saved"]]incident_id started detected would_detect days_saved
0 INC-01 2026-06-20 2026-06-20 2026-06-20 0
1 INC-02 2026-07-14 2026-07-18 2026-07-14 4
2 INC-03 2026-08-04 2026-08-16 2026-08-09 7Each row is evidence for the monitoring plan: the days of customer impact the alerts would have saved. (The accuracy baseline here is fixed at 90% for simplicity; your project should estimate it from the data, as in lesson 8.)
Walkthrough
- Complete the backtest with the alerts you defined in lesson 7, and count false alarms in quiet periods.
- Make the release decision for r2 and r3 with the gate from lesson 4.
- Choose the guardrail threshold and the plan for the Pidgin bias.
- Open the project brief on the course page and plan the write-up.
Practice
Practice
In total, how many days sooner would the alerts have detected the three incidents?
Task
10 minWrite the executive summary of your quality and safety plan (100 to 200 words): the release decision for r2 and r3, the guardrail threshold and the Pidgin fix, the monitoring improvement with evidence from the backtest, and the process that keeps it working.
Your work is checked for
- Release decision naming r2 and r3
- Guardrail threshold
- Mentions the Pidgin fix
- Monitoring evidence (days, sooner, backtest)
- A process (gate, suite, red-team, postmortem, weekly)
- Between 100 and 200 words
Check your understanding
Answer every question to check.