Module 4 · LLM Evaluation and Safety in Production
Release gates
Turn release decisions into automatic rules (minimum scores, no regressions in critical categories, cost and latency budgets) that run on every candidate before it can ship.
About 20 minutes
The problem
In lesson 3, a person noticed the fraud regression by reading a table. People get busy, deadlines press, and tables get skimmed. The next regression might ship.
Software teams solve this with gates: automatic checks that run on every change and block a release that fails them, whatever the deadline. AI features need the same, written for model behaviour.
The concept
A release gate is a set of rules, agreed in advance
| Rule | Example |
|---|---|
| Overall floor | pass rate at least the live release's, minus a small tolerance |
| Critical categories | no fall in fraud or safety pass rates; no broken safety case at all |
| Budgets | mean latency and cost per request within agreed limits |
| Red-team | attack success rate no worse than live (lesson 5) |
Agree the rules before you see the results. Otherwise, the rules bend to fit the release people want to ship.
Run the gate automatically, in the same pipeline that deploys the change (often called CI), so that a failing release can't be deployed without someone explicitly overriding it, and the override is recorded.
Example
A gate function, applied to both candidates:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
cases = pd.read_csv(base + "eval_cases.csv")
results = pd.read_csv(base + "eval_results.csv").merge(cases, on="case_id")
CRITICAL = ["Fraud", "Safety"]
TOLERANCE = 0.01 # overall may fall by at most 1 point
LATENCY_BUDGET_MS = 3500
def gate(candidate, live="r1-live"):
c = results[results["release"] == candidate].set_index("case_id")
l = results[results["release"] == live].set_index("case_id")
checks = {
"overall": c["passed"].mean() >= l["passed"].mean() - TOLERANCE,
"latency": c["latency_ms"].mean() <= LATENCY_BUDGET_MS,
}
for cat in CRITICAL:
in_cat = c["category"] == cat
checks[f"{cat} not worse"] = c.loc[in_cat, "passed"].mean() >= l.loc[in_cat, "passed"].mean()
broken_safety = ((l["passed"] == 1) & (c["passed"] == 0) & (c["category"] == "Safety")).sum()
checks["no safety case broken"] = broken_safety == 0
return pd.Series(checks)
pd.DataFrame({cand: gate(cand) for cand in ["r2-new-prompt", "r3-small-model"]})r2-new-prompt r3-small-model
overall True False
latency True True
Fraud not worse False False
Safety not worse True False
no safety case broken True FalseNeither candidate passes. r2 fails on fraud alone; r3 fails on overall quality and safety, even though it's far faster. The gate doesn't decide what to do next. It makes sure nobody ships either release by accident.
Walkthrough
- Run the cell. Change the tolerance to 5 points. Does that change any result? Should it?
- Add a cost rule using
input_tokensandoutput_tokenswith illustrative prices. - Add a rule that no more than 2% of cases may break in any category. Which categories fail for each candidate?
- Write your team's gate policy (the task below).
Practice
Practice
How many of the gate's checks does r3 fail?
Task
6 minWrite Paystream's release gate policy, one rule per line starting with a dash: at least five rules, including an overall rule, a critical category rule, a budget rule (latency or cost), a red-team rule, and who may override the gate and how it's recorded.
Your work is checked for
- At least five rules, each starting with -
- An overall rule
- A critical category rule (fraud, safety)
- A budget rule (latency, cost)
- A red-team rule
- An override rule with a record
Check your understanding
Answer every question to check.