Module 2 · LLM Evaluation and Safety in Production
Building a regression suite
Design the fixed set of test cases every release must pass (covering categories, difficulty, refusals and safety), and work out how precise its scores are, overall and per category.
About 25 minutes
The problem
Before a release, Paystream's team used to try "a few questions" and ship if the answers looked good. Different people tried different questions, and nobody could say whether a change made things better or worse.
A regression suite fixes that: the same set of cases, with the expected behaviour written down, run against every release. Paystream's suite has 400 cases. The question for this lesson is what makes a suite good, and how much a score on it can be trusted.
The concept
What a good suite covers
- Every category of real traffic, roughly in proportion, plus extra cases for high-stakes ones (fraud).
- Difficulty: easy, medium and hard cases, so improvements on hard cases show.
- Things the assistant must not do: out-of-scope questions it should decline, and safety cases (requests for other people's data, attempts to override its rules).
- Real wording: typos, Pidgin, urgency, mixed questions, taken from real traffic with personal data removed.
- Past failures: every incident and bug adds a case.
Each case needs a clear expected behaviour, written so that two graders would agree whether an answer passes.
How precise is a score?
A pass rate from a finite set of cases is an estimate. With n cases and pass rate p, a 95% confidence interval is roughly ± 2 × √(p(1 − p)/n). The Wilson interval is a better version for rates near 0 or 1. The key point: a category with 30 cases has a much wider interval than the whole suite of 400.
Example
import numpy as np
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
cases = pd.read_csv(base + "eval_cases.csv")
results = pd.read_csv(base + "eval_results.csv").merge(cases, on="case_id")
print(pd.crosstab(cases["category"], cases["difficulty"], margins=True))difficulty easy hard medium All
category
Account access 25 8 17 50
Cards 15 9 21 45
Fees 20 6 14 40
Fraud 24 10 16 50
Out of scope 17 5 8 30
Safety 16 1 13 30
Savings 13 7 10 30
Transfers 37 13 30 80
Verification 17 10 18 45
All 184 69 147 400Now the live release's score, with Wilson intervals:
def wilson(passed, n, z=1.96):
p = passed / n
centre = (p + z**2 / (2 * n)) / (1 + z**2 / n)
half = z * np.sqrt(p * (1 - p) / n + z**2 / (4 * n**2)) / (1 + z**2 / n)
return float(round(centre - half, 3)), float(round(centre + half, 3))
live = results[results["release"] == "r1-live"]
print("Overall:", round(live["passed"].mean(), 3), wilson(live["passed"].sum(), len(live)))
by_cat = live.groupby("category")["passed"].agg(["sum", "size"])
by_cat["rate"] = (by_cat["sum"] / by_cat["size"]).round(3)
by_cat["interval"] = [wilson(s, n) for s, n in zip(by_cat["sum"], by_cat["size"])]
by_catOverall: 0.918 (0.886, 0.941)
sum size rate interval
category
Account access 45 50 0.900 (0.786, 0.957)
Cards 42 45 0.933 (0.821, 0.977)
Fees 40 40 1.000 (0.912, 1.0)
Fraud 45 50 0.900 (0.786, 0.957)
Out of scope 28 30 0.933 (0.787, 0.982)
Safety 28 30 0.933 (0.787, 0.982)
Savings 26 30 0.867 (0.703, 0.947)
Transfers 74 80 0.925 (0.846, 0.965)
Verification 39 45 0.867 (0.738, 0.937)The overall score is pinned down to within a few points. Category scores are much looser: with 30 to 50 cases, a category's true pass rate could easily be several points either side of what you measured. That matters in the next lesson, when a candidate release's category scores move.
Walkthrough
- Run the cells. Which category has the widest interval, and why?
- Read ten cases. Are the expected behaviours specific enough for two people to agree?
- Work out roughly how many fraud cases you'd need for an interval of ± 5 points at a 90% pass rate.
- Write five new cases from the incidents in lesson 1 (the task below).
Practice
Practice
What is the live release's overall pass rate on the suite? As a percentage, one decimal place.
Task
8 minWrite five new regression cases from Paystream's incidents and risks, one per line in the form category | input | expected behaviour. Include at least one refusal or safety case and one in Pidgin.
Your work is checked for
- Five cases in the form category | input | expected
- A refusal or safety case
- A Pidgin input (abeg, wetin, dey, una, make I)
- Specific expected behaviours (a figure, an article or a step)
Check your understanding
Answer every question to check.