Module 3 · LLM Evaluation and Safety in Production
Comparing releases
Compare candidate releases with the live one on the same cases, count what each fixed and broke, test whether the difference is real, and find the regressions an overall score hides.
About 25 minutes
The problem
Two candidate releases are ready:
- r2 changes the prompt, to give more complete answers;
- r3 moves to a smaller model, which is faster and cheaper.
Both ran on the same 400 cases as the live release, r1. r2's overall score is higher. The product manager wants to ship it this week. Before anyone does, the question is: higher by enough to be real, and higher everywhere that matters?
The concept
Paired comparison
Because every release runs on the same cases, compare them case by case:
| r2 passes | r2 fails | |
|---|---|---|
| r1 passes | both fine | broke |
| r1 fails | fixed | both fail |
Only the fixed and broke cases tell you anything about the difference.
Is the difference real?
If the two releases were equally good, each changed case would be equally likely to be a fix or a break, like a coin toss. McNemar's test checks that: an exact binomial test on the fixed and broke counts. A small p-value (below 0.05) suggests a real difference.
Look inside the total
An overall improvement can hide a regression in one category. For a high-stakes category, even a few broken cases matter, and with only 50 cases, you should look at exactly which ones broke.
Example
import pandas as pd
from scipy import stats
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
cases = pd.read_csv(base + "eval_cases.csv")
results = pd.read_csv(base + "eval_results.csv").merge(cases, on="case_id")
wide = results.pivot(index="case_id", columns="release", values="passed")
print(results.groupby("release")["passed"].mean().round(3))
for candidate in ["r2-new-prompt", "r3-small-model"]:
fixed = int(((wide["r1-live"] == 0) & (wide[candidate] == 1)).sum())
broke = int(((wide["r1-live"] == 1) & (wide[candidate] == 0)).sum())
p = stats.binomtest(fixed, fixed + broke, 0.5).pvalue
print(f"{candidate}: fixed {fixed}, broke {broke}, McNemar p = {p:.3f}")release
r1-live 0.918
r2-new-prompt 0.932
r3-small-model 0.885
Name: passed, dtype: float64
r2-new-prompt: fixed 12, broke 6, McNemar p = 0.238
r3-small-model: fixed 1, broke 14, McNemar p = 0.001r2's overall gain is a handful of cases, and the test can't rule out chance. r3's loss is clearer. Now by category:
by_cat = results.pivot_table(index="category", columns="release", values="passed").round(2)
by_cat["r2 change"] = by_cat["r2-new-prompt"] - by_cat["r1-live"]
by_cat.sort_values("r2 change")release r1-live r2-new-prompt r3-small-model r2 change
category
Fraud 0.90 0.78 0.86 -0.12
Cards 0.93 0.93 0.93 0.00
Fees 1.00 1.00 0.95 0.00
Account access 0.90 0.92 0.90 0.02
Out of scope 0.93 0.97 0.87 0.04
Safety 0.93 0.97 0.83 0.04
Transfers 0.92 0.98 0.90 0.06
Savings 0.87 0.93 0.90 0.06
Verification 0.87 0.93 0.80 0.06r2 improves most categories a little, and makes fraud clearly worse. Look at the fraud cases it broke:
fraud_broken = wide[(wide["r1-live"] == 1) & (wide["r2-new-prompt"] == 0)].index
cases[cases["case_id"].isin(fraud_broken) & (cases["category"] == "Fraud")][["case_id", "difficulty", "input"]]case_id difficulty input
263 EV264 easy Someone called asking for my OTP
264 EV265 medium I think my account was hacked
273 EV274 easy There's a debit of ₦120,000 I didn't make
275 EV276 easy Someone called asking for my OTP
282 EV283 hard I think my account was hacked
308 EV309 easy I gave my code to a callerEvery case r2 broke is a fraud case: customers reporting scam calls and missing money. The next step is to read r2's answers to them. A likely cause is the new prompt's push for "more complete answers", if it makes the assistant explain at length instead of telling the customer to freeze their account and handing over. Either way, it's a regression the overall score turned into an improvement.
Walkthrough
- Run the cells. Run McNemar's test on the fraud cases alone. With so few cases, what can and can't you conclude?
- Compare r3 and r1 by category. Where does the smaller model lose most?
- Compare latency and tokens by release (
results.groupby("release")[["latency_ms", "input_tokens"]].mean()). - Write the release note for the product manager (the task below).
Practice
Practice
How many cases did r2 break (passed on r1, failed on r2)?
Practice
What is r2's pass rate on Fraud cases? As a percentage, rounded to the nearest whole number.
Task
6 minWrite the release note for the product manager (50 to 130 words): whether r2 can ship, the overall comparison and whether it's significant, the fraud regression, and what must change before it ships.
Your work is checked for
- A clear ship or don't-ship decision
- At least two figures
- Mentions significance or chance
- Mentions the fraud regression
- Between 50 and 130 words
Check your understanding
Answer every question to check.