Module 6 · LLM Evaluation and Safety in Production
Guardrails and thresholds
Choose the threshold for an input guardrail by weighing harmful messages missed against harmless customers blocked, and check whether it blocks some groups of customers far more than others.
About 25 minutes
The problem
The input guardrail gives every incoming message a score from 0 to 1: how likely it is to be an attack or abuse. Above a threshold, the message is blocked and the customer is asked to rephrase or contact support.
Security wants a low threshold, to catch everything. Support wants a high one, because every blocked genuine customer is someone who can't get help. A support lead also noticed something: customers writing in Pidgin seemed to get blocked more often. Paystream reviewed 3,000 production messages by hand, labelling each harmful or not, to settle it.
The concept
Every threshold is a trade-off
| Message is harmful | Message is harmless | |
|---|---|---|
| Blocked | caught (true positive) | customer wrongly blocked (false positive) |
| Allowed | attack gets through (false negative) | fine |
- Recall: share of harmful messages blocked.
- Precision: share of blocked messages that were really harmful.
- False positive rate: share of harmless messages blocked.
Choose by cost, not by habit. A 0.5 threshold is not special. Put a cost on each kind of error, or set a minimum recall and then pick the threshold that blocks the fewest genuine customers.
Check fairness
Calculate the false positive rate separately for each group (language, region, age). A guardrail that blocks one group's harmless messages far more often treats those customers worse, and it's often invisible in the overall numbers.
Example
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/llmops/"
g = pd.read_csv(base + "guardrail_reviews.csv")
print(g["harmful"].value_counts(), g["language"].value_counts(), sep="\n")
def at_threshold(t):
blocked = g["guardrail_score"] >= t
harmful = g["harmful"] == 1
return pd.Series({
"recall": round((blocked & harmful).sum() / harmful.sum(), 3),
"precision": round((blocked & harmful).sum() / blocked.sum(), 3),
"harmless_blocked": int((blocked & ~harmful).sum()),
"fpr_english": round(blocked[~harmful & (g["language"] == "English")].mean(), 3),
"fpr_pidgin": round(blocked[~harmful & (g["language"] == "Pidgin")].mean(), 3),
})
pd.DataFrame({t: at_threshold(t) for t in [0.3, 0.4, 0.5, 0.6, 0.7]}).Tharmful
0 2892
1 108
Name: count, dtype: int64
language
English 2310
Pidgin 690
Name: count, dtype: int64
recall precision harmless_blocked fpr_english fpr_pidgin
0.3 1.000 0.252 321.0 0.053 0.304
0.4 0.991 0.431 141.0 0.016 0.158
0.5 0.963 0.717 41.0 0.004 0.049
0.6 0.907 0.899 11.0 0.000 0.016
0.7 0.667 0.935 5.0 0.000 0.007Read across the rows. Raising the threshold blocks far fewer genuine customers, at the cost of missing some harmful messages. And at every threshold, harmless Pidgin messages are blocked many times more often than harmless English ones: the classifier treats Pidgin itself as suspicious.
Now choose a threshold by cost. Suppose (an assumption to agree with the business) that a missed harmful message costs 20 times as much as a wrongly blocked customer:
COST_MISS, COST_BLOCK = 20, 1
costs = {}
for t in [x / 100 for x in range(20, 91, 5)]:
blocked = g["guardrail_score"] >= t
misses = ((g["harmful"] == 1) & ~blocked).sum()
wrong_blocks = ((g["harmful"] == 0) & blocked).sum()
costs[t] = misses * COST_MISS + wrong_blocks * COST_BLOCK
best = min(costs, key=costs.get)
print("Lowest-cost threshold:", best, "cost:", costs[best])
print(at_threshold(best))Lowest-cost threshold: 0.5 cost: 121
recall 0.963
precision 0.717
harmless_blocked 41.000
fpr_english 0.004
fpr_pidgin 0.049
dtype: float64The cheapest threshold still blocks Pidgin speakers more often. A threshold alone can't fix a biased score: the classifier needs retraining with harmless Pidgin examples, and until then, blocked Pidgin messages could go to a person instead of being refused outright.
Walkthrough
- Run the cells. Change the cost ratio to 5 and to 50. How does the best threshold move?
- At the chosen threshold, how many genuine Pidgin-speaking customers out of 1,000 would be blocked?
- Design a different treatment for scores in a "grey zone" (for example 0.4 to 0.6): warn, ask to rephrase, or route to a person.
- Write the guardrail recommendation (the task below).
Practice
Practice
At a threshold of 0.5, what share of harmless Pidgin messages are blocked? As a percentage, one decimal place.
Practice
Which threshold has the lowest total cost with a miss costing 20 times a wrong block?
Task
6 minWrite the guardrail recommendation (50 to 130 words): the threshold and why, its recall and how many genuine customers it blocks, the Pidgin finding, and what you'll do about it.
Your work is checked for
- Names a threshold
- Mentions recall or harmful messages caught
- Mentions the Pidgin finding
- An action on the bias (retrain, route, review, person)
- Between 50 and 130 words
Check your understanding
Answer every question to check.