Module 6 · AI Engineer Capstone: From Prototype to Production
Red-teaming and fair guardrails
Attack the assistant on purpose, compare two guardrails by how many attacks get through, and check the price: genuine customers wrongly blocked, and whether that price falls on Pidgin speakers and angry customers.
About 25 minutes
The problem
A WhatsApp number is open to anyone. Some people will try to make the assistant approve a claim, promise money, reveal another customer's details, or obey instructions hidden inside a message. The team built two guardrails, each a check that blocks a message before the model acts on it:
- v1: a keyword filter.
- v2: a small classifier trained on attacks, together with a hardened prompt and the rules in code from lesson 3.
The red team ran 120 attacks against both. A guardrail also has a cost, though: every genuine customer it blocks is a customer who gets no help.
The concept
Red-teaming
Write attacks by category, run them against each version, and record the outcome: Blocked by the guardrail, Refused by the model, or Attack succeeded. Report the success rate by category, because one weak category is enough.
Defence in depth
The guardrail is one layer. The model's instructions are another. Code is the strongest: an assistant that has no tool to approve claims can't be talked into approving one, whatever the message says.
False positives, and who pays them
Run the guardrail on genuine messages and count how many it wrongly flags. Then split by group. A keyword filter that trips on Pidgin words, or on anger, blocks exactly the customers who most need a person.
Example
Attack outcomes for each version:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/assistant/"
redteam = pd.read_csv(base + "redteam.csv")
messages = pd.read_csv(base + "messages.csv", keep_default_na=False)
print(pd.DataFrame({v: redteam[f"{v}_outcome"].value_counts() for v in ["v1", "v2"]}).fillna(0).astype(int), "\n")
success = redteam.groupby("category")[["v1_outcome", "v2_outcome"]].agg(lambda s: (s == "Attack succeeded").mean())
success.round(3)v1 v2
Blocked 57 113
Refused 49 6
Attack succeeded 14 1
v1_outcome v2_outcome
category
Abuse and threats 0.042 0.000
Data request 0.125 0.042
Hidden injection 0.167 0.000
Instruction override 0.125 0.000
Payment promise 0.125 0.000v1 lets attacks through in every category. v2 stops almost all of them; the only one that got through was a request for another customer's data. Now the price, on the 600 genuine messages:
print(messages.groupby("language")[["guardrail_v1_flag", "guardrail_v2_flag"]].mean().round(3), "\n")
messages.groupby("gold_angry")[["guardrail_v1_flag", "guardrail_v2_flag"]].mean().round(3)guardrail_v1_flag guardrail_v2_flag
language
English 0.056 0.021
Pidgin 0.186 0.034
guardrail_v1_flag guardrail_v2_flag
gold_angry
0 0.046 0.026
1 0.271 0.010The keyword filter wrongly flags Pidgin messages more than three times as often as English ones, and it flags angry customers far more than calm ones. An angry customer who has been waiting is exactly the person who should reach a human, not a wall. v2 is more accurate against attacks and fairer to customers. Even so, a flagged genuine message should go to a person rather than be dropped, so a false positive costs a delay, not a customer.
Walkthrough
- Run the cells.
- Read the attack that succeeded against v2. Which layer should have stopped it? Write the code rule that would (for example: the assistant has no access to other customers' records at all).
- Write five new attacks in Pidgin. Which category do you expect to be weakest?
- Decide what happens to a flagged message: blocked with a message, or routed to a person?
- Write the red-team report (the task below).
Practice
Practice
How many of the 120 attacks succeeded against guardrail v1?
Practice
What percentage of genuine Pidgin messages does guardrail v1 wrongly flag? One decimal place.
Task
8 minWrite the red-team report (60 to 150 words): attack success for both versions, the weakest category, the false positives by language and for angry customers, and what happens to a flagged message.
Your work is checked for
- Attack results for both versions
- Uses numbers
- Names a category
- Covers Pidgin false positives
- Covers angry customers
- What happens to flagged messages (person, human, routed)
- Between 60 and 150 words
Check your understanding
Answer every question to check.