Module 10 · Generative AI Engineering
"Final project: Paystream's support assistant"
Plan your final project, a support assistant prototype with ticket triage and grounded answers, evaluated honestly and made safe, and set up the evaluation harness first.
About 20 minutes
The problem
Paystream's head of support wants a decision in a month: should the company launch an AI support assistant, and if so, which design? Your final project is the prototype and, more importantly, the evidence: how accurate it is, where it fails, what it costs, and what keeps it safe.
Experienced AI engineers build the evaluation harness first, before the feature itself. If you can't measure it, you can't improve it, and you can't tell the head of support whether it's ready.
The concept
The prototype has two parts
| Part | What it does | Evaluated with |
|---|---|---|
| Ticket triage | classifies each ticket into a category, with validation and a fraud safety net | accuracy and per-category recall on labelled tickets |
| Help answers | answers customer questions from retrieved help articles, with citations and refusals | retrieval hit rate, human grades, automatic citation checks, a checked LLM judge |
The evaluation harness
A single notebook section that, given a version of the system, produces the same table every time: accuracy, fraud recall, retrieval hit@3, correct-answer rate, refusal accuracy, failure rate and cost per 1,000 requests. Run it on every change and keep the history.
The decision
Launch, launch with limits (for example, triage only, with people answering), or don't launch yet, with the evidence for each.
Example
The start of an evaluation harness: one function that summarises triage quality for any column of predicted categories.
import pandas as pd
tickets = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/genai/tickets.csv")
def triage_report(df, predicted):
fraud = df[df["true_category"] == "Fraud or scam"]
return pd.Series({
"accuracy": round((df[predicted] == df["true_category"]).mean(), 3),
"fraud_recall": round((fraud[predicted] == "Fraud or scam").mean(), 3),
"tickets": len(df),
})
pd.DataFrame({col: triage_report(tickets, col) for col in ["small_model_category", "large_model_category"]}).Taccuracy fraud_recall tickets
small_model_category 0.798 0.723 900.0
large_model_category 0.893 0.892 900.0Every later version of the triage system gets a row in this table. Add a column for the classic baseline from lesson 5, and one for your own prompt if you have an API key.
Walkthrough
- Build the harness: triage report, retrieval hit rates and answer-grade summaries, in one place.
- Add the safety layers from lesson 9 and measure them: redaction coverage and injection flags.
- If you have an API key, run your own prompt version on a sample and add it to the table.
- Open the project brief on the course page and plan the write-up.
Practice
Practice
What is the small model's fraud recall across all tickets? As a percentage, one decimal place.
Task
8 minWrite the launch recommendation for the head of support (80 to 180 words): launch, launch with limits, or not yet, with the evidence (at least three numbers from your evaluations), the main risk, and the safeguards that would be in place.
Your work is checked for
- Makes a clear choice (launch, with limits, not yet)
- Uses at least three numbers
- Names a risk
- Lists safeguards (review, fraud, redact, validation, monitor)
- Between 80 and 180 words
Check your understanding
Answer every question to check.