Module 5 · Generative AI Engineering
Evaluating an LLM classifier
Measure an LLM's classifications against human labels, compare a small model, a large model and a classic machine learning baseline, and find where each one fails.
About 25 minutes
The problem
Paystream's team ran its triage prompt on 900 past tickets with two models: a small, cheap one and a large one. A support lead had already labelled every ticket with its true category. Now the question is: which approach should run in production? The large model? The small one? Or, as one analyst suggests, a simple classifier trained on the labelled tickets, with no LLM at all?
That last suggestion often surprises people, and it's exactly why you measure. An LLM is one tool among several. When you have labelled examples and stable categories, a classic model can be cheaper, faster and almost as accurate.
The concept
A labelled test set is non-negotiable
You can't judge an LLM feature without examples where the right answer is known. Build one early: a few hundred real inputs, labelled by people who know the business. Use it for every prompt and model change.
Measures
- Accuracy: share of tickets classified correctly.
- Per-category recall: of the tickets that really are fraud, how many did the model call fraud? For high-stakes categories, this matters more than overall accuracy.
- Confusion: which categories get mixed up with which.
The baseline: TF-IDF and logistic regression
TF-IDF turns each text into numbers by weighting the words it contains (common words like "the" count little; distinctive words like "reversal" count a lot). A logistic regression trained on those numbers is a fast, cheap text classifier. Train it on part of the labelled tickets and test it on the rest, alongside the LLMs on the same tickets.
When each wins
| Approach | Strengths | Weaknesses |
|---|---|---|
| Classic classifier | cheap, fast, consistent, runs anywhere | needs labelled examples; struggles with new phrasings and new categories |
| LLM | works with few or no examples; handles unusual wording; easy to add categories | costs per call; slower; outputs need validation; can be manipulated |
Example
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
tickets = pd.read_csv("https://academy.cloudtechanalytics.com/datasets/genai/tickets.csv")
print("Small model accuracy (all 900):", round((tickets["small_model_category"] == tickets["true_category"]).mean(), 3))
print("Large model accuracy (all 900):", round((tickets["large_model_category"] == tickets["true_category"]).mean(), 3))
train, test = train_test_split(tickets, test_size=0.3, random_state=42, stratify=tickets["true_category"])
vectoriser = TfidfVectorizer(ngram_range=(1, 2))
baseline = LogisticRegression(max_iter=2000).fit(vectoriser.fit_transform(train["text"]), train["true_category"])
test = test.assign(baseline_category=baseline.predict(vectoriser.transform(test["text"])))
for col in ["baseline_category", "small_model_category", "large_model_category"]:
print(f"{col} on the same {len(test)} test tickets:", round((test[col] == test["true_category"]).mean(), 3))Small model accuracy (all 900): 0.798
Large model accuracy (all 900): 0.893
baseline_category on the same 270 test tickets: 0.867
small_model_category on the same 270 test tickets: 0.789
large_model_category on the same 270 test tickets: 0.885On these tickets, a classic classifier trained on 630 labelled examples lands close to the large model and ahead of the small one. Now look at the category that matters most:
recall = test.groupby("true_category").apply(
lambda g: pd.Series({col: (g[col] == g.name).mean() for col in ["baseline_category", "small_model_category", "large_model_category"]}),
include_groups=False,
)
recall.round(2)baseline_category small_model_category large_model_category
true_category
Account access 0.96 0.86 0.88
Cards 0.84 0.90 0.87
Cash-out agent 0.78 0.78 0.83
Failed or pending transfer 0.98 0.78 0.92
Fees and charges 0.70 0.85 0.89
Fraud or scam 0.84 0.64 0.96
Savings 0.60 0.60 0.75
Verification and limits 0.91 0.77 0.91Overall accuracy hides differences by category. Which model misses the most fraud reports is the number the head of support will ask about first.
Walkthrough
- Run the cells. Build a confusion table for the large model:
pd.crosstab(test["true_category"], test["large_model_category"]). - Read ten tickets the large model got wrong. Are they mistakes, or genuinely ambiguous tickets (two issues, or too vague)?
- Train the baseline on only 100 tickets. How does it compare with the LLMs now?
- Decide which approach you'd recommend, and what would change your mind.
Practice
Practice
What is the large model's accuracy across all 900 tickets? As a percentage, one decimal place.
Practice
What is the TF-IDF baseline's accuracy on the 270 test tickets? As a percentage, one decimal place.
Task
6 minWrite a recommendation (50 to 130 words) on which approach should sort Paystream's tickets: the accuracies you compared, fraud recall, and at least two factors beyond accuracy (cost, speed, new categories, labelled data, manipulation).
Your work is checked for
- Recommends an approach
- Gives at least two accuracy figures
- Mentions fraud
- Names at least two other factors
- Between 50 and 130 words
Check your understanding
Answer every question to check.