Module 5 · AI Engineer Capstone: From Prototype to Production
Grounded answers and judging
Compare answers with and without retrieval, see how answer quality depends on retrieval hits, and test whether an LLM judge can replace human grading, including the one kind of error it misses.
About 25 minutes
The problem
A wrong answer about insurance is costly. A customer told "yes, you can use your own mechanic" may lose part of their claim. So the assistant must answer from the policy, and Shieldline needs a way to check answer quality every time the prompt or model changes. People graded the answers this time. Next time, the team wants an LLM judge to do it automatically. Can it be trusted?
The concept
Grades
| Grade | Meaning |
|---|---|
| Correct | Right, and supported by the policy |
| Partly correct | Right but incomplete |
| Wrong | Contradicts the policy |
| Unsupported | States something the policy doesn't say: made up, even if it sounds plausible |
Unsupported answers are the most dangerous kind. They sound confident and can't be traced to anything.
Answer quality depends on retrieval
Split the graded answers by whether retrieval found the right section. If most bad answers come from retrieval misses, fix retrieval first.
Judging the judge
Before an LLM judge replaces people, compare its grades with human grades on the same answers. Look at overall agreement, but especially at agreement on the grades that matter most. A judge that calls unsupported answers "Correct" is worse than useless for safety.
Example
Answers with and without retrieval:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/assistant/"
questions = pd.read_csv(base + "questions.csv")
questions["hit"] = [gold in ids.split(";") for gold, ids in zip(questions["gold_section_id"], questions["retrieved_ids"])]
ORDER = ["Correct", "Partly correct", "Wrong", "Unsupported"]
pd.DataFrame({
"without retrieval": questions["answer_no_rag_grade"].value_counts(),
"with retrieval": questions["answer_rag_grade"].value_counts(),
}).reindex(ORDER).fillna(0).astype(int)without retrieval with retrieval
Correct 38 57
Partly correct 12 12
Wrong 7 6
Unsupported 23 5Retrieval raises correct answers substantially and cuts unsupported ones sharply. Where do the remaining bad answers come from?
pd.crosstab(questions["hit"], questions["answer_rag_grade"]).reindex(columns=ORDER, fill_value=0)answer_rag_grade Correct Partly correct Wrong Unsupported
hit
False 2 1 4 4
True 55 11 2 1When retrieval finds the right section, the answer is almost always right. Most wrong and unsupported answers (8 of 11) follow a retrieval miss, although misses are only 11 of the 80 questions. So the retrieval fixes from lesson 4 are also the best answer fixes. In the meantime, instruct the model to say it will pass the question to a colleague when the retrieved sections don't contain the answer, and test that instruction.
Now the judge:
agree = (questions["judge_grade"] == questions["answer_rag_grade"]).mean()
print(f"The judge agrees with people on {agree:.1%} of answers")
pd.crosstab(questions["answer_rag_grade"], questions["judge_grade"], rownames=["people"], colnames=["judge"]).reindex(index=ORDER, columns=ORDER, fill_value=0)The judge agrees with people on 90.0% of answers
judge Correct Partly correct Wrong Unsupported
people
Correct 53 0 2 2
Partly correct 0 12 0 0
Wrong 0 0 5 1
Unsupported 3 0 0 2The overall agreement looks good. But read the Unsupported row: of the answers people marked unsupported, the judge called most of them Correct. It's fooled by exactly the confident, made-up answers that matter most. So the judge can't run alone. Use it to grade at scale, and add a code check that every sentence of an answer can be matched to a retrieved section, plus a monthly human sample focused on answers the judge passed.
Walkthrough
- Run the cells.
- Read the unsupported answers' questions. Which policy sections were they about, and were they retrieval misses?
- Write the instruction that tells the model what to do when the retrieved sections don't contain the answer.
- Design the ongoing check: how many answers people grade each month, and which ones.
- Write the answer-quality note (the task below).
Practice
Practice
How many of the 80 answers with retrieval did people grade Correct?
Practice
Of the answers people graded Unsupported, how many did the judge grade Correct?
Task
7 minWrite the answer-quality note (60 to 140 words): what retrieval does for answer quality, where the remaining bad answers come from, whether the LLM judge can be trusted, and the safeguard you'd add.
Your work is checked for
- Uses numbers
- Links bad answers to retrieval misses
- Covers the judge
- Names the judge's weakness (unsupported)
- A safeguard (human sample, check, pass to a colleague)
- Between 60 and 140 words
Check your understanding
Answer every question to check.