Module 1 · AI Engineer Capstone: From Prototype to Production
The brief and the design
Meet Shieldline's WhatsApp claims assistant, decide what the model does and what code decides, define what it must never do, and get to know the recorded test data.
About 25 minutes
The problem
This is the capstone of the AI Engineer track. You'll take one LLM feature all the way from design to production monitoring, using the habits from the track: validated outputs, rules in code, retrieval, evaluation, red-teaming, release gates and alerts.
The company is Shieldline Insurance, the motor insurer from the Business Analyst Capstone. Its claims were slow partly because customers sent incomplete information and then heard nothing. Most customers already message Shieldline on WhatsApp, so the head of claims has asked for an assistant that:
- reads a customer's first message about an incident and turns it into a structured claim (type, date, plate number, injuries, police report);
- tells the customer exactly which documents are still needed;
- answers policy questions from the policy wording;
- hands anything serious to a person.
The team has already run three configurations over 600 real messages and recorded every output, so you can evaluate them without an API key.
The concept
The model reads; code decides
An LLM is good at reading messy text, including Pidgin, and at writing a clear reply from a source. It is not reliable at enforcing rules. So split the work:
| The model does | Code does |
|---|---|
| Extract fields from the message | Validate every field, and reject or repair what's wrong |
| Answer questions from retrieved policy text | Decide which documents are needed, from the policy table |
| Suggest whether a person is needed | Escalate by rule: injuries, theft, large amounts, anger, or any output that fails validation |
What it must never do
Approve, reject or promise payment for a claim; quote an amount Shieldline will pay; reveal anything about another customer; or answer from memory when the policy doesn't say.
The test data
messages.csv holds 600 messages with gold labels (what a claims officer extracted) and the recorded outputs of three configurations:
| Configuration | What it is |
|---|---|
small_v1 | A small, cheap model with a short prompt |
large_v1 | A large model with the same prompt |
large_v2 | The large model with a JSON schema, three worked examples and today's date in the prompt |
Example
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/assistant/"
messages = pd.read_csv(base + "messages.csv", keep_default_na=False)
print(f"{len(messages)} messages")
print(messages["language"].value_counts().to_string(), "\n")
print(messages["gold_claim_type"].value_counts().to_string(), "\n")
print(f"Need a person (gold): {messages['gold_needs_human'].mean():.1%}")600 messages
language
English 482
Pidgin 118
gold_claim_type
Accident damage 250
Windscreen 180
Third party 90
Theft 80
Need a person (gold): 36.0%One message, and what each configuration made of it:
example = messages[messages["message_id"] == "MSG-0001"].iloc[0]
print(example["text"], "\n")
for config in ["small_v1", "large_v1", "large_v2"]:
print(f"{config}: {example[config + '_output']}")Good morning o, stone break my windscreen on 26/06/2026 for Wuse 2. Na Lexus RX 350, number na FKJ 471 KT. Mechanic talk say e go cost like ₦275k. Which document una need?
small_v1: {"claim_type":"Windscreen","incident_date":"2026-06-26","vehicle_reg":"FKJ 471 KT","injuries":false,"police_report":false,"needs_human":false}
large_v1: {"claim_type":"Windscreen","incident_date":"2026-06-26","vehicle_reg":"FKJ 471 KT","injuries":false,"police_report":false,"needs_human":false}
large_v2: {"claim_type":"Windscreen","incident_date":"2026-06-26","vehicle_reg":"FKJ471KT","injuries":false,"police_report":false,"needs_human":false}All three got this one essentially right, but notice large_v2 wrote the plate without spaces. Small differences like that matter when the plate is matched against Shieldline's policy records. The rest of this course measures them properly.
Walkthrough
- Download the dataset below and open
messages.csvin Colab. - Read twenty messages, including some in Pidgin. What makes them hard to extract?
- Draw the design: message in, model calls, the checks in code, and the three ways out (reply with a checklist, answer a question, hand to a person).
- List what the assistant must never do, and where in the design each rule is enforced.
- Write the design note (the task below).
Practice
Practice
What percentage of the 600 messages need a person according to the gold labels? One decimal place.
Practice
How many of the messages are in Pidgin?
Task
8 minWrite the design note (70 to 160 words): what the model does, what code decides, when a person takes over, and at least three things the assistant must never do.
Your work is checked for
- Says what the model does (extract, read, answer)
- Says what code decides (validate, rules, checklist)
- Says when a person takes over (escalate, hand over)
- Lists things it must never do
- Covers approving or promising payment
- Covers other customers' data
- Between 70 and 160 words
Check your understanding
Answer every question to check.