Module 2 · AI Engineer Capstone: From Prototype to Production
Extraction and validation
Parse and validate the model's JSON in code, repair what can safely be repaired, catch invented plate numbers without gold labels, and measure each configuration field by field and by language.
About 30 minutes
The problem
The model's output goes straight into Shieldline's claims system. A missing bracket crashes the pipeline. A plate number in the wrong format doesn't match any policy. And an invented plate number is worse than none: it attaches the claim to someone else's car. Before comparing models, write the code that stands between the model and the system.
The concept
Validate everything
| Field | Check |
|---|---|
| The whole output | Parses as JSON |
claim_type | One of the four allowed types |
incident_date | A real date, not after the message was sent |
vehicle_reg | Standard Nigerian format, ABC 123 DE, or null |
injuries, police_report, needs_human | True or false |
Repair only what's safe
A plate written FKJ471KT can be safely rewritten as FKJ 471 KT. A claim type the model made up can't be repaired; the message goes to a person.
Grounding checks in code
If the model returns a plate number, it should appear in the customer's message. You can check that with no gold labels at all, so the check works in production too.
Measure by field and by group
Overall accuracy hides which fields fail and for whom. Measure each field, and compare English with Pidgin.
Example
The validator:
import json
import re
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/assistant/"
messages = pd.read_csv(base + "messages.csv", keep_default_na=False)
CLAIM_TYPES = {"Windscreen", "Accident damage", "Third party", "Theft"}
PLATE = re.compile(r"^([A-Z]{3})\s*(\d{3})\s*([A-Z]{2})quot;)
def validate(raw, text, sent_at):
"""Parse one model output. Returns (record or None, list of problems)."""
try:
rec = json.loads(raw)
except json.JSONDecodeError:
return None, ["not valid JSON"]
problems = []
if rec.get("claim_type") not in CLAIM_TYPES:
problems.append("unknown claim type")
date = rec.get("incident_date")
if not (isinstance(date, str) and re.fullmatch(r"\d{4}-\d{2}-\d{2}", date) and date <= sent_at[:10]):
problems.append("bad or future date")
reg = rec.get("vehicle_reg")
if reg is not None:
match = PLATE.match(reg)
if not match:
problems.append("plate not recognised")
else:
rec["vehicle_reg"] = " ".join(match.groups()) # safe repair: standard spacing
if rec["vehicle_reg"].replace(" ", "") not in text.replace(" ", ""):
problems.append("plate not in the message")
for field in ["injuries", "police_report", "needs_human"]:
if not isinstance(rec.get(field), bool):
problems.append(f"{field} not true or false")
return rec, problems
rec, problems = validate('{"claim_type": "Theft", "incident_date": "2026-06-02", "vehicle_reg": "KJA482TL", '
'"injuries": false, "police_report": true, "needs_human": true}',
"My car was stolen yesterday. Plate KJA 482 TL.", "2026-06-03 10:00")
print(rec["vehicle_reg"], problems)KJA 482 TL []The plate was repaired and found in the message. Now run all three configurations through it, and compare each field with the gold labels:
def evaluate(config):
rows = []
for _, m in messages.iterrows():
rec, problems = validate(m[config + "_output"], m["text"], m["sent_at"])
rec = rec or {}
rows.append({
"valid_json": "not valid JSON" not in problems,
"passes_checks": not problems,
"claim_type": rec.get("claim_type") == m["gold_claim_type"],
"incident_date": rec.get("incident_date") == m["gold_incident_date"],
"vehicle_reg": (rec.get("vehicle_reg") or "") == m["gold_vehicle_reg"],
"invented_plate": m["gold_vehicle_reg"] == "" and rec.get("vehicle_reg") is not None,
"caught_by_check": "plate not in the message" in problems,
"language": m["language"],
})
return pd.DataFrame(rows)
results = {config: evaluate(config) for config in ["small_v1", "large_v1", "large_v2"]}
summary = pd.DataFrame({config: r.drop(columns="language").mean() for config, r in results.items()})
summary.round(3)small_v1 large_v1 large_v2
valid_json 0.932 0.955 0.997
passes_checks 0.887 0.947 0.992
claim_type 0.812 0.887 0.962
incident_date 0.873 0.923 0.973
vehicle_reg 0.900 0.953 0.993
invented_plate 0.042 0.007 0.003
caught_by_check 0.042 0.007 0.003large_v2 is better on every field. The schema and examples almost eliminated invalid JSON, and giving the model today's date fixed most relative dates such as "yesterday". Look at invented_plate: the small model invents a plate for a large share of the messages that don't contain one. The caught_by_check row shows the grounding check catches those inventions without needing any gold labels. By language:
pd.DataFrame({config: r.groupby("language")["claim_type"].mean() for config, r in results.items()}).round(3)small_v1 large_v1 large_v2
language
English 0.838 0.902 0.969
Pidgin 0.703 0.822 0.932Every configuration does worse on Pidgin, and the small model much worse. Since a fifth of customers write in Pidgin, that gap is part of the decision, not a footnote.
Walkthrough
- Run the cells.
- Print ten messages where
large_v2got the claim type wrong. Is there a pattern? - Look at the outputs that aren't valid JSON. Which could a simple repair fix (such as removing a "Here is the JSON:" prefix), and which should go to a person?
- Add a check of your own: for example, that a theft claim has
police_reportset. - Write the extraction report (the task below).
Practice
Practice
What percentage of small_v1 outputs invent a plate number for a message that doesn't contain one? (As a share of all 600 messages.) One decimal place.
Practice
What is large_v2's claim type accuracy on Pidgin messages? One decimal place.
Task
8 minWrite the extraction report (60 to 150 words): the configuration you'd use and why, its field accuracy with numbers, the gap for Pidgin, and what the validator catches.
Your work is checked for
- Names a configuration
- Uses numbers
- Covers Pidgin
- Covers the validator or checks
- Covers invented plates
- Between 60 and 150 words
Check your understanding
Answer every question to check.