Module 8 · AI Engineer Capstone: From Prototype to Production
Launch, monitoring and the presentation
Watch the assistant's first four weeks in production with control limits set from its first two weeks, catch a change nobody announced, trace it to its cause, and present the project to Shieldline's leadership.
About 30 minutes
The problem
large_v2 went live on 3 August 2026. For two weeks everything looked steady. Then the WhatsApp team switched on voice notes: customers can now speak their message, and a transcription service turns it into text before the assistant sees it. Nobody told the AI team. The release gate was passed on typed messages, so does the assistant still work?
daily.csv has the production metrics for the first 28 days, including the results of a daily audit: each day, a claims officer checks 40 random extractions against the original message.
The concept
Control limits from a stable period
Use the first two weeks as the baseline. For each metric, set an upper limit at the baseline mean plus three standard deviations. A day above it is very unlikely to be normal variation: investigate.
Leading and lagging signals
| Signal | Arrives | Shows |
|---|---|---|
| Invalid JSON rate | At once | The model is struggling with the input |
| Audit error rate | Daily | Extractions are wrong, even when valid |
| Thumbs down | Hours to days | Customers are unhappy |
| p95 latency | At once | Replies are slowing |
The audit is the most important, because invalid JSON can stay low while valid-looking extractions are wrong. But it's also the noisiest, because each day's audit is small.
From alert to cause
When a limit is breached, ask what changed on that day: a release, a new channel, a new kind of customer. Then compare the affected and unaffected messages.
Example
The metrics, with limits from the first 14 days:
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/assistant/"
daily = pd.read_csv(base + "daily.csv", parse_dates=["date"])
daily["invalid_rate"] = daily["invalid_json"] / daily["messages"]
daily["audit_error_rate"] = daily["audit_field_errors"] / daily["audited"]
daily["thumbs_down_rate"] = daily["thumbs_down"] / daily["messages"]
metrics = ["invalid_rate", "audit_error_rate", "thumbs_down_rate", "p95_latency_ms"]
baseline = daily.iloc[:14]
limits = baseline[metrics].mean() + 3 * baseline[metrics].std()
print("Upper limits:", limits.round(4).to_dict(), "\n")
for m in metrics:
breaches = daily.loc[daily[m] > limits[m], "date"]
first = breaches.min().date() if len(breaches) else "none"
print(f"{m:17} days above the limit: {len(breaches):2} first: {first}")Upper limits: {'invalid_rate': 0.0121, 'audit_error_rate': 0.1732, 'thumbs_down_rate': 0.0463, 'p95_latency_ms': 3726.4217}
invalid_rate days above the limit: 13 first: 2026-08-18
audit_error_rate days above the limit: 4 first: 2026-08-26
thumbs_down_rate days above the limit: 13 first: 2026-08-17
p95_latency_ms days above the limit: 0 first: noneThumbs down and invalid JSON break out first, a day or two after 17 August, and stay above their limits. The audit error rate, the most important signal, breaches only from 26 August, nine days later. With just 40 checks a day it's noisy, so its limit is wide and only a large rise crosses it. A bigger daily audit, or a weekly rate pooled from the daily checks, would detect the change sooner. Latency never breaches. What changed?
daily["week"] = (daily.index // 7) + 1
daily.groupby("week")[["voice_note_share", "invalid_rate", "audit_error_rate", "thumbs_down_rate", "p95_latency_ms"]].mean().round(3)voice_note_share invalid_rate audit_error_rate thumbs_down_rate p95_latency_ms
week
1 0.000 0.009 0.064 0.038 3141.143
2 0.000 0.007 0.064 0.039 3212.857
3 0.155 0.016 0.118 0.050 3181.286
4 0.317 0.023 0.182 0.063 3343.286Voice notes started in week 3, and every quality metric moved with their share. Transcripts have no punctuation, more Pidgin and more spoken fillers. The assistant was never tested on them, so the release gate didn't cover them. The response:
- Contain: route voice-note messages to a person (or ask the customer to confirm the extracted details) until fixed.
- Measure: add a few hundred voice-note transcripts, with gold labels, to the evaluation set.
- Fix and gate: change the prompt (or add a clean-up step) and rerun the full release gate, including the new transcripts.
- Prevent: any new input channel must go through the AI team before launch.
Walkthrough
- Run the cells.
- Plot the audit error rate by day with its limit line and mark day 15.
- Write the incident note: what happened, when it was detected, the impact, the cause and the actions.
- Plan the presentation for Shieldline's leadership: five slides at most.
- Open the project brief on the course page and plan your submission.
Practice
Practice
On how many of the 28 days was the audit error rate above its upper limit?
Practice
What was the average audit error rate in week 4? One decimal place, as a percentage.
Task
12 minWrite the executive summary for Shieldline's leadership (120 to 230 words): what the assistant does, the evidence it's ready (the gate), the safeguards, the cost, the voice-note incident and what you did, and the next steps.
Your work is checked for
- Says what the assistant does
- Uses numbers
- Cites the release gate
- Safeguards (escalation, guardrail, rules, never)
- Gives the cost in naira
- Covers the voice-note incident
- Next steps
- Between 120 and 230 words
Check your understanding
Answer every question to check.