Module 9 · AI Agents and Tool Use
Monitoring agents in production
Decide what to log, which numbers to watch every day, when to alert a person, and how to roll an agent out gradually, so that problems are found by your dashboard, not your customers.
About 20 minutes
The problem
Passing an evaluation once doesn't keep an agent safe. After launch, new kinds of requests arrive, a tool's data changes, the model provider updates the model, someone edits the prompt. Any of these can quietly change what the agent does.
Paystream's head of support asks for the operating plan: what will be watched, who looks at it, what triggers an alarm, and what happens when one goes off.
The concept
Log every step
The trace format from this course (run, step, tool, arguments, result, tokens, seconds), plus the prompt version, model version and final action. Without traces, you can't investigate a complaint or an incident.
Watch daily
| Measure | Why |
|---|---|
| Share of runs by final action | a sudden rise in case openings or hand-overs signals a change |
| Hand-over rate | the main safety valve; too high wastes staff, too low may mean over-confidence |
| Tool error and not_found rates | data or integration problems |
| Step-limit stops and repeated calls | loops |
| Tokens and seconds per run | cost and customer wait |
| Injection flags | attacks |
Sample and review
People review a random sample of runs every week, graded with the same labels as the evaluation set, so accuracy is measured on live traffic, not just the test set.
Roll out gradually
Start with shadow mode (the agent proposes, people act), then a small share of live requests, then more, with a kill switch that sends everything back to people instantly.
Example
A daily monitoring table from the traces (here, all the v2 runs treated as one day's traffic):
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/agents/"
runs = pd.read_csv(base + "runs.csv")
steps = pd.read_csv(base + "steps.csv")
def daily_report(version):
r = runs[runs["version"] == version]
s = steps[steps["run_id"].isin(r["run_id"])]
return pd.Series({
"runs": len(r),
"handover_rate": round(r["final_action"].isin(["escalate_human", "escalate_fraud"]).mean(), 3),
"case_rate": round((r["final_action"] == "open_transfer_case").mean(), 3),
"tool_error_rate": round((s["result"] == "error").mean(), 3),
"step_limit_stops": int((r["stop_reason"] == "max_steps").sum()),
"mean_steps": round(r["steps"].mean(), 2),
"p90_seconds": round(r["seconds"].quantile(0.9), 1),
})
pd.DataFrame({v: daily_report(v) for v in ["v1", "v2"]})v1 v2
runs 150.000 150.000
handover_rate 0.173 0.287
case_rate 0.393 0.187
tool_error_rate 0.009 0.007
step_limit_stops 3.000 0.000
mean_steps 2.870 3.070
p90_seconds 9.900 12.300Look at the case rate. v1 opened cases on a far larger share of requests than v2. If v2 were live and its case rate suddenly moved towards v1's, that alone would be a reason to investigate, before anyone had graded a single run.
Walkthrough
- Run the cell. Which measure would have caught v1's problems fastest?
- Set an alert threshold for each measure, based on v2's values.
- Write the operating plan (the task below).
- Decide who owns the kill switch, and how quickly it must work.
Practice
Practice
What share of v1 runs ended with open_transfer_case? As a percentage, one decimal place.
Task
8 minWrite the operating plan for running v2 live, one line each starting Logging:, Daily checks:, Alerts:, Review:, Rollout: and Kill switch:.
Your work is checked for
- A Logging line
- A Daily checks line naming measures
- An Alerts line with a threshold
- A Review line with a sample
- A Rollout line that starts small (shadow, pilot, %)
- A Kill switch line
Check your understanding
Answer every question to check.