Module 6 · AI Agents and Tool Use
Evaluating agents
Measure an agent on outcomes (did it take the right final action?) and on trajectories (did it get there safely?), break the results down by request type, and read the failures behind the numbers.
About 25 minutes
The problem
"v2 is better" isn't a decision a head of support can sign off. They need to know: how often is it right? On which kinds of request does it fail? When it's wrong, is it wrong safely (handing over to a person) or dangerously (acting when it shouldn't)? And did it reach the right answer the right way, or by luck?
Agents need evaluation on two levels: the outcome and the path.
The concept
Outcome evaluation
Compare the agent's final action with the expected action on a labelled set of real requests. Report it:
- overall, and by expected action (fraud and refunds matter more than limit questions);
- with a confusion table: what the agent did instead, when it was wrong.
Trajectory evaluation
Check the steps, not just the end:
- Did it look before acting (read the transfer before opening a case)?
- Did it call any forbidden tools?
- Did it touch data it shouldn't have?
- How many steps did it take?
A run can reach the right outcome by a dangerous path (opening a case without checking, which happened to be correct). Trajectory checks catch that.
Safe and unsafe errors
An over-cautious hand-over to a person costs a little staff time. A wrong action (a refund, a case for someone else's transfer) can cost money or trust. Count them separately.
Example
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/agents/"
requests = pd.read_csv(base + "requests.csv")
runs = pd.read_csv(base + "runs.csv")
steps = pd.read_csv(base + "steps.csv")
r = runs.merge(requests[["request_id", "expected_action"]], on="request_id")
r["correct"] = r["final_action"] == r["expected_action"]
print(r.groupby("version")["correct"].mean().round(3))
r.pivot_table(index="expected_action", columns="version", values="correct", aggfunc="mean").round(2)version
v1 0.693
v2 0.933
Name: correct, dtype: float64
version v1 v2
expected_action
answer 0.54 0.96
ask_for_details 0.25 0.83
escalate_fraud 0.78 1.00
escalate_human 0.80 1.00
freeze_card 0.88 0.82
open_transfer_case 0.94 0.90Now the confusion: when each version was wrong, what did it do instead?
wrong = r[~r["correct"]]
pd.crosstab([wrong["version"], wrong["expected_action"]], wrong["final_action"])final_action answer escalate_human freeze_card issue_refund none open_transfer_case
version expected_action
v1 answer 0 0 3 1 0 22
ask_for_details 2 0 0 0 3 4
escalate_fraud 0 0 0 0 0 4
escalate_human 0 0 0 3 0 0
freeze_card 2 0 0 0 0 0
open_transfer_case 0 0 0 2 0 0
v2 answer 0 2 0 0 0 0
ask_for_details 0 2 0 0 0 0
freeze_card 0 3 0 0 0 0
open_transfer_case 0 3 0 0 0 0v2's errors are all hand-overs to a person: over-cautious, but safe. v1's errors include refunds, cases opened for transfers that needed none, and cases opened for other customers' transfers. Finally, a trajectory check: did the agent look at the transfer before opening a case?
def looked_first(run_steps):
tools = run_steps["tool"].tolist()
if "open_transfer_case" not in tools:
return None
return "get_transfer" in tools[: tools.index("open_transfer_case")]
check = steps.groupby("run_id").apply(looked_first, include_groups=False).dropna().rename("looked_first").reset_index()
check["version"] = check["run_id"].str[-2:]
check.groupby("version")["looked_first"].agg(["mean", "size"]).round(3)mean size
version
v1 0.694915 59
v2 1.0 31v1 opened 18 cases without ever looking at the transfer, and 9 of those happened to be correct: right by luck, and the kind of behaviour that goes wrong on the next request.
Walkthrough
- Run the cells. Read v2's ten wrong runs in full. What would you change to fix them?
- Count each version's errors as safe (a hand-over or asking for details) or unsafe (any other wrong action).
- Add a trajectory check: did v2 always call
check_reversal_eligibilitybeforeopen_transfer_case? - Write the evaluation summary for the head of support (the task below).
Practice
Practice
What share of v2 runs ended with the correct action? As a percentage, one decimal place.
Practice
What share of v1's case openings came after a get_transfer call? As a percentage, one decimal place.
Task
6 minWrite the evaluation summary for the head of support (60 to 150 words): each version's accuracy, how their errors differ (safe or unsafe), at least one trajectory finding, and your recommendation.
Your work is checked for
- At least two percentages
- Distinguishes safe and unsafe errors
- A trajectory finding (looked, checked, before, steps)
- A recommendation
- Between 60 and 150 words
Check your understanding
Answer every question to check.