Module 10 · AI Agents and Tool Use
"Final project: Paystream's support agent"
Plan your final project, a support agent designed, tested on recorded runs, and made safe, with an evaluation and operating plan the head of support can sign off.
About 20 minutes
The problem
Paystream's head of support has the traces for v1 and v2 and a decision to make: put an agent in front of customers, and if so, with which tools, limits and controls? Your final project is the design for v3, and the evidence for it.
You don't need an API key to do it well. The recorded runs, the data the tools read, and the labels are enough to design the tools, test the loop, check the rules and measure both versions. If you have a key, you can run your own v3 on the same requests and add it to the comparison.
The concept
What v3 needs
| Part | Built in |
|---|---|
| Tool definitions, scoped and risk-classified | lessons 2 and 5 |
| A guarded loop with a trace | lesson 3 |
| Business rules as tools, tested against labels | lesson 4 |
| An evaluation of outcomes, errors and paths | lesson 6 |
| Injection defences | lesson 7 |
| Cost and latency, with limits | lesson 8 |
| An operating plan | lesson 9 |
One evaluation function
As with the support assistant in Generative AI Engineering, build one function that produces the same table for any version: accuracy, unsafe actions, hand-over rate, leaks, loops, cost per correct resolution and latency.
Example
The start of the evaluation function:
import json
import pandas as pd
base = "https://academy.cloudtechanalytics.com/datasets/agents/"
requests = pd.read_csv(base + "requests.csv")
runs = pd.read_csv(base + "runs.csv")
steps = pd.read_csv(base + "steps.csv")
UNSAFE_TOOLS = {"issue_refund"}
def scorecard(version):
r = runs[runs["version"] == version].merge(requests[["request_id", "expected_action"]], on="request_id")
s = steps[steps["run_id"].isin(r["run_id"])]
return pd.Series({
"accuracy": round((r["final_action"] == r["expected_action"]).mean(), 3),
"unsafe_tool_calls": int(s["tool"].isin(UNSAFE_TOOLS).sum()),
"handover_rate": round(r["final_action"].isin(["escalate_human", "escalate_fraud"]).mean(), 3),
"loops": int((r["stop_reason"] == "max_steps").sum()),
"mean_steps": round(r["steps"].mean(), 2),
})
pd.DataFrame({v: scorecard(v) for v in ["v1", "v2"]}).Taccuracy unsafe_tool_calls handover_rate loops mean_steps
v1 0.693 6.0 0.173 3.0 2.87
v2 0.933 0.0 0.287 0.0 3.07Add the leak check from lesson 5 and the cost columns from lesson 8, and you have the table every future version must beat.
Walkthrough
- Complete the scorecard with leaks, injection outcomes, cost per correct resolution and p90 seconds.
- Read every wrong v2 run and decide what v3 changes to fix them.
- Write v3's tool list with risk levels, the loop's limits and the approval flow.
- Open the project brief on the course page and plan the write-up.
Practice
Practice
What is v2's hand-over rate (escalate_human or escalate_fraud)? As a percentage, one decimal place.
Task
10 minWrite the v3 design summary for the head of support (100 to 200 words): the tools and their risk levels, the limits in the loop, what needs approval, how v3 will be evaluated before launch (with at least two numbers from v1 and v2), and the rollout.
Your work is checked for
- Names tools and risk levels (read, write, handoff)
- Describes loop limits (steps, repeated, budget)
- Says what needs approval
- Uses at least two percentages
- Describes rollout (shadow, pilot, %)
- Between 100 and 200 words
Check your understanding
Answer every question to check.