An AI agent doesn't just answer; it acts, by calling tools in a loop. That makes it useful and dangerous. In this course you take apart two versions of Paystream's support agent, recorded step by step on 150 real requests that a support lead labelled. You'll design tools scoped to the logged-in customer, move business rules out of the prompt into tested code, and write a loop with limits, repeated-call detection and a trace. You'll find the refunds v1 issued without approval and the other customers' transfers it read, evaluate both versions on outcomes and on the paths they took, and see how instructions hidden in transfer narrations hijacked v1. Finally, you'll work out cost per correct resolution, design monitoring and a gradual rollout, and write the design for a safer v3. Every number comes from running the code on the recorded runs; live model calls are optional and shown with the Anthropic SDK.