AI Agent Governance: Rules, Escalation and Audit Trails for Customer-Facing Agents
How to govern AI agents that talk to customers: approved content rules, escalation paths, logging, human-in-the-loop review and a monthly QA routine.
A voice agent script template organized by section, sample lines to adapt, the anti-patterns that break prompts, and a test-call plan that runs before launch.
A voice agent script is not a monologue the agent reads aloud. It is the policy the model consults on every turn: who it is, what it may do, how it speaks, what happens when a call goes sideways, and which words trigger a transfer. Get the structure right and the agent sounds composed; skip it and you get a polite robot that talks too much and still books the wrong day.
This playbook gives you a section-by-section template, sample lines you can adapt, the anti-patterns that reliably break voice prompts, and a test-call plan - simulation first, real calls second - so failures happen in testing rather than in front of a customer.
| Section | What goes in it | Common mistake |
|---|---|---|
| Identity and role | Who the agent is, who it works for, its one job | Writing a biography the caller never hears |
| Scope and guardrails | What it can and cannot handle; hard limits | Leaving scope implicit until a caller finds the edge |
| Speaking style | Tone, brevity, spoken forms for numbers and dates | Chat-style formatting rules copied from a text bot |
| Call flow | The happy path, step by step, one question at a time | Embedding three branches into every step |
| Escalation | Exact triggers and phrases that move the call to a human or fallback | Making the human path a last resort |
| Data capture | Which fields to collect, confirm and write back | Asking for everything at once, confirming nothing |
| Examples | Happy path, edge case, tool failure | Skipping examples because the instructions sound clear |
Vapi's guide to voice prompts organizes production prompts as identity, response guidelines, guardrails, context, workflow and examples, and notes that each section exists because text-chat habits fail on the phone. The two extra rows above - escalation and capture - are where service businesses get burned most, so they deserve their own sections.
Write yours against this order. Keep it to about a page before examples.
Two habits keep this template healthy. First, refer to tools by what they do ("book the appointment"), never by internal names that could leak into speech. Second, when something fails, revise the instruction that should have prevented it rather than appending a new prohibition.
Illustrative phrasing, not a script to copy word for word. Adjust names and policies.
Greeting and intent:
"Thanks for calling Bellweather Plumbing. This is Avery. Are you calling about a new job, an existing appointment, or something urgent?"
AI disclosure when asked:
"Yes - I'm an AI assistant. I can book this right now and the office will confirm, or I can get you to a person. Which would you prefer?"
One-question collection:
"What's the best mobile number for the technician to text?"
Read-back before booking:
"So that's Thursday the ninth, between one and three, at the Oak Street address - did I get that right?"
Booking confirmation:
"You're on the schedule for Thursday, April ninth, between one and three. You'll get a text confirmation in a moment."
Tool failure fallback:
"I'm having trouble reaching the schedule right now. I can take your details and have the office call you back within the hour, or transfer you now."
Emergency recognition:
"That sounds urgent, so I'm getting you to a live person right now. If you smell gas, please step outside first."
After-hours choice:
"The office is closed, but I can book you for tomorrow morning or pass a message to the on-call technician. Which is better?"
Transfer with context:
"Let me get you to the team - I'll pass along everything you've told me so you don't have to repeat it."
Closing:
"You're all set. Anything else before we hang up?"
Two layers: graded simulations that run before launch, then live calls that prove the stack. Retell's simulation testing plays the caller with an AI, grades each run pass or fail against success criteria, and mocks functions so a test never creates a real booking. Rerun the suite after every script or flow change, and judge a scenario by its pass rate across runs, since the simulated caller and the grader are also models.
| # | Scenario | What it tests | Pass criteria |
|---|---|---|---|
| 1 | Clean booking | Happy path end to end | Booking created with correct day, time and contact details |
| 2 | Reschedule | Calendar write and confirmation | Old slot released, new slot confirmed aloud |
| 3 | After-hours call | Availability and fallback | Caller never hits a dead end; booking or message offered |
| 4 | Wrong number or spam | Fast exit | Call ends politely and quickly, no lead created |
| 5 | Interruption mid-sentence | Turning-taking and barge-in | Agent stops, listens, responds to the new question |
| 6 | Noisy call or poor audio | Repeat handling | Agent asks conversationally, transfers after two failures |
| 7 | Strong accent from your market | Recognition limits | Booking completes or transfer happens fast, no mis-booked address |
| 8 | "Am I talking to a robot?" | Disclosure | Clear answer, human offered, call continues naturally |
| 9 | "I want a person now" | Escalation speed | Transfer within one turn, context passed along |
| 10 | Angry caller about price | Boundaries and de-escalation | No invented discounts, transfer or callback offered |
| 11 | Emergency phrasing | Urgency recognition | Immediate transfer with the safety line spoken first |
| 12 | Tool failure | Fallback behavior | Message captured, callback promised, no dead air |
| 13 | Correction mid-call | State updates | Read-back uses the corrected detail only |
| 14 | Silent caller | Turn management | One check-in, then follow the scripted path |
How to run it:
Listen to the first fifty to one hundred calls in full before trusting any dashboard. Once a week, review transfers and abandoned calls, and add each failure to the test suite. Keep the prompt short as the suite grows by pushing stable answers into the knowledge base and letting tools do heavy lifting. Disclosures, consent lines and recording rules deserve the same discipline as the sales script; the AI agent governance checklist covers those requirements, and the lead qualification questions guide is the reference for what the intake section should ask. If the agent will text callers after the call, run those message templates past the SMS compliance rules first.
This is the same discipline we apply in the AI voice receptionist workflow on client builds, and the voice agent guide goes deeper on platform mechanics. If you would rather not maintain the machinery, the AI voice agent service exists for exactly that.
How long should an AI voice agent script be?
Short enough to follow and short enough to keep latency low. Platform guidance says keep prompts concise because length costs time to first response and can reduce instruction-following. Aim for a page of structured sections plus examples; if it passes roughly 1,000 words or covers more than five tools, split the work into states or sub-agents.
Should the script disclose that callers are talking to an AI?
Disclose at least when asked, and check your obligations. Retell's Agent Handbook ships an AI Disclosure When Asked preset enabled by default, which is a reasonable floor. Build the disclosure line into the script and the caller experience stays honest without turning the greeting into a disclaimer.
How do you test a voice agent before going live?
Run simulated test cases that grade pass or fail, then place real test calls. Retell's simulation testing lets an AI play the caller and checks success criteria, with function mocks so tests do not create real bookings. After simulation passes, run a dozen live calls covering noisy audio, interruptions and upset callers, and rerun the suite after every script change.
What makes voice scripts fail?
Text-chat monologues, piles of do-not rules added after each failure, multiple questions per turn, vague tool descriptions, and treating the prompt as a security boundary. Long negative lists also prime the model toward the banned behavior, which is why platform guides recommend short positive instructions over exhaustive prohibitions.
If you want the script written, tested and maintained for you, start at the Praktivo funnel or book a call with a list of your most common call types. We build the agent, run the test plan and hand over a maintained system - see how the service is scoped and what the voice receptionist workflow looks like when it ships.
How to govern AI agents that talk to customers: approved content rules, escalation paths, logging, human-in-the-loop review and a monthly QA routine.
An honest comparison of AI appointment setters and human schedulers: what each does well, where each falls short, and the hybrid setup that wins.
What an AI receptionist does for a small service business in 2026, what it really costs, where it fails, and a 30-day rollout plan you can run yourself.
More articles: browse the full Praktivo blog.