Skip to content
Workflows Resources Case Studies Pricing About
Playbooks

AI Voice Agent Script Template: Write It and Test It

A voice agent script template organized by section, sample lines to adapt, the anti-patterns that break prompts, and a test-call plan that runs before launch.

By Ahmad TawfikPublished 10 min read

A voice agent script is not a monologue the agent reads aloud. It is the policy the model consults on every turn: who it is, what it may do, how it speaks, what happens when a call goes sideways, and which words trigger a transfer. Get the structure right and the agent sounds composed; skip it and you get a polite robot that talks too much and still books the wrong day.

This playbook gives you a section-by-section template, sample lines you can adapt, the anti-patterns that reliably break voice prompts, and a test-call plan - simulation first, real calls second - so failures happen in testing rather than in front of a customer.

Key takeaways

  • Production voice prompts follow a structure: identity, guardrails, speaking style, call flow, escalation, data capture and examples. Vapi's prompting guide frames it as six required sections; phone work adds explicit escalation and capture rules.
  • Voice has constraints text does not: every token adds latency, long answers become monologues, and turn-taking replaces scrolling.
  • Keep prompts short and state each instruction once; Retell's guidance warns that longer prompts reduce reasoning quality and increase response latency.
  • Never pile "do not" rules on top of every failure. Fix the instruction that caused the failure instead, and prefer short positive rules over long banned lists.
  • Disclose AI when asked - it is the default expectation on modern platforms and the honest baseline for callers.
  • Test with graded simulations first, then live calls. Retell's simulation testing lets an AI play the caller and score each scenario pass or fail, with function mocks so tests never create real bookings.

The anatomy: seven sections that do the work

SectionWhat goes in itCommon mistake
Identity and roleWho the agent is, who it works for, its one jobWriting a biography the caller never hears
Scope and guardrailsWhat it can and cannot handle; hard limitsLeaving scope implicit until a caller finds the edge
Speaking styleTone, brevity, spoken forms for numbers and datesChat-style formatting rules copied from a text bot
Call flowThe happy path, step by step, one question at a timeEmbedding three branches into every step
EscalationExact triggers and phrases that move the call to a human or fallbackMaking the human path a last resort
Data captureWhich fields to collect, confirm and write backAsking for everything at once, confirming nothing
ExamplesHappy path, edge case, tool failureSkipping examples because the instructions sound clear

Vapi's guide to voice prompts organizes production prompts as identity, response guidelines, guardrails, context, workflow and examples, and notes that each section exists because text-chat habits fail on the phone. The two extra rows above - escalation and capture - are where service businesses get burned most, so they deserve their own sections.

The template, filled in structure

Write yours against this order. Keep it to about a page before examples.

  1. Identity and role. One or two sentences. Example shape: you are a booking assistant for a plumbing company; your job is to book service visits, answer scheduling questions and route everything else.
  2. Scope and guardrails. What the agent may do (book, reschedule, share approved answers), what it must never do (quote prices, promise arrival times, give advice), and the privacy line: no payment card numbers, no ID numbers, no health details.
  3. Speaking style. Short turns, one or two sentences, contractions, no lists or formatting in speech, numbers and dates spoken naturally, names and phone numbers read back for confirmation.
  4. Call flow. Numbered steps: greeting and intent, collect the minimum details, check availability, offer at most two options, confirm, book, close. One question per turn. If the caller volunteers information early, do not ask again.
  5. Escalation. The exact triggers, phrased as observable conditions: the caller asks for a person, the caller mentions an emergency, the request falls outside scope, the same question was misunderstood twice, or a required tool fails twice. Name the action for each trigger: transfer, message, or text-back link.
  6. Data capture. The fields to collect, when to read each back, and what to send to the CRM at the end. Ambient platforms can capture fields incrementally as the call runs, so a dropped call still leaves data behind.
  7. Examples. At least three short transcripts: a clean booking, an edge case such as no availability, and a tool failure with the fallback spoken aloud. Vapi's guide recommends exactly these three shapes.

Two habits keep this template healthy. First, refer to tools by what they do ("book the appointment"), never by internal names that could leak into speech. Second, when something fails, revise the instruction that should have prevented it rather than appending a new prohibition.

Sample lines to adapt

Illustrative phrasing, not a script to copy word for word. Adjust names and policies.

Greeting and intent:

"Thanks for calling Bellweather Plumbing. This is Avery. Are you calling about a new job, an existing appointment, or something urgent?"

AI disclosure when asked:

"Yes - I'm an AI assistant. I can book this right now and the office will confirm, or I can get you to a person. Which would you prefer?"

One-question collection:

"What's the best mobile number for the technician to text?"

Read-back before booking:

"So that's Thursday the ninth, between one and three, at the Oak Street address - did I get that right?"

Booking confirmation:

"You're on the schedule for Thursday, April ninth, between one and three. You'll get a text confirmation in a moment."

Tool failure fallback:

"I'm having trouble reaching the schedule right now. I can take your details and have the office call you back within the hour, or transfer you now."

Emergency recognition:

"That sounds urgent, so I'm getting you to a live person right now. If you smell gas, please step outside first."

After-hours choice:

"The office is closed, but I can book you for tomorrow morning or pass a message to the on-call technician. Which is better?"

Transfer with context:

"Let me get you to the team - I'll pass along everything you've told me so you don't have to repeat it."

Closing:

"You're all set. Anything else before we hang up?"

Anti-patterns that break voice scripts

  • The text-chat monologue. Multi-sentence paragraph answers work in chat and fail out loud. Vapi's guide recommends one to two sentences per turn; callers remember less than you think.
  • The do-not pile-up. A new "never say X" rule after each failure bloats the prompt, and long negative lists can prime the very behavior they ban. Retell's guide says to fix the underlying instruction instead.
  • Three questions at once. Name, address and problem in one turn guarantees a confused answer. Collect one field, confirm, move on.
  • Length as thoroughness. Longer prompts cost time to first response and can reduce instruction-following. Stable detail belongs in the knowledge base, not the prompt.
  • Prompt as security boundary. Verification the agent must not fake belongs in server-side tools. Prompts are probabilistic; a determined caller can talk around them.
  • Vague tool descriptions. If the agent calls the wrong tool or none at all, fix the tool description before editing the script.
  • No examples. Instructions tell the model what you want; examples show it. Include the happy path, an edge case and an error recovery.

The test-call plan

Two layers: graded simulations that run before launch, then live calls that prove the stack. Retell's simulation testing plays the caller with an AI, grades each run pass or fail against success criteria, and mocks functions so a test never creates a real booking. Rerun the suite after every script or flow change, and judge a scenario by its pass rate across runs, since the simulated caller and the grader are also models.

#ScenarioWhat it testsPass criteria
1Clean bookingHappy path end to endBooking created with correct day, time and contact details
2RescheduleCalendar write and confirmationOld slot released, new slot confirmed aloud
3After-hours callAvailability and fallbackCaller never hits a dead end; booking or message offered
4Wrong number or spamFast exitCall ends politely and quickly, no lead created
5Interruption mid-sentenceTurning-taking and barge-inAgent stops, listens, responds to the new question
6Noisy call or poor audioRepeat handlingAgent asks conversationally, transfers after two failures
7Strong accent from your marketRecognition limitsBooking completes or transfer happens fast, no mis-booked address
8"Am I talking to a robot?"DisclosureClear answer, human offered, call continues naturally
9"I want a person now"Escalation speedTransfer within one turn, context passed along
10Angry caller about priceBoundaries and de-escalationNo invented discounts, transfer or callback offered
11Emergency phrasingUrgency recognitionImmediate transfer with the safety line spoken first
12Tool failureFallback behaviorMessage captured, callback promised, no dead air
13Correction mid-callState updatesRead-back uses the corrected detail only
14Silent callerTurn managementOne check-in, then follow the scripted path

How to run it:

  1. Build the cases in your platform's test tool and mock every function that touches production.
  2. Run the suite before launch and fix failures by editing instructions, not by adding prohibitions.
  3. When simulation passes, place live test calls from at least two phones, including one noisy location.
  4. Have someone play a frustrated caller and someone play an accented caller; record both.
  5. Fix, then rerun the full suite - a fix for one case commonly breaks another.
  6. After launch, convert every real failed call into a new test case so the suite grows from reality.

After launch: the maintenance loop

Listen to the first fifty to one hundred calls in full before trusting any dashboard. Once a week, review transfers and abandoned calls, and add each failure to the test suite. Keep the prompt short as the suite grows by pushing stable answers into the knowledge base and letting tools do heavy lifting. Disclosures, consent lines and recording rules deserve the same discipline as the sales script; the AI agent governance checklist covers those requirements, and the lead qualification questions guide is the reference for what the intake section should ask. If the agent will text callers after the call, run those message templates past the SMS compliance rules first.

This is the same discipline we apply in the AI voice receptionist workflow on client builds, and the voice agent guide goes deeper on platform mechanics. If you would rather not maintain the machinery, the AI voice agent service exists for exactly that.

FAQ

How long should an AI voice agent script be?

Short enough to follow and short enough to keep latency low. Platform guidance says keep prompts concise because length costs time to first response and can reduce instruction-following. Aim for a page of structured sections plus examples; if it passes roughly 1,000 words or covers more than five tools, split the work into states or sub-agents.

Should the script disclose that callers are talking to an AI?

Disclose at least when asked, and check your obligations. Retell's Agent Handbook ships an AI Disclosure When Asked preset enabled by default, which is a reasonable floor. Build the disclosure line into the script and the caller experience stays honest without turning the greeting into a disclaimer.

How do you test a voice agent before going live?

Run simulated test cases that grade pass or fail, then place real test calls. Retell's simulation testing lets an AI play the caller and checks success criteria, with function mocks so tests do not create real bookings. After simulation passes, run a dozen live calls covering noisy audio, interruptions and upset callers, and rerun the suite after every script change.

What makes voice scripts fail?

Text-chat monologues, piles of do-not rules added after each failure, multiple questions per turn, vague tool descriptions, and treating the prompt as a security boundary. Long negative lists also prime the model toward the banned behavior, which is why platform guides recommend short positive instructions over exhaustive prohibitions.

Next step

If you want the script written, tested and maintained for you, start at the Praktivo funnel or book a call with a list of your most common call types. We build the agent, run the test plan and hand over a maintained system - see how the service is scoped and what the voice receptionist workflow looks like when it ships.

Frequently asked questions

How long should an AI voice agent script be?
Short enough to follow and short enough to keep latency low. Platform guidance says keep prompts concise because length costs time to first response and can reduce instruction-following. Aim for a page of structured sections plus examples; if it passes roughly 1,000 words or covers more than five tools, split the work into states or sub-agents.
Should the script disclose that callers are talking to an AI?
Disclose at least when asked, and check your obligations. Retell's Agent Handbook ships an AI Disclosure When Asked preset enabled by default, which is a reasonable floor. Build the disclosure line into the script and the caller experience stays honest without turning the greeting into a disclaimer.
How do you test a voice agent before going live?
Run simulated test cases that grade pass or fail, then place real test calls. Retell's simulation testing lets an AI play the caller and checks success criteria, with function mocks so tests do not create real bookings. After simulation passes, run a dozen live calls covering noisy audio, interruptions and upset callers, and rerun the suite after every script change.
What makes voice scripts fail?
Text-chat monologues, piles of do-not rules added after each failure, multiple questions per turn, vague tool descriptions, and treating the prompt as a security boundary. Long negative lists also prime the model toward the banned behavior, which is why platform guides recommend short positive instructions over exhaustive prohibitions.
Keep reading

Related articles

Get Your AI Automation Plan