Write an AI Agent SOP: The One-Page Template for Owners
Every agent task needs a written page covering purpose, triggers, limits, approvals, and review. Includes a filled inbox example plus a blank template to copy.
Not every agent action needs a human click. Learn the four approval tiers, who approves what, and how to set response times your team can keep.
An AI agent without approvals is an intern with your credit card and no curfew. It might do fine work for weeks, then send the wrong message to two hundred people on a Friday night. Approvals are how you keep the speed of automation without handing over judgment on the decisions that matter.
The good news is that every major agent platform now ships approval controls. ChatGPT Dots lets actions be allowed, blocked, or set to require approval. Claude Cowork uses per-task approvals with an admin-controlled auto-approve mode. Meta Muse keeps working in the background and returns when approval is needed, for example before purchases. This guide turns those platform features into a workflow your team can actually run.
Every action your agent can take belongs in one of four tiers. Write the list down before you switch the agent on, because defaults decide what happens at 11 pm when nobody is watching.
Auto-run covers reversible, low-risk, low-visibility actions: drafting a message, summarizing a thread, scoring a lead, updating an internal note. Notify covers actions worth seeing but not worth stopping: a booked appointment, a sent confirmation, a filed call summary. The agent acts, and a person gets a message with enough context to spot a problem. Approve first covers anything hard to undo or visible to customers: quotes over a threshold, refunds, bulk messages, deletions, published posts. Blocked covers what the agent must never do: change passwords, move money beyond a cap, export full customer lists, or touch payroll.
ChatGPT Dots maps to this model directly, with actions allowed, blocked, or requiring approval, and background research running read-only. That read-only default is worth copying everywhere: let the agent gather and draft freely, but require a click before it changes a system of record. If your platform supports an Activity View or equivalent, use it during the first two weeks to watch what the agent attempts, then set the tiers from evidence rather than guesses. Our agent SOP template gives you a one-page place to record these tiers, and the first-90-days plan shows how to loosen them gradually.
| Tier | Meaning | Examples | Review habit |
|---|---|---|---|
| Auto-run | Acts alone | Draft replies, summaries, lead scoring, internal notes | Weekly log scan |
| Notify | Acts, then tells you | Booked appointments, confirmations, filed summaries | Daily glance |
| Approve first | Waits for a click | Quotes over threshold, refunds, bulk sends, deletions | Every request decided in target time |
| Blocked | Never allowed | Password changes, payroll, full-list exports | Quarterly re-confirm |
An approval workflow with no named approver is a suggestion. For each workflow, name one approver who understands the consequence and one backup for days off. In a ten-person service business this is usually simple: the owner or office manager approves money and customer-facing messages, a lead technician approves job-scope promises, a bookkeeper approves invoice actions.
Then set a response target, sometimes called an SLA, that your staffing can keep. Thirty minutes during staffed hours is realistic for most shops; evenings and weekends need a different rule, such as queue customer-visible actions until morning while letting internal drafts continue. Publish the target where the team sees it, and track whether approvals actually land inside it. A workflow that waits four hours for a click is a workflow the team will start bypassing, which is worse than a slower tier designed honestly.
Escalation needs one sentence in the SOP: if the approver does not respond in the target time, the action waits, and the customer gets a holding message. Never escalate by auto-approving. The safe fallback for an unanswered approval is delay plus communication, not silent execution. If after-hours coverage matters to you, the speed-to-lead SMS workflow shows how immediate acknowledgment plus a morning decision beats either silence or unsupervised action.
You do not need to take our word for the pattern; it is visible in how vendors build. OpenAI Dots, announced at Dev Day on September 29, 2026, runs each agent on its own cloud computer with connections to more than 4,000 apps, and lets users allow, block, or require approval per action. Sensitive tasks such as changing a password always stay with the user, credentials can be used without exposing passwords to the model, and Dots can be stopped if systems detect malicious instructions. An Activity View lets users watch and redirect work in progress.
Claude Cowork, Anthropic's workspace agent for analysts, lawyers, account staff, and marketers, takes a similar approach with per-task approvals and an Automatically approve mode that administrators can control. It runs tasks in the cloud, works with local files and connectors like Slack and Google Drive, and can take browser actions. Admin controls include role-based access, spend limits, and activity streaming. Meta Muse, launched September 8, 2026, keeps working in the background after the app closes and returns when approval is needed, for example before purchases.
The lesson for a buyer: ask every vendor where approvals live, whether they are per action or per task, who can change them, and where the log goes. If approvals are all-or-nothing, or if only the person who built the agent can change them, that is a limitation to price into your decision. The internal assistant service follows the tiered pattern described here, with supervised actions and reviewable records.
Start with one workflow, not five. Lead follow-up or inbox triage is ideal because the actions are familiar and the stakes are clear. List every action the agent will take: read new leads, draft a reply, send the reply, book a time, update the CRM, notify the owner. Assign each to a tier. Drafting and CRM notes are auto-run. Booking a standard appointment is notify. Sending a custom quote or a discount is approve first. Exporting the list or changing the calendar setup is blocked.
Next, define the approval message. An approver needs four things: what the agent wants to do, the exact content or amount, the customer and context, and one-tap approve or reject. A message that says approve requested action is not reviewable; a message that shows the draft text, the customer record, and the proposed time is. Rejections should loop back with a reason so the agent learns the boundary.
Then test with ten realistic cases before going live, including two adversarial ones: a refund above the threshold and a bulk send to the whole list. Confirm both wait for approval and that the holding message to the customer reads naturally. Run the workflow supervised for two weeks, review the approval log weekly, and only then widen auto-run to the actions that proved boring. This sequence, narrow start, evidence, gradual widening, is also how Anthropic frames Cowork rollout guidance with its maturity model, and it works at any company size.
The most common failure is approval fatigue. If approvers get forty pings a day, they start tapping approve without reading, and the workflow becomes theater. The fix is tiering discipline: move genuinely low-risk actions down to notify, batch non-urgent approvals into two daily windows, and keep approve-first for the short list that deserves attention. Fewer, better approvals beat blanket coverage.
The second failure is the missing backup. One approver goes on holiday, requests pile up, and someone disables approvals to unblock the queue. Every workflow needs a named backup with the same context access.
The third failure is silent auto-approve. An admin switches on blanket auto-approval to stop the pings, nobody records the change, and the agent runs unsupervised for months. Treat auto-approve like a permission: scoped to named actions, time-boxed for the pilot, logged, and re-confirmed quarterly. small businesses already worry about accuracy and reliability of AI, with 73 percent of owners naming it a top concern in a 2026 Small Business Majority survey of 222 owners, so keeping a visible human gate on consequential actions is both good operations and good customer trust.
What is a human-in-the-loop approval workflow?
It is a rule that decides which agent actions run automatically, which notify a person, and which wait for explicit approval. ChatGPT Dots supports exactly this pattern: actions can be allowed, blocked, or set to require approval, with background research running read-only until a person signs off.
Which agent actions should always need approval?
Anything hard to undo: sending money, issuing refunds, deleting records, contacting a full customer list, changing credentials, or publishing in your name. Claude Cowork uses per-task approvals for this reason, and OpenAI keeps sensitive tasks like changing a password with the user rather than the agent.
Who should approve agent actions in a small business?
The person closest to the consequence, usually the owner or office manager for money and customer-facing messages, and a technician or bookkeeper for job-specific actions. Name one approver and one backup per workflow, set a response target such as 30 minutes in staffed hours, and log every decision.
How do auto-approve modes stay safe?
Scope them narrowly by action, amount, and customer segment, keep them read-only by default, and review the log weekly. Claude Cowork offers an Automatically approve mode that administrators can control, which is the right pattern: convenience switched on deliberately, for low-risk actions, with an audit record.
Pick one agent workflow and sort its actions into the four tiers this week. If you want a second pair of eyes on the tiering before you go live, book a call or start with the free six-step AI automation plan.
Every agent task needs a written page covering purpose, triggers, limits, approvals, and review. Includes a filled inbox example plus a blank template to copy.
A week-by-week rollout plan for your first 90 days with an AI agent: access, training, review cadence, and clear rules for when to expand or stop.
Every AI agent action should leave a trace. Learn what to log, how long to keep it, and a lightweight audit trail setup that fits a small business.
More articles: browse the full Praktivo blog.