Write an AI Agent SOP: The One-Page Template for Owners
Every agent task needs a written page covering purpose, triggers, limits, approvals, and review. Includes a filled inbox example plus a blank template to copy.
A week-by-week rollout plan for your first 90 days with an AI agent: access, training, review cadence, and clear rules for when to expand or stop.
Most AI agent projects do not fail on technology; they fail on pacing. Too much access too soon, no named owner, no review rhythm, and five tasks launched before the first one works. This plan gives your first agent ninety disciplined days: one task tamed in month one, access and training hardened in month two, and a measured expand-or-stop decision in month three. Follow it as written the first time; improvise on the second agent.
Start by writing the task card and naming the owner. The card states the task, its trigger, the single app involved, the expected output, the written target with a number, and the review schedule. The owner is the manager closest to the task, with authority to change instructions and access. Announce both to the team in plain terms: what the agent will do, what it will not touch, and who to tell when something looks wrong.
Configure access narrowly. The agent gets a scoped connection to one app, approvals required before anything reaches a customer or changes a record permanently, and a readable activity log. This mirrors, at small scale, what larger organizations formalize: Anthropic's enterprise guidance for tools like Claude Cowork describes staged rollout with role-based access, spend limits, and activity streaming, and Cowork itself runs cloud tasks with per-task approvals and connector scoping. You do not need enterprise software to copy the discipline; you need the settings turned on and an owner who checks them.
Run supervised tests in week two using real examples, including at least one awkward case such as an angry customer or an ambiguous request. The correct behavior for confusion is a flagged handoff, never a confident guess. Keep every correction in a one-page log with the date, what happened, and the brief sentence you changed. If the log shows the same failure three times, the task is too broad: halve it and continue. End day 14 with a go-or-narrow decision, written down and signed by the owner.
Move to live work with supervision: the agent handles real volume daily, and the owner reviews every output before it counts. This is the most labor-intensive fortnight of the quarter, and skipping it is the most common cause of later failure. The review burden should visibly shrink by week four; if it does not, the brief or the scope is wrong, not the reviewer.
Hold the weekly fifteen-minute review without exception. Same agenda each time: what the agent handled, what needed correction, whether any approval setting should change, and whether access should widen by one small step or stay put. Loosen exactly one restriction at a time, such as letting drafts go to internal staff without approval while keeping customer messages gated, and watch the log for two weeks before the next loosening. Restrictions removed in bundles cannot be evaluated; restrictions removed singly teach you what each one was protecting.
By day 30, compute the first numbers: volume handled, correction rate, and hours the owner and staff spent reviewing versus hours the task used to take. You are not proving ROI yet; you are checking direction. Corrections falling, review time falling, and output quality steady means month two can begin. Anything else means another two weeks on the same task with a narrower brief.
Month two has three parallel tracks. First, harden what month one built: confirm the agent's connections use dedicated credentials rather than anyone's personal login, review the access list and remove anything unused, and verify the log retention covers at least the quarter. Larger companies put this under agent identity management, giving each agent its own identity with lifecycle controls; Microsoft's Agent 365 control plane, generally available since May 2026, packages exactly that for organizations. Your version is a checklist and a calendar reminder, but it enforces the same idea: every agent has an identity, a scope, and an off switch.
Second, train the team. Show staff what the agent does, how to spot its mistakes, and how to escalate. Address the worry directly: the World Economic Forum's 2026 entry-level work report found only 16 percent of organizations have fully redesigned roles and processes around AI, which means most teams are improvising. Do not improvise. Tell each affected person what changes about their week, what stays theirs, and how their judgment overrules the agent. Teams that understand the agent correct it; teams that fear it route around it, and then you are paying for a system nobody uses.
Third, and only if task one holds its target for four consecutive weeks, add task two through the full cycle again: narrow access, approvals on, five supervised tests, two weeks of review. Task two should share the same app or data area as task one so your supervision compounds rather than scatters. If task one wobbles, spend month two stabilizing it instead. Adding a second shaky task to a first shaky task produces one convincing demonstration that agents do not work, which helps no one.
| Signal at day 60 | Verdict | Action |
|---|---|---|
| Target met 4 weeks, corrections rare, staff asking for more | Expand | Add task two, keep weekly reviews |
| Target mostly met, corrections cluster on one confusion | Fix | Narrow that case, retighten one approval, retest |
| Target missed twice after fixes, or review costs exceed savings | Stop or replace | Pause the agent, keep the log, pick an easier task |
| Nobody uses the output, or the owner never reviews | Reset ownership | Name a new owner or admit the timing is wrong |
Month three converts the log into a decision. Measure the same three numbers as day 30, plus one more: outcome quality as judged by the downstream consumer, such as booking rates on drafted follow-ups or error rates on entered records. Compare against the written target from day one, not against the vendor's marketing. A target of follow-up within one hour either happened or did not; the log knows.
Write the verdict as a one-page memo with four sections: what was attempted, what the numbers show, what it costs in money plus review time, and the recommendation. The recommendation has three honest options. Expand means adding a third task or widening hours, with the same cycle. Fix means one more month on current scope with a specific change named. Stop means pausing the agent, archiving the brief and log, and revisiting in a quarter. Stopping a failing pilot is a successful quarter; funding it for a year would have been the failure.
Whatever the verdict, keep two artifacts: the final brief and the correction log. The brief is the seed of your agent playbook: write the task's operating rules with our AI agent SOP template and design the handoff gates with our approval workflows guide. The log is training material for staff and evidence for the next budget conversation. Teams that keep these compound their learning; teams that rely on memory restart every project from zero.
Close the quarter by reviewing the whole picture with anyone affected: what changed about roles, what training is still missing, and what the next quarter's candidate task is. Then protect the review habit that got you here. The weekly fifteen minutes can become thirty minutes biweekly once quality is stable, but it should never become nothing. Unsupervised agents do not stay good; they stay unexamined.
Who should own the AI agent in a small business?
Name one person with authority to change its instructions and access, ideally the manager closest to the task. That owner runs the weekly fifteen-minute review, keeps the correction log, and makes the expand-or-stop calls. Agents without a named owner drift: Anthropic's enterprise rollout guidance stresses staged maturity and assigned responsibility for exactly this reason.
How many tasks should an AI agent handle in the first 90 days?
One task for the first month, a second task in month two only if the first passes its target, and at most three by day 90. Each addition repeats the same cycle of narrow access, enforced approvals, supervised testing, and weekly review. Teams that add five tasks in week one spend the quarter debugging instead of saving time.
When should we expand an AI agent versus shut it down?
Expand when the agent hits its written target for four consecutive weeks with a clean correction log and staff asking for more. Pause and fix when corrections cluster around one confusion, and stop when the target is missed twice after fixes or the weekly review costs more than the time saved. Decide with numbers from your log, not enthusiasm.
What access should an AI agent have by day 90?
Only what earned tasks require: scoped accounts per app, approvals before anything customer-facing or irreversible, and a complete activity log you actually review. Larger firms formalize this with agent registries and identity controls, such as Microsoft Agent 365 at 15 dollars per user per month, but the discipline of least privilege and logged approvals is the same at any size.
Name your owner and write your task card this week, then follow the day-by-day plan above. If you want help scoping the task or setting the target, book a call and bring your numbers. For the broader automation sequence around the agent, the free six-step plan at the homepage funnel puts it in order. When you are ready to systematize approvals, our onboarding automation service shows how supervised rollouts run at team scale.
Every agent task needs a written page covering purpose, triggers, limits, approvals, and review. Includes a filled inbox example plus a blank template to copy.
Calculate AI agent ROI before you buy: hours saved times loaded pay rates, plus error and delay costs, with a worked example for a ten-person firm.
Set up your first AI agent in about an hour: pick one task, connect one app, set approvals, test safely, and review results before expanding further.
More articles: browse the full Praktivo blog.