Put Approval Gates Around AI Agents Before Deployment
Matthew Leo · Published October 11, 2026 · Guides
An AI agent becomes an operational risk when it can turn an uncertain answer into a real action. Sending an email, filing a form, changing a customer record, deploying code or approving a payment creates consequences that a chat response does not.
The right control is not “keep a human in the loop” as a slogan. An organization needs to specify which actions require approval, what the approver will see, which identity performs the action and what evidence remains afterward.
Canada’s Office of the Superintendent of Financial Institutions makes this boundary unusually clear in its 2026 bulletin on generative and agentic systems. OSFI recommends limits on autonomy, human accountability for material decisions, monitoring and approval checkpoints for high-risk actions. The guidance is aimed at federally regulated financial institutions, but the control design is useful well beyond banking.
List actions, not job titles
Start with an inventory of every action the agent can request or perform. “Customer-service agent” is too broad. The list should name verbs and destinations:
- read a customer record;
- draft a reply;
- send a message to an external address;
- change an address or account setting;
- issue a refund;
- upload a document to a government or partner portal;
- run code, change infrastructure or deploy software;
- delete data or close an account.
Record the tool or application, data involved, maximum value or scope, affected person and whether the action can be reversed.
Use four permission levels
A practical policy can divide actions into four levels.
- Read: The agent may retrieve approved information but cannot change anything.
- Draft: The agent may prepare a message, form or change for a person to review.
- Approve then execute: A named person must approve the exact action and arguments before a separate execution layer carries it out.
- Prohibited: The agent cannot perform the action, even with ordinary user approval. This tier can cover irreversible, legally sensitive or unusually high-value tasks.
Do not let the agent decide which tier applies to its own action. Put the policy in an independent authorization layer that receives the proposed action and enforces the decision.
Make approval specific
An approval screen should show the destination, complete payload, amount, permissions, data leaving the organization and whether the action can be undone. “Allow this task?” is not meaningful if the task contains several hidden steps.
Bind the approval to those exact arguments and give it a short expiry. If the destination, amount, attachment or command changes, require a new decision. Never treat approval of a plan as blanket approval of every later action the agent invents.
Give the agent its own identity
Do not run an agent under a shared administrator account. Use a dedicated service identity with the minimum permissions required for the approved workflow. Prefer short-lived credentials, restrict which applications and records it can reach, and prevent one connected service from silently becoming a bridge to another.
Separate development, testing and production identities. Test credentials should not work against real customer systems or public filing portals.
Keep an independent audit record
For every proposed and completed action, record:
- the agent and version;
- the initiating user or event;
- the requested tool and exact arguments;
- the policy decision;
- the human approver, when required;
- the result returned by the external service;
- related data access; and
- the time and correlation identifier.
Store this record outside the agent’s control. The agent’s own explanation can help an investigation, but it is not proof of what happened or why.
Test failure paths before production
Build realistic test copies of forms, email systems and business applications that cannot reach real people or records. Exercise ambiguous instructions, missing confirmation pages, repeated retries, timeouts, changed destinations, malicious content in retrieved documents and attempts to work around blocked tools.
Anthropic’s October 9 incident report provides a concrete reason: one evaluation submitted a sensitive form on a real police website, while other tests found ways around access restrictions. The lesson is not that every agent will do the same thing. It is that live connections turn an evaluation error into an external event.
Plan the stop button and the aftermath
Operators need a way to revoke the agent’s credentials, stop queued work and block outbound connections without waiting for the vendor. Define who can activate that response and how the organization will preserve logs.
Contracts should set notification deadlines for unintended actions, access to evidence, responsibility for third-party impacts and support for credential rotation or rollback. Review these controls whenever the underlying model, tools, permissions, data sources or workflow changes.
A final pre-deployment check
- Every action has a named permission tier.
- High-impact actions require approval of exact arguments.
- Policy enforcement is outside the agent.
- The agent uses a dedicated least-privilege identity.
- Testing cannot reach live recipients or real filing systems.
- Logs are independent, complete and reviewable.
- An operator can stop execution and revoke access quickly.
- A named owner reviews incidents and control changes.
Mapletechie’s analysis of Windows agent containers explains one useful containment layer. Approval gates solve a different problem: deciding which real-world actions should be allowed after the agent has already produced a plausible plan.
Tags: AI agents, AI governance, cybersecurity, Canada