AI Security
Human Approval for AI Agents: Designing Safe Control Points
A practical framework for placing meaningful human approval controls around consequential AI-agent actions without causing approval fatigue.
Human approval is most useful at a risk transition: the point where an agent moves from understanding or preparing work to causing an effect that is hard to reverse, broad in scope, externally visible, privileged, or insufficiently evidenced. Requiring a click for every tool call creates rubber-stamping. Allowing every action to proceed because a model sounds confident creates unsafe autonomy. The design goal is a reviewable control point that is proportional to the action.
This page is about designing that control point. AI Agent Permissions: Least-Privilege Authorization by Design explains whether a principal is allowed to act; approval supplies an additional, bounded human decision when the policy requires it. The wider AI Agent Security guide covers the surrounding system boundaries.
Autonomy and approval are a continuum
Not every capability deserves the same friction. A useful progression is:
| Autonomy level | Typical action | Default control |
|---|---|---|
| 0 — recommend | Explain, plan, or draft | No execution; show evidence and uncertainty |
| 1 — read-only | Search an authorized record or inspect status | Authorization, scope, logging |
| 2 — bounded/reversible | Save a draft, tag a ticket, stage a change | Limits, undo, notification or sampled review |
| 3 — external side effect | Send a message, create an order, change a shared workflow | Risk-based pre-action approval |
| 4 — privileged/destructive | Delete data, deploy production, change access or credentials | Strong approval, narrow execution, independent audit |
This is a starting model, not a universal rulebook. A low-value internal message may be safely automated; a read operation against sensitive medical or payroll data can be consequential. The action, object, audience, and deployment context determine the control.
Decide risk with more than model confidence
Model confidence is neither consent nor authorization. It can inform a user-facing explanation, but it cannot prove that the proposed recipient, resource, or effect is correct. Evaluate deterministic facts alongside evidence quality:
- Reversibility: Can the exact effect be reliably undone, and by whom?
- Financial or operational impact: Does it move money, consume a scarce resource, or interrupt service?
- External communication: Will it reach a customer, supplier, regulator, or public audience?
- Privilege and sensitivity: Does it alter access, credentials, production state, or protected data?
- Blast radius: Is this one object, one tenant, or a fleet-wide change?
- Novelty and uncertainty: Is the destination new, the request ambiguous, the evidence weak, or the outcome unusual?
NIST's AI RMF treats risk as context-specific and calls for defined roles and responsibilities in human-AI configurations. That supports a risk-based design, not a blanket requirement to approve harmless reads. OWASP's agentic guidance similarly reinforces that consequential actions need controls outside a model's text output.
The Human Approval Decision Matrix
Use an explicit matrix before a tool is exposed to an agent:
| Action type | Reversibility | Privilege / blast radius | Evidence needed | Approval requirement |
|---|---|---|---|---|
| Summarize authorized data | No system change | Low | Source references | None; log retrieval |
| Create a private draft | Usually reversible | Low | Target owner and draft body | Policy limits; notify if needed |
| Send an external message | Recall is unreliable | Moderate | Recipient, body, purpose, attachments | User approval for novel or sensitive sends |
| Modify production configuration | Rollback may be partial | High | Diff, environment, change ticket, impact | Named operator or change approver |
| Delete account or data | Often irreversible | High | Exact objects, retention rule, legal hold state | Strong pre-action approval; often two-person |
| Change roles or credentials | Enables further authority | High | Principal, new scope, reason, expiry | Security/admin approval, short execution window |
The matrix should become deterministic policy: object classification, action type, amount, environment, tenant, and destination are better inputs than an LLM's self-assessed risk label. A model can propose a risk category; the policy service must independently calculate or validate it.
Pre-action approval is not post-action review
Post-action monitoring is essential for detection, correction, and learning, but it cannot make a payment, credential rotation, public disclosure, or destructive change safe after the fact. Require approval before the policy enforcement point releases a high-impact tool call.
For bounded reversible work, an after-the-fact notification, periodic sampling, or a rollback queue may be more usable than a blocking dialog. Document which actions are intentionally autonomous and why. This turns "human in the loop" from a slogan into an operational decision.
Give the reviewer a decision, not a vague prompt
An approver needs enough context to judge the actual transaction. The approval screen or signed record should show:
- requested action and human-readable effect;
- initiating user, agent/workload identity, and delegated authority if any;
- tenant, target resource, recipient or destination;
- proposed parameters, diff, amount, or message body as applicable;
- data classification, policy/risk tier, evidence and stated reason;
- expiry, alternatives, and whether the action is reversible.
"Allow the agent to continue?" is not meaningful approval. If a reviewer cannot see the recipient, amount, environment, or changed parameters, they cannot judge the consequential effect. Keep sensitive values appropriately masked while preserving enough detail to make the decision.
Bind approval to the exact proposed action
An approval record should name the policy version, action, resource identifiers, normalized parameters or a digest, tenant, requesting principal, approving principal, issuance time, and expiry. The enforcement point validates that record immediately before execution.
Do not let the model fabricate a success message, edit an approval object, or convert an approval for one recipient into an approval for another. Parameter changes, resource changes, scope expansion, a different agent identity, or an expired request require re-evaluation. Use one-time request identifiers and record consumption to prevent replay. For long queues, re-check current authorization and resource state at execution time.
This integrity property is why approval belongs outside model context. The model may prepare the proposal and summarize evidence; a separate workflow, policy service, and executor control approval state and the credentialed call. AI Agent Identity explains how the initiating user, workload, and delegated authority remain distinguishable in that path.
Avoid approval fatigue with tiering and sensible defaults
People faced with constant low-value prompts learn to click through them. Reduce fatigue by removing prompts that add no decision value: use read-only authorization, typed tools, budgets, destination allowlists, and reversible drafts for routine work. Aggregate related low-risk actions only when their aggregate effect remains bounded.
Reserve interrupts for meaningful transitions: a first-time external recipient, unusual volume, privileged environment, sensitive data, material value, destructive operation, or evidence conflict. Calibrate from actual approval outcomes and near misses, not simply from how many dialogs a product displays.
Step-up, separation, and emergency paths
Risk can warrant a stronger approver rather than another identical confirmation. A normal user may approve an email draft; a manager may approve a material commitment; a security owner may approve a credential or production access change; two independent approvers may be justified for high-value irreversible actions. Do not over-engineer a two-person workflow for ordinary tasks.
Break-glass workflows are exceptional. They should require a named emergency reason, a narrow time window, heightened logging and alerting, and mandatory retrospective review. They are not a bypass button the agent can invoke when a normal approval is denied.
Approval across agents
If Agent A delegates to Agent B, approval remains attached to explicit authority and the exact action, not to a conversational claim that "someone approved it." B must verify the approver, actor chain, scope, expiry, and request binding before executing. It should not silently combine several small approvals into broader authority.
A defensible approval architecture
Agent → action proposal → deterministic risk/policy service → low risk: execute under policy / high risk: approval queue → human decision → signed, bound approval record → policy enforcement point → tool execution
The risk component may use classification or heuristics, but deterministic policy owns the final gate. The executor checks current permission, approval binding, tenant, destination, and limits immediately before acting. It returns an observable result to the agent without granting the agent possession of the approval authority or the underlying credential. See Secure Credential Handling for AI Agents for that custody boundary.
Audit evidence and testing
Record who proposed, approved, denied, executed, or cancelled an action; the target, parameters or digest, policy version, risk factors, time, expiry, and outcome. Preserve the relation between approval and actual execution, including failed executions. Do not log secrets or hidden reasoning as a substitute for evidence.
Test approval integrity as part of AI Red Teaming: modified parameters after approval, replayed records, stale requests, cross-tenant IDs, a different executor, an agent attempting to claim approval, partial failures, and emergency-path misuse.
Design checklist
- Inventory actions by reversibility, privilege, audience, sensitivity, impact, and blast radius.
- Define autonomous, notify-only, pre-approval, and step-up tiers in deterministic policy.
- Require pre-action approval for irreversible, externally visible, privileged, or unusually uncertain effects.
- Present reviewer-ready transaction details, evidence, and risk—not a generic continuation prompt.
- Bind approval to actor, tenant, target, normalized parameters, policy version, and a short expiry.
- Enforce and consume the approval outside the model immediately before tool execution.
- Reauthorize when any consequential parameter or resource changes.
- Use separation of duties and break-glass controls only where the consequence justifies them.
- Reduce fatigue with bounded autonomous actions, alerts, and sampled review.
- Log proposal, decision, execution, and outcome; test replay and bypass resistance.