AI Security
Secure Tool Calling for LLM Applications
Implementation guidance for making model-proposed tool calls bounded, validated, authorized, and observable.
An LLM-generated function call is a proposal, not an authorized operation. The application must validate the proposal, identify the caller, apply policy, and decide whether a side effect may occur. Secure tool calling turns model flexibility into bounded capabilities without handing the model a universal credential.
Model proposes, application decides
A model can select a tool and produce arguments, but it cannot grant itself permission. Keep the enforcement point in trusted application code. The request should carry a session, tenant, principal, and correlation ID. The gateway resolves current state, validates arguments, checks policy, and invokes an adapter only after all gates pass.
This distinction matters because a syntactically valid call may target another tenant, an unauthorized record, an excessive amount, or a dangerous destination. Treat model output as untrusted input at every hop.
Build an explicit tool registry
Maintain a registry of known tools with stable names, owners, schemas, capabilities, risk tier, allowed environments, and data classification. Expose only the tools required for the task. A registry entry should say whether an operation reads, drafts, writes, sends, deletes, or changes privileges. Tool descriptions help the model choose, but descriptions and annotations are metadata, not policy.
Prefer narrow tools such as create_ticket_draft over a generic HTTP client or shell. Separate read, preview, and commit operations so the model cannot silently cross an effect boundary.
Validate every argument
Use strict schemas and then apply business validation. Check types, ranges, enum values, resource identifiers, tenant ownership, formats, quantities, dates, and allowed destinations. Resolve IDs against trusted state rather than trusting names generated by the model. Reject unknown fields and oversized strings. Recalculate prices, recipients, or permissions from server-side records.
Validation should be deterministic and observable. Record the normalized argument digest and rejection reason without logging secrets. A model confidence score is not a substitute for validation or authorization.
Authorize at the action boundary
For every side-effecting call evaluate principal, action, resource, and context. Include tenant, purpose, session, current state, rate or spend limit, and delegated authority. AI Agent Permissions explains least-privilege decisions; this page focuses on the tool-call gateway that applies them.
Do not reuse one universal agent credential. Secure Credential Handling for AI Agents describes brokered, tool-specific credentials. A mail tool should not receive a database token, and a read-only reporting tool should not inherit admin scope.
Allowlisting and risk tiers
Allowlist capabilities, resource classes, and destinations. A useful baseline is read-only, reversible write, external side effect, and privileged or destructive. Low-risk calls can execute under fixed policy. Writes may require stronger validation, idempotency, or a user confirmation. Destructive, financial, production, or security changes should pause for the approval controls in Human Approval for AI Agents.
Idempotency and transaction boundaries
Agents retry when they see timeouts or ambiguous errors. Use idempotency keys bound to the task and intended operation so a retry does not duplicate an irreversible effect. Where possible split prepare, validate, authorize, and commit. Return a durable operation ID and status rather than asking the model to guess whether a commit occurred. Support compensating actions or rollback for reversible changes.
Output is untrusted input too
Tool output becomes model input and may contain hostile text, secrets, or misleading claims. Parse responses against an expected schema, limit size, redact sensitive fields, and label provenance. Do not let an error string become a new instruction. Indirect Prompt Injection covers the broader problem. Apply destination-specific encoding before output reaches HTML, SQL, a shell, or another API.
Errors, rate limits, and resources
Return safe error classes such as denied, invalid, unavailable, or needs-approval. Do not expose stack traces, bearer tokens, internal URLs, or full backend responses to the model. Apply per-agent and per-tenant rate, concurrency, spend, payload, and time limits. Circuit-break repeated denials and runaway loops. An external operator should be able to disable a tool or revoke its credential.
Secure tool gateway architecture
Use this path:
LLM → Tool Proposal → Schema Validation → Identity/Authorization → Risk Policy → Approval if needed → Credential Broker → Tool Adapter → Target → Output Validation → LLM
The gateway owns identity and policy. The credential broker injects a short-lived token only into the adapter. The adapter translates a narrow operation into the target API and strips unnecessary response fields. Telemetry records proposal, decision, approval, execution, and outcome with one correlation ID.
Tool security matrix
| Tool type | Default exposure | Key controls | Evidence |
|---|---|---|---|
| Read-only search | Allow narrowly | Tenant filter, result cap, redaction | Query, resource IDs |
| Draft message | Allow with policy | Recipient allowlist, preview | Draft digest |
| Send message | Approval or restricted | Exact recipient/body binding, idempotency | Approval and send ID |
| Database write | Deny by default | Typed command, row policy, transaction | Before/after change |
| File or code execution | Isolated | Sandbox, quotas, artifact scan | Code hash, exit status |
| Privilege change | Strong approval | Separation of duties, step-up auth | Approver, policy version |
Human approval and replay resistance
Approval must bind to exact tool, parameters, target, identity, and expiry. If a model changes the amount, recipient, resource, or command after approval, require a new decision. Prevent replay with one-time operation IDs, nonces, and server-side state. The model must not be able to fabricate, edit, or mark an approval as complete.
Practical checklist
- Treat every call as an untrusted proposal.
- Keep a registry of typed, owned, risk-classified tools.
- Validate schema, ranges, tenant, resource, and business invariants.
- Authorize principal, action, resource, purpose, and limits in trusted code.
- Expose the minimum capabilities and separate read, draft, and commit.
- Broker tool-specific short-lived credentials; never pass master secrets to the model.
- Use idempotency keys, transaction boundaries, and safe retries.
- Validate and redact tool output before returning it to the model.
- Require bound approval for irreversible, external, privileged, or costly effects.
- Apply rate, concurrency, time, and spend limits with an external stop control.
- Log decisions and outcomes without secrets, and test denial and replay paths.
Related protocol analysis
Use Model Context Protocol Security when the tool boundary is implemented through MCP clients, servers, prompts, resources, and elicitation.
Sources
- OWASP Agentic AI Threats and Mitigations
- OWASP Top 10 for LLM Applications
- NIST AI Risk Management Framework
- IETF OAuth 2.0 and token security standards
- Provider function-calling documentation for interface semantics only
Ownership and change control
Every tool needs an owner who can explain its data contract, failure behavior, and revocation path. Review schema changes like API changes: compare permissions, affected resources, and error semantics; run authorization and replay tests before rollout. Keep development and production registries separate. A tool that silently broadens a response field or accepts a new identifier can create a security regression even when the model prompt is unchanged.
Test the negative path
Exercise wrong-tenant IDs, expired approvals, duplicate idempotency keys, malformed arguments, oversized results, revoked credentials, unavailable targets, and policy changes during a workflow. Confirm the gateway denies safely, returns a bounded error, emits evidence, and does not retry with broader authority. Test tool output containing instructions or secrets, and verify redaction before it reaches the model. These tests measure the application boundary rather than the model’s willingness to comply.