AI Security

Secure Multi-Agent Systems: Trust, Delegation, and Message Boundaries

A practical trust graph for securing delegation and message boundaries between autonomous agents.

By PermsAI Editorial Team
Secure Multi-Agent Systems: Trust, Delegation, and Message Boundaries featured image

A multi-agent system is secure only when every relationship between agents has an explicit trust boundary. An orchestrator, planner, specialist, critic, and tool-facing worker may all be authenticated, yet a valid identity does not make every message true or every delegation authorized. The design question is: which agent is speaking, what authority was delegated, what data may cross the edge, and which policy enforces the result? This complements AI Agent Security, while AI Agent Identity focuses on principal identity and AI Agent Permissions on authorization decisions.

The Multi-Agent Trust Graph

Use this graph as a review artifact: User → Orchestrator Agent → Specialist Agent → Tool Gateway → Resource. At each edge record identity, delegated authority, message trust, policy, and evidence. The user may authorize a task; the orchestrator may delegate a narrow subtask; the specialist may propose an action; the gateway decides whether that action is allowed; and the resource applies its own authorization. No edge should inherit unlimited authority from the previous one.

Identity is necessary but not sufficient

Give each agent and service a distinct workload identity. A receiver should verify the sender, audience, tenant, request identifier, and freshness. Also preserve the initiating human or service principal. A message signed by Agent A proves origin, not that its claims are correct or that A was allowed to ask for a database export. Keep identity and authorization separate.

Delegation and authority attenuation

Delegation should reduce or preserve scope, never silently expand it. Bind a grant to task, tenant, resource, action, audience, time, and budget. If an orchestrator can read tickets but a specialist only needs one ticket, the delegated grant should name that object. Reauthorize at the tool gateway rather than trusting a conversational statement such as “the user approved this.” Expiry and revocation must work while a workflow is running.

A confused deputy appears when a powerful agent uses its authority for a less-trusted caller. A transitive chain can make this worse: A delegates to B, B asks C, and C receives more authority than A intended. Carry a bounded delegation record and require each hop to attest what it received and what it is passing onward.

Message boundaries

Treat inter-agent messages as structured data, not executable instructions. Validate schema, size, encoding, allowed fields, sender, recipient, audience, timestamp, nonce, and correlation ID. Reject unknown capabilities and stale or replayed requests. Separate facts, evidence, proposed actions, and policy decisions in the message model. A specialist’s output should not be able to overwrite the orchestrator’s system policy or inject arbitrary tool parameters.

Use explicit content labels and provenance for retrieved documents and tool results. A message can be authentic and still contain untrusted text. Keep untrusted content in a data field and prevent downstream agents from interpreting it as a control directive without a policy decision.

Shared memory and state

Shared state is another trust boundary. Scope records by tenant, user, task, and purpose; apply read and write authorization independently; and preserve writer, source, version, and expiry. Do not let every agent write durable instructions. AI Memory Security explains why reusable context needs provenance and revocation. For short-lived state, use task identifiers and clear it when the workflow ends.

Architectural patterns

A hierarchical design centralizes policy in an orchestrator, which simplifies audit but creates a high-value coordinator. Peer-to-peer designs can reduce bottlenecks but make trust discovery, revocation, and transitive delegation harder. A hybrid pattern uses a coordinator for identity and policy while allowing specialists to exchange bounded evidence. Choose based on failure containment, not novelty.

Keep tool access behind a gateway. Agents should request typed operations; the gateway checks current identity, delegated scope, resource ownership, rate and spend limits, and required approval before issuing credentials. A specialist that only drafts an email should not inherit the sender’s mailbox token. High-impact actions can require the human approval controls described in Human Approval for AI Agents.

Failure containment

Assume one agent will be compromised, confused, or unavailable. Limit retries, fan-out, message size, data volume, and execution time. Circuit-break repeated denials, unexpected destinations, privilege requests, or cross-tenant identifiers. Design partial failure explicitly: a coordinator should not interpret silence as approval, and a failed specialist should not cause an uncontrolled fallback with broader authority.

Audit chains

Record the initiating principal, every agent identity, delegation identifier and scope, message digest, policy version, tool request, approval, resource, result, and state change. Correlate the chain with a stable workflow ID. Logs should be tamper-resistant and should not contain secrets or unnecessary payloads. This evidence lets responders distinguish a malicious sender, a forged message, a policy bug, and an honest downstream failure.

Multi-agent security checklist

  1. Draw every agent, gateway, store, and resource as a node and every message as an edge.
  2. Assign distinct identities and verify sender, audience, tenant, freshness, and correlation.
  3. Bind delegation to exact task, resource, action, duration, and budget.
  4. Attenuate authority at each hop; prevent confused-deputy and transitive expansion.
  5. Use schemas, nonces, replay checks, provenance, and separate data from instructions.
  6. Authorize shared-memory reads and writes independently with version and expiry.
  7. Keep tools behind a policy gateway and isolate credentials per capability.
  8. Add approval for irreversible, external, privileged, or high-blast-radius actions.
  9. Limit fan-out, retries, time, spend, and data volume; provide an external stop control.
  10. Preserve an end-to-end audit chain and test compromised-agent recovery.

Sources

Review a delegation edge before production

For each edge ask five questions: who is the sender, who is the receiver, which credential authenticates it, what authority is being transferred, and what evidence will remain? Then test denial paths. Remove the receiver’s grant and confirm queued work stops; alter the resource and confirm the request is rejected; replay an old message and confirm the nonce fails; substitute a different tenant and confirm the gateway refuses it. These tests are more informative than a diagram that shows only arrows.

Agent-to-agent protocols should make the safe path easier than an ad hoc HTTP call. Define canonical serialization, version negotiation, error semantics, and explicit capability names. Return policy denials as structured results so an orchestrator cannot mistake them for a transient model failure and retry with a broader agent. Where an agent needs another agent’s output, prefer a signed assertion of a narrow fact or proposal over forwarding an entire conversation.

Human operators need a kill switch that does not depend on the coordinator behaving correctly. Disable an identity, revoke a delegation, pause a queue, or block a destination independently. Preserve enough evidence to investigate without allowing a compromised agent to edit its own history. The goal is composable autonomy: each agent can be useful, but no single message can smuggle authority across the whole graph. Version these contracts and review them whenever a new agent, tool, tenant, or shared store is introduced. Security ownership must be named for both the sending and receiving side. Record unresolved assumptions as residual risk and schedule review. That makes delegation accountable rather than merely convenient.