AI Security

Model Context Protocol Security: Threats and Defensive Design

A current defensive guide to MCP hosts, clients, servers, tools, resources, trust boundaries, consent, and least privilege.

By PermsAI Editorial Team
Model Context Protocol Security: Threats and Defensive Design featured image

Model Context Protocol (MCP) standardizes how an LLM application connects to external data and tools. The current specification describes hosts, clients, and servers communicating with JSON-RPC; servers can expose resources, prompts, and tools, while clients may offer sampling, roots, and elicitation. The protocol makes integration composable, but it does not decide whether a server is trustworthy or whether an operation is appropriate for a user. Those remain host and application responsibilities.

What MCP changes architecturally

A host application initiates connections. An MCP client inside that host manages a connection to a server. The server provides capabilities and returns data or performs operations. This separates integration plumbing from the host’s model orchestration, yet creates a new trust edge. A server may be local to a developer machine, a remote service, or an intermediary that reaches another system.

The official specification explicitly treats tools as arbitrary code-execution paths and tool descriptions or annotations as untrusted unless obtained from a trusted server. It also emphasizes user consent, data privacy, and clear authorization UI. These are protocol security principles, not a guarantee that every client implements them correctly.

MCP trust boundary map

Use this conceptual map when adopting an integration:

User → Host/Application → Model → MCP Client → MCP Server → Tool/Resource → External System

At each edge record identity, trust level, credentials, authorization decision, untrusted data, and audit evidence. The host authenticates the user and decides what data may leave the application. The client authenticates or establishes a channel to the server according to the deployment. The server’s tool or resource remains a separate authority; a successful protocol exchange does not grant it unrestricted access to the host.

Server identity and trust

Evaluate who operates a server, how its software is obtained, what update path it uses, and which data it can access. A local server can read files or invoke processes with the user’s OS privileges; a remote server can receive data over a network and may have its own downstream credentials. Do not equate “installed locally” with trusted or “remote” with malicious. Require explicit registration, pinned configuration, narrow roots or resources, and an owner.

Treat discovery results, names, descriptions, examples, and annotations as metadata that can be wrong or manipulated. The model may use descriptions to select a tool, but the host should maintain an independent allowlist and risk classification. Changes to a server’s advertised capabilities should trigger review.

Authorization and consent

MCP does not replace application authorization. Before exposing a resource, the host must check user, tenant, purpose, and data classification. Before invoking a tool, validate arguments and apply the action policy described in Secure Tool Calling for LLM Applications. High-impact operations should require a meaningful approval showing target, effect, and recipient; the model must not manufacture consent.

Do not blindly forward the user’s bearer token to every server. Use audience-restricted credentials, a broker, or a server-specific identity. Secure Credential Handling for AI Agents covers custody and rotation. AI Agent Permissions covers the authorization decision; MCP adds an integration boundary where that decision must be rechecked.

Data exposure and context flow

Resources can contain private documents, code, configuration, or personal data. The host should minimize what it sends, apply tenant and object filters before context assembly, and make data sharing visible to the user. Tool results are untrusted input to the model and may include instructions, secrets, or misleading content. Preserve provenance and validate structured responses. A server should not be able to silently turn a resource read into broader host access.

Sampling, roots, and elicitation deserve explicit policy. If a server asks the client to sample another model, inspect what prompt, context, and result the server can see. If it asks about filesystem roots, return only approved boundaries. If it requests user information, show a clear purpose and scope. Features should be disabled when the application cannot explain or control their data flow.

Local versus remote servers

Local transports can reduce network exposure but increase sensitivity to process, filesystem, and environment privileges. Run local servers with a dedicated OS identity, restricted working directory, limited environment variables, and an update process. Remote servers need endpoint authentication, transport protection, audience checks, egress policy, and monitoring. In both cases isolate credentials and assume the server can be compromised after installation.

Supply-chain and change risk

An MCP server is software plus configuration, schemas, and often dependencies. Pin versions where practical, verify provenance, review updates, and stage changes before production. Inventory which hosts connect to which server and which tools or resources were available at a given time. Revoke a server identity or remove it from the allowlist when compromise is suspected. This complements AI Supply Chain Security without turning MCP into a generic supply-chain article.

Logging and accountability

Capture host, user, client, server identity, capability, request ID, target resource, authorization result, approval, result class, and downstream request ID. Redact tokens, personal data, and full resource contents. Correlate MCP events with the model session and tool gateway so investigators can answer which server supplied context and which server caused a side effect. AI Agent Observability provides the broader evidence model.

Defensive adoption checklist

  • Identify the server operator, source, version, transport, and update path.
  • Inventory every advertised tool, resource, prompt, sampling, root, and elicitation capability.
  • Maintain an independent allowlist and risk tier; do not rely on descriptions as policy.
  • Verify server identity and audience; isolate local processes and remote egress.
  • Apply user, tenant, object, purpose, and least-privilege authorization before data access.
  • Keep master credentials out of model context and avoid blind token forwarding.
  • Obtain explicit consent for data sharing and high-impact tool calls.
  • Validate arguments and outputs; treat server data as untrusted context.
  • Pin and review updates, record capability changes, and retain rollback options.
  • Log identities, decisions, approvals, resource IDs, and outcomes with redaction.
  • Test compromised-server, revoked-credential, cross-tenant, and excessive-capability scenarios.

What MCP does not guarantee

The protocol can standardize message shape and capability negotiation, but it cannot enforce your tenant model, business authorization, credential custody, sandbox, or incident response. A compliant server may still be over-privileged, compromised, deceptive, or unsuitable for sensitive data. Conversely, a carefully governed integration can be useful when the host keeps the final decision and evidence.

Sources

Review the integration as a product boundary

Document the server’s intended users, data classification, support owner, incident contact, and maximum acceptable blast radius. Test the first connection in a staging tenant with synthetic data, then compare observed requests with the declared capability list. Review every update, including dependency and configuration changes, because a server can become more powerful without a protocol version changing. Keep a removal plan that disables the client, revokes credentials, and preserves evidence.

A defensible operating posture

The safest MCP deployment makes the host the final policy authority. The model can discover useful capabilities, but it cannot enlarge the allowlist, approve its own request, or bypass tenant filters. Servers receive only the context and credentials needed for a particular operation. Clear consent, narrow tools, isolated credentials, validated outputs, and tamper-resistant telemetry turn a flexible protocol into a reviewable integration. Reassess that posture whenever a new server, transport, feature, or downstream system is introduced.