AI Security
LLM Security Risks: A Practical Guide for Developers
A developer-focused architecture guide to LLM trust boundaries, authorization, retrieval, output validation, tool safety, secrets, and monitoring.
An LLM security program protects the application around the model, not just the prompt sent to it. A large language model is a probabilistic component that may receive instructions, user input, retrieved records, external documents, and tool responses, then produce text that another system may act on. Those inputs do not share one trust level, and the output is not automatically safe because it came from the model.
The practical design rule is simple: keep authentication, authorization, data isolation, validation, and transaction controls deterministic and outside the LLM. Prompts can guide behavior, but application and infrastructure controls must decide what data the model can see and what actions the system can perform.
Start with an LLM application threat model
An LLM application is a chain of trust boundaries. Mapping the data crossing each boundary is more useful than asking whether the model itself is secure.
| Data or component | Typical trust level | Principal security question |
|---|---|---|
| System and developer instructions | Privileged but potentially exposable | Does the application rely on secrecy of the prompt as a control? |
| User input | Untrusted | Can it alter instructions, request another user's data, or reach a dangerous sink? |
| Retrieved records | Mixed trust | Was access authorized for this user, and can document text manipulate the model? |
| Web pages, email, and uploaded files | Untrusted content | Are embedded instructions being confused with application policy? |
| Tool responses | Authenticated source, untrusted payload | Can returned text influence later actions or disclose excess data? |
| Model output | Untrusted | Is it validated for the specific destination before display or execution? |
| Tools and target systems | High impact | Are permissions scoped and every consequential action independently authorized? |
Threat modeling should identify the protected assets, actors, entry points, data stores, tools, and downstream sinks. It should also record which tenant or user owns every retrieved object. The model's context window is not a security boundary: if sensitive data enters it, assume the model may reproduce or transform that data in its response.
Prompt injection is an architecture problem
Prompt injection occurs when untrusted content influences the model as though it were a trusted instruction. The content may be supplied directly by a user or indirectly through a document, web page, email, database field, or tool result. The model cannot reliably distinguish intent merely because one string was labeled “data” in a prompt.
The dedicated prompt injection guide explains attack paths and layered defenses. At the application level, the important consequence is that no prompt instruction should be the only control protecting data or authorizing an action. Input filtering may remove known patterns, but natural language has too many equivalent forms and legitimate documents may contain instruction-like text. Reduce the impact with scoped data access, isolated content processing, constrained tools, output validation, and confirmation for high-impact operations.
Sensitive information disclosure
Sensitive information can leak because the application supplies too much context, retrieval crosses an authorization boundary, logs capture confidential prompts, or the model is asked to reveal hidden instructions. Relevant assets include personal data, proprietary documents, access tokens, internal identifiers, and customer-specific context.
Minimize disclosure before inference. Retrieve only fields the authenticated principal may access; enforce tenant filters in the data layer; redact unnecessary sensitive values; and avoid placing secrets in system prompts. A denial instruction such as “never reveal this token” does not make a token safe to disclose to the model. Where the use case requires sensitive context, define retention, provider, geographic, and logging requirements explicitly, and test responses for cross-user leakage.
Treat model output as untrusted
Improper output handling turns model text into an injection path for another component. The correct validation depends on the destination:
- Rendered HTML should be escaped or sanitized by a maintained policy before insertion into a page.
- Database queries should use parameterized operations and a restricted data-access layer, not generated SQL executed verbatim.
- Shell commands should not be passed to a general shell. Prefer a fixed operation interface with typed arguments and strict allowlists.
- API requests should be created from validated schemas, constrained destinations, and server-side authorization.
- URLs, filenames, and document formats require destination-specific canonicalization and validation.
Validation after generation is necessary even when the prompt requests a safe format. Structured output or a tool schema improves parsing but does not prove that an action is authorized or semantically appropriate.
Excessive agency expands the blast radius
An assistant that only proposes text has a different risk profile from a system that can send messages, modify records, call APIs, run tools, or execute transactions. More autonomy combines uncertain model behavior with real credentials and side effects. The AI agent security guide covers agent-specific permissions, memory, isolation, and approval design in depth.
For an LLM application, expose the smallest useful action set. Separate read and write operations, scope credentials to one service and tenant, limit transaction value or batch size, and require a fresh authorization check when a proposed action becomes an actual request. High-impact actions should present the user with the exact target and effect before confirmation; approval of a broad goal is not approval of every later tool call.
Authorization belongs outside the model
An LLM can help classify a request, but it should not be the policy enforcement point. “Should this user see this record?” is an authorization decision that must be derived from authenticated identity, resource ownership, roles, attributes, and current policy.
Enforce access twice where appropriate: when constructing context and again at the action or data sink. This prevents a model from retrieving an unauthorized object and also prevents generated output from bypassing the normal service boundary. Denials should fail closed. Tool gateways should receive the authenticated principal and a narrowly defined proposed action, then evaluate policy without trusting the model's explanation of why it is allowed.
Secure retrieval-augmented generation
RAG adds a retrieval system and a corpus to the threat model. Common failures include retrieving documents across tenants, indexing untrusted content without provenance, returning more context than necessary, and allowing poisoned document text to influence tool use.
Apply authorization before retrieval results enter the prompt, not after the model answers. Preserve document owner, tenant, source, classification, and ingestion time as metadata; filter by those attributes in the query path. Separate corpora when policy requires stronger isolation. Treat chunks from public sites and user uploads as untrusted, show citations so users can inspect provenance, and avoid granting action authority based solely on retrieved text. Embedding similarity is a relevance signal, not an access-control decision.
Constrain tools, plugins, and functions
Tool schemas are interfaces, not security boundaries by themselves. A robust gateway validates names and arguments, restricts destinations, checks authorization, applies rate and spend limits, and records the final request and result. It should reject unknown fields and ambiguous identifiers rather than letting the model repair them silently.
Credentials belong in the gateway or target service, not in model context. Prefer short-lived, task-scoped credentials where the platform supports them. Do not reuse a powerful backend service account merely because it simplifies integration. Allowlists should be specific enough to express permitted operations; an allowlisted HTTP client that can reach arbitrary URLs still creates a broad server-side request surface.
Manage supply-chain dependencies
The application depends on more than a model provider. Libraries, orchestration frameworks, plugins, embedding services, vector stores, external APIs, model artifacts, and prompt or agent templates can all change behavior or expose data.
Maintain an inventory of providers and components, pin and scan software dependencies, review update channels, and define what data each external service receives. Verify model and artifact provenance when self-hosting. Contractual assurances do not replace technical controls: service credentials, egress restrictions, tenant isolation, and fallback behavior still need testing. A provider outage or model update should fail safely rather than silently routing sensitive work to an unapproved service.
Keep secrets out of prompts
Do not place an API key, database password, or signing secret in a prompt unless the model truly must reproduce it—which is rarely a sound design. Store secrets in a managed server-side facility, inject them only into the component that needs them, rotate them, and scope them to the minimum operations.
When a model proposes a tool call, the gateway should attach the appropriate credential after authorization. This separates reasoning context from execution authority. If a credential is exposed, short lifetime, narrow scope, and auditable use reduce the incident's blast radius.
Logging must preserve both evidence and privacy
Useful observability records the authenticated actor, model and configuration, retrieved document identifiers, policy decisions, tool requests, confirmations, outcomes, and correlation IDs. Monitor unusual retrieval breadth, repeated denials, unexpected tool sequences, sharp changes in token or cost use, and attempts to reach disallowed destinations.
Indiscriminate prompt logging can create a second sensitive database. Apply minimization, redaction, access controls, retention limits, and tenant separation to telemetry. Preserve enough evidence to reconstruct consequential actions without storing every confidential document verbatim. Incident response procedures should support revoking tool credentials, disabling risky actions, locating affected requests, and notifying data owners where required.
A trust-boundary control ownership model
Prompt engineering contributes to security, but it cannot own controls that require deterministic guarantees. This table provides a review test for common risks.
| Risk | Prompt contribution | Application control | Infrastructure or service control |
|---|---|---|---|
| Prompt injection | Mark trusted instructions and untrusted content | Isolate content, constrain actions, validate outputs | Sandbox processors and restrict network access |
| Data disclosure | Ask the model not to reveal sensitive context | Minimize and authorize retrieval; redact responses | Encrypt data and isolate tenants |
| Unauthorized action | Ask for confirmation | Enforce policy on principal, resource, and action | Use scoped credentials and target-side checks |
| Unsafe output | Request a schema | Parse, validate, encode, and parameterize for the sink | Apply browser, database, and runtime protections |
| Excessive tool impact | Describe allowed behavior | Allowlist tools, set limits, require approval | Sandbox execution and restrict egress |
| Weak auditability | Request explanations | Emit structured decision and action events | Protect logs and enforce retention policy |
If the only entry in the application or infrastructure columns is “none,” the design is relying too heavily on the model.
Secure reference architecture
A practical request-to-action path keeps policy decisions visible:
- User and authentication: establish a principal using the application's normal identity controls.
- Authorization and policy layer: decide which data and actions that principal may access before prompt construction.
- Input and context construction: label provenance, retrieve only authorized records, minimize sensitive fields, and separate untrusted documents.
- LLM inference: use an approved model and explicit configuration; treat the response as a proposal.
- Output validation: parse against a schema and apply destination-specific validation and encoding.
- Tool gateway: authorize the exact proposed operation, enforce allowlists and limits, attach scoped credentials, and request human confirmation when needed.
- Target system: apply its own access controls, integrity constraints, and transaction semantics.
- Audit and monitoring plane: correlate identity, retrieval, model, policy, and tool events while protecting logged data.
This architecture contains failures. A successful injection may influence a proposal, but it still encounters data authorization, output validation, tool policy, scoped credentials, and the target system's controls.
Developer security checklist
Before release, confirm that:
- every context source and output destination has an assigned trust level;
- retrieval applies user and tenant authorization before returning chunks;
- model output is untrusted and validated for its exact sink;
- authorization is deterministic and rechecked at consequential operations;
- tools use narrow schemas, allowlists, limits, and scoped credentials;
- sensitive actions show their exact effect and require appropriate confirmation;
- secrets remain server-side and outside prompts and logs;
- dependencies, models, providers, and external services are inventoried;
- telemetry supports investigation without becoming a sensitive-data warehouse;
- tests include injection, cross-tenant access, malformed output, denied actions, and tool failures.
Related standards tracking
For release-specific changes rather than the evergreen risk map, use OWASP GenAI Top 10 Updates.
Sources
- OWASP Top 10 for Large Language Model Applications
- OWASP: LLM01 Prompt Injection
- OWASP: LLM02 Sensitive Information Disclosure
- OWASP: LLM03 Supply Chain
- OWASP: LLM05 Improper Output Handling
- OWASP: LLM06 Excessive Agency
- OWASP: LLM08 Vector and Embedding Weaknesses
- NIST SP 800-218A: Secure Software Development Practices for Generative AI
- MITRE ATLAS