AI Security
Prompt Injection Attacks: How They Work and How to Defend Against Them
A threat-model-driven guide to direct and indirect prompt injection, attack paths, impact containment, tool authorization, output validation, testing, and monitoring.
Prompt injection is an attack in which untrusted content changes how a language-model application follows instructions. The content can come directly from a user or indirectly from a page, document, email, image-derived text, retrieved record, or tool result. The core security problem is not an impolite phrase. It is that natural-language instructions and data share a probabilistic interpreter, while the surrounding application may give the model access to private context or consequential tools.
OWASP classifies prompt injection as LLM01:2025 and states that foolproof prevention is unclear. Effective defense therefore combines model-level resistance with conventional controls that limit what a manipulated model can read, request, and cause.
Prompt injection, jailbreaking, and ordinary bad input
Prompt injection changes an application’s intended behavior by introducing competing instructions. A jailbreak is a related attempt to bypass a model’s safety behavior, often to elicit prohibited content. The categories overlap, but application security needs the distinction: a model that produces a disallowed answer is different from a model that uses an authorized business tool for an unauthorized purpose.
Ordinary malformed input can also cause wrong output without being adversarial. The same containment controls help. Authorization, data minimization, typed tool interfaces, output validation, and transaction approval reduce harm whether the initiating cause is an attacker, ambiguity, or model error.
Direct and indirect prompt injection
A direct injection is supplied through an interface the attacker controls, such as a chat message or API field. It may try to replace the task, reveal private instructions, manipulate a tool call, or convince the model that policy no longer applies.
An indirect injection is planted in content the application fetches on the user’s behalf. Examples include hidden or visible text in a web page, a shared document, an email thread, a support ticket, a source-code comment, metadata, OCR output, or an indexed knowledge-base record. The user may never see the instruction that influences the model.
Indirect injection is especially important because trusted automation can retrieve attacker-controlled material. The attack-chain analysis below follows that material from content placement through retrieval and action.
Why agents increase the impact
In a text-only assistant, injection may alter the answer or disclose context. In an agent, the same manipulation can influence a sequence of tool calls. The agent may search private data, choose an external destination, update memory, retry around a failure, or delegate the hostile objective to another component.
Impact depends on reachable authority, not only model susceptibility. A frequently fooled summarizer with no secrets and no write tools may create low operational risk. A rarely fooled agent with production credentials, broad network access, and automatic execution can create high risk. This leads to the most useful design question: if the model follows an injected instruction, which independent control prevents the consequence?
A prompt-injection attack path
Treat the attack as a chain with six stages:
| Stage | Attacker requirement | Defensive breakpoint |
|---|---|---|
| 1. Placement | Control content the system may consume | Source policy, quarantine, provenance |
| 2. Selection | Cause that content to be uploaded or retrieved | Permission-aware retrieval and ranking tests |
| 3. Interpretation | Make data influence the model as instruction | Context separation, model hardening, instruction detection |
| 4. Capability | Reach private data or a useful tool | Least privilege and data minimization |
| 5. Authorization | Obtain approval for unsafe arguments | Deterministic policy and transaction validation |
| 6. Consequence | Exfiltrate, modify, execute, or persist | Egress controls, sandboxing, output handling, monitoring |
A defense is stronger when it breaks several stages. A filter aimed only at stage three leaves the system exposed when a new phrasing bypasses it.
Common attack paths and impacts
Retrieval and browsing
A research agent may interpret text from a page as a new task. A RAG system may retrieve a poisoned record whose wording influences answers to many users. In both cases, ingestion and retrieval must preserve provenance and enforce access policy before content reaches the model.
Email, tickets, and collaborative documents
An attacker can send content to a mailbox or queue that an agent routinely processes. The sender is untrusted even when the mailbox and retrieval connector are approved. Authentication of the connector proves where the data came from, not that the content is safe to obey.
Tool results and multi-step workflows
Tool output is another untrusted input. A fetched error message, repository file, calendar description, or API response can contain language that changes the next model decision. Every hop must preserve provenance and authority; trust must not increase merely because content passed through an internal tool.
Memory and feedback
If manipulated output becomes durable memory, a one-time attack can affect future tasks. Store typed facts with source, subject, tenant, timestamp, and expiry rather than copying free-form instructions into memory. Require review for policy-like or cross-task state.
Likely consequences include sensitive-data disclosure, cross-tenant access, unsafe external messages, altered transactions, malicious code or queries reaching an interpreter, and poisoned memory. System-prompt disclosure may expose useful design information, but the deeper failure is storing secrets or authorization logic in a prompt.
Why simple string filters fail
Exact-match blocks are easy to evade through paraphrase, spacing, translation, encoding, misspelling, token splitting, or multimodal content. A model-based detector can recognize broader patterns, but it is still a probabilistic component that may share blind spots with the target model. False positives also matter: a security analyst may legitimately ask the system to summarize an injection example.
Filtering remains useful as one signal. Use it to quarantine suspicious content, lower tool authority, request review, or add telemetry. Do not let a “clean” classifier result grant access the user and agent did not otherwise possess.
Defense in depth
Separate instructions and untrusted data
Maintain distinct fields for operator policy, user intent, retrieved evidence, and tool results. Preserve source identifiers and trust labels through context assembly. Tell the model that retrieved text is evidence, not authority. This improves robustness and auditability, but because the same model still processes the combined context, separation is not a hard security boundary.
Apply least privilege before inference
Do not place unnecessary records, secrets, or credentials in context. Scope retrieval by authenticated user, tenant, and object policy before semantic similarity is evaluated. Give the agent only the tools required for the current task. The AI agent security guide explains the surrounding architecture; its permission model must be implemented in trusted code, not inferred from model output.
Authorize every tool call outside the model
A model may propose an action, but deterministic code must decide whether it is allowed. Validate a typed schema and then check object ownership, tenant, recipient, destination, quantity, price, purpose, and current policy. Recalculate sensitive values from trusted state. A syntactically valid tool call can still be unauthorized.
Separate tools by effect. Reading a message, drafting a response, and sending it should be different capabilities. Avoid generic shells and unrestricted database or browser tools when a narrow API can perform the task.
Restrict destinations and isolate execution
Use default-deny egress with explicit destination and protocol allowlists. Resolve redirects and final destinations before sending data. Block cloud metadata and internal control-plane addresses unless required. Run code and file transformations in disposable sandboxes with minimal filesystem mounts, short timeouts, and no production secrets.
Validate output for its destination
Model output is untrusted input to the next component. Parse structured output against a schema, then enforce business invariants. Escape text for HTML; sanitize permitted Markdown; parameterize database queries; review generated code before isolated execution; and reject URLs that violate destination policy. Validation must be context-specific—there is no universal “sanitize AI output” function.
Use allowlists and approvals deliberately
Allowlists should name permitted operations and destinations, not merely blocked words. Approval should present the complete, immutable transaction: who will receive what data, which object will change, and why. Require confirmation for irreversible, externally visible, privileged, or unusually costly actions. Do not overwhelm reviewers with prompts for routine low-risk steps.
Monitor the action path
Record source IDs, model and policy versions, proposed tools, validated arguments, authorization results, approvals, destinations, and outcomes. Alert on new domains, repeated denied actions, unusual data volume, cross-tenant identifiers, privilege escalation, and attempts to disable logging. Keep evidence outside the agent’s control.
A concrete defensive example
Consider a procurement assistant that reads supplier pages and prepares purchase requests. A hostile page contains an instruction to choose a different recipient and attach internal pricing.
A weak design sends the page, internal records, and a purchasing credential to one model and executes its output. A stronger design retrieves only public supplier content; labels the page as untrusted; gives the model a draft-only tool; resolves product and price from approved catalogs; checks the supplier and recipient against trusted account data; requires a buyer to approve the exact request; and sends through a broker with a short-lived credential. Even if the model repeats the hostile instruction, it cannot select an unauthorized recipient or access the pricing attachment.
That is the central pattern: tolerate some model failure while preventing unauthorized consequence.
Testing prompt-injection defenses
Build tests from real data paths, not only a list of jailbreak strings. Place benign and adversarial instructions in pages, PDFs, email, tickets, OCR text, tool responses, and memory candidates. Vary language, formatting, visibility, position, and retrieval rank. Test single-step and long-running workflows.
Define success in terms of security properties:
- no secret or unauthorized record enters model context;
- no cross-tenant object is retrieved;
- no tool executes outside the requester’s policy;
- no unapproved destination receives data;
- no untrusted instruction becomes durable memory;
- unsafe output does not reach an interpreter;
- the system stops safely and preserves evidence.
Include benign documents that discuss attacks so detection does not destroy utility. Track false positives, false negatives, task completion, cost, and latency. Re-run the suite after changing the model, system prompt, retriever, tool, policy, or approval UI. The AI model benchmarks guide explains why the harness, scoring rule, and tested configuration must accompany any reported result.
Developer checklist
- Inventory every user-controlled and externally controlled content source.
- Identify private data, write tools, interpreters, and outbound channels.
- Apply user, tenant, and object authorization before retrieval.
- Keep credentials out of prompts and model-visible memory.
- Expose narrow, typed, effect-specific tools.
- Validate tool arguments and business invariants in trusted code.
- Restrict egress and isolate code, browser, and file processing.
- Bind approval to complete transaction details.
- Treat tool results and model output as untrusted at every hop.
- Store memory with provenance, scope, expiry, and review controls.
- Monitor action decisions and preserve tamper-resistant evidence.
- Test indirect injection and ordinary failure after every material change.
Related defense evidence
Use Prompt Injection Research Review for the time-bounded evidence on whether detector, separation, authorization, and containment defenses generalize beyond the evaluations that tuned them.