AI Security
Indirect Prompt Injection: Threat Model and Defenses
A threat-boundary guide to indirect prompt injection through external content, with attack-path analysis, containment architecture, testing, and developer controls.
Indirect prompt injection occurs when an AI application retrieves attacker-controlled or otherwise untrusted content and the model interprets text inside that content as instructions. The attacker may never speak to the model directly. A poisoned web page, email, document, search result, ticket, repository file, API response, or RAG chunk can cross into model context and influence what the system says or does.
The durable defense is not a better blacklist of hostile phrases. It is an architecture that keeps external content untrusted, restricts what the model can access, authorizes every consequential tool call outside the model, validates outputs for their destination, and records enough evidence to detect and contain failures.
Direct versus indirect prompt injection
In a direct prompt injection, the person interacting with the application puts the conflicting instruction into their own input. In an indirect injection, the instruction arrives through material the application fetched or received while completing a legitimate task. OWASP's LLM01:2025 guidance uses websites and files as representative external sources and notes that injected content need not be obvious to a human if the model still parses it.
The distinction concerns delivery, not a separate model primitive. Both attacks exploit the difficulty of reliably separating instructions from data inside a language-model context. The broader prompt injection guide covers both forms. This page follows the indirect path from hostile content to consequence.
The trust-boundary problem
A useful application may combine three classes of text:
- trusted application instructions describing the task and limits;
- an authenticated user's request and parameters;
- untrusted material retrieved from external systems.
Those classes can be labelled or placed in separate message fields, but they ultimately influence one model computation. A delimiter, warning, or system prompt can help the model behave correctly; it does not create the kind of hard isolation provided by an authorization check, process boundary, or network policy.
The security mistake is therefore to treat retrieved text as trusted merely because the application selected it. Relevance is not trust. A search engine can return a compromised page, a help desk can receive an attacker-written ticket, and an internal knowledge base can contain a document uploaded by a user with narrower privileges than the eventual reader.
Common external injection surfaces
| Surface | How hostile content enters | Consequence depends on |
|---|---|---|
| Web pages and search results | Indexed text, page metadata, or extracted content | Whether the system only summarizes or can browse, disclose, and act |
| PDFs and office documents | Parsed text, annotations, tables, or OCR output | Parser isolation, document provenance, and available data/tools |
| Email and messages | Sender-controlled body, quoted threads, attachments, or links | Mailbox scope and permissions to reply, forward, or fetch URLs |
| Tickets and issues | Public or customer-authored descriptions and comments | Repository, support, and deployment tool authority |
| RAG chunks | Poisoned or insufficiently governed knowledge sources | Retrieval probability, access controls, context construction, and model response |
| Third-party APIs and tools | Remote response fields presented back to the model | Destination trust, schema validation, and the next tool decision |
| Repositories | README files, comments, issues, test output, or generated artifacts | Filesystem scope, secrets exposure, and code-execution controls |
Risk depends on both source controllability and the model's authority; a read-only summarizer has a smaller blast radius than an agent holding mail or deployment credentials.
The Indirect Prompt Injection Attack Path
PermsAI's attack-path model separates six stages so defenders can place and test controls before the final effect.
| Stage | Trust boundary and failure | Defensive breakpoint | Residual risk to test |
|---|---|---|---|
| 1. Content control | An attacker can influence a page, document, message, repository, or API field | Source allowlists, ownership rules, provenance, integrity checks, quarantine | Compromised trusted sources and authorized-but-malicious contributors |
| 2. Ingestion or retrieval | Hostile content is parsed, indexed, fetched, or ranked for a task | Parser sandboxing, type limits, permission-aware retrieval, risk signals | Hidden text, OCR, stale indexes, and content that wins retrieval naturally |
| 3. Context construction | External text is merged with instructions, user data, or secrets | Minimize context, label source and trust, isolate sensitive fields, exclude unnecessary secrets | Models may still follow clearly separated hostile instructions |
| 4. Interpretation | The model treats data as an instruction or changes its plan | Model safeguards, explicit task constraints, injection detectors as signals | Novel wording, multilingual content, encoding, and false positives |
| 5. Action request | The influenced model proposes a tool call, URL, query, or message | Typed schemas, deterministic authorization, allowlisted destinations, rate and scope limits | Permitted tools chaining into a higher-impact effect |
| 6. Consequence | Data leaves, state changes, or a user receives manipulated output | Output validation, transaction approval, sandboxing, egress policy, rollback, alerting | Low-and-slow leakage, misleading advice, and activity below thresholds |
Retrieval, model influence, tool availability, authorization, and final effect are separate test outcomes. The 2024 Rag and Roll study found that high malicious-document rank did not automatically produce a reliable end-to-end attack; its measurements apply only to the tested configurations.
A defensive example
Consider a support assistant asked to summarize a newly filed ticket. The ticket text claims that completing the task requires reading an unrelated customer's record and sending the result to an external address. A vulnerable design gives the model a broad customer database tool and a general email function, then trusts its natural-language judgment.
A safer design retrieves only records already authorized for the requesting support workflow. The model can propose a response draft, but cannot select arbitrary customer IDs or send mail. A policy gateway validates the ticket, tenant, destination, and action; external sending requires a separately presented approval. The injected text may still distort the draft, so the system marks its source, validates destinations and sensitive fields, and alerts on unusual instructions. Containment does not depend on perfectly recognizing the malicious sentence.
Why phrase filtering is insufficient
Blocking a phrase such as “ignore previous instructions” can catch a known demonstration, but it does not define the vulnerability. An attacker can express the same goal indirectly, split it across fields, use another language or representation, or write text that looks like ordinary task content. Legitimate documents can also discuss prompt injection and trigger false positives.
Semantic classifiers and content-risk scores are still useful. They can reduce exposure, route material for review, or disable high-risk capabilities. They should be treated as probabilistic signals with measured false-positive and false-negative rates, not as permission to bypass downstream controls. OWASP states that foolproof prompt-injection prevention is unclear and recommends layered impact reduction, including least privilege, separation of external content, output validation, and human approval for high-risk operations.
Impact follows authority and data access
Indirect injection can corrupt a summary or recommendation even when no tool is present. With sensitive context, it can pressure the model to reveal private data in its response. With tools, it may influence messages, records, queries, code changes, purchases, or deployment actions. If generated output is embedded into HTML, SQL, shell input, or another interpreter without destination-specific validation, a model-level failure can become a conventional application vulnerability.
Durable influence requires a concrete write path and later retrieval. Preserve provenance and prevent unreviewed external text from becoming high-trust memory or policy.
Build defense in layers
Control ingress and provenance
Inventory every external content source and identify who can modify it. Preserve source URI or identifier, owner, retrieval time, integrity information, parser used, and trust classification. Quarantine unsupported formats and parse complex files in an isolated service. Sanitization can remove active file content and normalize text, but cannot safely delete every natural-language instruction without also destroying useful data.
For retrieval systems, apply the requester's tenant and object permissions before candidates enter model context. Relevance ranking must operate only over the authorized set. Monitor sudden source changes, duplicated chunks, unusual hidden text, and documents engineered to dominate results.
Minimize and separate context
Include only the data needed for the current task. Keep secrets, credentials, and unrelated conversation history out of context. Clearly delimit retrieved passages, attribute each passage, and instruct the model to use them as evidence rather than authority. This improves task clarity and supports citations, but it remains a soft model control.
Where practical, let a tool-less component extract typed facts, validate them deterministically, then pass only the approved structure forward. Extracted values still require validation.
Enforce actions outside the model
The model may propose an action; application code must decide whether it is allowed. Use narrow, typed tools and verify the authenticated subject, agent identity, action, resource, tenant, destination, constraints, and current state. The AI agent permissions guide provides a complete authorization model.
Default-deny unknown tools, parameters, and destinations. Separate read, draft, and execute capabilities. Issue short-lived, scoped credentials only after policy approval, and let the target service repeat resource-level checks. Least privilege does not stop model influence, but it turns many injected requests into denied proposals.
Validate output and contain execution
Treat model output as untrusted for its next destination. Validate structured fields against schemas and business invariants. Encode text for HTML and other rendering contexts. Never pass free-form output directly to a shell, SQL engine, template interpreter, or unrestricted URL fetcher. The LLM security risks guide explains these application trust boundaries.
Run browsers, parsers, generated code, and file transforms in disposable environments with minimal filesystem and network access. Require transaction-specific human approval for destructive operations, external communications, privilege changes, or unusual data movement. Approval should show the exact arguments and expire if they change.
Monitor and test the whole path
Log source identifiers, retrieved chunks, model and policy versions, proposed tool calls, authorization decisions, destinations, approvals, and results without indiscriminately retaining secrets. Alert on repeated denials, cross-tenant identifiers, new destinations, unusual retrieval patterns, unexpected write requests, and high-volume output.
Test each breakpoint and the full path across representative sources, formats, retrieval ranks, models, tools, and attacker budgets. Report retrieval, model influence, action request, policy decision, and effect separately; never treat one study's success rate as universal.
Developer checklist
- Classify every model input as trusted instructions, authenticated user data, or untrusted external content.
- Inventory who can modify each web, mail, document, API, repository, and RAG source.
- Preserve provenance and authorization metadata through parsing and retrieval.
- Apply tenant and object permissions before content reaches model context.
- Exclude credentials and unrelated sensitive data from prompts and memory.
- Treat delimiters and detectors as soft signals, not security boundaries.
- Give the model narrow typed tools rather than generic network, shell, or database access.
- Reauthorize every proposed action and destination outside the model.
- Separate read, draft, and execute; require exact approval at high-impact boundaries.
- Validate model output for the next interpreter and constrain egress.
- Prevent untrusted content from becoming durable high-trust memory automatically.
- Log the retrieve-to-consequence chain and make workflows stoppable.
- Test retrieval, influence, policy enforcement, and consequence as separate outcomes.
The AI agent security guide places these controls within the broader identity, memory, isolation, and incident-response architecture.
Sources
- OWASP LLM01:2025 Prompt Injection
- OWASP Agentic AI — Threats and Mitigations
- MITRE ATLAS
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems