AI Security

RAG Security: Protecting Retrieval-Augmented Generation Systems

An end-to-end RAG security architecture for source governance, retrieval authorization, tenant isolation, poisoned content, safe output, and audit evidence.

By PermsAI Editorial Team
RAG Security: Protecting Retrieval-Augmented Generation Systems featured image

RAG security protects the data and decisions that surround retrieval-augmented generation: which sources enter the knowledge base, who may retrieve each item, how retrieved text is treated in model context, and where generated output is allowed to go. A vector store can improve relevance, but similarity search is not authorization and retrieved text is not trusted instruction.

A secure RAG system authorizes before retrieval, preserves provenance, treats chunks as untrusted data, validates output for its destination, and logs the connected policy, sources, model, and result.

RAG architecture in security terms

NIST defines RAG as pairing a generative model with a separate retrieval system or knowledge base. A query identifies relevant information, which is provided to the model as context for its response. In a typical implementation, the flow is:

Data sources → ingestion and parsing → chunking and embeddings → index or vector store → query and retrieval → context builder → LLM → output or action

A document can be legitimate but unauthorized, a relevant chunk can be malicious, and an accurate quotation can still violate access policy. Security must cover the pipeline, not only the model endpoint.

Map the trust boundaries

Start with identities, assets, and transitions:

BoundarySecurity questionEvidence required
Source dataWho owns the source, who can modify it, and is ingestion allowed?Source ID, owner, classification, integrity and approval
Ingestion pipelineCan parsers or transforms be abused, and what metadata survives?Parser version, scan result, extracted types, provenance record
Index or vector storeAre chunks isolated, encrypted where appropriate, and removable?Tenant, document ID, ACL reference, index version, retention state
QueryWhich authenticated subject and tenant requested this task?User/workload identity, session, purpose and query policy
RetrievalIs every candidate authorized before it becomes context?Policy decision, applied filters, selected document and chunk IDs
Context builderAre trust labels preserved and sensitive data minimized?Context manifest, token selection, source attribution
ModelWhat configuration processed the context?Provider/model version, system policy and evaluation version
Output consumerIs the response merely displayed or used by another interpreter or tool?Validation result, destination, approval and action outcome

This inventory prevents “the RAG system” from becoming one opaque component with no accountable control owner.

Retrieval authorization is not similarity search

The fact that a document exists in a vector store does not mean the current user may retrieve it. Embedding proximity answers “which items appear semantically relevant?” It does not answer “which principal may access this resource under these conditions?”

Build the authorized candidate set from authenticated context and authoritative policy. Scope by tenant, user or group, document ownership, classification, workflow purpose, and current resource state. Then rank only candidates the requester is allowed to see. Where the store supports metadata filters, use them as one enforcement layer, but validate sensitive resource access in application code or the system of record as well.

Filtering after retrieval is weaker. Unauthorized text may already have crossed into logs, caches, traces, or model context before a late check removes it from the displayed answer. Pre-retrieval or retrieval-time enforcement also reduces side channels from result counts and metadata. Fail closed if tenant, ownership, or policy data is missing.

The AI agent permissions guide provides a deterministic model for subject, action, resource, context, constraints, and approval. Apply the same principle to a retrieval action: the model may formulate a query, but it cannot grant access to the matching documents.

Prevent cross-tenant leakage

In multi-tenant RAG, tenant identity must come from the authenticated session or workload identity—not the prompt and not a model-selected parameter. Bind each chunk to a canonical tenant and source resource. The retrieval service should reject candidates whose tenant and ACL do not match the authoritative request context.

Separate indexes and shared indexes with mandatory filters have different operational tradeoffs; neither is automatically secure. Test caches, hybrid search, rerankers, background jobs, backups, and administrative paths.

Use known document identifiers from another tenant in negative tests. Verify that the API returns no content or revealing metadata, that the model context does not contain the chunk, and that an audit event records the denied attempt. OWASP's current vector and embedding guidance specifically identifies unauthorized access, data leakage, and cross-context leakage as RAG risks and recommends fine-grained access control and logical partitioning.

Poisoned content and indirect prompt injection

A RAG knowledge base can be poisoned intentionally through a malicious upload, compromised source, public page, support ticket, or insider modification. It can also accumulate unsafe content accidentally through stale, low-quality, or wrongly classified sources. Once retrieved, text may mislead the answer, impersonate policy, or attempt to control the model.

This is indirect prompt injection: hostile instructions travel inside content rather than through the user's direct request. Do not assume RAG neutralizes prompt injection. OWASP states that retrieval and fine-tuning do not fully mitigate it, while Rag and Roll demonstrates why retrieval probability and downstream model influence must be evaluated as separate stages.

Control both halves. At ingestion, restrict sources, preserve ownership, scan and parse safely, detect suspicious changes, and require review for high-trust collections. At generation, label retrieved passages as untrusted evidence, minimize context, constrain model capabilities, and keep authorization and tool execution outside the model. A poisoned chunk may still influence text, so least privilege must limit the consequence.

Preserve data provenance and lifecycle

Every chunk should remain traceable to its source document and security metadata. Record source identity, owner, tenant, classification, ingestion time, content version or integrity digest, parser and embedding versions, approval state, and retention policy. Keep a stable document ID even if chunks are regenerated.

Provenance supports explanation, incident response, and reliable update or deletion of derived chunks, embeddings, caches, and summaries. “Internal” is not an instruction-trust label: an internal wiki may be broadly editable. Track confidentiality and content-control risk separately.

Secure the ingestion pipeline

Authenticate source connectors, authorize collections, restrict file types and sizes, isolate complex parsers, and limit their network and filesystem access. For sensitive collections, quarantine, parse, classify, review when warranted, then promote to a serving index. Re-ingestion must not loosen permissions. Sanitization can reduce parser and active-content risk, but natural-language instructions can look like valid prose; a scanner or phrase filter does not make content trusted for model obedience.

Protect the index and vector store

Apply ordinary datastore controls: authenticate services, authorize query and administrative operations, restrict network reachability, encrypt according to the threat model, protect backups, rotate credentials, patch dependencies, and monitor exports. Embeddings are not safely anonymous by default; minimize stored data and protect metadata. ACL, tenant, and deletion changes must propagate to every serving replica and cache, and rollback must not restore stale access.

Build context from authorized, attributable data

Include only authorized chunks needed for the question, preserve source identifiers, and ensure citations resolve to resources the user may still open. Separate application instructions from evidence and tell the model not to treat passages as policy or commands, while recognizing this is a behavioral control. Exclude credentials and unrelated sensitive history; validate any extracted structured fields before use.

Treat output according to its destination

A RAG answer can disclose unauthorized facts, reproduce hostile content, or emit unsafe text. Recheck citation authorization, apply application-specific sensitive-data controls, encode HTML, validate structured output, and never pass free-form text directly to SQL, a shell, a template engine, or a privileged API. The LLM security risks guide covers these boundaries. External actions require deterministic policy and transaction-specific approval; retrieved evidence cannot manufacture authority.

Log enough to investigate without creating another leak

Correlate subject, tenant, privacy-preserving query reference, policy decision, filters, retrieved IDs and scores, context manifest, model version, output validation, and action. Protect logs with access control, retention, redaction, and integrity; copying raw queries and contexts into broad analytics creates another disclosure surface. Alert on cross-tenant identifiers, denied-document probes, sudden retrieval concentration, bulk access, repeated policy failures, and stale indexes.

The RAG Security Control Map

PermsAI's control map assigns an asset, control, and auditable evidence to each lifecycle phase.

PhaseMain asset and threatRequired controlEvidence to log
IngestSource integrity; malicious or unauthorized documentsConnector authentication, source allowlist, parser isolation, classification and approvalSource, owner, digest, parser/scan result, decision
IndexChunks, embeddings, metadata; leakage or stale permissionsTenant/resource binding, access control, encryption as required, versioned update/deleteDocument/chunk IDs, tenant, ACL reference, index version
RetrieveAuthorized knowledge; cross-user or cross-tenant disclosureAuthenticated scope, policy evaluation, permission-aware filters, defense-in-depth resource checkSubject, purpose, filter, decision, selected IDs and scores
GenerateContext and model behavior; injection or unsupported outputProvenance-aware context, minimization, trust labels, model safeguardsContext manifest, model/prompt version, citations
Act or outputUser data and target systems; disclosure or unsafe executionDestination validation, rendering safety, tool authorization, approval, egress limitsValidation, destination, approval, action and outcome

The map is useful during design reviews because every phase has a control owner and a testable artifact. A team cannot claim that a model prompt compensates for a missing retrieval authorization decision.

A secure RAG request path

A defensible architecture looks like this:

  1. Authenticate the user or workload and derive tenant and session scope.
  2. Authorize the search purpose and construct mandatory resource filters.
  3. Retrieve and rank only within the authorized candidate set.
  4. Recheck sensitive resources and attach provenance to selected chunks.
  5. Build a minimal context that labels each passage as untrusted evidence.
  6. Ask the model to answer with attributable sources, without credentials or unnecessary data.
  7. Validate citations, sensitive output, structure, and rendering for the destination.
  8. Route any proposed action through a separate authorization and approval gateway.
  9. Record the connected policy, retrieval, context, generation, and outcome evidence.

NIST IR 8579 documents security decisions for an internal NCCoE RAG chatbot, including access controls and validation filters, but explicitly describes a point-in-time prototype rather than implementation guidance. Use it as a transparent case study, not a universal blueprint.

RAG security checklist

  • Inventory every source, owner, modifier, confidentiality class, and content-control risk.
  • Authenticate connectors and authorize ingestion into each collection.
  • Sandbox parsers and restrict file types, size, network, and filesystem access.
  • Preserve tenant, document, ownership, ACL, integrity, and lifecycle metadata per chunk.
  • Treat embeddings and metadata as protected data, not anonymous by default.
  • Build the retrieval candidate set from authenticated tenant and resource permissions.
  • Test cross-user and cross-tenant access with known forbidden document IDs.
  • Keep retrieved passages untrusted and exclude credentials and unnecessary secrets from context.
  • Detect and review poisoned or suspicious sources without relying on phrase filters alone.
  • Validate citations and source authorization before returning sensitive answers.
  • Encode and validate output for its destination; never execute free-form text directly.
  • Gate tool actions with deterministic authorization and meaningful approval.
  • Propagate permission changes and deletion to indexes, replicas, caches, and serving paths.
  • Log policy and source identifiers while minimizing sensitive query and content retention.
  • Re-run ordinary and adversarial tests after source, parser, embedding, model, or policy changes.

For the wider attack model, see prompt injection attacks. The dedicated indirect-injection guide explains how hostile content moves from retrieval to consequence, while this page owns the complete RAG data and authorization lifecycle.

Related empirical evidence

Use RAG Security Research Review for the current evidence on poisoning, retrieval manipulation, leakage, provenance, and the limits of proposed defenses.

Sources