AI News
RAG Security Research Review: Poisoning, Leakage, and Provenance
Current RAG security evidence separates poisoning, leakage, access control, and provenance, then maps research claims to defensible controls.
Current RAG security evidence: separate the trust boundaries
Evidence cutoff: 2026-09-17. Retrieval-augmented generation is not one attack surface. A corpus can be poisoned; a ranking can be manipulated; retrieved text can attempt indirect prompt injection; a caller can receive data outside its entitlement; and a team can lose the origin and version evidence needed to investigate. Calling all of these “RAG poisoning” hides different assumptions and different controls.
This research review complements RAG Security, which owns evergreen architecture. Here the focus is empirical claim scope. A successful result in a public benchmark says something about the evaluated retriever, embedding model, corpus, generator, attacker knowledge, and success metric. It does not by itself establish compromise of another organization’s private corpus or access-controlled deployment.
What selected research establishes
Zou, Geng, Wang, and Jia’s PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models (arXiv preprint, February 12, 2024) formally studies targeted knowledge corruption. Its reported setting allows an attacker to inject a small number of texts per target question into a large knowledge base and evaluates selected defenses. The headline result is evidence that retrieval can be a practical influence channel under that ingestion assumption; it is not a recipe or a rate for managed corpora with controlled write access.
Provably Secure Retrieval-Augmented Generation by Zhou, Feng, and Yang (arXiv preprint, August 1, 2025) proposes encryption and formal confidentiality/integrity claims under a stated computational model. Formal guarantees are valuable where their model holds, especially for authorized access to stored content and embeddings. They do not establish that retrieved claims are true, benign, relevant, or safe to obey as instructions.
Security and Privacy in Retrieval-Augmented Generation by Palanisamy, Chalapathi, Hassija, and Buyya (arXiv survey preprint, June 24, 2026) maps poisoning, membership and index inference, leakage, and deployment variants. It is a survey, not a new benchmark result. ToxicRAG (arXiv preprint, September 2026) is newer single-shot poisoning research; as of the cutoff it should be treated as preliminary and evaluated against its disclosed corpus and metrics before transfer claims.
RAG SECURITY EVIDENCE MAP
| Layer | Poisoning/leakage/provenance question | Defensive evidence to request |
|---|---|---|
| Source | Who supplied the document and under what authority? | Source identity, admission policy, signature or accountable owner. |
| Ingestion | Was content transformed, accepted, or quarantined correctly? | Validation logs, scanner outcomes, ingestion time, immutable version. |
| Storage/index | Can an unauthorized principal read, alter, or infer indexed content? | Access policy tests, encryption scope, index change history. |
| Retrieval | Did the caller receive only eligible documents and was ranking manipulated? | Entitlement-before-retrieval test, query/result trace, anomaly review. |
| Context assembly | Is untrusted prose labeled as data rather than instructions? | Source labels preserved into prompt construction and adversarial test traces. |
| Generation | Does the answer expose unnecessary sensitive content or overstate evidence? | Citation checks, redaction tests, output review sampling. |
| Tool/action boundary | Can retrieved prose cause a side effect beyond task authority? | Scoped tool policy, allowlist, approval and execution logs. |
The map makes a crucial distinction. Corpus poisoning assumes a path to add or alter material. Retrieval manipulation may instead affect ranking or queries. Indirect injection requires content to influence instruction following. Sensitive leakage can occur even with perfectly benign content when retrieval ignores caller entitlements. Provenance failure means the system cannot reconstruct origin or transformations; it does not itself prove maliciousness.
RAG RESEARCH EVIDENCE MATRIX
| Study | Threat | System setup | Attack assumptions | Metric | Defense | Transfer evidence | Limitation |
|---|---|---|---|---|---|---|---|
| PoisonedRAG (2024 preprint) | Targeted corpus corruption | RAG knowledge base and selected LLM/retrieval settings | Ability to inject selected texts | Target-answer success | Evaluated filters/defenses | Limited to declared datasets and configurations | Write access, corpus and target assumptions bound the result. |
| Provably Secure RAG (2025 preprint) | Confidentiality/integrity of stored/retrieved data | Encrypted storage and embedding design | Computational adversary in formal model | Formal security plus benchmark experiments | Encryption and authorized access | Security theorem transfers only with model assumptions | Does not establish semantic truth, relevance, or injection resistance. |
| RAG security/privacy survey (2026 preprint) | Taxonomy including poisoning and leakage | Literature across centralized, on-device, federated, hybrid RAG | Varies by cited work | Survey synthesis | Architectural/algorithmic/cryptographic families | Identifies gaps, not a single replicated rate | Heterogeneous studies are not directly comparable. |
| ToxicRAG (2026 preprint) | Single-shot knowledge poisoning | Declared benchmark RAG setup | Manipulated content admission | Reported task-specific success | Proposed/evaluated defenses | New evidence only | Preprint and setup-specific; independent replication absent at cutoff. |
A study should record corpus source, retriever, embedding model, reranker, generator, dataset, attacker capability, manipulated-document count or rate, success criterion, baseline, samples, retries, and whether adaptation was allowed. Without those fields, a number cannot be compared safely. Do not infer transfer across embedding models or ranking policies merely because the generated answer looks similar.
Leakage and access control are separate questions
RAG can leak through index inference, retrieved passages, logs, citations, or generated summaries. An ingestion filter cannot correct a retrieval authorization mistake. Check authorization before retrieval, not only after the model has seen a passage. Then constrain context assembly to the approved result set and make redaction decisions auditable.
Test with paired principals: one entitled and one denied, identical query wording, and a documented expected result. Verify no sensitive identifier appears in retrieved chunks, model context, output, trace, cache, or retry artifact for the denied principal. Repeat after index rebuilds, connector changes, and document reclassification. AI Memory Security is relevant because durable context can become another retrievable store.
Provenance is an evidence system, not a truth oracle
A provenance record can make origin and history auditable. It does not automatically mean a source is accurate, benign, or appropriate for every use. The record below is a minimal design target, not a standard schema.
RAG PROVENANCE RECORD
| Field | Meaning | Verification question |
|---|---|---|
| document/source ID | Stable origin reference | Can a reviewer locate accountable source ownership? |
| version | Immutable or versioned content identity | Was the exact retrieved revision preserved? |
| ingested_at | Time and pipeline event | Which policy/configuration admitted it? |
| hash/signature | Integrity evidence where applicable | Did content change after validation? |
| access classification | Entitlement label | Was it checked before retrieval? |
| chunk lineage | Segment and transformation history | Which source span produced this context? |
| retrieval event | Query, rank, policy decision, timestamp | Why did this caller see this chunk? |
| citation/output trace | Answer claims connected to evidence | Can a reader distinguish source text from synthesis? |
This record helps containment: suspicious chunks can be quarantined, affected answers found, and policies retested against the same version. It also helps quality review, but it cannot determine whether a signed source contains misleading prose. Combine provenance with trusted-source admission, change control, anomaly monitoring, and human review for high-impact domains.
Defense evidence and practical priorities
Trusted-source policies reduce the number of unaccountable write paths. Ingestion validation and continuous corpus monitoring make changes visible. Access control before retrieval confines data exposure. Reranking, classifiers, and sanitization may reduce specific manipulation outcomes, but sanitation alone does not solve semantic attacks: fluent malicious text can remain semantically persuasive. Instruction/data separation limits how retrieved prose is interpreted, while least privilege and action authorization limit consequences if it still influences a model.
Evaluate defenses as systems. Use a clean baseline and a controlled, authorized test corpus; measure retrieval quality, task utility, false positives, authorization failures, and investigation completeness alongside attack-oriented metrics. Include static and adaptive variants where safe. Preserve configurations so a later model or embedding update can be compared. Indirect Prompt Injection and AI Agent Permissions cover the adjacent context-to-action boundary.
As of 2026-09-17, the selected literature supports a bounded conclusion: RAG adds distinct ingestion, retrieval, authorization, and provenance surfaces; selected studies demonstrate vulnerability under their assumptions and several defense directions. It does not support the claim that any filter, encryption layer, or provenance field universally prevents poisoning, leakage, or instruction influence.
Evaluation protocol: verify both security and usefulness
Treat every defense as a measurable change to a retrieval pipeline. Pin the corpus snapshot, ingestion configuration, chunking algorithm, embedding and reranker versions, generator, system instructions, tool policy, and evaluator. Use an explicit success predicate: a wrong answer, an unauthorized disclosure, a target answer, an unsafe action, and a missing provenance trace are not interchangeable. Measure ordinary retrieval relevance and answer quality at the same time, because a defense that simply removes useful documents can create a misleading security score.
For adaptive testing, disclose what the evaluator knows about the ranking policy and detector. For access tests, use approved synthetic identities and no live sensitive records. Keep a trace for every sampled answer so reviewers can distinguish the retrieval result from model synthesis. Re-run the suite after a corpus update, model migration, or connector change. This does not create universal robustness; it creates a change-detection discipline that can catch regressions before they become claims about safety.
Entitlement-aware retrieval
A document’s presence in a corpus is not proof that it is retrievable by every principal. Filtering only after generation is too late: the model, trace, cache, or citation candidate may already contain material outside the caller’s entitlement. Apply access policy before or during candidate retrieval, carry the principal and policy decision through reranking and context assembly, and record the reason a document was included or excluded.
This is distinct from relevance. A highly ranked passage may be useful yet unavailable to this caller; an authorized passage may be low quality or untrusted. Evaluation should therefore pair the same query with entitled and denied principals, verify that no denied chunk enters the context or retry path, and measure retrieval utility for allowed users. The generator’s refusal is a supplementary control, not the primary confidentiality boundary.
Provenance validation limits
Known origin, known version, and known transformation lineage are evidence of history, not evidence that content is truthful, safe, authorized, or non-malicious. A signed supplier document can still be stale; an intact internal version can be misclassified; a traceable chunk can still contain an instruction that should remain data. Provenance becomes useful when combined with admission controls, access policy, retrieval telemetry, and a citation/output trace that lets reviewers reconstruct how a claim reached an answer.
During review, ask four separate questions: who supplied this version; was it admitted under a current policy; was this principal entitled to retrieve it; and did the answer accurately represent it? The answers can differ. This separation avoids treating provenance labels as a universal content-safety assertion while preserving the evidence required to quarantine a bad document and identify affected answers.
Sources
- Wei Zou, Runpeng Geng, Binghui Wang, Jinyuan Jia, PoisonedRAG, February 12, 2024, arXiv preprint. Source. Targeted corruption evidence; corpus-write and setup assumptions limit transfer.
- Pengcheng Zhou, Yinglun Feng, Zhongliang Yang, Provably Secure Retrieval-Augmented Generation, August 1, 2025, arXiv preprint. Source. Formal confidentiality/integrity model; semantic safety is out of scope.
- Balamurugan Palanisamy et al., Security and Privacy in Retrieval-Augmented Generation, June 24, 2026, arXiv survey preprint. Source. Taxonomy and trade-offs; not one comparable experiment.
- ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks, September 2026, arXiv preprint. Source. New preliminary evidence; replication and transfer remain open.
- NIST, AI Risk Management Framework, official guidance, accessed 2026-09-17. Source. Risk-management context; does not prescribe one RAG architecture.
- OWASP, LLM Top 10 / GenAI Security Project, official project material, accessed 2026-09-17. Source. Threat/control context; not empirical proof of a specific defense.