AI Security

AI Memory Security: Preventing Context and Memory Poisoning

A practical guide to preventing context and memory poisoning through controlled writes, provenance, scoped retrieval, expiry, and recovery.

By PermsAI Editorial Team
AI Memory Security: Preventing Context and Memory Poisoning featured image

AI memory security is the practice of treating persistent or reusable context as a protected data and decision boundary. An agent should not treat a remembered item as true, authorized, or safe merely because it was retrieved. The key question is not whether a system "has memory," but who can write which kind of state, who can read it, how long it survives, and how the team can correct or revoke it.

This differs from ordinary prompt injection. Indirect Prompt Injection focuses on untrusted content influencing an active interaction. Memory security focuses on the extra persistence step: untrusted or incorrect context is accepted, stored, and later reused. RAG Security covers retrieval architecture broadly; this article concentrates on durable agent memory.

Memory is an architectural taxonomy, not one feature

Products use the word memory for several different mechanisms. These categories are useful for threat modeling, not universal standards:

Memory mechanismTypical persistenceMain security question
Conversation historySession or accountWho can reopen or summarize it?
Summary memoryCross-sessionDid compression lose provenance or introduce a false claim?
User profileLong-livedIs it correct, consented to, and scoped to the user?
Vector-store memoryLong-lived retrievalWhich source and tenant produced the retrieved item?
Episodic/task memoryJob or workflowCan a stale plan or result be reused after conditions change?
Scratchpad/stateShort-lived executionCan untrusted tool output become control data?
Structured key/value memoryVariableWhich schema, writer, and retention rule govern each field?

The same database can contain several of these forms. They should not all receive the same write policy, access scope, retention period, or trust label.

Why persistence changes the threat model

Without persistence, malicious content may influence one model call and then leave the active context. With persistent memory, a bad item can survive the session, be retrieved later for a different request, and shape a future recommendation or tool proposal. The security path is:

untrusted input → accepted as memory → stored → later retrieved → future behavior influenced

Research on memory poisoning describes this durable influence channel; its reported experimental attack rates are not a guarantee about every deployment. The durable architectural point is stronger: a write accepted today can affect a user or agent after the original source and interaction are no longer visible.

Potential sources include malicious user content, indirect injection in a document or web page, compromised tool output, a poisoned knowledge source, a faulty summarizer, a cross-tenant cache error, and model-generated state that was never verified. Do not need an attacker to directly access the database before treating this as a real trust boundary.

Make memory writes an authorized operation

The model should not automatically persist arbitrary prose because it labels it "important." A memory write is a state-changing action with an owner, scope, and future consequences. Define:

  • which principals may create, update, supersede, or delete each memory type;
  • which sources are eligible for durable storage;
  • which fields require a schema, validation, or explicit user confirmation;
  • what tenant, user, project, or task scope attaches to the item;
  • what trust status the item receives at creation; and
  • when a reviewer, service, or deterministic rule must approve the write.

A preference explicitly supplied by the authenticated user may be eligible for a narrow profile field. A claim extracted from an untrusted webpage should remain attributed evidence, quarantined, or ephemeral—not become an instruction or trusted user fact. When a memory can influence a consequential action, the write policy should be at least as deliberate as the later action policy. AI Agent Permissions describes the separate authorization decision.

Read authorization prevents privacy and tenant failures

Retrieval needs its own policy. User A's memory must not become User B's context; Tenant A's item must not be retrieved for Tenant B; a general worker should not read an executive's profile merely because semantic similarity is high. Apply identity, tenant, user, project, purpose, classification, and allowed-memory-type filters before similarity search, ranking, or context assembly.

Relevant is not equivalent to authorized. A retrieved memory is also not automatically true or a safe instruction. Pass only the minimum authorized fields to the model, keep the item identifier and provenance available to the application, and apply destination-specific authorization again when the model proposes an action.

The Memory Trust Table

Memory sourcePersistenceTrust statusWrite authorizationRead scopeExpiry and audit
User-confirmed preferenceLong-livedConfirmed but revisableAuthenticated user or approved support workflowSame user/tenantRevision history, review date
Application-owned policy factVersionedControlledRestricted publisher pipelineAuthorized workloadsVersion, signer, rollout, rollback
Retrieved document claimShort or attributedUntrusted evidenceIngestion policy onlyAuthorized task scopeSource, retrieval time, TTL
Tool resultTask-scopedContext-dependentTyped tool and schemaSame job/tenantCorrelation ID, expiry
Model-generated summaryReviewableDerived, not authoritativeValidator or confirmation gateNarrow original scopeInputs, model/version, supersession

Trust labels should affect how the system uses an item. An untrusted retrieved note may help propose a question or point to a source; it should not instruct a tool executor, widen access, or silently overwrite a confirmed profile.

Validate before persistence, but do not promise a magic filter

Structured schemas, type checks, allowed memory classes, length limits, source allowlists, duplicate detection, and user confirmation can reduce accidental writes. Deterministic rules can reject a profile update that contains a tool directive or a tenant identifier inconsistent with the authenticated session.

Those controls do not solve semantic poisoning by themselves. A well-formed, plausible false statement can still be harmful. Preserve provenance and trust status, require corroboration or confirmation for high-impact facts, and avoid converting a source's instructions into durable operational policy. Treat every unverified memory as data, not control flow.

Expiry, conflict, correction, and integrity

Some memory should expire quickly: task notes, volatile status, temporary access decisions, and external claims. Other items need a revision schedule rather than perpetual retention. Attach TTL, review date, version, and a clear owner. Expiration is both a security and correctness control: stale instructions and former roles should not silently remain influential.

When new information conflicts with an existing item, do not blindly append another paragraph. Preserve a versioned history, mark the older claim superseded or disputed, record why, and choose a retrieval rule that exposes the conflict rather than presenting an arbitrary winner as fact. For sensitive stores, protect append-only audit history, integrity checks, and access-controlled update paths so a compromised agent cannot rewrite its own evidence.

Retrieve safely and preserve observability

At retrieval time, log the memory ID, source, writer, tenant, user scope, trust status, timestamp, consumer, query purpose, and policy decision. Keep the raw sensitive content out of broad logs. A relevance score is a selection signal, not a truth score or permission grant.

Use retrieved memory as untrusted context when its status or source warrants it. Separate quoted evidence from application instructions; limit how much context is injected; and do not let a memory item specify credentials, authorization headers, approval results, or arbitrary tool parameters. The system still needs policy enforcement outside the model, as explained in the broader AI Agent Security guide.

The PermsAI AI Memory Security Lifecycle

INGEST → VALIDATE → AUTHORIZE WRITE → STORE WITH PROVENANCE → AUTHORIZE READ → RETRIEVE → USE AS UNTRUSTED CONTEXT → REVIEW / EXPIRE / REVOKE

At each stage, preserve a stable memory ID and correlation ID. Ingest identifies source and owner; validation checks shape and allowed class; write authorization binds a scope; storage retains provenance and version; read authorization filters before retrieval; use does not elevate the item's trust; and review, expiry, or revocation removes its influence. This lifecycle is PermsAI's architectural synthesis, not a claim that any single standard mandates the exact sequence.

Incident response for poisoned or leaked memory

Teams should be able to answer: Which memory item is suspect? Who or what wrote it? Which source and policy admitted it? Which sessions or agents retrieved it? Which actions followed? Can it be disabled without deleting forensic evidence?

Quarantine the item from retrieval, revoke related grants or tools if needed, preserve a protected record, identify consumers through retrieval logs, correct or supersede the item, invalidate derived summaries and caches, and review downstream actions. Test these operations before an incident. The absence of a memory delete button is a serious operational gap.

Practical checklist

  1. Inventory every session, profile, vector, summary, task, and structured memory store.
  2. Define owner, tenant, user, purpose, schema, trust status, and retention for each memory type.
  3. Authorize writes independently; do not let a model persist arbitrary text by default.
  4. Filter read access before vector search and context assembly.
  5. Store provenance, writer, source, interaction, version, confidence/status, and timestamps.
  6. Treat relevance as neither truth nor permission; keep untrusted memory out of control paths.
  7. Use TTL, revision, conflict handling, and controlled deletion or revocation.
  8. Separate confirmed user facts from extracted claims and model-derived summaries.
  9. Log write, read, retrieval, supersession, and invalidation events without logging secrets.
  10. Exercise cross-tenant, stale-memory, poisoning, correction, and incident-containment tests.

Sources