Web Security
Cache Security for Personalized AI Responses
Design tenant-aware LLM caches that preserve personalization without leaking responses, retrieval results, tool data, or stale permissions.
An AI response can be perfectly generated and still be unsafe to reuse from a cache. A cache hit proves only that a key matched stored bytes; it does not prove that the current requester may see those bytes, that the tenant still owns the underlying data, or that the policy used to create the answer is still current. Secure LLM caching therefore treats every cached result as data that must pass the same identity, authorization, freshness, and deletion rules as a live response.
Why an AI cache is a security boundary
Traditional HTTP caches usually reason about a URL, method, headers, and freshness. AI applications add hidden inputs: tenant, user role, retrieved documents, memory, tool results, model, prompt template, safety policy, locale, and feature flags. If any of those affect the output but are absent from the cache key or authorization check, one user can receive another user’s personalized answer. This is a confidentiality failure, not merely a stale-content bug.
The application must make a deliberate decision about what is cacheable. A public, identical answer may be shared broadly. A response containing private retrieval results, account state, or tool output normally needs a principal- or tenant-scoped cache, or no shared cache at all. Read RAG security for the upstream data boundary and multi-tenant LLM security for tenant context that must survive retrieval and storage.
Cache taxonomy and ownership
Name each cache and its owner before adding a TTL:
- Response or completion cache: stores a model answer for a prompt and configuration.
- Prompt/output fragment cache: reuses deterministic template sections or embeddings of prior work.
- Retrieval cache: stores document IDs, chunks, or ranked results from a RAG query.
- Embedding cache: reuses a vector for identical source content and model version.
- Tool-result cache: stores an API response returned to an agent.
- Conversation or summary cache: holds state derived from a user’s messages.
- Browser, CDN, or HTTP cache: may retain rendered pages or API responses outside the application.
- Application or database cache: memoizes authorization, configuration, or personalization data.
For each one, document the data classification, allowed sharing scope, eviction owner, encryption boundary, and deletion process. A cache entry should have an owner or scope that can be checked, not just a string key that happens to contain a user ID.
The cache decision path
Use this gate for every lookup:
AI CACHE TRUST DECISION
Request identity and tenant → classify requested data → build key from all security-relevant inputs → check current authorization → check policy/model/source versions → evaluate freshness and revocation → return, recompute, or reject
The cache lookup is an optimization after the trusted application has resolved the principal and tenant. It is not an authorization service. For a personalized answer, the server should verify that the requester still has access to the conversation, documents, and tools represented by the entry. A hit can be discarded and recomputed when a role, membership, document ACL, or safety policy changes.
Design keys around security context
Start with a canonical representation rather than concatenating arbitrary strings. Depending on the cache, the key may need a normalized request, tenant identifier, principal or audience class, resource IDs and versions, model and provider, system/prompt-template version, policy version, retrieval filter, locale, and feature flags. Include a schema version so a change in key semantics cannot collide with old entries.
Do not put secrets or raw private prompts in a key that appears in metrics or a shared cache. Hash sensitive components with a documented algorithm and preserve a separate reference for debugging. A hash prevents casual disclosure but does not grant permission; the lookup still needs authorization.
Choose scope deliberately. A global cache is suitable only for content whose answer and authorization are genuinely identical for every requester. Tenant-scoped entries support shared workspace knowledge while preserving isolation. Principal-scoped entries are safer for personalized chats, recommendations, and account data. If proving equivalence is difficult, do not share the entry.
RAG, memory, and authorization changes
Caching retrieval results is especially risky because semantic similarity is not ownership. Apply tenant and document filters before querying the vector store, and include the effective filter or policy version in the cache identity. When a source document is reclassified, removed, or shared with a different group, invalidate affected retrieval and response entries. Do not assume that deleting the raw file automatically deletes chunks, embeddings, summaries, or cached answers.
The same rule applies to memory. A summary generated while a user had access may outlive that access. AI chat session security explains why a conversation or session identifier is not a continuing grant. On cache hit, re-check current conversation and memory permissions instead of trusting the authorization decision stored with an old result.
HTTP and browser cache controls
Use HTTP caching directives that match content sensitivity. Private personalized responses should normally be marked so shared intermediaries do not reuse them; public responses may be cacheable only when their content and authorization are truly public. Use Vary for request headers that legitimately change representation, but do not rely on it to encode an entire authorization model. Review framework defaults for server components, edge caches, and data-fetch libraries.
Browser caches can retain HTML, JSON, images, and streamed responses after logout. Decide whether sensitive routes should be revalidated or prevented from storage, and clear client state on logout where appropriate. A CDN or reverse proxy must receive an explicit cache policy; otherwise a default may be broader than the application expects. Treat service-worker caches as another data store with an owner and deletion path.
Tool results and provider caching
Tool-result caches require both data and action review. A weather lookup may be shareable for a short period; an account balance, authorization decision, or pending transaction usually is not. Never replay a cached result as proof that an action is still permitted. Before a write, execute a fresh authorization and idempotency check.
Model providers may offer prompt or prefix caching with provider-specific retention and isolation semantics. Treat those claims as an implementation dependency: understand what is retained, who can address it, how deletion works, and whether tenant data can be co-located. Do not place credentials or unnecessary personal data in a prompt merely because a provider says a cache is private.
Poisoning and freshness
An attacker who can influence a cache key, source, or write path may poison later answers. Validate inputs before storing, bind entries to provenance, and prevent untrusted users from selecting another tenant’s namespace. Prompt injection in retrieved material can become durable if its output is cached; keep source IDs and trust labels so reviewers can invalidate affected entries.
TTL is only one freshness control. Define invalidation triggers for ACL changes, tenant membership changes, model or prompt-policy releases, source updates, tool revocation, account deletion, and incident response. Short TTLs reduce exposure but do not replace event-driven invalidation for urgent revocation. Versioned namespaces make broad invalidation predictable when a policy changes.
Encryption at rest protects a stolen cache volume; it does not fix a wrong key or an over-broad reader. Restrict cache service access, isolate environments, rotate keys, and keep management operations auditable. Apply secure credential handling so cache backends do not receive long-lived provider secrets in application data.
Cache security matrix
| Cache | Main exposure | Safe default | Required evidence |
|---|---|---|---|
| Personalized response | cross-user disclosure | principal-scoped or disabled | key dimensions and authorization result |
| Tenant RAG results | cross-tenant retrieval | tenant filter before lookup | tenant, policy version, source IDs |
| Embeddings | stale or unauthorized index data | versioned, tenant-aware namespace | source hash, model version, deletion status |
| Tool result | replayed private or unsafe state | short TTL; no write authorization reuse | tool, principal, freshness, outcome |
| Conversation summary | context leakage after revocation | conversation ownership check on hit | conversation ID, owner, expiry |
| Browser/CDN | shared intermediary exposure | explicit private/public directives | response headers and purge record |
| Configuration/policy | stale enforcement | versioned and centrally invalidated | policy version and rollout event |
Observe, test, and delete
Emit cache-hit and miss events with a pseudonymous principal, tenant, cache type, key hash, policy/model version, freshness decision, and outcome. Do not log full prompts, tokens, or retrieved documents by default. Alert on hits across unexpected tenants, sudden miss/hit changes after an ACL event, and cache namespaces accessed by an unapproved workload.
Security tests should attempt cross-user and cross-tenant reads, role revocation after population, document deletion, prompt-template changes, locale collisions, stale tool results, browser back-button access after logout, and concurrent invalidation. Test both positive sharing cases and denials. Verify that a purge reaches raw cache, CDN, browser/service-worker, retrieval, embedding, and derived response stores; deletion is a workflow, not a single Redis command.
Practical checklist
- Inventory every AI, HTTP, browser, CDN, and provider cache.
- Classify entries and choose public, tenant, principal, or no sharing.
- Build canonical keys containing every security-relevant input and a schema version.
- Re-authorize cache hits against current tenant, resource, and conversation policy.
- Filter RAG data before retrieval and preserve provenance and source versions.
- Define ACL, model, policy, deletion, and incident invalidation triggers.
- Keep tool-result caches from authorizing writes or replaying secrets.
- Restrict cache service access, encrypt where appropriate, and redact telemetry.
- Test cross-scope leakage, revocation, stale data, and complete purge propagation.
Sources
OWASP Web Cache Poisoning guidance, RFC 9111 HTTP Caching, MDN HTTP caching, and framework cache documentation provide the baseline. The AI Cache Trust Decision and matrix are PermsAI’s security synthesis.