Web Security
Multi-Tenant Security for LLM and RAG Applications
A practical architecture for preventing cross-tenant leakage across LLM prompts, RAG, memory, caches, tools, credentials, jobs, logs, and billing.
Multi-tenant security for LLM and RAG applications is the discipline of keeping every identity, dataset, conversation, memory item, cache entry, credential, job, log, and billable action attached to the correct tenant. A tenant identifier in a prompt is not an authorization boundary. The server must resolve and enforce tenant scope before data reaches a model or a tool.
Tenant context propagation chain
A reliable chain is Login → Membership → Tenant Context → Authorization → Data and RAG → Tool → Logging and Billing. At each stage record the tenant source, trusted enforcement point, and common failure.
| Stage | Tenant source | Trusted enforcement point | Common failure |
|---|---|---|---|
| Login | Authenticated account | Session verifier | Trusting a name in prompt |
| Membership | Server membership record | Membership policy | Accepting browser tenant ID |
| Context | Selected workspace validated against membership | Request middleware | Context overwritten by model |
| Data/RAG | Object ownership and tenant columns | Query/retrieval layer | Filtering after retrieval |
| Tool | Principal, tenant, action, resource | Tool gateway | Shared credential |
| Logs/billing | Server correlation and usage meter | Telemetry service | Model-supplied tenant tag |
Identity and request binding
A user may belong to several tenants, so the active workspace is a server-validated choice, not a permanent property of the user. Resolve it from an authenticated membership and bind it to the request, conversation, upload, background job, and tool call. Never accept a tenant ID from the model as proof of membership. If a client sends a workspace identifier, use it only to select a candidate and then verify membership and status in trusted code.
Keep user identity, tenant identity, and agent/workload identity distinct. The agent may act for a user, but its workload credential should be scoped to the tenant and operation. AI Agent Identity explains how to preserve attribution, while AI API Authorization covers the decision at an API boundary.
Database and object ownership
Every tenant-owned object should carry enforceable ownership or a relation from which ownership can be derived: conversations, documents, jobs, memories, tool configurations, API connections, and invoices. Apply the tenant and object authorization check on every read and write, including internal calls from an orchestrator. A conversation ID or document ID is only a lookup key until ownership is checked.
Prefer server-derived tenant scope in queries. Do not let a request supply both tenant_id and object_id and assume they agree. Resolve the object, compare its owner, and deny mismatches. Database row-level security can add a second boundary where appropriate, but it does not remove the need for application policy, migrations, and negative tests.
RAG and vector isolation
RAG is a data authorization problem before it is a relevance problem. Apply tenant and object permissions before keyword or vector search; semantic similarity never implies permission. Preserve tenant metadata through ingestion, chunking, embedding, indexing, query, and retrieval. If a vector store cannot enforce reliable metadata filters, use separate indexes or a retrieval service that can.
Do not assemble context and filter it afterward. A forbidden chunk may already have influenced a summary or model response. Keep provenance such as document ID, owner, classification, and retrieval time. Re-test after permission changes, document deletion, re-indexing, and embedding-model changes. See RAG Security for retrieval-specific controls.
Conversations and memory
A session or conversation ID must not authorize access by itself. Verify owner and tenant on every load, stream, export, and continuation. Memory is durable state: write only approved fields, record writer and source, and enforce read scope before retrieval. Tenant A memory must never become Tenant B context. Invalidate summaries and embeddings when source permissions change. AI Memory Security covers provenance, expiry, and poisoning controls.
Prompt and context leakage can occur through system context, retrieved chunks, summaries, tool results, error messages, and caches. Minimize fields before model invocation and mark external text as untrusted evidence. Never use the model to decide whether a record belongs to the current tenant.
Cache boundaries
Caches are a frequent cross-tenant leak because the response may look identical while permissions differ. Include tenant, user or role when needed, resource scope, model/prompt version, and policy version in response-cache keys. Retrieval, embedding, and tool-result caches need equivalent scope. Do not cache a sanitized or summarized response under a key that can serve a different tenant. Test cache hits, misses, invalidation, and permission changes explicitly.
Credentials and tools
Tenant A must not receive Tenant B credentials. Keep credential references in a server-side broker, bind them to tenant and audience, and issue only the scope needed for one tool call. Avoid one shared provider key when it prevents attribution, quota enforcement, or revocation. Secure Credential Handling for AI Agents provides the custody pattern.
Every tool call should evaluate principal, tenant, action, resource, purpose, and current state. Expose narrow tools instead of a universal database or cloud credential. Read, draft, and commit capabilities should be separate. A model proposal cannot grant access; the gateway and downstream API must enforce it.
Jobs, storage, logs, and billing
Tenant context must survive queues, workers, async ingestion, embedding jobs, and agent runs without storing unlimited authority in a payload. Put a tenant reference and task ID in the job, then reauthorize with a current workload identity when the worker starts. Reject missing, disabled, or mismatched tenant context.
Documents and uploads in object storage need tenant-scoped IDs or paths, authorization on download, and short-lived signed access where appropriate. Do not assume an unguessable filename proves ownership. Derived artifacts and exports need the same scope.
Logs can contain prompts, retrieved text, identifiers, and tool parameters. Tag events with tenant and correlation IDs, restrict access, redact sensitive content, and set retention. Billing and quota meters must attribute tokens, model calls, tool calls, and storage from server-side records, not model-provided metadata. A usage event without a trusted tenant is incomplete.
Deletion and lifecycle
Tenant deletion should cover the database, vector index, memory, caches, object storage, derived files, queues, and observability copies. Immediate physical deletion may not be possible for backups or retention archives; document the lifecycle, mark data inaccessible promptly, and prove the remaining retention path. Revoke tenant credentials and disable jobs during the process.
AI multi-tenant isolation matrix
| Layer | Tenant-owned asset | Primary failure | Required control | Verification |
|---|---|---|---|---|
| DB | Rows and relations | Object leakage | Server ownership checks | Negative API tests |
| RAG | Chunks and indexes | Cross-tenant retrieval | Filter before ranking | Retrieval isolation tests |
| Memory | Facts and summaries | Wrong context | Scoped read/write | Memory boundary tests |
| Cache | Responses and embeddings | Key collision | Tenant-aware keys | Hit/miss tests |
| Storage | Uploads and exports | Guessable access | Scoped IDs and auth | Download tests |
| Tools | API connections | Credential crossover | Tenant-bound broker | Credential tests |
| Jobs | Ingestion and runs | Lost context | Reauthorize worker | Queue tests |
| Logs | Prompts and actions | Insider leakage | Redaction and ACLs | Access review |
| Billing | Usage and quotas | Misattribution | Server meter | Reconciliation |
Testing and incident response
Maintain positive and negative tests for every tenant-owned endpoint. Attempt cross-tenant object access, retrieval, memory reads, cache reuse, signed downloads, tool calls, background jobs, and log access. Include a user who belongs to two tenants and a tenant that has just been disabled. Run tests after schema, index, cache, prompt, model, or worker changes.
When leakage is suspected, identify affected tenant, objects, sessions, caches, credentials, jobs, and logs. Revoke credentials, disable tools, invalidate caches and memory, quarantine derived artifacts, preserve evidence, notify the correct owners, and rerun isolation tests. Observability guidance in AI Agent Observability helps reconstruct the chain.
Practical checklist
- Resolve tenant from authenticated membership, never model text.
- Bind tenant to session, conversation, stream, job, tool, and billing event.
- Enforce ownership before retrieval, ranking, context assembly, and rendering.
- Include tenant scope in cache keys and invalidation.
- Keep credentials and object storage access tenant-bound.
- Reauthorize workers with current authority.
- Redact and restrict tenant-sensitive logs.
- Cover two-tenant users and disabled-tenant cases.
- Test deletion across primary and derived stores.
- Rehearse revocation and leakage response.
Sources
- OWASP ASVS and API Security guidance
- OWASP GenAI and multi-tenant security guidance
- NIST AI Risk Management Framework
- Database and vector-store access-control documentation
- Cloud object-storage and identity security documentation
Multi-tenancy is not a prompt convention. It is a chain of server-enforced bindings that must remain intact from login to retrieval, action, evidence, and deletion.