Web Security

Web Security for AI Applications: Architecture and Checklist

A practical architecture and checklist for securing web applications that combine LLMs, RAG, uploads, streaming, tools, and multi-tenancy.

By PermsAI Editorial Team
Web Security for AI Applications: Architecture and Checklist featured image

An AI web application is still a web application. It needs sound authentication, authorization, session management, input handling, output encoding, API protection, tenant isolation, secret management, logging, and secure deployment. Adding a model does not retire those controls; it adds new trust boundaries where language, retrieved data, memory, and tool proposals cross into software that can change state.

This guide gives a practical architecture review method for a web app that combines chat, retrieval-augmented generation (RAG), file uploads, streaming responses, and agent tools. Treat the model as an untrusted decision component. The server, policy layer, and destination systems remain responsible for enforcing security.

AI web application security architecture

A representative request path is:

Browser → Web Frontend → Application/API → Authentication + Authorization → AI Orchestration

The orchestration layer connects to Model, RAG/Data, Memory, and a Tool Gateway, which then reaches external systems. Each arrow is a trust boundary. Browser input is untrusted. The API establishes the user and tenant. The orchestrator assembles context but must not turn text into authority. The tool gateway validates and authorizes actions before a downstream service receives them.

A useful architecture review asks five questions at every hop: what identity is present, what data crosses, what can be changed, which policy is enforced, and what evidence is retained?

Authentication and sessions

Authenticate users with the application’s normal session or token mechanism. Do not let a model infer identity from a name in a prompt. Bind every conversation, upload, retrieval request, and tool call to the authenticated session and tenant on the server. Cookies should use appropriate secure attributes and session tokens should not be placed in model context.

Session state needs explicit ownership. A conversation identifier supplied by a browser is only a lookup key after the server verifies its owner and tenant. Regenerate or reauthorize sessions when privilege changes. Expire idle sessions and revoke them when an account is disabled. For streaming responses, check authorization when the stream is opened and ensure the stream cannot be redirected to another user’s conversation.

Authorization is not a model decision

Server-side authorization must govern documents, database objects, tools, actions, and tenant data. The model can propose “retrieve invoice 42” or “send this message,” but it cannot grant itself permission. Evaluate the principal, action, resource, tenant, purpose, and current state in trusted code. AI Agent Permissions develops this decision boundary.

Use object-level checks on every request, including internal calls from an orchestrator. Avoid an API that accepts a tenant ID or resource ID and trusts the client value. Resolve ownership from the authenticated principal, then apply the narrowest policy. For high-impact actions, bind approval and credentials to the exact operation.

Multi-tenant data boundaries

Tenant isolation must hold in identity, storage, queries, retrieval, memory, caches, tool credentials, and logs. Apply the tenant filter before keyword or vector search; similarity is not authorization. A memory record written by Tenant A must never become context for Tenant B. Namespaces and encryption can help, but the decisive control is a server-side authorization binding.

Do not use one shared provider key for all tenants if it prevents attribution or revocation. Keep per-tenant credential references and enforce resource scope at the tool gateway. Correlation IDs and audit records should include tenant identity without exposing private content. RAG Security and AI Memory Security cover these boundaries in depth.

Traditional web threats still apply

Prompt injection is an additional application threat, not a replacement for XSS, CSRF, IDOR/BOLA, SQL injection, SSRF, insecure deserialization, or broken access control. Use the OWASP ASVS and Web Security Testing Guide as the baseline. Validate input on the server, parameterize database operations, protect state-changing browser requests with CSRF defenses where cookie authentication makes them relevant, and enforce CORS deliberately rather than broadly.

Rate-limit login, chat, retrieval, upload, and tool endpoints. Bound request size, model tokens, concurrent streams, tool calls, retries, and downstream spend. Return generic errors to clients while preserving diagnostic detail in protected logs.

RAG, memory, and prompt injection

RAG connectors introduce a data authorization boundary. Authenticate the requester before retrieval, filter by tenant and object permissions, and preserve document provenance. Retrieved text is evidence, not an instruction. A document can contain an indirect prompt injection that attempts to redirect the model; the gateway must still validate any proposed action. See Prompt Injection Attacks and LLM Security Risks.

Memory is durable application state. Define which fields may be written, who may write them, what source and timestamp are recorded, and when they expire. Read authorization should be at least as strict as write authorization. Treat summaries and embeddings as derived data that can be invalidated when source permissions change.

File uploads and ingestion

Use ordinary upload security before AI processing: authenticate the uploader, enforce size and count limits, validate content type and file signatures, quarantine uploads, scan where appropriate, and store them outside executable web paths. Archives need traversal and decompression-bomb defenses. Images, PDFs, and office files may contain active content or malicious parser edge cases.

Do not send an upload directly to a model or retriever. Assign an object ID, tenant, classification, provenance, and processing status. Extract text in an isolated worker, cap CPU and memory, and keep the original immutable. Review whether extracted content can influence tools or memory. A failed scan should block downstream indexing rather than silently degrading to “trusted.”

API, tools, and agent actions

Model and tool endpoints need authentication, authorization, schema validation, rate controls, and safe error handling. The browser should call a server API; it should not hold provider keys or database credentials. Secure Credential Handling for AI Agents explains brokered custody.

Expose an allowlisted registry of narrow tools. Validate types, ranges, resource ownership, destinations, and business rules before execution. Separate read, draft, and commit capabilities. Use idempotency keys and transaction boundaries for effects that may be retried. Route high-impact actions to approval controls described in Human Approval for AI Agents. Secure Tool Calling for LLM Applications provides the gateway pattern.

Output, streaming, and browser rendering

Treat model output as untrusted data. Plain text can be rendered as text; rich Markdown, links, HTML-like fragments, and structured data need explicit parsing and validation. Do not insert generated HTML into the DOM through raw HTML escape hatches without context-appropriate sanitization. Secure Rendering of LLM Output is the dedicated output guide.

Streaming does not lower the trust requirement. Escape and classify each chunk, handle delimiters safely, and do not execute partial output as JavaScript, SQL, shell, or a workflow command. Keep client state changes behind typed server APIs. Security headers, including a carefully designed Content Security Policy, provide defense in depth; they do not sanitize model text.

AI web security checklist

LayerPrimary riskRequired boundary/controlVerification/evidence
BrowserXSS, token theft, unsafe DOM updatesEscaped components, secure cookies, CSPXSS tests and header review
Web/APIBOLA, CSRF, injection, abuseAuth, object policy, schemas, limitsASVS/WSTG tests
IdentityConfused principal or over-privilegeUser and workload identity, scoped claimsNegative authorization tests
Tenant/DataCross-tenant retrieval or memoryTenant-bound queries, namespaces, ownership checksIsolation integration tests
Model/ContextPrompt injection, leakageUntrusted labels, minimization, provenanceAdversarial context tests
Tools/ActionsUnsafe side effectsAllowlist, validation, approval, idempotencyReplay and transaction tests
OutputXSS or interpreter injectionSafe rendering and destination validationRendering and schema tests
ObservabilityUntraceable action or secret leakCorrelation IDs, redaction, retentionReconstruction exercise

Deployment and dependency boundaries

Secure deployment still matters when the model is the newest component. Separate development, evaluation, and production credentials and networks. Pin application and model dependencies, review changes, and maintain an inventory of model versions, retrieval indexes, parsers, and tool adapters. Verify artifacts before promotion and keep a rollback path. A compromised package, parser, or model can undermine otherwise careful prompts.

Protect internal service endpoints from model-chosen URLs and enforce egress policy at the network boundary. Do not expose cloud metadata services or administrative interfaces to an execution worker. Set CPU, memory, queue, and timeout limits so an expensive prompt or recursive tool loop cannot exhaust shared capacity. These are ordinary platform controls applied to an AI-shaped data flow.

Observability and incident response

Record the initiating user, tenant, agent identity, session/task ID, model and policy versions, retrieved object IDs, tool proposal, authorization result, approval binding, destination, result, and state change. Minimize full prompts and documents; never log bearer tokens or API keys. AI Agent Observability describes an evidence chain.

Prepare response actions before launch: revoke sessions and credentials, disable a tool, quarantine a document, invalidate memory, isolate a tenant, stop runaway streams, preserve logs, and roll back state. Test these actions with a tabletop exercise. A control that cannot be observed or revoked is incomplete.

Architecture review checklist

  • Define browser, API, orchestration, data, model, tool, and external trust boundaries.
  • Authenticate every user and bind sessions, conversations, uploads, and streams to tenant ownership.
  • Enforce resource and action authorization outside the model.
  • Filter RAG and memory reads before relevance ranking; record provenance and expiry.
  • Quarantine and scan files before extraction and indexing.
  • Keep provider and tool secrets server-side in a broker or secret manager.
  • Expose narrow typed tools with validation, budgets, idempotency, and approval gates.
  • Render output safely and keep displayed text separate from executable actions.
  • Apply normal web controls: CSRF where applicable, XSS defenses, secure headers, and API limits.
  • Capture attributable, redacted evidence and rehearse revocation and recovery.
  • Retest after model, prompt, retriever, dependency, policy, or tenant changes.

Focused implementation guides

Apply this boundary model with Content Security Policy for AI-Generated Interfaces, File Upload Security, Webhook Security, Denial-of-Wallet Defense, Cache Security, and Secrets Management.

Sources

  • OWASP Application Security Verification Standard (ASVS)
  • OWASP Web Security Testing Guide (WSTG)
  • OWASP API Security Top 10
  • OWASP GenAI and Cheat Sheet guidance
  • NIST AI Risk Management Framework

The central design rule is simple: let the model help interpret and compose, but keep identity, authorization, secrets, rendering, and side effects in independently testable application controls.