AI Security
Threat Modeling AI Applications: A Practical Method
A repeatable worksheet for threat modeling LLM, agent, RAG, memory, and model-supply-chain boundaries.
Threat modeling an AI application is a repeatable way to discover where data, authority, and model-influenced decisions cross trust boundaries. It is not a checklist of fashionable risks and it does not make traditional security obsolete. The method here adapts established threat modeling to LLM, RAG, agent, memory, and model-supply-chain behavior. Use it before launch, after material changes, and when incidents reveal a missing assumption. LLM Security Risks gives the broad risk context; this page provides a reusable working method.
1. Define the system and business action
Write one sentence describing what the system does and what outcome matters. “An assistant uses a model” is too vague. State whether it answers questions, edits records, sends messages, deploys code, or recommends a decision. Identify users, tenants, operators, service owners, and the consequences of error. Define what is in scope: model provider, prompts, retrieval, tools, memory, identity, external services, logs, and deployment pipeline.
2. Inventory assets and principals
List assets that need confidentiality, integrity, availability, or accountability: personal data, documents, credentials, prompts, model artifacts, datasets, memory, tool actions, approval records, and audit logs. Name principals separately: human user, agent workload, orchestrator, tool gateway, data store, model provider, administrator, and attacker. For each principal record authority and tenant scope. Identity is not authorization; the Agent Permissions guide covers the decision at the action boundary.
3. Draw data and action flows
Create a diagram or table for user input, system instructions, retrieved documents, model calls, tool requests, tool results, memory writes, outputs, and downstream actions. Mark trust boundaries where data changes owner, parser, privilege, tenant, or execution environment. Include asynchronous queues, caches, retries, human approval, and model or dataset promotion. An arrow to “the model” is not enough: show what enters context and what can leave it.
4. Mark AI-specific influence
For every model-influenced step ask whether output can disclose data, choose a resource, construct an argument, write memory, call a tool, or delegate work. Treat prompts and retrieved text as inputs with different trust levels, not as a single instruction stream. Indirect Prompt Injection covers untrusted content influencing an active call; RAG Security covers retrieval boundaries. Memory persistence, model artifacts, and multi-agent edges deserve their own rows.
5. Enumerate threats with useful lenses
Use STRIDE-style questions where they help: can a sender be spoofed, can data be tampered with, can an action be repudiated, can secrets or private records be disclosed, can a tool or model be denied service, or can authority be elevated? Add AI cases: prompt or context injection, cross-tenant retrieval, unsafe output interpretation, excessive tool scope, poisoned data or artifacts, memory poisoning, confused deputy delegation, and runaway loops. MITRE ATLAS helps map observed AI techniques; OWASP provides risk guidance. Do not assume a model confidence score is a control.
6. Map controls to the boundary
Controls should sit where the threat occurs. Enforce identity and authorization at the gateway, before retrieval, before memory writes, and before tool execution. Keep credentials out of model context, validate output for its destination, use typed tools, isolate execution, require approval for high-impact effects, and pin model artifacts. Preventive controls should have detective and responsive partners: logs, alerts, rate limits, circuit breakers, revocation, and rollback.
The PermsAI AI Threat-Modeling Worksheet
Use one row per component or flow.
| Component | Asset | Input source | Trust level | Authority | Threat | Control | Evidence | Residual risk |
|---|---|---|---|---|---|---|---|---|
| Retrieval index | Tenant documents | Upload and ingestion jobs | Mixed | Read for tenant | Poisoning or leakage | Provenance, filters, review | Source ID, policy log | Stale or false data |
| Model context | Prompt and records | User, system, retrieval | Mixed | No direct authority | Injection or disclosure | Minimize, label, redact | Context ID, access decision | Model may still misread |
| Tool gateway | API capability | Agent proposal | Controlled boundary | Narrow action scope | Excessive or forged call | Typed schema, policy, approval | Request digest, decision | Logic defects |
| Memory store | Profile/task state | User, tool, model summary | Scoped | Write/read by policy | Poisoning or cross-tenant read | Provenance, TTL, isolation | Memory ID, writer, retrieval | Semantic falsehood |
| Artifact registry | Weights and adapters | Suppliers and builds | Quarantined | Promote/deploy roles | Substitution or theft | Digest, signature, release gate | Manifest, scan, deploy ID | Unknown upstream risk |
Add owner, review date, affected tenant, test case, and severity fields in your working copy. The worksheet is PermsAI’s synthesis; standards provide principles, while your team supplies system-specific facts.
7. Prioritize and define evidence
Rank threats by impact, likelihood, exposure, and detectability. A low-likelihood credential leak may outrank a frequent harmless hallucination because the blast radius is larger. Define acceptance criteria: which tenant filter must hold, which tool arguments are rejected, what approval evidence is required, how quickly a token or model can be revoked, and what alert proves containment. Every control should have an owner and a test.
8. Test the model and the system
Turn threats into safe, repeatable tests. Try untrusted documents, wrong-tenant identifiers, stale approvals, altered tool arguments, revoked credentials, poisoned memory records, malformed outputs, repeated retries, and compromised artifact digests. Test the surrounding application, not only model responses. AI Red Teaming explains how to build regression evidence without relying on a single prompt.
9. Revisit after change or incident
A new tool, model, provider, dataset, memory field, tenant, or deployment environment changes the threat model. Compare the new data flow with the previous one, update worksheet rows, rerun affected tests, and record accepted residual risk. After an incident, preserve the original model and logs, identify the failed boundary, add a control and a test, then verify rollback and recovery.
Practical checklist
- State the business action, users, tenants, and consequence of error.
- Inventory assets, principals, authority, and external dependencies.
- Draw data, action, memory, approval, and promotion flows.
- Mark trust boundaries and model-influenced decisions.
- Enumerate spoofing, tampering, disclosure, denial, elevation, and AI-specific threats.
- Map preventive, detective, and responsive controls to each boundary.
- Record evidence, owner, severity, residual risk, and review date.
- Test denial, isolation, revocation, replay, poisoning, and unsafe-output paths.
- Revisit after material changes and every incident.
Sources
- NIST AI Risk Management Framework
- NIST SP 800-154: Data-Centric Threat Modeling
- MITRE ATLAS
- OWASP Top 10 for LLM Applications
- OWASP Agentic AI Threats and Mitigations
Make the worksheet operational
Store the worksheet with the service threat model and link each row to an issue, test, policy, or runbook. During review, ask a second engineer to challenge the trust assumptions and attempt to trace a harmful output to the first boundary that should have blocked it. If the team cannot name the evidence for a control, treat that control as unverified. Residual risk is acceptable only when a named owner understands the consequence, monitoring, and recovery path.
The method scales from a small assistant to a platform with many agents. Start with the highest-impact action and the most sensitive asset, then add detail where flows cross tenants, identities, interpreters, or deployment environments. A concise, maintained model is more useful than an elaborate diagram that no one updates.