AI Security
AI Security Controls Matrix: From Threat to Verification
An operational matrix connecting AI threats to controls, evidence, response, verification, and ownership.
A security control is useful only when a team can show that it operates. “We have a policy” is weaker than a repeatable authorization test, protected telemetry, or an incident drill demonstrating that the policy blocked the intended threat. This reference maps major AI threats to preventive, detective, responsive, and verification activities, with an owner for each control.
How to use the matrix
Start with Threat Modeling AI Applications. Identify the business action, sensitive assets, principals, tenant boundaries, model influence, tools, and external effects. Choose controls according to autonomy, data sensitivity, privilege, side effects, deployment environment, and recovery requirements. A read-only assistant does not need the same gates as an agent that changes production.
Preventive controls reduce likelihood or impact. Detective controls produce signals and evidence. Responsive controls contain, revoke, restore, and investigate. Verification proves that the first three classes are configured and effective. No single class is sufficient.
PermsAI AI Security Controls Matrix
| Threat | Preventive control | Detective evidence | Response | Verification | Control owner |
|---|---|---|---|---|---|
| Direct prompt injection | Separate instructions/data; narrow tools | Injection signal, denied call | Disable capability, preserve session | Adversarial regression | App security |
| Indirect injection | Source labels, retrieval filters, output validation | Source ID and unusual tool path | Quarantine source, revoke grant | Poisoned-document tests | Data + app |
| Excessive agency | Step, time, spend, and tool budgets | Loop and sequence telemetry | Circuit breaker, stop workflow | Runaway-agent test | Platform |
| Authorization failure | Principal/action/resource policy | Decision and reason logs | Revoke identity, rollback | Negative authorization tests | IAM |
| Credential exposure | Brokered short-lived credentials | Access and redaction events | Rotate/revoke secret | Leakage and secret scans | IAM |
| RAG leakage | Tenant/object filters before retrieval | Query, document, tenant IDs | Remove item, investigate | Cross-tenant retrieval tests | Data layer |
| Memory poisoning | Typed writes, provenance, TTL | Writer and retrieval events | Invalidate memory and summaries | Poison/correction tests | App + data |
| Model/data poisoning | Lineage, holdouts, release gate | Drift and behavior slices | Quarantine, rollback artifact | Reproducible rebuild | ML platform |
| Supply-chain compromise | Digests, signatures, quarantine | Registry/deployment audit | Revoke artifact or key | Provenance review | Platform |
| Unsafe tool calls | Schemas, allowlists, transaction policy | Arguments, decision, result | Disable tool, revoke grant | Fuzz and replay tests | Tool gateway |
| Generated-code execution | Ephemeral sandbox and quotas | Command and resource logs | Destroy sandbox, invalidate output | Escape/resource tests | Infrastructure |
| Cross-tenant leakage | Tenant-bound identity and storage | Tenant mismatch alerts | Isolate tenant, purge exposure | Isolation integration tests | Architecture |
| Missing audit evidence | Correlation IDs, append-only records | Completeness and tamper signals | Preserve forensic copy | Reconstruction exercise | SecOps |
Controls versus evidence
For every rule define its input, decision, enforcement point, owner, and durable record. “Tool calls require least privilege” should identify policy version, principal, resource, result, and a test proving a wrong-tenant call is denied. Evidence must be minimized and redact secrets, but it must remain sufficient to reconstruct what happened.
Ownership and placement
Application owners maintain context assembly, output handling, and business rules. IAM owns identity, delegation, credential lifecycle, and revocation. The model layer owns evaluation and prompt versioning but not authorization. Tool gateways own schemas, capability allowlists, and adapters. Data owners govern retrieval and memory scope. Infrastructure owns sandbox, network, runtime, and resource limits. Security operations owns detection, retention, incident response, and exercises.
Assign one accountable owner even when several teams implement a control. Record dependencies: a tool gateway may rely on IAM claims, a retrieval filter on tenant identity, and incident response on observability fields. Ownership gaps are themselves a risk.
Risk-based selection
Choose controls proportionate to consequence. A summarizer over public text may need input handling, output encoding, and logging. An agent that sends money, modifies production, or reads regulated records needs strong identity, narrow scopes, pre-action approval, transaction binding, isolation, and tested rollback. Excessive controls can create operational risk if reviewers are flooded or recovery is impossible.
Document residual risk, assumptions, and expiry. Reassess when autonomy, data sources, tools, tenants, model providers, or environments change. The matrix is a decision aid, not a certification checklist.
Verification methods
Unit tests exercise policy functions and schema validation. Integration tests prove tenant filters, credential brokers, tool adapters, approval gates, and transaction boundaries work together. Adversarial tests place hostile instructions in user input, documents, tool results, and memory candidates. Configuration review checks identities, routes, storage ACLs, retention, and alerts. AI Red Teaming describes how to scope and retest realistic paths.
Use Threat Modeling AI Applications to connect each control to an asset and boundary. Retest after model, prompt, retriever, tool, policy, dependency, or tenant changes. Store test version, environment, inputs, expected decision, observed result, and reviewer. A green test that cannot be reproduced is weak evidence.
Evidence quality
Prefer evidence that is complete, attributable, time-bounded, tamper-resistant, and useful to an investigator. A dashboard screenshot is less durable than an event with a correlation ID and policy version. A policy file is less persuasive than a denied request captured in an automated test. Avoid logging raw secrets or full sensitive prompts merely to create the appearance of evidence.
Practical workflow
- Define the business action and consequence of error.
- Identify assets, principals, tenants, tools, and trust boundaries.
- Select preventive controls at each boundary.
- Define telemetry proving decisions and outcomes.
- Write response actions: stop, revoke, quarantine, rollback, notify.
- Create verification tests, owners, cadence, and acceptance criteria.
- Run tests and record residual risk.
- Revisit after changes and incidents.
Keep the matrix alive
Store each row with an owner, policy version, test link, last result, and next review date. When a control fails, mark it degraded and track containment rather than leaving a green status unchanged. During an incident, compare observed evidence with the expected row to find the first failed boundary. A compact matrix covering identity, retrieval, tools, credentials, output, and response is more valuable than dozens of unchecked theoretical entries.
Practical checklist
- Name the threat and affected asset, not only a generic category.
- Separate preventive, detective, responsive, and verification work.
- Place enforcement outside the model at identity, data, tool, and runtime boundaries.
- Assign accountable owners and dependencies.
- Define evidence fields with minimization, access, and retention.
- Test wrong-tenant, replay, revocation, poisoning, leakage, and runaway paths.
- Exercise incident response and rollback, not only prevention.
- Reassess after material system or provider changes.
Related change tracking
Use AI Regulation and Security Standards: Technical Change Log for dated standards changes, and Year in AI Security for the year-to-date synthesis of evidence, incidents, and enduring controls.
Sources
- NIST AI Risk Management Framework
- NIST SP 800-53 security and privacy controls
- OWASP Agentic AI Threats and Mitigations
- MITRE ATLAS
- OWASP Application Security Verification Standard
Verification should include operating evidence, not only design review. Sample authorization decisions, inspect redaction, confirm alerts arrive, and perform a tabletop exercise that revokes a credential while an agent is running. Track failed-test age and unowned rows. These measures reveal whether a control is dependable under pressure and turn the matrix into a living security program. Keep rejected controls and accepted exceptions visible so risk decisions remain reviewable.