AI Security

AI Security Controls Matrix: From Threat to Verification

An operational matrix connecting AI threats to controls, evidence, response, verification, and ownership.

By PermsAI Editorial Team
AI Security Controls Matrix: From Threat to Verification featured image

A security control is useful only when a team can show that it operates. “We have a policy” is weaker than a repeatable authorization test, protected telemetry, or an incident drill demonstrating that the policy blocked the intended threat. This reference maps major AI threats to preventive, detective, responsive, and verification activities, with an owner for each control.

How to use the matrix

Start with Threat Modeling AI Applications. Identify the business action, sensitive assets, principals, tenant boundaries, model influence, tools, and external effects. Choose controls according to autonomy, data sensitivity, privilege, side effects, deployment environment, and recovery requirements. A read-only assistant does not need the same gates as an agent that changes production.

Preventive controls reduce likelihood or impact. Detective controls produce signals and evidence. Responsive controls contain, revoke, restore, and investigate. Verification proves that the first three classes are configured and effective. No single class is sufficient.

PermsAI AI Security Controls Matrix

ThreatPreventive controlDetective evidenceResponseVerificationControl owner
Direct prompt injectionSeparate instructions/data; narrow toolsInjection signal, denied callDisable capability, preserve sessionAdversarial regressionApp security
Indirect injectionSource labels, retrieval filters, output validationSource ID and unusual tool pathQuarantine source, revoke grantPoisoned-document testsData + app
Excessive agencyStep, time, spend, and tool budgetsLoop and sequence telemetryCircuit breaker, stop workflowRunaway-agent testPlatform
Authorization failurePrincipal/action/resource policyDecision and reason logsRevoke identity, rollbackNegative authorization testsIAM
Credential exposureBrokered short-lived credentialsAccess and redaction eventsRotate/revoke secretLeakage and secret scansIAM
RAG leakageTenant/object filters before retrievalQuery, document, tenant IDsRemove item, investigateCross-tenant retrieval testsData layer
Memory poisoningTyped writes, provenance, TTLWriter and retrieval eventsInvalidate memory and summariesPoison/correction testsApp + data
Model/data poisoningLineage, holdouts, release gateDrift and behavior slicesQuarantine, rollback artifactReproducible rebuildML platform
Supply-chain compromiseDigests, signatures, quarantineRegistry/deployment auditRevoke artifact or keyProvenance reviewPlatform
Unsafe tool callsSchemas, allowlists, transaction policyArguments, decision, resultDisable tool, revoke grantFuzz and replay testsTool gateway
Generated-code executionEphemeral sandbox and quotasCommand and resource logsDestroy sandbox, invalidate outputEscape/resource testsInfrastructure
Cross-tenant leakageTenant-bound identity and storageTenant mismatch alertsIsolate tenant, purge exposureIsolation integration testsArchitecture
Missing audit evidenceCorrelation IDs, append-only recordsCompleteness and tamper signalsPreserve forensic copyReconstruction exerciseSecOps

Controls versus evidence

For every rule define its input, decision, enforcement point, owner, and durable record. “Tool calls require least privilege” should identify policy version, principal, resource, result, and a test proving a wrong-tenant call is denied. Evidence must be minimized and redact secrets, but it must remain sufficient to reconstruct what happened.

Ownership and placement

Application owners maintain context assembly, output handling, and business rules. IAM owns identity, delegation, credential lifecycle, and revocation. The model layer owns evaluation and prompt versioning but not authorization. Tool gateways own schemas, capability allowlists, and adapters. Data owners govern retrieval and memory scope. Infrastructure owns sandbox, network, runtime, and resource limits. Security operations owns detection, retention, incident response, and exercises.

Assign one accountable owner even when several teams implement a control. Record dependencies: a tool gateway may rely on IAM claims, a retrieval filter on tenant identity, and incident response on observability fields. Ownership gaps are themselves a risk.

Risk-based selection

Choose controls proportionate to consequence. A summarizer over public text may need input handling, output encoding, and logging. An agent that sends money, modifies production, or reads regulated records needs strong identity, narrow scopes, pre-action approval, transaction binding, isolation, and tested rollback. Excessive controls can create operational risk if reviewers are flooded or recovery is impossible.

Document residual risk, assumptions, and expiry. Reassess when autonomy, data sources, tools, tenants, model providers, or environments change. The matrix is a decision aid, not a certification checklist.

Verification methods

Unit tests exercise policy functions and schema validation. Integration tests prove tenant filters, credential brokers, tool adapters, approval gates, and transaction boundaries work together. Adversarial tests place hostile instructions in user input, documents, tool results, and memory candidates. Configuration review checks identities, routes, storage ACLs, retention, and alerts. AI Red Teaming describes how to scope and retest realistic paths.

Use Threat Modeling AI Applications to connect each control to an asset and boundary. Retest after model, prompt, retriever, tool, policy, dependency, or tenant changes. Store test version, environment, inputs, expected decision, observed result, and reviewer. A green test that cannot be reproduced is weak evidence.

Evidence quality

Prefer evidence that is complete, attributable, time-bounded, tamper-resistant, and useful to an investigator. A dashboard screenshot is less durable than an event with a correlation ID and policy version. A policy file is less persuasive than a denied request captured in an automated test. Avoid logging raw secrets or full sensitive prompts merely to create the appearance of evidence.

Practical workflow

  1. Define the business action and consequence of error.
  2. Identify assets, principals, tenants, tools, and trust boundaries.
  3. Select preventive controls at each boundary.
  4. Define telemetry proving decisions and outcomes.
  5. Write response actions: stop, revoke, quarantine, rollback, notify.
  6. Create verification tests, owners, cadence, and acceptance criteria.
  7. Run tests and record residual risk.
  8. Revisit after changes and incidents.

Keep the matrix alive

Store each row with an owner, policy version, test link, last result, and next review date. When a control fails, mark it degraded and track containment rather than leaving a green status unchanged. During an incident, compare observed evidence with the expected row to find the first failed boundary. A compact matrix covering identity, retrieval, tools, credentials, output, and response is more valuable than dozens of unchecked theoretical entries.

Practical checklist

  • Name the threat and affected asset, not only a generic category.
  • Separate preventive, detective, responsive, and verification work.
  • Place enforcement outside the model at identity, data, tool, and runtime boundaries.
  • Assign accountable owners and dependencies.
  • Define evidence fields with minimization, access, and retention.
  • Test wrong-tenant, replay, revocation, poisoning, leakage, and runaway paths.
  • Exercise incident response and rollback, not only prevention.
  • Reassess after material system or provider changes.

Related change tracking

Use AI Regulation and Security Standards: Technical Change Log for dated standards changes, and Year in AI Security for the year-to-date synthesis of evidence, incidents, and enduring controls.

Sources

  • NIST AI Risk Management Framework
  • NIST SP 800-53 security and privacy controls
  • OWASP Agentic AI Threats and Mitigations
  • MITRE ATLAS
  • OWASP Application Security Verification Standard

Verification should include operating evidence, not only design review. Sample authorization decisions, inspect redaction, confirm alerts arrive, and perform a tabletop exercise that revokes a credential while an agent is running. Track failed-test age and unowned rows. These measures reveal whether a control is dependable under pressure and turn the matrix into a living security program. Keep rejected controls and accepted exceptions visible so risk decisions remain reviewable.