AI Security

AI Red Teaming: A Practical Guide to Security Evaluations

A practical method for threat-led AI red teaming, safe test design, model-versus-system evidence, severity, remediation, and regression testing.

By PermsAI Editorial Team
AI Red Teaming: A Practical Guide to Security Evaluations featured image

AI red teaming is structured adversarial testing of an AI model or AI-enabled system to discover security and safety failures before an attacker, accident, or production edge case does. A useful red team does more than collect provocative prompts. It starts from a threat model, defines an observable failure, exercises the complete application path, captures evidence, and turns every validated finding into an engineering change and regression test.

The central question is not simply whether a model can produce an unsafe answer. It is whether the deployed system permits an unwanted disclosure, decision, tool call, state change, or external effect under realistic conditions.

What AI red teaming is—and is not

AI red teaming complements several adjacent practices rather than replacing them.

PracticePrimary questionTypical evidence
Ordinary QADoes the feature work for expected users and inputs?Requirements, test cases, expected outputs
BenchmarkingHow does a defined model or system perform on a repeatable task set?Dataset, harness, configuration, metric
Penetration testingCan conventional software, identity, network, or configuration weaknesses be exploited?Technical path, affected asset, demonstrated impact
AI red teamingHow can an adversarial actor or hostile context cause an AI component or workflow to violate its intended security or safety boundaries?Threat scenario, model behavior, control decisions, resulting system effect

These practices overlap. A red-team finding may expose an ordinary API authorization bug; a benchmark can become one part of a regression suite; and a penetration test can include prompt injection against an agent. The difference is emphasis: AI red teaming deliberately explores failures created or amplified by models, probabilistic behavior, untrusted context, and tool-mediated workflows.

MITRE ATLAS provides a living knowledge base of adversary tactics and techniques for AI-enabled systems. It is useful for scenario coverage and shared language, but it is not a certification and should not replace a system-specific threat model.

Scope the evaluated system

A result is meaningful only when readers know what was tested.

  • Model-level testing evaluates behavior at a model interface: harmful assistance, sensitive-data reproduction, instruction following, robustness, or misuse capability under a stated configuration.
  • Application-level testing includes system instructions, retrieval, memory, output rendering, business rules, and user access controls.
  • Agent and tool testing follows multi-step planning, tool selection, credentials, authorization, retries, and real or simulated side effects. The AI agent security guide describes these boundaries.
  • Deployment and infrastructure testing covers the surrounding API, identity, network, storage, sandbox, logging, and administrative surfaces.

Do not report a model-only test as proof that the product is secure. Conversely, a system-level denial may contain unsafe model output successfully. Record both outcomes.

Begin with a threat model

Start with assets, actors, trust boundaries, and impact. Assets might include customer data, model weights, credentials, payment authority, private documents, production code, or the integrity of a decision. Actors may be unauthenticated users, ordinary tenants, insiders, compromised data suppliers, or attackers controlling a web page the system retrieves.

Map how input crosses boundaries: user to application, retrieval source to context, model to tool gateway, gateway to target system, and output to browser or downstream interpreter. Define preconditions rather than assuming unlimited attacker access. A public chatbot, an authenticated support assistant, and an internal deployment agent require different tests.

NIST's AI Risk Management Framework is voluntary guidance for incorporating trustworthiness into the design, development, use, and evaluation of AI systems. Its risk-management framing supports an iterative program, but it does not supply a ready-made exploit list for a particular application.

What to test

Choose scenarios from the actual architecture and consequences:

  • direct and indirect prompt injection;
  • disclosure of secrets, personal data, proprietary context, or cross-tenant records;
  • missing object-, function-, or tenant-level authorization;
  • excessive agency, unsafe tool parameters, action chaining, and approval bypass;
  • RAG poisoning, unauthorized retrieval, and provenance failures covered in the RAG security guide;
  • unsafe rendering or model output passed to SQL, shells, templates, URLs, or APIs;
  • model misuse and capability risks under a defined access level and attacker budget;
  • monitoring blind spots, incomplete audit trails, ineffective revocation, and recovery failures.

The broad prompt injection guide and LLM security risks guide supply deeper control context. A red-team plan should reference those threat classes without restating them as generic prose.

Design a test that can produce evidence

A reproducible case contains:

  1. Objective: the security property being challenged.
  2. Threat scenario: actor, goal, asset, entry point, and assumed access.
  3. Preconditions: model version, system configuration, tenant, tools, data, credentials, and environment state.
  4. Inputs and variations: prompts, documents, retrieved content, tool responses, or action sequences.
  5. Success criteria: an observable model or system outcome, defined before execution.
  6. Evidence: context, output, policy decision, tool call, logs, and resulting state.
  7. Severity and confidence: impact, exploitability, repeatability, and detection with uncertainties stated.
  8. Mitigation and retest: control owner, expected blocking point, regression case, and acceptance result.

For example, “the model discussed another tenant” is ambiguous. A better objective is: an authenticated Tenant A user must not cause a Tenant B document identifier or content to enter retrieval results, model context, response, logs visible to that user, or a downstream tool call. The test specifies known forbidden documents and inspects every stage.

Separate model failure from system failure

Suppose a hostile document tells an assistant to email a private record externally. There are several distinct outcomes:

  • the document is retrieved;
  • the model follows it and proposes the action;
  • the tool gateway accepts or denies the request;
  • the target service authorizes or rejects it;
  • data actually leaves the controlled environment.

An unsafe model response is a model-behavior finding. An executed unauthorized action is a system-security failure with greater demonstrated impact. A denied call is evidence that a control worked, although teams should still measure how often the model generates dangerous requests and whether denial creates operational or social-engineering risks.

Report each stage. Collapsing them into one “attack success rate” obscures which control failed and which engineering team owns the fix.

Combine human and automated red teaming

Expert humans are strong at forming new hypotheses, understanding business logic, noticing ambiguous harm, chaining subtle weaknesses, and adapting to feedback. Their work is costly and difficult to reproduce at scale.

Automation can generate variations, replay a corpus across versions, explore languages and formats, and detect regressions. But automated attackers may repeat the generator model's blind spots, and model-based judges may misclassify nuanced outcomes. A large run is not broad coverage if every case tests the same assumption.

A hybrid approach works best: experts discover and refine high-value scenarios; engineers encode stable cases; automation expands controlled variations; and humans review uncertain or high-impact results. Anthropic describes a similar progression from qualitative expert work toward standardized automated evaluations, while noting that different red-team methods have distinct advantages and limitations.

Use a safe test environment

Run consequential tests in an isolated environment with synthetic or minimized data, scoped credentials, fake destinations, reversible state, and explicit resource limits. Agent tests should use instrumented tool adapters or controlled external targets. Deny production egress unless a narrowly reviewed test requires it.

A sandbox must resemble the relevant production boundary closely enough to expose real failures. Record differences such as unavailable tools, substitute data, or stricter network policy. Never test unauthorized third-party systems. Establish stop conditions for unexpected data access, cost, persistence, or external communication.

Capture an evidence bundle

Preserve the test ID, threat hypothesis, evaluator identity, timestamp, model and application versions, system instructions where appropriate, user input, retrieved source IDs, context manifest, model output, proposed tool call, authorization result, approval, logs, and final state change. Use a correlation ID across services.

Minimize secrets and personal data. Store sensitive evidence under restricted access and retention rules; redact exported reports without destroying the facts needed for reproduction. Hidden chain-of-thought is not a dependable audit record. Observable inputs, decisions, tool events, and state are.

A practical severity model

PermsAI's rubric avoids treating a vivid transcript as severity by itself.

DimensionQuestion
ImpactWhat confidentiality, integrity, availability, safety, financial, or trust consequence occurred or was credibly reachable?
PrivilegeWhat authentication, role, tool, credential, or tenant position was required?
ExploitabilityHow much access, knowledge, control of content, and adaptive effort did the attacker need?
RepeatabilityDoes the result reproduce across fresh runs and representative configurations?
DetectionWould existing telemetry alert, and could responders reconstruct the path?
Human interactionDid success require an informed approval, a misleading approval, or none?

Record demonstrated impact separately from plausible worst case. Severity can rise when a modest model failure crosses a powerful permission boundary; it can fall when a deterministic control reliably contains the behavior. Do not mechanically reuse CVSS when model sampling, context, and business process dominate the finding.

The AI Red-Team Evidence Loop

PermsAI's lifecycle makes red teaming an engineering feedback system:

THREAT MODEL → DESIGN TEST → EXECUTE SAFELY → CAPTURE EVIDENCE → TRIAGE → MITIGATE → RETEST → REGRESSION SUITE

At threat model, select an asset and credible actor. At design, define preconditions and success criteria. At execution, preserve configuration and bound side effects. At evidence, distinguish model influence, control decision, and consequence. At triage, assign severity, confidence, and an owner. At mitigation, change the narrowest effective control, including any affected AI supply-chain component, application, identity, or infrastructure boundary. At retest, use the original case plus nearby variants. At regression, run stable cases on every material model, prompt, policy, retrieval, tool, or dependency change.

The loop prevents two common failures: a dramatic demonstration with no reproducible artifact, and a one-time fix that silently regresses. The AI benchmark guide explains why harness, configuration, and inference settings must accompany any reported metric.

Common red-team mistakes

Avoid testing only a base model when production adds retrieval and tools; collecting jailbreak screenshots without defined impact; granting unrealistic administrator access; changing multiple variables without recording them; treating one successful sample as universal reliability; reporting only average rates while missing severe slices; using the same model as attacker, target, and unquestioned judge; or fixing a prompt without testing the authorization boundary.

Red teaming is not a compliance stamp. No finite suite proves absence of failure. Its value is evidence about named threats and controls under documented conditions.

Practical checklist

  • Define the asset, actor, boundary, and unacceptable outcome.
  • State whether the target is a model, application, agent, or full deployment.
  • Record model, prompt, retrieval, tool, policy, and environment versions.
  • Predefine success criteria for model behavior and system consequence separately.
  • Use synthetic data, scoped credentials, controlled destinations, and stop conditions.
  • Combine expert exploration with reproducible automated regression cases.
  • Capture retrieved context, tool proposals, authorization decisions, logs, and state changes.
  • Score demonstrated impact, privilege, exploitability, repeatability, detection, and human interaction.
  • Assign every accepted finding a control owner and deadline.
  • Retest the exact case and meaningful variants after mitigation.
  • Run the regression suite after any material system change.
  • Document coverage gaps and residual uncertainty.

Sources