Benchmarks
AI Security Benchmarks: What Attack Success Rates Mean
How to interpret AI security benchmark attack-success rates with clear denominators, attacker budgets, defenses, utility, uncertainty, and evidence.
Security benchmark numbers are useful only when the measurement contract is visible. “Attack success rate” is not a universal property of a model. It is an observed proportion under a defined attacker, target surface, policy, evaluator, budget, and system configuration. A low rate can reflect strong controls, a weak attack set, a narrow surface, or a scorer that misses unsafe outcomes. A high rate can reflect realistic pressure, broad exposure, or an overly permissive test harness. Treat the number as evidence about a system under conditions, not as a permanent security label.
Define attack success before counting it
Write the success predicate before running the benchmark. Decide which security property is being tested: unauthorized data disclosure, policy-violating tool use, unsafe code execution, cross-tenant access, or another concrete outcome. A refusal, a safe completion, a blocked side effect, and a harmless but awkward response should not be collapsed into one category.
A basic rate is:
ASR = successful security violations / eligible attack trials.
That formula is intentionally incomplete. The denominator must say which trials were eligible, whether repeated attempts count separately, and whether a failed setup or unavailable tool is excluded. The numerator must define what evidence proves a violation. If the benchmark changes either rule between systems, the comparison is not fair.
Security results should be paired with benign utility. A defense that blocks every request may produce a low ASR while making the product unusable. Report safe-task success, refusal quality, latency, and cost separately from attack outcomes. This is the same discipline used in AI model benchmarks: a score has meaning only alongside its task definition and conditions.
Model-only versus system-level claims
A model-only test sends prompts to a fixed interface and observes text. A system-level test includes the application, retrieval layer, memory, tool gateway, credentials, authorization checks, browser or runtime, and monitoring. The second is usually closer to deployment risk because many serious failures require a chain of components.
State exactly which layer was tested. If tools were disabled, do not claim the production agent is secure. If a gateway blocked an action before the model saw it, credit the control to the gateway and report the blocked decision. If the model produced a dangerous proposal but trusted code denied it, that is a control success with a model-layer warning, not a model refusal.
The dimensions behind an attack success rate
Attack surface determines what can be reached: chat instructions, retrieved documents, memory, tool arguments, file parsers, browser actions, or external APIs. Map each surface and state whether the test is direct, indirect, authenticated, tenant-aware, or side-effecting.
Attacker knowledge and adaptivity matter. A fixed public prompt set measures one exposure. An adaptive attacker who sees outputs, learns error messages, and chooses the next attempt measures a different exposure. Disclose whether the attacker knows the system prompt, tool schemas, policy, model family, or prior results. Never compare a fixed set against an adaptive campaign without labeling the difference.
Budget is another security variable. Bound attempts, conversation turns, tokens, wall-clock time, tool calls, parallel sessions, and human effort. A rate after ten tries is not equivalent to a rate after ten thousand tries. Report both the budget and the opportunity for selection, such as best-of-N attempts. Include the cost of reaching the target surface when that cost is material.
The defense configuration must be versioned. Record model and system-prompt versions, filters, retrieval policy, authorization policy, tool permissions, rate limits, and logging. A benchmark rerun after a policy change is a new measurement. Keep the original result for trend analysis rather than silently replacing it.
Attack success rate deconstruction
| Dimension | Question to answer | Evidence to retain |
|---|---|---|
| Target property | What exact violation counts as success? | Written predicate and examples of safe/unsafe outcomes |
| Denominator | Which trials were eligible and why? | Trial manifest and exclusions |
| Attacker | Who or what generated attempts? | Generator, knowledge and adaptivity settings |
| Budget | How many attempts, turns, tokens and tool calls? | Per-run budget ledger |
| Surface | Which model, RAG, memory, tool or API boundary was reachable? | Capability and permission snapshot |
| Defense | Which filters, policies and approvals were active? | Configuration hash and policy version |
| Scorer | How was the outcome judged and by whom? | Scoring code, judge rubric and adjudications |
| Utility | What benign behavior remained available? | Matched safe-task results |
| Uncertainty | How stable is the estimate? | Sample size, interval and slice breakdown |
This table prevents a single percentage from hiding the actual experiment.
Choose realistic security tasks
Build tasks from a threat model, not from whatever prompts are easiest to generate. Include direct and indirect untrusted content where the product accepts it, but keep examples defensive and do not publish bypass strings. For tool-using systems, test unauthorized reads, unsafe writes, and requests that should require approval. For retrieval, test tenant and document boundaries. For memory, test whether a prior instruction can improperly influence a later authorized task.
Use a clean baseline and a defended configuration. The baseline helps estimate attack difficulty; it is not a production recommendation. Test multiple model families or versions only when the surrounding scaffold is held constant. If the scaffold changes, publish a separate system-level comparison.
Separate discovery from confirmation. Exploratory red teaming can find candidate failures. A benchmark should then replay a frozen, reviewed set with deterministic eligibility and independent confirmation of each claimed violation. AI red teaming is valuable for finding cases, while Prompt Injection Attacks and Indirect Prompt Injection provide threat context. Neither page turns a prompt list into a complete security metric.
Scoring, judges, and false outcomes
Human or automated judges need an explicit rubric. A judge that rewards any policy deviation may over-count harmless wording. A judge that looks only for a refusal may miss a partial data leak or a side effect that happened before the final answer. Prefer observable evidence: authorized resource checks, tool-gateway decisions, sandbox events, and output classification. Use independent review for disputed cases and retain adjudication notes.
Track false positives and false negatives. A false positive marks a safe result as an attack success and can make a control look worse than it is. A false negative misses a real violation and creates false confidence. Sample both accepted and rejected cases for manual review, especially near the decision threshold. If a judge model is used, report its version, prompt, calibration process, and agreement with human labels.
Do not treat refusal rate as security rate. Refusal may be unnecessary on a harmless task, or it may occur after sensitive information was already disclosed. Likewise, an application can allow useful model output while enforcing authorization and side-effect controls outside the model.
Comparability across studies
Two studies are comparable only when their major dimensions align. At minimum compare target property, attack surface, attacker budget, adaptivity, defense configuration, scorer, task distribution, and reporting unit. Note whether trials are independent and whether the same conversation or environment is reused.
Security benchmark comparability matrix
| Dimension | Study A | Study B | Comparable? |
|---|---|---|---|
| Target violation | Exact predicate | Exact predicate | Only if equivalent |
| Surface and permissions | Same tools and scope | Same tools and scope | Required for system claims |
| Attacker budget | Attempts, turns, tokens | Attempts, turns, tokens | Normalize or report separately |
| Adaptivity | Fixed or adaptive | Fixed or adaptive | Do not mix silently |
| Defense state | Versioned policy/config | Versioned policy/config | Required |
| Scorer | Rubric and adjudication | Rubric and adjudication | Check agreement |
| Utility baseline | Matched benign tasks | Matched benign tasks | Required for trade-off claims |
| Uncertainty | Sample and interval | Sample and interval | Report before ranking |
A benchmark may still be valuable when studies are not comparable; simply limit the claim to the measured protocol.
Report a security result card
| Field | Required statement |
|---|---|
| System scope | Model, application, tools, data and permissions tested |
| Threat property | Exact violation and protected asset |
| Trials | Count, task slices and exclusions |
| Attacker | Knowledge, adaptivity, generator and human effort |
| Budget | Attempts, turns, tokens, time and best-of-N rules |
| ASR | Numerator, denominator, estimate and interval |
| Utility | Benign-task success, refusals, latency and cost |
| Safety evidence | Gateway decisions, authorization logs and side-effect checks |
| Errors | False positives, false negatives and adjudication process |
| Limitations | Untested surfaces, drift, contamination and environment gaps |
Publish this card with every release so a reader can audit the claim.
Statistical and operational cautions
Use task-level uncertainty methods that match the design. If multiple attempts share a conversation, environment, or attack generator, they are correlated; do not present them as independent samples. Report slices by model, surface, privilege, task type, and severity. A small but severe violation may deserve more attention than a larger count of low-impact wording deviations.
Watch for contamination and benchmark overfitting. Keep held-out tasks or rotate reviewed cases. If prompts, policies, or model weights were tuned against the set, label the result as development performance. NIST evaluation guidance recommends recording probes and structured evidence; follow that pattern so trend charts show what changed.
Security is also an operational property. Record whether a violation was detected, contained, reversed, and attributed. A system with occasional model-layer failures but reliable authorization, sandboxing, and incident controls can have lower real-world risk than a system with a slightly better refusal score and broad credentials. Use AI security controls and the agentic measurement guidance in Agentic Benchmarks as adjacent engineering references once those controls are live.
ASR claim checklist
Before publishing an attack-success number, confirm:
- The protected property and violation predicate are written in advance.
- The denominator, exclusions, retries, and best-of-N rules are explicit.
- Model, scaffold, tools, permissions, data, policies, and environment are versioned.
- Direct, indirect, authenticated, tenant, and side-effect surfaces are labeled.
- Attacker knowledge, adaptivity, generator, and resource budget are reported.
- Benign utility and refusal quality are measured alongside ASR.
- Automated judges are calibrated, disputed cases are adjudicated, and error rates are sampled.
- Sample sizes, intervals, correlated trials, slices, and limitations are visible.
- Every reported violation has observable evidence and a reproducible case identifier.
Attack success rates are decision-supporting measurements, not proof of safety or insecurity by themselves. Use them to prioritize fixes, retest after every material change, and combine them with authorization tests, threat modeling, monitoring, and incident response. That is how a benchmark becomes a security control rather than a misleading score.
Sources
- NIST AI Measurement and Evaluation
- NIST AI Risk Management Framework
- MITRE ATLAS
- OWASP Generative AI Security Project
- OWASP Top 10 for LLM Applications
These sources provide measurement, risk-management, adversarial-technique, and application-security context; benchmark results remain limited to their declared protocol.