Vulnerability Intelligence
Writing High-Quality AI Vulnerability Reports
How to give vendors enough evidence to reproduce, assess, and fix an AI vulnerability without overstating the claim.
A useful AI vulnerability report is an evidence package, not a dramatic description of what a model might do. Its job is to let a vendor reproduce the behavior, determine whether a security boundary is crossed, identify the affected versions, and coordinate a fix without exposing real customer data. This article presents a practical structure for AI vulnerability disclosure, including model and application context, safe reproduction, probabilistic results, impact claims, and retest evidence. It complements AI Vulnerability Intelligence, How to Read an AI Security Advisory, Prompt Injection Vulnerability Disclosures, and CVSS for AI Vulnerabilities: Uses and Limits. It is a reporting method, not a promise that every unusual model response is a vulnerability or that a CVE will be assigned.
Start with a precise claim
Open with a one-paragraph executive summary that a triage engineer can route quickly. State the affected product or service, the security property that fails, the required access level, the observed consequence, and whether the report is reproducible. Use a neutral claim such as: “A workspace member can cause the retrieval layer to include another workspace’s document in a generated answer when a specific index configuration is enabled.” Avoid conclusions that the evidence cannot support, such as claiming arbitrary account takeover when you only observed disclosure of synthetic content.
A strong summary answers five questions:
- What component is affected: web application, model endpoint, RAG index, agent tool, memory store, or integration?
- What security boundary is crossed: user, tenant, role, data classification, or execution authority?
- Who can trigger it and what prerequisites are required?
- What did the reporter actually observe?
- What action should the vendor take first: reproduce, disable a feature, rotate a secret, or preserve logs?
Describe the environment so reproduction is possible
AI behavior depends on configuration. Record the product name, deployment type, version or commit, model family and revision, provider or serving runtime, relevant system settings, retrieval/index configuration, tool definitions, feature flags, authentication role, tenant/workspace arrangement, and date tested. Include whether the run used streaming, a particular temperature or sampling setting, a safety policy, a custom prompt template, or a connector.
Separate version scope into confirmed affected, confirmed not affected, fixed, and untested. Pin model and repository revisions where possible; a mutable alias can change between the first report and vendor reproduction. Record the client and API versions when serialization, browser rendering, or request signing matters.
State the expected and actual boundary
Write the security invariant before the reproduction steps. Examples include “a member can retrieve only documents belonging to the active tenant,” “tool calls require the caller’s current authorization,” or “an untrusted document is treated as data, not privileged instructions.” Then contrast it with the actual result. This keeps the report focused on a broken control rather than on an interesting output.
Describe the trust boundary in application terms: principal, tenant, resource, model context, tool gateway, or external system. Model text is evidence about behavior; it is not proof of ownership or authorization. If the issue involves prompt injection, state which untrusted content entered context and which trusted action followed. If it involves a model file or dependency, state whether the loading path, repository provenance, or learned behavior is the subject of the claim.
Give safe, minimal reproduction steps
Use synthetic accounts, dummy secrets, and non-sensitive documents. Create the smallest fixture that demonstrates the issue and identify the role of each account. Number every step, including setup, request, model or tool configuration, and observation. Quote only the relevant input; do not attach live credentials, customer records, private prompts, or internal URLs.
A safe reproduction might create Tenant A and Tenant B with documents named 'A-test-001' and 'B-test-001', grant a normal member access only to Tenant A, run the documented query, and show that the response contains the B marker. Explain how to reset the fixture. For a tool issue, use a harmless endpoint that records a marker instead of performing a destructive action. For a file or model issue, provide hashes and a minimal sample only when the recipient has a safe intake channel.
Separate preconditions from the trigger. Preconditions can include membership, a feature flag, an index refresh, a connector permission, or a model revision. A vendor should be able to tell which condition is necessary and which is incidental.
Report expected versus actual results
Use a short comparison:
| Expected | Actual | |
|---|---|---|
| Identity | Request runs as the signed-in member | Same member obtains data outside the membership scope |
| Context | Retrieval is filtered to the active tenant | A result from another tenant enters model context |
| Action | Gateway denies an unauthorized tool call | Tool adapter accepts the model-proposed resource |
| Evidence | Audit record names the decision and resource | Log lacks enough identifiers to reconstruct the event |
Tie the actual result to an observable artifact: response excerpt with synthetic markers, request/response IDs, a redacted trace, or a screenshot of a permission state. Do not substitute a model’s statement that it accessed something for server-side evidence that it did.
Handle probabilistic behavior honestly
Many AI findings are not deterministic. Report the number of trials, successes, failures, and success rate. Record model revision, sampling settings, context size, retrieved documents, tool state, account role, and test dates. If the result changed after a retry, say so. “Observed in 7 of 20 runs with temperature 0.2” is more useful than “works sometimes.”
Explain whether failures were safe denials, unrelated model answers, timeouts, or missing retrieval. If a specific phrase, document order, or conversation history changes the result, describe it without publishing a bypass recipe. Include a minimal deterministic variant if one exists. A vendor can then decide whether to reproduce with the same model, an instrumented stub, or a policy test.
Classify the claim and impact
Label each conclusion as demonstrated, potential, or speculative. A demonstrated impact has evidence in the test fixture. A potential impact is a reasonable consequence that was not exercised, such as possible exposure of a larger dataset after confirming one unauthorized record. A speculative impact is a hypothesis that needs a separate test. This classification prevents severity inflation and helps the recipient prioritize.
Identify confidentiality, integrity, and availability effects separately. State whether the reporter crossed a tenant boundary, changed data, invoked a privileged action, exposed a credential, consumed excessive resources, or merely produced an unsafe-looking answer. Connect the claim to the relevant prompt-injection disclosure guidance when untrusted instructions are involved, but do not call every prompt-policy failure a product vulnerability.
A CVSS score can communicate a comparable severity estimate when the required facts are known. Use CVSS for AI Vulnerabilities: Uses and Limits as context, document the vector and assumptions, and keep the score subordinate to the evidence. A score does not prove exploitability, affected scope, or business impact. A CWE mapping can add vocabulary when a conventional weakness fits, but do not force a label where the boundary failure is genuinely application-specific.
Evidence that protects both sides
The report should preserve enough evidence for verification while minimizing disclosure. Include timestamps with timezone, test account and tenant aliases, product/model revisions, request or run IDs, hashes of attachments, and redacted inputs and outputs. Explain every redaction and keep the unredacted material in the reporter’s controlled evidence store. Never place API keys, session cookies, signing secrets, customer data, or private system prompts in a normal issue tracker.
When a recording is useful, prefer a short capture using synthetic data and show the authorization state before and after. Treat attachments as untrusted files on both sides; verify their hashes and do not require a recipient to execute a supplied script.
AI VULNERABILITY REPORT EVIDENCE BUNDLE
Claim → Environment → Preconditions → Reproduction → Boundary → Impact → Evidence → Version scope → Mitigation/fix → Retest
Use this chain as a completeness test. If the report jumps from a claim directly to impact, the vendor cannot tell which configuration produced it. If it omits version scope, a fix cannot be confirmed. If it lacks a retest, closure is only an assertion.
Coordinate disclosure and ownership
Send the report through the vendor’s documented security contact, PSIRT, security.txt endpoint, bug-bounty intake, or CNA process. CERT/CC guidance describes coordinated vulnerability disclosure as a process of receiving, validating, coordinating, and communicating a vulnerability; it does not guarantee a particular timeline or outcome. A CVE request is not the same as a CVE assignment: the request may be submitted to a CNA or the CVE Program, while assignment depends on the responsible authority and its criteria.
Ask for a tracking identifier and record response dates. Vendor states might be received, needs-more-information, accepted, disputed, fixed, duplicate, or not-applicable. A disputed report should remain evidence-based; do not silently change the claim to win an argument.
Verify the fix and its scope
A vendor’s patch note is not the same as a retest. Re-run the original minimal fixture against the fixed version, then test the negative control that should still work. Check adjacent paths: alternate API endpoints, streaming versus non-streaming, cached versus fresh retrieval, another role, and a restarted worker where state may persist. Record the exact fixed version, model or index revision, configuration, date, trials, and results.
If the fix is partial, report what remains and narrow the affected scope. If the issue is mitigated by a configuration change, say whether the safe setting is enabled by default and whether upgrades preserve it. Close only when the original security invariant holds and the evidence bundle is complete.
PERMSAI AI VULNERABILITY REPORT TEMPLATE
- Summary: product, affected security property, concise demonstrated impact.
- Environment: deployment, versions/commits, model/provider, configuration, tenant and role.
- Scope: confirmed affected, not affected, fixed, and untested versions.
- Prerequisites: accounts, permissions, feature flags, data fixtures, and tool state.
- Reproduction: numbered, minimal, safe steps with expected and actual results.
- Frequency: trials, successes, settings, date, and relevant context.
- Boundary analysis: principal, tenant, resource, trust boundary, and failed control.
- Evidence: redacted outputs, IDs, hashes, traces, and attachment manifest.
- Impact classification: demonstrated, potential, speculative; confidentiality, integrity, availability.
- Taxonomy: CWE or CVSS vector when justified, with assumptions.
- Mitigation and retest: workaround, fixed version, negative control, residual risk.
- Disclosure: recipient, tracking ID, dates, response state, and requested coordination.
EVIDENCE STRENGTH LADDER
This is a PermsAI editorial framework, not an official CERT or CVE classification:
- LEVEL 0 — CLAIM ONLY: an assertion without a reproducible observation.
- LEVEL 1 — OBSERVATION: one recorded behavior with environment details.
- LEVEL 2 — REPRODUCIBLE BEHAVIOR: the behavior repeats under stated conditions.
- LEVEL 3 — SECURITY BOUNDARY IMPACT: evidence shows unauthorized disclosure, change, execution, or resource use across a defined boundary.
- LEVEL 4 — VERSIONED / ACTIONABLE VULNERABILITY PACKAGE: scope, safe reproduction, evidence, mitigation, owner coordination, and retest are all documented.
Use the ladder to identify missing work, not to manufacture severity. A Level 2 behavior may deserve engineering attention even when the security impact is still uncertain.