Vulnerability Intelligence

Prompt Injection Vulnerability Disclosures: An Evidence Framework

An evidence framework for distinguishing model behavior, product weakness, boundary failure, and actionable prompt-injection vulnerability disclosures.

By PermsAI Editorial Team
Prompt Injection Vulnerability Disclosures: An Evidence Framework featured image

Prompt injection reports often use “vulnerability” to describe several different things. A model may follow an attacker’s instruction, a product may expose a control that was supposed to hold, or an application may have placed untrusted text on the privileged side of a boundary. These are not interchangeable claims. This framework helps security teams collect evidence, communicate scope, and decide when a disclosure supports remediation or a formal vulnerability record. It complements prompt injection attacks and indirect prompt injection without turning either page into an exploit manual.

Start with the security property

The core question is: what security property was supposed to hold, and which implemented boundary failed? “The model did something surprising” is a useful observation, not yet a vulnerability classification. State the intended property in testable terms: a tenant must not read another tenant’s document, a tool must not send without authorization, model output must render as inert text, or a secret must not leave a broker-controlled destination.

Then identify the component that owns the property. It may be a model, system prompt, retrieval pipeline, tool gateway, API authorization layer, browser renderer, memory store, or network control. A prompt can influence behavior, but it does not by itself authenticate a user, enforce object ownership, or authorize a transaction. A bypass of “never reveal the system message” may show instruction-following weakness; it does not automatically prove a software access-control failure.

Distinguish the claim types

Use precise labels before discussing severity:

  • Model behavior: an output is unsafe, biased, or inconsistent under certain inputs. The remedy may be evaluation, product policy, or model improvement.
  • Product weakness: a feature makes an unsafe behavior easy or lacks a documented guardrail, even if no independent boundary is crossed.
  • Application design flaw: the integrator trusted model output, retrieval text, or a conversation identifier where server-side validation was required.
  • Security-boundary failure: an implemented control such as authorization, tenant isolation, secret handling, or output encoding is bypassed.
  • Software vulnerability: a patchable defect in code or service logic that causes a security property to fail under defined conditions.
  • Actionable disclosure: a report with reproducible evidence, affected component and scope, impact, and a remediation or coordination path.

These categories can overlap. A prompt injection may expose a product weakness, while the downstream unauthorized tool call is an application authorization failure. AI vulnerability intelligence covers asset and advisory triage; this page focuses on evidence for prompt-injection disclosures.

PROMPT-INJECTION DISCLOSURE CLASSIFIER

OBSERVATION → STATE INTENDED SECURITY PROPERTY → IDENTIFY TRUST BOUNDARY → CLASSIFY INPUT (DIRECT / INDIRECT / RETRIEVED / TOOL / MEMORY) → REPRODUCE WITH CONTROLLED CONTEXT → RECORD MODEL, PRODUCT, VERSION, AND CONFIGURATION → TEST WHETHER A DETERMINISTIC CONTROL CROSSES → MEASURE DATA OR ACTION IMPACT → CHECK PATCHABILITY AND AFFECTED COMPONENT → SEPARATE PROVIDER, PRODUCT, AND APPLICATION OWNERS → ASSIGN CLAIM STRENGTH → DISCLOSE RESPONSIBLY AND RETEST THE FIX

The flow is an editorial framework from PermsAI, not an official CVE rule. It prevents a persuasive transcript from becoming an unsupported claim about every deployment.

Describe the prompt and context without overfocusing on tricks

Record whether the input was a direct user message, retrieved page, document, email, tool output, memory record, or system-prompt extraction attempt. Include the relevant context boundary and what the application did with the result. A single successful trial is often insufficient because model outputs vary with model version, temperature, context ordering, safety settings, tool availability, and retries.

Use minimal, sanitized demonstrations. Do not include real secrets, destructive commands, private tenant data, or instructions that make exploitation easier. Preserve hashes or redacted excerpts so reviewers can identify the tested artifact without distributing sensitive content.

Reproducibility and deterministic controls

Record product and model identifiers, deployment mode, policy configuration, retrieval sources, memory state, enabled tools, user privilege, and number of trials. Explain whether the result was consistent, intermittent, or dependent on a particular context. A hosted model may change without an application release; a provider-side safety update can alter the result.

Evidence becomes stronger when the behavior crosses an independent, deterministic control. For example, a model merely says it will call deleteInvoice; a tool gateway then validates arguments and denies the request. That is not an authorization bypass. If the gateway executes a resource the principal cannot access, the report has evidence of a separate boundary failure. If a retrieved document causes a cross-tenant record to enter context because filtering was absent, the data layer is part of the defect.

Claim-strength ladder

LevelClaimMinimum evidenceAppropriate language
0Model anomalyrepeatable or well-described output deviation“behavior observed under these conditions”
1Instruction bypasssystem or application policy was ignored in output“prompt-policy bypass”
2Product weaknessproduct feature predictably enables unsafe behavior“product design weakness; scope documented”
3Security-boundary impactan independent control was crossed“unauthorized disclosure or action demonstrated”
4Actionable vulnerabilitypatchable component, affected versions, reproducibility, impact, and owner“candidate vulnerability; coordinate disclosure”

This ladder is PermsAI’s communication aid, not a CVE severity scale. Do not call every level-1 jailbreak a CVE. Conversely, a prompt is a valid attack vector when it reliably reaches a patchable boundary such as authorization, tenant isolation, secret handling, or unsafe rendering.

Define the impacted boundary

Ask which property failed and whose code enforces it:

  • Authentication: did an unauthenticated actor obtain a session or act as another principal?
  • Authorization: did the server perform an action on an object the principal could not use?
  • Tenant isolation: did context, retrieval, memory, cache, or logs cross a tenant boundary?
  • Confidentiality: did a secret, private prompt, document, or credential leave its intended scope?
  • Tool policy: did a gateway accept an unapproved tool, destination, argument, or side effect?
  • Rendering: did generated content become executable markup or script in a browser?
  • Network or file policy: did a fetcher or parser reach a prohibited destination or path?

The report should show the transition from input to impact, not only the final transcript. Connect tool-action questions to secure tool calling and prompt secrecy to system prompt security.

What counts as impact

Classify the observed effect: secret disclosure, unauthorized read or write, cross-tenant access, privilege escalation, unsafe tool invocation, network-policy bypass, persistent memory poisoning, stored XSS, or availability/resource abuse. State whether the result was simulated, a test tenant, a real production asset, or a provider-controlled environment. Never use real customer data to increase convincingness.

An instruction to “ignore policy” with no data or action impact may still deserve product hardening, but its security claim should remain limited. A model-generated URL that is rejected by an SSRF gateway is evidence of attempted influence, not an SSRF vulnerability. A URL that causes the server to contact an internal service despite an enforced destination policy is materially different.

Patchability, ownership, and scope

Identify the component that can fix the property: model provider, model runtime, orchestration framework, tool gateway, data service, browser renderer, or integrating application. A provider’s model behavior may not have a versioned patch, while an application’s missing authorization check does. If the issue crosses products, split the report into claims with separate owners and evidence.

CVE assignment depends on the CVE Program and the responsible CNA’s scope and policy. There is no universal rule that prompt injection is always eligible or always ineligible. A CVE record normally describes a software or service vulnerability with affected scope; an advisory may instead document a product weakness or design guidance. Do not promise an identifier. Include the evidence a CNA or vendor needs to decide.

Evidence package for a disclosure

Provide a concise timeline, affected product and version, deployment mode, user privilege, model and policy configuration, input class, number of trials, expected property, observed result, independent control crossed, impact, and safe reproduction description. Attach redacted logs, request IDs, hashes, screenshots, or test-tenant identifiers. Distinguish “can influence,” “did influence,” and “caused an unauthorized effect.”

If a vendor responds, preserve its acknowledgement, severity or scope decision, mitigation, fix version, and communication dates. Re-test the exact property after the fix, including the path that previously failed. A change in model behavior is not proof that a deterministic application boundary is repaired.

Responsible disclosure

Use the vendor’s security contact or coordinated-disclosure process. Minimize impact: test in an owned environment, avoid real data, stop when the security property is demonstrated, and do not publish credentials or operational bypasses. Give maintainers enough time to validate and remediate, while agreeing on a safe public timeline. If no patch is possible, document compensating controls and residual risk rather than overstating closure.

PermsAI prompt-injection evidence matrix

ClaimMinimum evidenceWhat it provesWhat it does not prove
Jailbreakrepeatable policy-violating output in defined model/configbehavior under tested conditionssoftware vulnerability or unauthorized access
Prompt-policy bypasssystem/application instruction ignoredinstruction hierarchy weaknessauthentication or authorization failure
Indirect injectionretrieved/tool content changes agent behavioruntrusted data can influence control flowimpact if tools and data gates hold
Tool invocationmodel proposes or reaches a toolcapability exposure or action pathpermission bypass unless gateway allows it
Secret disclosurecontrolled secret appears outside intended scopeconfidentiality failure in tested boundarydisclosure of other tenants or production secrets
Cross-tenant accesstest tenant retrieves another tenant’s datatenant isolation failureroot cause without data-flow evidence
Unauthorized side effectserver performs forbidden write/send/fetchsecurity-boundary impactbroad product-wide exploitability
CVE-worthy claimpatchable component, versions, conditions, impact, evidenceactionable vulnerability candidateguaranteed CVE assignment or severity

Avoid common overclaims

Do not treat a hidden system prompt as an access-control mechanism, a refusal failure as remote code execution, a successful transcript as proof of every model version, or a provider mitigation as a fix for your application’s authorization. Do not infer cross-tenant impact from a prompt that mentions another tenant; verify the actual data path. Keep model output, product behavior, application design, and patchable software defects separate.

Practical disclosure checklist

  • State the intended security property and boundary owner.
  • Classify the input source and preserve minimal, redacted evidence.
  • Record product, model, version, configuration, tools, retrieval, memory, and user privilege.
  • Repeat safely and describe consistency, prerequisites, and hosted-service drift.
  • Test whether deterministic authorization, tenant, secret, network, rendering, or tool controls were crossed.
  • Separate observed behavior from impact and from a CVE eligibility claim.
  • Identify affected scope, patchability, vendor response, and compensating controls.
  • Disclose responsibly, avoid real data, and retest the exact property after remediation.

Sources

CVE Program, NVD, OWASP Prompt Injection, OWASP GenAI, vendor security advisories, and credible primary research inform this framework. The classifier, claim ladder, and evidence matrix are PermsAI’s editorial synthesis.