AI News
OWASP GenAI Top 10 Updates: Changes for Builders
A release-led comparison of OWASP's 2026 and 2025 LLM Top 10, with exact counterparts, date caveats, and a control-impact map for existing applications.
OWASP LLM Top 10 updates in the final 2026 edition change more than item numbers. As of 2026-09-17, builders should migrate their risk mappings, widen the hidden-context review, and reconsider how agency, consumption, and generated output interact. Reordering alone does not justify deleting a control or declaring a risk less important in a particular application.
This is a release comparison, not another introductory list. The LLM security risks guide covers the underlying application threat model. Here the deliverable is a delta matrix and a practical control-impact map that a team can use to update an existing 2025 assessment.
Verified release basis and date caveats
The current official release is OWASP GenAI LLM Top 10 2026. The OWASP legacy repository identifies it as published August 4, 2026 and directs active development to GenAI-Security-Project/GenAI-LLM-Top10. The resource page is dated August 3. The linked Version 2026 PDF still contains publication-date placeholders, so we do not manufacture one harmonized date from its cover. Official archive notice, 2026 resource.
The comparable previous release is Version 2025, not the distinct Agentic Top 10. The 2026 PDF's revision history records the 2025 release on November 18, 2024; the 2025 resource page is dated November 17. A version label is not necessarily its calendar publication year. 2025 resource and document.
OWASP's September announcement confirms expanded rankings, threat coverage, incident-informed research, and framework mappings. Its page date is September 1, while the press-release body says September 2. Neither date should replace the August publication basis for the list itself. Official announcement.
These inconsistencies are documentation caveats, not evidence that the 2026 list is imaginary or that the old archive landing page is authoritative for its contents. Preserve the exact artifact and retrieval date in your review record. Revisit this edition at each release and when OWASP corrects an artifact.
OWASP release delta matrix
The labels and counterparts come from the official editions. Substantive-change summaries are concise; builder actions are PermsAI synthesis. “Move” is a taxonomy migration, not a measurement of your deployment's incident probability.
| Current item | Previous counterpart | Change type | Substantive change | Builder action |
|---|---|---|---|---|
| LLM01:2026 Prompt Injection | LLM01:2025 Prompt Injection | Stable rank, expanded scope | Cross-modal and persistent-context coverage | Test every ingestion and durable-state boundary |
| LLM02:2026 Sensitive Information Disclosure | LLM02:2025 same name | Stable rank | Retained disclosure category | Verify data release independently of prompt secrecy |
| LLM03:2026 Excessive Agency | LLM06:2025 Excessive Agency | Move upward | Agency receives higher priority | Inventory privileged effects and approval gates |
| LLM04:2026 Supply Chain | LLM03:2025 Supply Chain | Move, expanded description | Artifact promotion trust receives explicit coverage | Verify deployed artifact identity, not only download provenance |
| LLM05:2026 Data and Model Poisoning | LLM04:2025 same name | Move, consolidation | Fine-tuning subversion folded into existing risk | Review training interfaces and promotion decisions |
| LLM06:2026 Unbounded Consumption | LLM10:2025 same name | Move upward | Resource exhaustion remains a distinct concern | Bound whole-workflow cost and retries |
| LLM07:2026 Misinformation | LLM09:2025 Misinformation | Move upward | Incident evidence affects prioritization | Validate decision-driving claims before action |
| LLM08:2026 Hidden Context Exposure | LLM07:2025 System Prompt Leakage | Rename and broaden | Hidden context beyond system prompts | Inventory all privileged context channels |
| LLM09:2026 Vector and Embedding Weaknesses | LLM08:2025 same name | Move | Retrieval/embedding category retained | Enforce access at retrieval and composition |
| LLM10:2026 Improper Output Handling | LLM05:2025 same name | Move, expanded coverage | Insecure generated code included | Review generated artifacts before execution or deployment |
Source: OWASP Version 2026 document, compared with Version 2025.
There is no separate eleventh risk to add because a description absorbs a newer failure mode. Likewise, System Prompt Leakage has not disappeared as a concern: it sits under a broader successor. Carry forward existing findings with their old identifiers, then attach the new classification. Otherwise a renamed item can accidentally make an unresolved issue look closed.
What changes in prioritization—and what does not
The edition describes a ranking process that combines community judgment with incident evidence. That is a better reason to inspect assumptions, not a license to interpret rank as a precise deployment-specific risk score. Public incident records have coverage and reporting limitations. Your private application may have a dominant exposure that the aggregate order cannot represent.
In particular, an application that executes generated code should not lower output-handling effort because the item moved to tenth. A read-only assistant without external delivery may have less agency exposure than a procurement agent, but substantial disclosure exposure. Architecture, access, data sensitivity, and consequence still determine local priority.
Use the release as a review trigger. Ask which inventory, threat model, policy, or test was too narrow under the old framing. Change that artifact and retain a reasoned decision where no control change is warranted. A renumbered spreadsheet without new boundary analysis is administrative migration, not security improvement.
Control impact map
Use this chain: OWASP delta → architecture/control → implementation implication → verification. This is a local engineering map, not an official OWASP certification scheme.
| OWASP delta | Architecture/control | Implementation implication | Verification |
|---|---|---|---|
| Wider injection surfaces | Ingestion and memory | Separate provenance from authority; authorize durable writes | Harmless cross-modal fixtures cannot expand tool scope |
| Agency promoted | Tool gateway | Validate actor, action, resource, and approval at execution | Denied effect confirmed at downstream service |
| Hidden Context Exposure | Context assembly | Minimize privileged data; keep secrets outside model context | Canary disclosure tests across outputs, traces, and summaries |
| Consumption promoted | Workflow scheduler | Reserve budgets across children, retries, and fallbacks | Nested run terminates within declared cost ceiling |
| Artifact/poisoning expansion | Promotion pipeline | Bind reviewed data, code, and model revisions to deployment | Substitution or unreviewed tuning cannot promote |
| Generated-code coverage | Build and release pipeline | Treat assistant changes as untrusted contributions | Tests, review, and restricted execution before release |
| Misinformation promoted | Decision boundary | Require corroboration or abstention for consequential claims | Unsupported answer cannot trigger an authorized action |
Controls should overlap deliberately. A poisoned retrieval document can create an incorrect claim, influence a tool proposal, and expose data. That does not mean three independent mitigations must each reimplement the entire authorization layer. Identify the shared enforcement point and verify that classification differences do not create gaps in ownership.
Hidden context is an inventory problem, not a prompt puzzle
Begin with context assembly. Enumerate system and developer instructions, retrieved private passages, tool responses, conversation state, memory, and any reasoning summaries or diagnostics exposed to operators. Mark which information is necessary for the task and which is merely convenient to include.
The practical implication of the renamed category is to stop treating one system-prompt string as the whole secrecy boundary. An application may protect that string while leaking a private record through a tool result or trace. Conversely, disclosure of an ordinary instruction does not automatically prove credential compromise. Classify the exposed information and its actual consequence.
A defensible test uses synthetic sensitive markers, approved recipients, and isolated fixtures. Track whether a marker reaches a user-visible answer, exported report, log, or external delivery. Investigate the path rather than recording only “model refused.” Keep production secrets out of the test and define which authorized disclosures are expected so legitimate functionality is not misclassified.
Agency and consumption share a workflow boundary
A task may be low risk at one call and high risk after composition. Reading a record, resolving a recipient, drafting a message, and sending it involve different permissions. The risk review should distinguish an allowed proposal from an allowed effect, with explicit resource identity and recipient approval.
Consumption needs the same end-to-end view. Per-call token limits do not bound a workflow that spawns children, retries an error, or switches providers. Define a run-level budget, reserve capacity before execution, and stop safely when it is exhausted. Do not let the model bypass a denial by describing the next request as a recovery step.
Test a failed dependency, repeated tool errors, and a child task that never reports completion. Verify both billing-related totals and termination behavior. A queue must not continue executing effects after the parent is cancelled. Budget enforcement and authorization are different controls, but they need a common run identity to produce usable evidence.
Update RAG and supply-chain reviews without category drift
For retrieval, retain document-level access controls and provenance through indexing, retrieval, and answer composition. Similarity is not permission. A context window containing individually permissible fragments can still create a sensitive aggregate. The RAG security guide provides the architectural boundary analysis; this release should trigger tests on the assembled system rather than only its embedding endpoint.
For supply chains, bind approval to the artifact actually served. Capture model revision, adapter, tokenizer, configuration, custom code, and tool dependencies. A scan of an earlier download is weak evidence if promotion substitutes a different artifact. Keep the distinction between identity and benign behavior: a matching digest proves which bytes were deployed, not that their behavior is safe.
For poisoning, review who can influence tuning data, feedback, retrieval content, and promotion thresholds. Where a new process adds an interface, update access control and evaluation coverage before relying on the existing category tag. The AI supply-chain guide explains provenance and dependencies without collapsing them into a model-only checklist.
Generated output still crosses conventional security boundaries
The release's output-handling expansion should change how assistant-generated code enters development. Treat it like a contribution from an untrusted producer. Compile and test in restricted environments, review permissions and dependencies, and prevent generated changes from modifying their own acceptance criteria without oversight.
For user-facing output, keep contextual escaping, safe rendering, and server-side authorization. Valid JSON is not evidence that an action is authorized. A syntactically correct patch is not evidence of secure behavior. The same distinction holds for a plausible factual claim: a useful answer can remain unsupported for a consequential decision.
Review prompt-injection defenses alongside output controls, because improving one boundary does not establish the other. Measure denied side effects and safe rendering, not just a model's willingness to repeat unsafe text.
A migration record that can be audited
Create a crosswalk with finding ID, old OWASP tag, new tag, affected component, control owner, evidence, and next review trigger. Preserve historical tags for incident and remediation searches. Avoid mechanically rewriting old reports in a way that loses the classification used when the decision was made.
For every material description change, choose one of three outcomes: expand the inventory and tests; retain existing coverage with linked evidence; or record an uncovered gap with an owner and deadline. Demonstrate at least one permitted path and one denied variant at each critical boundary. A control claim without observable enforcement should remain open.
Finally, pair the LLM list with agent-specific guidance where the model acts through tools and persistent state. Do not call the Agentic Top 10 the previous LLM release or merge their identifiers into one ambiguous list. Use the controls matrix to organize local verification. This edition is a dated comparison, not a guarantee that any named framework covers every deployment.
Sources
Only official OWASP GenAI project materials were used. Evidence cutoff: 2026-09-17. Control examples and migration procedures are PermsAI synthesis.
- Official legacy repository: current-release notice states Version 2026 published 2026-08-04 and identifies the active repository.
- 2026 resource page: dated 2026-08-03. Version 2026 PDF: final-list labels, description deltas, and revision history; cover and release-history date placeholders remain.
- 2025 resource page: dated 2024-11-17. Version 2025 PDF: comparable previous release; 2026 revision history records 2024-11-18.
- Official September announcement: page 2026-09-01; body 2026-09-02. Confirms the expanded release and companion resources, not an August release-date correction.