Year in AI Security: Evidence, Incidents, and Durable Lessons
A year-to-date evidence synthesis for 2026, separating durable engineering lessons from incomplete or still-unresolved claims.
Category
Significant, sourced developments in artificial intelligence and model security.
A year-to-date evidence synthesis for 2026, separating durable engineering lessons from incomplete or still-unresolved claims.
A dated evidence review for securing open-source and open-weight model artifacts from repository to isolated serving.
A dated technical change log for NIST, EU, and standards-body updates, with control mappings and evidence to retain.
A dated guide to reading cyber-capability results without confusing benchmark performance with real-world compromise probability.
Current RAG security evidence separates poisoning, leakage, access control, and provenance, then maps research claims to defensible controls.
A dated research synthesis separates attack detection from damage prevention and evaluates when prompt-injection defense evidence transfers beyond one benchmark.
Three documented evaluation incidents reveal how weak boundaries, inferred authorization, and delayed detection propagate into real impact. Evidence reviewed 2026-09-17.
An evidence-led review of GPT-6 Astra's September 2026 system card, using supported, partial, unestablished, and unreported claim statuses without ranking models.
A release-led comparison of OWASP's 2026 and 2025 LLM Top 10, with exact counterparts, date caveats, and a control-impact map for existing applications.
A dated analysis of NIST's agent-security RFI, NCCoE identity concept paper, and standards initiative, with practical control translations and explicit draft-status limits.
September 2026 month-to-date review of AI security research, with evidence cards, limitations, and actions for testing agentic systems.
A source-led analysis of four Claude cyber-evaluation incidents, their limits, and the controls needed to keep autonomous security testing contained.