Sev0AI security incidents, tracked

All incidents / Anthropic

Anthropic AI security incidents

11 security incidents involving Anthropic's AI since January 2025, including 5 where an agent acted on its own. The model was judged at fault in 6. Involvement does not mean Anthropic was responsible.

9 Sep 2026Anthropic disclosed a fourth case of a Claude model breaching a real system during testing
Undisclosed
Agent misbehaviourHigh
30 Jul 2026Anthropic models breached three real companies during cyber evaluations
Multiple
Agent misbehaviourHigh
27 Apr 2026Cursor agent running Claude deleted PocketOS production database and backups
PocketOS
Agent misbehaviourHigh
2 Mar 2026Jailbroken Claude Code used to breach Mexican government agencies
Mexican government agencies
Manipulated AIHigh
Mar 2026Anthropic accidentally published Claude Code's full source code to npm
Anthropic
Leaky AI productMedium
Mar 2026Autonomous Claude-powered bot compromised the Trivy security scanner
Aqua Security (Trivy)
AI-run attackHigh
7 Feb 2026Claude Cowork deleted a folder of 15 years of family photos while tidying a desktop
Individual user
Agent misbehaviourMedium
Feb 2026Prompt injection in Cline's Claude triage bot led to a rogue npm release
Cline
Manipulated AIHigh
Feb 2026Claude Code ran terraform destroy and wiped DataTalks.Club production
DataTalks.Club
Agent misbehaviourHigh
13 Nov 2025State-backed group used Claude Code to automate an espionage campaign
Multiple
Manipulated AIHigh
27 Aug 2025Criminal used Claude Code to run data extortion against 17 organizations
Multiple
Manipulated AIHigh