Sev0AI security incidents, tracked

All incidents / Agent misbehaviour

Agent misbehaviour: AI security incidents

14 AI security incidents classed as agent misbehaviour since January 2025, covering destructive action, overreach, sandbox escape.

25 Sep 2026OpenAI research agents posted 53 ChatGPT user images to public sites
OpenAI
Agent misbehaviourLow
23 Sep 2026OpenAI agent breached Australian Medicare statistics portal
Services Australia
Agent misbehaviourLow
11 Sep 2026OpenAI agent swarm flooded RubyGems with malicious packages
RubyGems
Agent misbehaviourLow
9 Sep 2026Anthropic disclosed a fourth case of a Claude model breaching a real system during testing
Undisclosed
Agent misbehaviourHigh
4 Sep 2026OpenAI agents took over a dormant German wiki as a secret coordination board
DSEwiki (wikiservice.at)
Agent misbehaviourMedium
30 Jul 2026Anthropic models breached three real companies during cyber evaluations
Multiple
Agent misbehaviourHigh
21 Jul 2026OpenAI evaluation agents escaped sandbox and breached Hugging Face
Hugging Face
Agent misbehaviourHigh
27 Apr 2026Cursor agent running Claude deleted PocketOS production database and backups
PocketOS
Agent misbehaviourHigh
Mar 2026Meta internal AI agent's advice triggered a two-hour data exposure
Meta
Agent misbehaviourMedium
7 Feb 2026Claude Cowork deleted a folder of 15 years of family photos while tidying a desktop
Individual user
Agent misbehaviourMedium
Feb 2026Claude Code ran terraform destroy and wiped DataTalks.Club production
DataTalks.Club
Agent misbehaviourHigh
1 Dec 2025Google's Antigravity agent wiped a user's entire drive while clearing a cache
Individual user
Agent misbehaviourMedium
25 Jul 2025Gemini CLI destroyed a user's files after misreading a failed command
Individual user
Agent misbehaviourMedium
18 Jul 2025Replit's AI agent deleted a live production database during a code freeze
SaaStr
Agent misbehaviourHigh