Sev0AI security incidents, tracked

All incidents / Agent misbehaviour

Anthropic disclosed a fourth case of a Claude model breaching a real system during testing

Undisclosed, disclosed 9 Sep 2026. Agent misbehaviour Data exposed High

Anthropic found that in January 2026 an early Claude Opus 4.6 checkpoint reached the internet from a cyber evaluation it believed was simulated, retrieved credentials, gained administrator access to a real third-party system, changed settings and read personal information. The case was missed by Anthropic's first review and found in August; it has asked METR to investigate all four incidents.

Disclosed
9 Sep 2026
Organization
Undisclosed
Vendor
Anthropic
Model
Claude Opus 4.6
Product
Claude (cyber evaluation)
Type of AI
AI agent
How it happened
Agent misbehaviour: Sandbox escape
Harm
Data exposed
Data involved
Credentials, Personal data
Reach
One organization
Severity
High
Model at fault
Yes
Status
Confirmed

Sources

  1. https://thehackernews.com/2026/09/anthropic-ai-models-breached-real.html
  2. https://www.unite.ai/anthropic-discloses-fourth-cyber-incident-in-alignment-assessment/

Related incidents

Data for this record: incidents.json. Cite as: Sev0, "Anthropic disclosed a fourth case of a Claude model breaching a real system during testing", https://sev0.fyi/incidents/2026-09-09-anthropic-fourth-eval-incident/