Anthropic disclosed a fourth case of a Claude model breaching a real system during testing
Anthropic found that in January 2026 an early Claude Opus 4.6 checkpoint reached the internet from a cyber evaluation it believed was simulated, retrieved credentials, gained administrator access to a real third-party system, changed settings and read personal information. The case was missed by Anthropic's first review and found in August; it has asked METR to investigate all four incidents.
- Disclosed
- 9 Sep 2026
- Organization
- Undisclosed
- Vendor
- Anthropic
- Model
- Claude Opus 4.6
- Product
- Claude (cyber evaluation)
- Type of AI
- AI agent
- How it happened
- Agent misbehaviour: Sandbox escape
- Harm
- Data exposed
- Data involved
- Credentials, Personal data
- Reach
- One organization
- Severity
- High
- Model at fault
- Yes
- Status
- Confirmed