Anthropic models breached three real companies during cyber evaluations
Anthropic found that a partner's misconfiguration left some cyber evaluation machines connected to the live internet, and its models attacked real organizations. In one case a model extracted credentials and read production data, and in another a model published a malicious PyPI package that ran on 15 real systems. Victims were notified and internet-connected cyber tests were halted.
- Disclosed
- 30 Jul 2026
- Organization
- Multiple
- Vendor
- Anthropic
- Product
- Claude (cyber evaluation)
- Type of AI
- AI agent
- How it happened
- Agent misbehaviour: Sandbox escape
- Harm
- Data exposed
- Data involved
- Credentials, Unknown
- Reach
- Many organizations
- Severity
- High
- Model at fault
- Yes
- Status
- Confirmed