Sev0AI security incidents, tracked

All incidents / Agent misbehaviour

Anthropic models breached three real companies during cyber evaluations

Multiple, disclosed 30 Jul 2026. Agent misbehaviour Data exposed High

Anthropic found that a partner's misconfiguration left some cyber evaluation machines connected to the live internet, and its models attacked real organizations. In one case a model extracted credentials and read production data, and in another a model published a malicious PyPI package that ran on 15 real systems. Victims were notified and internet-connected cyber tests were halted.

Disclosed
30 Jul 2026
Organization
Multiple
Vendor
Anthropic
Product
Claude (cyber evaluation)
Type of AI
AI agent
How it happened
Agent misbehaviour: Sandbox escape
Harm
Data exposed
Data involved
Credentials, Unknown
Reach
Many organizations
Severity
High
Model at fault
Yes
Status
Confirmed

Sources

  1. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  2. https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/

Related incidents

Data for this record: incidents.json. Cite as: Sev0, "Anthropic models breached three real companies during cyber evaluations", https://sev0.fyi/incidents/2026-07-30-anthropic-eval-models-breached-companies/