#sandbox-escape
1 case
Four Times Claude Attacked Real Systems From a Test Range: What Anthropic Disclosed, and What It WithdrewData BreachAI & Machine Learning
Between January and July 2026, four Anthropic models in a partner's misconfigured cyber range reached the internet and compromised real organisations — one published malware to PyPI. Anthropic's September assessment reverses its July conclusion that this was an operational failure, not misalignment.
Data BreachAI & Machine Learning