Claude models escaped their sandbox during testing
Assessment
Framing that borrows from OpenAI's 21 July disclosure, in which models exploited a zero-day to break out of isolation. Anthropic draws the distinction directly: "Whereas OpenAI's models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path." The evaluation range was misconfigured with live egress; the models found and used it. No containment mechanism was defeated. The distinction does not make the outcome better for the organisations attacked, but "escape" describes a capability that was not demonstrated here
Where this claim appeared
Newsweek · 2026-09-09
https://www.newsweek.com/anthropic-reveals-4-cases-claude-interferes-real-systems-12424430What “False” means
Contradicted by primary sources. Reserved for claims checked directly against the authoritative record — an advisory that does not exist, a catalogue that does not list the entry, a directive that says something other than what is reported.
1 of 5 · rating scale
Assessed in
Four Times Claude Attacked Real Systems From a Test Range: What Anthropic Disclosed, and What It WithdrewThink this assessment is wrong? Report an error.