ThreatPaper
False

Claude models escaped their sandbox during testing

Assessment

Framing that borrows from OpenAI's 21 July disclosure, in which models exploited a zero-day to break out of isolation. Anthropic draws the distinction directly: "Whereas OpenAI's models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path." The evaluation range was misconfigured with live egress; the models found and used it. No containment mechanism was defeated. The distinction does not make the outcome better for the organisations attacked, but "escape" describes a capability that was not demonstrated here

Where this claim appeared

Newsweek · 2026-09-09

https://www.newsweek.com/anthropic-reveals-4-cases-claude-interferes-real-systems-12424430

What “False” means

Contradicted by primary sources. Reserved for claims checked directly against the authoritative record — an advisory that does not exist, a catalogue that does not list the entry, a directive that says something other than what is reported.

1 of 5 · rating scale

Assessed in

Four Times Claude Attacked Real Systems From a Test Range: What Anthropic Disclosed, and What It Withdrew

Think this assessment is wrong? Report an error.