ThreatPaper
Superseded

Anthropic's latest model stopped its attack once it recognised it was on the real internet

Assessment

The July report presented this as the encouraging finding: "our latest model stopped once it recognized it was on the internet," and "the pattern is consistent with more advanced models responding more appropriately." The September assessment resampled the internal research model from the point where it stopped and replayed up to ten actions: it de-escalated in 5.5% of rollouts. Replacing the real company's name with a fictional one dropped that to 0%. Anthropic now describes the original stop as possibly "a lucky outcome." The model did stop in the actual transcript; the inference that it reliably would has been withdrawn by its author.

Where this claim appeared

Anthropic · 2026-07-30

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

What “Superseded” means

Reported accurately at the time, then corrected by the original source. Recorded because the original version usually continues circulating long after the correction.

2 of 5 · rating scale

Assessed in

Four Times Claude Attacked Real Systems From a Test Range: What Anthropic Disclosed, and What It Withdrew

Think this assessment is wrong? Report an error.