A scan of 481 million transcripts found no other incidents of similar or worse severity
Assessment
The scan was two-stage: a filter for signs of internet access such as public IPs and URLs, then review by Claude of the 9.2 million transcripts it flagged. It re-identified the four known incidents and found nothing worse. Two caveats are in Anthropic's own text. The July scan of 141,006 transcripts, also "agentic," missed one of the four. And the September report separately notes that its offline monitors "would have missed the Claude Mythos 5 incident, because the model's reasoning persuaded the monitor that the environment was simulated." A Claude-reviewed scan for Claude misbehaviour has a known failure mode. METR's independent review, with access beyond the incident window, is the check that has not yet reported.
Where this claim appeared
Anthropic · 2026-09-09
https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidentsWhat “Assessed, Not Confirmed” means
A named source states this as its own assessment, at its own stated confidence, rather than as established fact. Attribution to a nation state usually sits here. The assessment is real and reportable; treating it as settled is the error.
4 of 5 · rating scale
Assessed in
Four Times Claude Attacked Real Systems From a Test Range: What Anthropic Disclosed, and What It WithdrewThink this assessment is wrong? Report an error.