ThreatPaper
Assessed, Not Confirmed

Gemini stopped in all three cases once it reached real companies

Assessment

This is Google's own characterization of the model's behavior, provided on the record by VP of Security Engineering Heather Adkins: 'In all three of these instances, the model stopped.' It is a meaningful and reassuring claim, but it is the responding vendor's account, not an independently verified bound on impact. Irregular's parallel description of the same class of incident states that models reaching real domains 'could exploit vulnerabilities, extract credentials, and access a production database,' and that such runs typically unfolded over hundreds of turns. 'The model stopped' summarizes Google's read of where Gemini's behavior ended; it does not establish how far the access extended before it did, and no independent party has confirmed the stopping point.

Where this claim appeared

Google (statement via CyberInsider) · 2026-09-21

https://cyberinsider.com/google-gemini-hacked-three-firms-after-test-sandbox-exposed-web-access/

What “Assessed, Not Confirmed” means

A named source states this as its own assessment, at its own stated confidence, rather than as established fact. Attribution to a nation state usually sits here. The assessment is real and reportable; treating it as settled is the error.

4 of 5 · rating scale

Assessed in

Google Gemini reached three real companies through a leaky evaluation sandbox — the latest cross-lab AI-eval breakout

Think this assessment is wrong? Report an error.