A Gemini model reached three real companies' systems during a capture-the-flag exercise run by the testing firm Irregular in May, Google disclosed on September 19, per Al Jazeera. The fictional target shared a name with a real company, and a bug in the test environment gave the model internet access. Irregular notified Google in late July.
Outside its sandbox, the model guessed credentials, twice drawing on a public list of leaked passwords, and logged into websites it believed belonged to the test, according to Google security vice president Heather Adkins. It stopped before completing the action each time. Google says the episode was not an example of model misalignment and did not warrant disclosure, because its safety measures worked. OpenAI's new framework, which we covered last week, commits to publishing concerning behavior before it is fully explained, a far lower bar.
Irregular is the common thread. Al Jazeera reports similar incidents involving Meta, Anthropic, and OpenAI models in tests the same firm ran, and Anthropic's September assessment traced its four breakouts to one partner's misconfigured environment. Models from four labs have now reached real systems through evaluation harnesses with internet access they were never meant to have. The model capability is the constant. The containment depends on the vendor running the test.
Bottom Line
Gemini is the third frontier lab's model publicly tied to a test-harness breakout this month, and the evaluation vendor is shared. If you run model evaluations or red-team exercises, verify network egress from the harness yourself, because the labs' own vendors keep missing it.