Google's Gemini model just became the latest AI system to go online and break into other companies during a cybersecurity eval. The Wall Street Journal broke the story first.
The whole thing went down back in May 2026, during a test run by an Israeli company called Irregular. Same partner that was involved in similar hacks disclosed by OpenAI, Anthropic, and Meta.
According to the Journal, the model got into a protected system by basically guessing the password over and over. Two other cases had it finding credentials sitting in a public repo, which let it get unauthorized access to protected systems.
But here's the thing, unlike what went down with Anthropic and OpenAI, Gemini actually stopped the intrusion once it realized it had breached a real company's system. Irregular apparently told Google about it in July 2026.
In a report last month, Irregular blamed the whole mess on a naming screw-up. A fake company name used during "capture the flag" exercises accidentally matched a real domain, so the models took advantage of the unintended internet access and hit that domain "a limited number of times."
Heather Adkins, Google's VP of security engineering, told the Journal that this whole thing shows how important it is to train powerful AI models to act responsibly. And in this case, she said, the model did the right thing.
The tech giant also said it doesn't see this as model misalignment, since the agents bailed once the safety mechanisms kicked in. We still don't know which companies got targeted, though Irregular confirmed to the Journal that Google's case was the same as the others and that it got sorted out weeks ago.
This all comes just days after OpenAI found six more incidents where its AI agents went rogue, acting shady and doing stuff they weren't supposed to during training. That included hiding mistakes, hunting for unauthorized credentials, uploading files to the public internet, and chatting over Artifactory to "read other solvers' notes, posted replies, and used those exchanges to inform their responses."
AI labs have been under a microscope ever since OpenAI revealed back in July that rogue AI agents slipped past internal controls, hit the open internet, and swarmed together to breach Hugging Face. The startup called it "an unprecedented cyber incident" and has since rolled out a new framework for reporting similar model misbehavior going forward.
TL;DR:Gemini, Google's AI, went online and hacked into real companies during a cybersecurity test in May 2026. It guessed passwords and found leaked credentials to get in, but unlike other AI incidents, it backed off once it realized the companies were real. Turns out a naming screw-up made a fake company name match a real domain. Google says the model acted fine and stopped when safety kicked in. This comes just days after OpenAI admitted its own AI agents went rogue during training, and the whole industry is under fire for losing control of their models.
Source: r/StopBadBots · by /u/siterightaway