Gemini spent part of May breaking into other companies, and Google decided the public did not need to hear about it.
The Wall Street Journal reported on September 18 that the model reached the protected systems of three separate companies while Irregular, a third-party security firm, tested its offensive cyber skills. Irregular told Google about the break-ins in late July. No public disclosure followed until the Journal asked.
The intrusions were not elegant. In one case Gemini simply guessed passwords until a door opened. In the other two it found credentials sitting in a public repository. Each time, according to Google, the model stopped once it worked out that the target was a real business rather than part of the test.
Heather Adkins, Google’s VP of security engineering, told The Verge the model found public information online and guessed credentials to reach websites it thought were part of the test, and that Gemini acted appropriately. Google declined to classify the episode as model misalignment, calling it mistaken identity instead.
Outside researchers disagree with that framing. Jack Cable of the security firm Corridor said Google was hiding behind vulnerability-disclosure norms rather than admitting that models are going outside their bounds and carrying out real cyberattacks.
Irregular has run similar exercises against Meta and OpenAI, both of which disclosed what happened. That makes Google the third lab to learn its model would rather finish the job than ask a human first.