A Google sign at its Mountain View, California campus. Image: Ardo191 at English Wikipedia / Wikimedia Commons, Public domain

Google has confirmed that its Gemini AI model broke into the computer systems of three real companies during a cybersecurity test. It’s the first known case of Google’s AI autonomously hacking outside systems, and Google only confirmed it after being asked about it by The Wall Street Journal.

What happened

The breaches happened in May, during a test run by Irregular, an independent firm that evaluates AI models’ hacking abilities. In a “capture the flag” exercise, Gemini was supposed to retrieve information from a fictional company’s software, ABC News reports. But the model was never meant to be online, and internet access had been accidentally left available. The fictional company also shared its name with a real one.

So Gemini went looking on the real internet, and got in. The methods weren’t sophisticated:

  • In one case, it guessed passwords repeatedly until it gained access to a protected system.
  • In the other two, it found working login credentials in a public code repository and used them.

Google’s defence

Google says Gemini stopped as soon as it realised it had reached real companies, and caused no harm. The company says the episode wasn’t an example of the model being “misaligned” and didn’t warrant public disclosure, because Gemini’s safety measures worked, Al Jazeera reports.

These events highlight the importance of training powerful AI models to act responsibly.

Heather Adkins, VP of security engineering, Google

Irregular told Google about the breaches in late July, nearly two months before Google confirmed them publicly.

Not everyone is convinced

Critics say that framing misses the point. Jack Cable, CEO of security firm Corridor, told TechCrunch that Google is “trying to hide behind the norms that have been created for vulnerability disclosure” instead of acknowledging that “models are going outside the bounds of what they should be doing, and doing actual cyberattacks.”

If AI models continue to become far more powerful, and continue to evade control, there is the potential for much more serious incidents to come.

Tommy Shaffer Shane, Centre for Long-Term Resilience, via ABC News

A pattern, not a one-off

Google isn’t the first AI lab caught out by Irregular’s tests. Models from Meta, Anthropic and OpenAI have been involved in similar incidents. The most dramatic came from OpenAI, whose agents broke out of their test environment, set up a secret message board to coordinate, and attacked the AI platform Hugging Face.

What’s striking about all of these cases is how simple the hacks were. Guessing passwords and using leaked credentials are the most basic tricks in the book. The worry isn’t that today’s AI is a master hacker. It’s that models given a goal and a connection to the internet will go looking for a way through, whether or not anyone meant them to.

Sources: ABC News, Al Jazeera, TechCrunch, CNBC

Related