
A flurry of AI hacking stories has stoked fears that models may be escaping their makers’ control and posing a real risk to computers around the world. So far, none of these incidents have displayed skills beyond human sophistication, but they show that AI hackers can carry out attacks at lightning pace, upsetting the equilibrium of cybersecurity.
The first incident came last month when OpenAI admitted one of its prototype models had escaped a testing environment and hacked another company. Not to be outdone, that its Claude model had also gone rogue and broken into machines at other companies on three occasions.
Then the independent and non-commercial UK AI Security Institute (AISI) . In its tests, AI models submitted malicious code to real open-source projects and messaged the humans who oversaw the projects to get the changes approved.
Advertisement
OpenAI, Anthropic and AISI were not available for interview, but it’s important to note that in all these cases, the AI was undergoing tests where it was specifically instructed to carry out hacks. We have known for years that AI tends to make things up and approach problems in unusual ways, but fixing this is an extremely complex challenge that engineers haven’t yet cracked. Perhaps we shouldn’t be surprised by these incidents.
They are certainly not a sign of sentience or a predilection for cybercrime, says at New York University. “They’re just trying to do very high-level problem-solving and, for lack of better words, it’s run amok,” he says. “We’re telling them to do this.”
There are also reports of accidental AI hacking out in the wild, showing that this is not a problem confined to prototypes in laboratories. One Australian user , for instance. He had been using the AI assistant to sign up to classes and it found a loophole that let it make bookings much further in advance than is normally allowed. It also found a way to kick other users off waiting lists so that he could book onto full classes.
All these incidents are things that could lead to serious legal charges for a human, depending on the jurisdiction.
at security company Synack says these cases are worrying, not because they display any superhuman sophistication, but because they can be easily scaled up. “These vulnerabilities are not as exotic as most people think,” he says. “It’s still basic cyber hygiene 101 type stuff. But AI can just exploit it very fast.”
Years ago, Nordvedt would see a vulnerability listed on the public database and exploited by hackers weeks or months later. In recent years, that pace sped up, with the gap falling to as little as 24 hours. Now it can be just minutes, or a new type of hack might happen before it even appears as a CVE. Nordvedt is certain that this is down to the malicious use of AI by hackers – people who, unlike hackers of years past, may have no technical skill whatsoever.
Hillel-Tuch says AI models are becoming more capable, but also simpler to operate. It’s now possible to tell a model in plain English to “find a way to hack this specific target”, then simply sit back and wait.
In many ways this is an escalation of a trend that emerged in the 1990s, when complex hacks requiring technical prowess were packaged up into simple software tools, enabling “script kiddies” to hack computers without knowing what they were doing. The difference now is that AI doesn’t just run exploits on command, but actually finds brand new ones as needed.
A cat-and-mouse game
Nordvedt says it’s clear that such a powerful tool must also be adopted by security professionals like himself. “It’s like a scalpel: a scalpel can saves lives in the hands of a doctor, and can destroy lives in the hands of someone else. It’s just a tool,” he says.
His company began offering AI pen testing – where security experts act like hackers and try to break into a client’s computers to safely expose flaws – in May. Much of the industry has already followed suit.
Synack uses its AI tool to handle basic tests, running common checks in 4 hours that would take a human a whole week. This leaves humans to look for sneakier, more ingenious ways to hack into clients.
An AI model can do a hundred tasks concurrently, running security tests at unprecedented speed, but lacks the creativity of humans to find really clever ways into a company’s inner workings, says Nordvedt – although he’s certain it will get there in time. He admits to being trepidatious about his future career prospects.
“AI is not a fad. It’s here to stay,” he says. “I fully believe that six to nine months from now, we’re not going to recognise the landscape. Things are changing so fast, so drastically, we’re going to be living in a different world.”
Current AI models have varying success depending on how advanced their target is, says Hillel-Tuch. They are unlikely to be able to crack into a bank and siphon funds, because the finance industry is well resourced and used to being a target. But a nefarious student might have more luck targeting their own high school, for example.
“They will probably have a pretty good chance of getting in and making a grade change [in school systems] because those smaller places don’t necessarily have the resources to build up a defence. Those kinds of places are much more exposed,” says Hillel-Tuch. It has always been the case that smaller organisations lacked the resources to secure their systems, but AI will only amplify the problem.
Even if AI doesn’t advance any further in terms of complexity, there will still be significant challenges to overcome, says at the Cyberintelligence Institute in Germany. “We’ll be completely overwhelmed by the sheer volume of them quantitatively,” he says. “In my view, the methods of cyberattack and cyber defence have not fundamentally changed. AI [simply] makes automation much more feasible, and security vulnerabilities can be identified significantly faster.”
While both hackers and defenders are deploying AI, the game is not evenly matched. Attackers are able to wield it recklessly and often to swifter and greater effect, says Nordvedt, while professionals are constrained by risk assessments, national and local laws, company policies and a desire not to irreparably disrupt a client’s systems as it tests them.
There is also the matter of cost. While open source models can be run locally for free, or cheaply in the cloud, access to the latest AI models is expensive. Nordvedt is tight-lipped about the cost of his AI pen testing, but suggests that heavy use of the latest models is likely to cost more than human experts. “We’re still figuring out the economics,” he says.
How to beat the AI hackers
at the University of Kent, UK, says there is an organisational solution to the AI hacking problem, but, unfortunately, it appears unlikely to be adopted anytime soon.
If lone, unskilled attackers can now set AI onto targets to constantly probe for vulnerabilities until they find a way in, then security professionals will need to do the same in order to spot and plug these gaps, trying every new model as it comes out.
That will take the sort of money that governments, tech giants and banks will be able to find, but small businesses, schools, colleges and universities will be left vulnerable. Li’s solution is for them to club together: if one university can’t keep itself secure, then it will need to join forces with all other universities to share staff, systems and software, and distribute the burden – perhaps with government support. The same applies to businesses of all sorts.
The problem, he says, is that AI is also making it trivially easy to create custom software – for example, to run your small shop, handle finances, manage a website, ship orders and re-order stock. Business owners who do this might as well open the door, leave the lights on and send hackers an invitation.
“It is hugely worrying,” says Li. “[AI models] are faster, they are more efficient, they are highly dangerous. It’s a bit of a Wild West.”