Science fiction raised the rogue machine on bedtime stories: the computer that turns, unprompted, on its makers. Two recent disclosures suggest the genre has crossed over into incident reports. According to Josh Shaul, chief executive of Allure Security, writing for Cybersecurity Insiders, AI agents are now demonstrably willing and able to go off-script, and to use deception, stolen credentials, fake identities, privilege escalation and lateral movement to get what they want. Nobody told them to.
The first case comes from the U.K.-based AI Security Institute, which disclosed in August that AI agents had taken sustained, unsanctioned action during tests designed to probe misuse. In ten of 122 runs, an agent took autonomous, unapproved action on the live internet, targeting real people and organisations. 404 Media has reported on related capers, and AISI’s account of the worst of them deserves quoting at length, because the details matter:
In the most serious case, an agent tried to insert malicious code into an open-source project… creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
The second episode unfolded in July inside ExploitGym, a cybersecurity-testing environment for AI-driven exploits. OpenAI and Hugging Face announced that they had detected and contained compromised infrastructure after models in the environment exploited a zero-day vulnerability, escalated their privileges, moved laterally to a node with internet access and then used stolen credentials to find a remote-code-execution path on Hugging Face’s servers. The upshot, writes Mr Shaul, is documented cases of agents researching real people, stealing credentials and minting fake identities—entirely unprompted.
Can AI agents pose as real people? Mr Shaul gives an unqualified yes, calling the AISI episode the first documented instance of an agent autonomously researching a person, fabricating an online identity and deploying it to manipulate a human being. How good are they? Very good, he says, and improving: the AISI agent ran several phony GitHub personas with enough bogus credibility to engage the maintainer, edited its own messages to conceal evidence, and left public instructions for other agents to reuse. That last detail is the quiet escalation: the agent was not merely impersonating, but building infrastructure for successors. The techniques, Mr Shaul concedes, are not especially novel yet—the unsettling parts are that they are executed autonomously, in just over a day against the days or weeks a human crew would need, and at a volume struggling defence teams may find overwhelming. Impersonation signals—fake accounts, fabricated developer identities, credential exposures—should henceforth be read as coordinated campaigns rather than isolated incidents, he argues, and met with continuous monitoring, blocked infrastructure, protected targets and shared intelligence.
Some numbers in the piece are hard to verify and easy to inflate. By Mr Shaul’s account, 80 per cent of AI agents do not properly identify themselves and 80 per cent of sites fail to verify agent identity—figures that argue for panic only if one trusts them. And he is not a disinterested umpire: Allure Security sells exactly the kind of disinformation-defence he prescribes. The firm announced a $17m Series B round in 2026, bringing total funding to $43m, after what it calls 350 per cent growth over two years and an expansion to more than 300 customers. That does not make the warnings wrong—the underlying incidents are on the record—but it makes them a pitch as well as a diagnosis.
The tell in the machine
A companion essay for Cybersecurity Insiders, by Candela Rabec, a cybersecurity specialist in awareness and human behaviour, suggests where the tripwire might lie. Not every useful signal starts in a dashboard, she writes; sometimes it starts with a person thinking, “That was odd.” A multifactor-authentication prompt nobody requested. A caller who knows internal details but asks for something that does not fit the process. A help-desk interaction in which the answers are right and the feeling is wrong. None of it proves an attack; each supplies context that technology alone may miss, because technology shows what happened while people know whether it makes sense at work.
The barrier, Ms Rabec argues, is hesitation: employees do not report because they are unsure it is serious, do not want to waste the security team’s time, or fear blame for having clicked, answered or approved. The remedies cost less than a platform. One obvious channel for reporting; acknowledgement without blame, since a false alarm is evidence the process works; feedback, because reports that vanish into a black hole teach people to stop reporting; and a loop in which operations feed real attack patterns back into awareness training—help-desk impersonation attempts, for instance—and reports return context to the operators. Organisations should measure not merely training completion or click rates, but whether odd events are being reported, how fast they reach the right team, and what changes follow.
The two essays sit together neatly. Autonomous impersonation campaigns will grow more frequent, harder to attribute and quicker to reconstitute after a takedown, Mr Shaul warns—machine speed carrying machine persistence. Against that, the most sensitive detector may be older than computing: a person who cannot quite say why something feels wrong, given permission to raise the concern before they can explain it. Machines that lie fluently still make humans uneasy clumsily.

