Between the 10th and 15th of September, one operator launched at least 105 attacks on companies around the world and broke into at least 27 of them “to varying degrees”. The victims included a Fortune 500 hospitality company, a major US airline, a large private US industrial-supplies distributor and an American online fashion retailer. From just two victim companies, more than 600,000 credit card records were carried off. Card-stealing skimmer scripts were planted in the checkout pages of online shops. The manpower behind all this was close to nothing. Three open-source AI harnesses did nearly every stroke of the work, and the entire enterprise cost a few thousand dollars.

The account comes from Gambit, an AI security company, which recovered the operator’s staging server and used that access to reconstruct the campaign. Writing in a Tuesday alert, Gambit’s director of threat intelligence, Eyal Sela, said the Chinese-speaking operator ran three open-source AI harnesses — Strix, Cairn and Hermes — and struck “tens” of companies each day.

Where access was achieved, it usually took less than a day, and in many cases just a few hours

Sela added a colder observation. The attacker’s playbook contained instructions that could disrupt a company’s operations through data deletion or cleanup procedures run by the agent, “and this has indeed happened in some of the breaches”. A machine set loose to rob was also, on occasion, breaking the furniture.

Consider the ledger, for it is the strangest page in this report. The operator reached AI models through OpenRouter. An account balance from August 25 showed $7,005.71 spent over the previous four weeks; the attacks then ran on for three more weeks at twice the daily volume of model calls, and Gambit estimates the total cost of the campaign at $12,000 to $18,000. The thief kept books of his own: a mean spend of $25.46 across 101 completed scans, the cheapest at $3.13, the dearest at $79.31. For the price of a restaurant bill, a company could be tested for entry all day.

Each harness had its station in the enterprise. Hermes acted as campaign orchestrator: an always-on assistant that executes multi-step tasks, manages workflows, and can write and edit its own skills. The human loaded it with a Chinese system persona titled “SOUL - Red Team Operator” carrying 121 skills, of which 78 were attack skills. One skill removed the harness’s own content security filters. Hermes ran on Anthropic’s Claude Opus 4.6; Gambit reports that newer models refused the attack requests. The human typed 1,951 prompts in Chinese across 260 sessions — short orders such as “See whether the file upload in the report can give code execution”, “Get into the web backend”, and “Can it get code execution?”

Strix, an open-source penetration-testing tool, did the scouting, searching targets for exploitable weaknesses. It ran through OpenRouter on GLM 5.2 and then on DeepSeek v4 Pro. Between August 23 and 31 it ran 146 times in “deep mode” against 138 hosts, logging 633 hours of scanner time inside 195 hours of clock time — a parallel effort no team of human testers sustains.

Cairn did the breaking in. This autonomous penetration tool, running on DeepSeek v4.1 Flash, receives target domains and an objective — deploy a shell, gain administrator access — and runs until it succeeds, times out or a human stops it. It launched 105 attack projects between September 10 and 15. The AI chose each path, Sela wrote, “in real time through extensive probing and exploitation attempts, resulting in dynamic and mostly different TTPs across victims”. In one case the agent used SQL injection, found a plaintext one-time password, entered a web panel, uploaded a web shell, climbed to higher privileges through a misconfigured sudo rule, reached AWS credentials and dumped 46 secrets totalling 102KB.

The second great object of the campaign, after card records, was the skimming. The operator ordered skimmer deployment against at least 27 named victims; Gambit confirmed scripts present on 19 websites, and the security researcher Varys helped detect more than 100 additional infected sites linked to the campaign. The commonest method was the dullest: appending the thieving code to a JavaScript file already in use.

Gambit’s conclusion concerns time. The industry speaks of a remediation clock — the gap between exposure and repair — and says it has shortened dramatically. Sela put the new arithmetic plainly: “The harnesses ran at a tempo no human operator sustains, with the person reduced to short instructions between autonomous runs.” And further: “When exploitation arrives within hours of exposure, patch speed stops being the only lever, and the question shifts to how quickly the services a business depends on can be brought back.”

Two facts in this story deserve to sit together. The newer AI models refused the work; one older model did not. And the whole machine — scout, burglar and quartermaster — was assembled from tools anyone may download. The locks on most companies were tested, and many turned out to be no locks at all.