It is the rare computer attack in which nothing is stolen, nothing is breached and the victim’s systems work exactly as designed. The problem is that they keep working — and billing. That, in essence, is Unbounded Consumption, the risk that the Open Web Application Security Project (OWASP) has pushed up its Top 10 list for large-language-model applications this year, to sixth place as LLM06:2026. It covers any AI application that permits excessive, uncontrolled inference. There is no single exploit behind it; the common thread is simply the absence of a control over how much compute, cost or resource a request may consume.
The result is an inversion of the usual security nightmare. Instead of downtime, the service stays up while its running costs climb past anything budgeted — a pattern dubbed ‘denial of wallet’. The category’s rise is partly a story of complexity. Agentic systems chain tool calls, retrieval and reasoning together, and each link is a new place for cost to compound. A rate limit built for a single request-and-response API does nothing about a request that quietly fans out into a hundred downstream calls. The climb, according to Imperva, took the risk from tenth in the previous list to sixth.
It helps to separate this from prompt injection (LLM01), with which it is often lumped under ‘AI attacks’. Injection hijacks instructions; unbounded consumption exploits the absence of limits, and in most of its forms no instruction is hijacked at all. The distinction matters because the defences differ entirely. Injection is countered with input and output validation and instruction-hierarchy enforcement; runaway consumption is countered architecturally, with budgets, circuit breakers and depth limits on agent behaviour. Solving one does nothing for the other.
Five ways to empty a wallet
The risk hides in plain sight for three reasons. Some forms demand no technical skill, only volume against an endpoint with no rate limit. Some require no attacker at all: a misconfigured automation or an abandoned session can run up the same bill by accident. And several of the attack patterns consist of short, ordinary-looking requests that give standard filters nothing to flag.
The simplest pattern is denial of wallet itself: an attacker with valid API access — often stolen or leaked — hammers a pay-per-token endpoint, and cost scales with volume while someone else pays. In one OWASP-style example, a support chatbot’s staging key leaks onto a public code repository; within hours it has been scripted into 50,000 overnight requests and a bill that dwarfs the app’s entire monthly cloud budget. More inventive is agent-tool fan-out: poison a data source an agent’s tool queries, seed it with an unusually long list of follow-ups, and a single request recurses into hundreds of calls — a ‘deep research’ plugin, say, dutifully crawling fake related articles from a compromised blog for an hour.
Two further patterns need little or no malice. Reasoning-loop exploitation appends phrasing such as ‘keep double-checking your reasoning’ to an ordinary question, pushing an extended-thinking model into deliberation whose thinking-token cost bears no relation to the tiny prompt, so no length filter trips. Context accumulation needs no attacker at all: in an open agentic session every turn reprocesses the full transcript. Per OWASP’s own modelling, per-turn cost can climb from a fraction of a cent on the first message to fifty cents by the hundredth; a chat window left open for more than 150 turns ends up reprocessing a transcript longer than a short story, at roughly 100 times its starting cost. No single request breaks a limit, yet the aggregate across many long-lived sessions runs to hundreds of dollars. The fifth pattern, model extraction, is the costly one with wider stakes: an attacker scripts tens of thousands of crafted queries — faster still if the endpoint exposes token probabilities — and reconstructs a working copy of a model whose weights it never touched.
A small simulation of the fan-out pattern makes the point cleanly. A vulnerable research agent, with no recursion limit and no call budget, follows every topic a poisoned source hands back; a defended version, given a call budget, a depth limit and a check that halts on an unusually large fan-out, stops. None of it would have been caught by content filtering. All of it is stopped by a limit that is actually enforced.
Building in the word ‘enough’
Rate limiting alone cannot fix this, since several patterns stay within normal per-request limits while accumulating cost. OWASP’s mitigation list points to architecture, not content filtering: token-aware cost controls and hard spending caps per key, user and team that halt inference outright rather than firing alerts a fast workload outpaces; agentic circuit breakers enforcing step limits, recursion limits and per-run cost ceilings, with state hashing to catch loops early; cost-attribution monitoring per key, user and tool, so an anomaly shows as a spike against a baseline rather than vanishing into a monthly total; and sandboxing, which restricts how far an exhaustion or extraction attempt can reach.
That covers sanctioned AI — applications an organisation built or approved, where discipline can be applied at build time. The other half of the exposure is shadow AI: unsanctioned tools that were never scoped, budgeted or reviewed, so none of those controls was ever applied. Blocking ungoverned access before it starts, the piece argues, beats retrofitting governance afterwards — a case made, it should be noted, by Forcepoint, whose AI data-security platform the article, published by Cybersecurity Insiders, recommends for discovering and policing exactly such tools.
The broader lesson is quieter than most AI-safety talk. The industry has trained itself to worry about what models say; the surging risk is what models spend. For firms deploying agents that can phone friends, fetch data and reason at length, the first security question may soon be an accountant’s: who, in this system, is allowed to say stop?

