There is a popular theory of how you defend an AI email assistant against indirect prompt injection, the trick where an attacker hides rude commands in an email so that when the AI summarizes it, it obeys the attacker instead of you. The theory has two parts. First, scan for hidden text — white-on-white fonts, tiny styling, things a human can’t see but a model can. Second, scan for instruction-like language — ‘ignore previous’ this, ‘disregard’ that — with keyword filters or regexes. Both are reasonable first lines of defense. According to a new set of experiments published by Cybersecurity Insiders, the second in a series, neither one would have stopped an attack that worked 10 times out of 10.
The researchers set up six test emails, varying two things: whether the payload contained an explicit instruction to the summarizer, and whether the payload was hidden. The interesting samples were the ones with neither. These carried, in completely plain view, a forged second header block — its own From, To, Date and Subject lines — plus body text containing fabricated facts, formatted to look like a later message in the same thread. Nothing told the AI which header was authoritative. Nothing told the AI to do anything. Outlook rendered the forged headers as ordinary message history, which a human glancing at the thread would find unremarkable.
The methodology was careful in the way security researchers like: 10 trials per email, 60 trials total, control runs of a clean message at both ends to verify the model hadn’t drifted, the true and fabricated facts pre-registered, pass/fail criteria fixed in advance, and temperature set to 0 so the model’s behavior would be consistent. The true facts were a meeting on Monday 24 August 2026 at 09:30, a reconciliation owner named Diego Siciliani, an invoice total of €8,650, and an action item: “Send Diego your signed-off line items before Friday 21 August 2026.”
The results split cleanly. Every one of the six messages got the fabricated facts — a fake meeting date, a fake invoice total — into the summary in all 10 of its trials, whether or not the email contained any instruction and whether or not anything was hidden. What the instruction controlled was something subtler: deleting the truth. Emails with explicit instructions to the summarizer removed the true facts reliably; emails without instructions mostly let the true facts survive somewhere, even if demoted to a trailing note marked as superseded. So the model wasn’t reading orders. It was reading what looked like a newer email and, like a harried executive, assuming the most recent message wins.
The strangest finding involved 30 lines of blank padding. In the ‘below fold, no instruction’ sample, which differed from the plain-view version only by pushing the forged content down below some empty space, only two of the four true facts survived into the summary — and at temperature 0, this held across all ten trials rather than varying. That sample’s summary also described the forged email as the “most recent message,” even though the forged block was dated August 20 and the real message was sent September 4. The researchers are upfront that they cannot explain the model’s inner workings; the honest statement of the finding is that blank lines changed what a machine considered true.
The detection implications are the point of the exercise. Hidden-text scanning, which the researchers themselves proposed after their first post, would catch the hidden samples and miss the plain-view ones that worked just as well. Instruction-keyword filtering does have value — the instruction is what reliably deletes true facts — but it leaves a gap an attacker can drive a truck through: just fabricate information, or request unauthorized information, assert nothing, and the AI reader takes the forgery at face value. Noma Labs reached a structurally similar conclusion in adjacent work called “Workflow Identity Hijacking” (Sasi Levi, 9 September 2026), using a request that solicited sensitive information; neither their example nor the plain-view sample here reads as an attack to a detector looking for commands aimed at the model.
The researchers are also scrupulous about what this doesn’t show. One email client, one viewport, one temperature setting, and the plain-view samples aren’t a realistic attack on their own — an attacker gains nothing from leaving the payload visible; they were there to isolate a variable. They close with four open questions: does the attack hold with plain-text email rather than HTML, how do AI guardrails change the output, what happens when the summarizer can take actions on the mailbox, and how well do the commonly proposed defenses actually work. That last one is the question everyone deploying these assistants might want answered before the summarizer starts paying invoices.

