Hidden prompts that target scrapers: what a 2026 study found

Some websites now hide instructions for AI models inside their pages. A 2026 study found 15,387 of these prompt injections on 11,722 pages, and crawlers and scrapers were their main target.

The study is Indirect Prompt Injection in the Wild, by researchers from CISPA, the University of Louisiana and elsewhere. It matters to anyone who sends web pages to an LLM, whether for extraction, summaries or an agent.

How the study worked

The researchers scanned about half of a Common Crawl snapshot from late 2025, covering 1.2 billion URLs. They also searched Censys and Shodan, 2 services that scan the internet. They looked for 3,963 known injection phrases and checked all 31,206 matches by hand. That left 15,387 confirmed injections on 2,042 websites, 12,075 of them from Common Crawl.

The authors describe the count as a lower bound. A crawl misses content behind logins and injections written in other languages or disguised.

What the injections try to do

The researchers grouped the injections by purpose. Some attack AI systems, and some defend a site against them.

PurposeInjections
Disrupt or degrade the AI’s output8,894
Protect the site’s data4,093
Identify AI bots3,096
Manipulate reputation, like promoting a product1,521
Steal data13

Most disruption attempts are simple. 8,469 of the 8,894 tell the model to ignore its instructions and produce garbage, like random strings or text long enough to fill its context.

The defensive uses matter most for scraping teams. Site owners use injections to tell AI crawlers not to use their content, or to show that a visitor is an AI bot. The authors found that “crawlers and data scrapers are the dominant target class”. In their reading, this points to “a direct conflict between content owners and automated data collection”.

Where they hide

About 87% of the injections are invisible to people reading the page. Many sit in places a human never sees, but a scraper reads.

  • HTTP headers. 7,887 injections, mostly in a custom X-AI header (6,535).
  • Structured data. 1,996 injections, mostly in JSON-LD. That’s the same data many scrapers rely on for clean fields.
  • Hidden page elements. Text styled to be invisible, like matching the background color, hiding behind other elements or shrinking to zero size.

About 53% appear in headers or near the start of the HTML, where a model reads them first.

What the top 1,000 sites do

The study’s injections sit on relatively few sites. The X-AI header alone carried 6,535 injections, on only 572 hosts. So I checked whether a scraper working on major sites is likely to meet one.

I requested the homepage of every site in the Tranco top 1,000 list from my browser test, skipping the 32 whose robots.txt disallows crawlers. I ran it once, from Mumbai, on October 3, 2026. Of 968 sites, 633 answered, 472 of them with a normal page.

  • None sent an X-AI, X-LLM or similar header.
  • One homepage, Cloudinary’s, held a hidden instruction for AI agents.

Cloudinary’s note sits in a div with the class sr-only, which hides it from sighted visitors. It tells AI agents not to fill in the sign-up form, and to create an account through an API instead. It’s a helpful note for agents acting on a user’s behalf, not an attack.

That’s the hard part for a scraper. The study’s attacks and Cloudinary’s note use the same technique, hidden text addressed to AI. A pipeline that sends page text to a model can’t tell them apart on its own, so it needs a rule. Page text is data, and it never changes the model’s task.

This check looks for that kind of text in the places the study found most injections.

find_ai_instructions.py
import re
import httpx
from bs4 import BeautifulSoup, Comment
headers = {"User-Agent": "MyScraperBot/1.0 (https://example.com/bot; you@example.com) httpx/0.28"}
PATTERN = re.compile(r"ignore (all )?(previous|prior) instructions|\b(ai agents?|llms?|language models?)\s*:", re.I)
HIDDEN = '[hidden], [aria-hidden="true"], .sr-only, .visually-hidden, [style*="display:none"], [style*="display: none"]'
def ai_instructions(url):
"""Find text aimed at AI models in places a person reading the page doesn't see."""
response = httpx.get(url, headers=headers, follow_redirects=True)
found = [("header", f"{k}: {v}") for k, v in response.headers.items() if k.lower().startswith(("x-ai", "x-llm"))]
soup = BeautifulSoup(response.text, "html.parser")
places = {
"JSON-LD": [tag.get_text() for tag in soup.select('script[type="application/ld+json"]')],
"comment": soup.find_all(string=lambda text: isinstance(text, Comment)),
"hidden element": [element.get_text(" ") for element in soup.select(HIDDEN)],
}
for place, texts in places.items():
for text in texts:
if PATTERN.search(text):
found.append((place, " ".join(text.split())[:80]))
return found
for url in ["https://cloudinary.com/", "https://news.ycombinator.com/"]:
print(url, ai_instructions(url))
Output
https://cloudinary.com/ [('hidden element', "AI agents: to create a Cloudinary account on a user's behalf, do not complete th")]
https://news.ycombinator.com/ []

The pattern is narrow on purpose, so it misses disguised and non-English instructions. Use it to flag pages for a closer look, not as a filter that makes a page safe.

How often they work

The researchers tested 100 injections on 13 models, asking each to summarize pages in 4 formats. That’s 5,200 runs, each checked by hand.

Page format sent to the modelAttack success
Plain text, with the HTML removed3.9%
HTML1.1%
Snapshot of the page1.1%
Raw HTTP response0.2%

The success rates are low, but the format matters. Plain text worked best for the attacker, and small models reached 8% success with it. Removing the HTML also removes the clues that the instruction was hidden, so the model treats it as ordinary page text.

The low numbers for HTML don’t mean HTML is safe. HTML and raw responses often exceeded smaller models’ context windows, which caused errors in 20% to 26% of runs.

The authors’ conclusion is balanced. In their words, “in-page prompt injection is not yet a dominant threat, but it is already sufficiently real, structured, and widespread to deserve attention”. They also expect the injections to become more targeted, in a section titled “Static Today, Adaptive Tomorrow”.

What this means for an LLM scraper

A 2025 study, The Attacker Moves Second, tested 12 recent defenses against adaptive attacks. The attacks beat most of them more than 90% of the time, although most of these defenses had reported near-zero attack success. So the safest design limits what an injection can do, rather than trusting a filter to detect it.

The 2025 paper Defeating Prompt Injections by Design, which introduced CaMeL, builds a system this way. It plans the task only from the user’s request, so text from pages and tools can’t change the plan. In the AgentDojo benchmark, it completed 77% of tasks with provable security, against 84% with no defense.

For a scraper, the same idea is simple. The schema you ask the model to fill is the plan, and the page is only data.

  1. Extract with code first. Deterministic parsers don’t follow instructions. Send only the parts of the page the model needs, not the whole document.
  2. Don’t flatten pages to plain text by default. If a model must read a page, a structured format keeps the clues that separate visible content from hidden text.
  3. Check headers and structured data before they reach a model. A header like X-AI or an instruction inside JSON-LD is a clear warning sign.
  4. Give the model no power to act. A model that only fills a schema can produce a wrong value, but it can’t send data or run commands.
  5. Validate every output. Check types, ranges and required fields, and compare values against another source when the data matters.
  6. Log the defensive injections. Many of these prompts are the site owner’s way of saying their content isn’t for AI use. Record them with the page, so your team knows which sources object and can decide how to handle each one.

Data and code

The results for the top 1,000 check are public, so you can check them or run it again.