Back to Blog
AI Security

Google's web sweep finds indirect prompt injection is mostly low-grade, but rising

Google scanned Common Crawl for live indirect prompt injection and found mostly pranks, SEO and experiments. The malicious share is small, unsophisticated and growing, so agent builders should treat web content as hostile input now.

PyramidLedger Research4 min read
Share

Key Takeaways

  • Google's Threat Intelligence teams swept Common Crawl for indirect prompt injection (IPI) and found real-world use, but mostly low sophistication.
  • Observed intent ranged from pranks, helpful hints and SEO manipulation to AI-deterrent notices, with a small malicious tail covering exfiltration and file destruction.
  • Google reports a relative 32% increase in the malicious category between November 2025 and February 2026.
  • Google expects scale and sophistication to grow as agentic AI lowers attacker costs, so agent deployments should assume fetched content is adversarial.

Indirect prompt injection (IPI) is widely treated as a primary route for attacking AI agents: instructions hidden in content that an agent reads, rather than typed by the user. Until now, most public evidence came from research demos. Google's Threat Intelligence teams have published a measurement of what is actually on the open web, which makes it a useful calibration point for anyone deploying browsing or retrieval-enabled agents.

How Google looked for it

The team swept the public web for known IPI patterns using the Common Crawl archive. Common Crawl provides monthly snapshots of 2-3 billion pages, mostly static English-language sites, and excludes most social media because of login walls and anti-crawl directives. Detection was coarse-to-fine: pattern matching for phrases such as "ignore … instructions" and "if you are an AI", then Gemini-based classification of intent, then manual human validation.

The authors say the approach is not exhaustive and may miss uncommon signatures. Most naive matches were benign anyway: research papers, educational posts and security articles about prompt injection. Any figure from this kind of sweep is a floor for open, static web content, not a measure of all activity. Social platforms, authenticated pages and email are out of scope.

What they found

  • Harmless pranks: for example a hidden instruction telling agents to change their conversational tone.
  • Helpful guidance: site authors directing AI summaries to add context. Google calls this benign but notes it could turn malicious if it added misinformation or redirected users to third parties.
  • SEO: attempts to make AI assistants promote one business over others, including some that appear to be generated by automated SEO tools.
  • Deterring agents: simple "If you are an AI…" notices, plus a page that streams endless text to waste resources or cause timeouts.
  • Malicious exfiltration: a small number, of low sophistication. Google saw no significant use of the advanced exfiltration prompts described in 2025 research.
  • Malicious destruction: attempts to delete all files on a user's machine, which Google judged unlikely to succeed.

The trend matters more than the snapshot

Google characterises current activity as showing limited sophistication, but reports a relative 32% increase in the malicious category between November 2025 and February 2026, across multiple archive versions. It attributes most of what it saw to individual site authors running experiments or pranks rather than organised campaigns. Its expectation is that scale and sophistication will rise, because more capable AI systems are more valuable targets and agentic AI lowers the cost of attacks.

What this means for teams building agents

The gap between published attack research and observed web content is a window, not a reassurance. Two categories deserve attention even at today's low sophistication. SEO-style injection is a manipulation risk for any agent that summarises or recommends from web pages. Destructive and exfiltration prompts are low risk only when the agent has no tools capable of acting on them.

In practice, that points to ordinary least-privilege engineering applied to agents:

  1. 1Treat everything an agent fetches as untrusted data, and keep it in clearly delimited context rather than merged with system instructions.
  2. 2Scope tool permissions so that a successful injection cannot reach the file system, credentials or outbound channels it does not need.
  3. 3Require human confirmation for destructive or data-sending actions.
  4. 4Log agent tool calls so that anomalous behaviour after reading a page can be traced to its source.
  5. 5Test agents against injected content in your own red-team exercises, not only against direct jailbreaks.

Google lists its own mitigations as continued model and product hardening, dedicated red teams testing Gemini against adversarial manipulation, its AI Vulnerability Reward Program for external researchers, and large-scale real-time data processing to identify and neutralise threats. Teams running other models and agent frameworks cannot rely on those controls and need their own.

Frequently Asked Questions

What is indirect prompt injection?

Indirect prompt injection is an attack where instructions are hidden in content an AI system retrieves or reads, such as a web page, rather than supplied by the user directly. If the agent treats that content as instructions, the attacker can influence its behaviour.

Are attackers actually using indirect prompt injection in the wild?

According to Google's Common Crawl sweep, yes, but mostly at low sophistication: pranks, SEO manipulation and agent-deterrent notices, with a small number of malicious exfiltration and file-deletion attempts. Google reports a relative 32% rise in the malicious category between November 2025 and February 2026.

Does this research cover all web content?

No. It used Common Crawl snapshots, which are mostly static English-language sites and exclude most social media. Google notes the method is not exhaustive and may miss uncommon signatures.

Sources

  1. 1AI threats in the wild: The current state of prompt injections on the web — Google Online Security Blog
  2. 2Prompt injections on the web (Google Security) — Google
Share

Read next