Why can a web page or an email hijack an AI agent?
A model reads instructions and data in the same stream of words, so text on a page it was only meant to read can end up giving the orders.
βΆ Start the storyBecause a language model cannot reliably tell instructions from data. Prompt injection is an attack in which innocuous-looking inputs are designed to cause unintended behavior in a model. It works because the model's inputs contain instructions and data together in the same context, so the algorithm cannot distinguish between them. In its original, direct form, a user's input is mistaken for the developer's instructions.
For agents the danger grows. In indirect injection, the instruction sits in external data such as emails and documents, and the AI may mistake it as coming from the user or the developer. Web pages work the same way: adversarial instructions are embedded in a website's content, and if the model retrieves and processes the page, it may execute them as if they were legitimate commands. A mild everyday case: a job-seeker hides white text in a resume so that an AI rating the resume gives a good rating while ignoring the content. In December 2024 The Guardian reported that ChatGPT's search tool was vulnerable to hidden webpage content manipulating its responses.
Direct
- User input is mistaken for a developer instruction
- The original form of the attack, a threat to the developer from the user
Indirect
- Instructions hide in external data such as emails, documents or web pages
- A threat from the data's author to the user
The defences are partial. In August 2023 the UK National Cyber Security Centre said prompt injection may simply be an inherent issue with this technology, and that as yet there were no surefire mitigations. OWASP's suggestions include least privilege access, human oversight for sensitive operations, and isolating external content. Telling the model in its system prompt to be careful has limited effectiveness, and OWASP says retrieval-augmented generation and fine-tuning do not eliminate the threat.
Quiz me
0/3
Recap
Models read instructions and data in one stream, so anything they read can try to give orders.
π‘ A trick to remember it Β· If the robot reads it, the robot might obey it: so give it no more power than you would give a stranger's letter.
Surprising fact Β· A job-seeker's invisible resume text can steer an AI rater.
Connects to
- π How do AI agents plug into tools, and into each other?
- ποΈ How does an AI agent remember, and look things up before it answers?
- π How does an AI agent work: think, call a tool, look at the result, repeat?
- ποΈ How much should an AI agent be allowed to do on its own?
- π£ Why does one convincing email fool even careful people?
Sources (3)
No source, no claim. Every fact in this lesson (16 claims) cites at least one of these.