Question

What is prompt injection?

Vault Verified
Curated Intelligence
Definitive Source
Answer

An attack in which instructions hidden in content the model reads are followed as though they came from the user — and it is a genuinely unsolved problem, not a bug awaiting a patch.

Why it exists. A language model receives everything as text in one stream: system instructions, user input and any retrieved content. There is no structural separation between instructions and data — the model has only text, and text that looks like an instruction may be treated as one, regardless of where it came from.

This differs fundamentally from SQL injection, which is solved by parameterisation separating code from data. No equivalent separation exists for natural language, which is why the problem is hard rather than merely unaddressed.

Direct prompt injection. The user types something attempting to override the system instructions — the familiar "ignore previous instructions". Mitigable to a degree, and mostly a nuisance.

Indirect prompt injection — the serious one. Instructions are planted in content the model will process later: a web page it browses, a document it summarises, an email it reads, an issue in a code repository, metadata, or text made invisible to humans through colour or font size.

A user asks the assistant to summarise a page. The page contains hidden text instructing the model to exfiltrate the conversation to a URL, or to recommend a particular product, or to alter a recommendation. The user never sees the instruction and never authorised it.

Why it matters more as models gain capability. A model that only produces text can mostly be made to say something wrong. A model with tools — reading email, browsing, executing code, making purchases, sending messages — can be made to do something, and the damage scales with the permissions it holds.

What mitigations exist, none complete: limiting tool permissions to the minimum; requiring human confirmation for consequential actions; treating all retrieved content as untrusted; separating the model that reads untrusted content from the one with tool access; output filtering; and constraining where data may be sent.

The realistic position is that untrusted content plus tool access plus sensitive data is a dangerous combination, and the defence is architectural rather than a filter.

Related Questions