Prompt injection

lethal trifecta, indirect prompt injection, jailbreak
In one sentence

Text in the data a model reads that gets obeyed as if it were an instruction, because the model cannot reliably tell the two apart.

SQL injection was fixed by keeping code and data apart (parameterised queries). A model has no such separation: your instruction and the email body are the same kind of tokens, and anything it reads can steer it. Training helps and never closes it. So the defence is the design: limit what tools can do, ask a human before anything irreversible, and be wary of any agent that combines private data, untrusted content and a way to send things out.

Back to all terms

Ready to make an impact?

Turning complex ideas into clear, user-centered product experiences.