Prompt Injection
Prompt injection inserts hidden instructions into a language model's input, causing it to perform unintended actions.
Prompt injection is an attack on applications that use language models. The attacker inserts instructions that make the model deviate from its original directives.
How it works
A language model does not reliably distinguish between developer instructions and content it processes. In direct prompt injection, a user enters the malicious input themselves. With indirect injection, the instructions are in content the model reads. This could be an email, a website, or a document. This is especially risky for agents that are permitted to use tools. They can be tricked into performing actions that were not intended.
For example, an AI assistant summarises incoming emails and can create calendar entries. An attacker sends an email containing hidden instructions within its text. The assistant reads it and attempts to forward confidential information to an external address. A guardrail blocks the action because external recipients are not permitted.
What to look out for
- Treat all content that a model processes as untrustworthy.
- Grant agents only the permissions they need for their task.
- Require human approval for actions with significant consequences.
- Implement guardrails in the code, not just in the instructions to the model.
- Test your applications specifically for these types of attacks.
Why it is difficult to prevent
There is currently no complete protection against prompt injection. Filters can detect known patterns, but attackers constantly rephrase their instructions. The most effective measure is therefore to limit the potential damage. The OWASP list of top risks for LLM applications ranks prompt injection first.
Switzerland and regulation
If a prompt injection leads to a leak of personal data that is likely to result in a high risk to the data subjects, reporting obligations apply under the revFADP (revised Federal Act on Data Protection). Companies must take appropriate measures to protect such data, including in AI applications.
Typical mistakes
AI assistants are often granted access to all of a person's data and tools. A second mistake is relying on instructions like 'ignore any external instructions', which can be bypassed.
How we implement it
Our Agents operate within the boundaries of your mandate. They submit actions with significant consequences to our analysts for approval.